System

A system using voice analysis and natural language processing to assess employee stress and provide personalized advice and soothing music addresses the challenge of ineffective mental health communication, improving productivity and comfort in the workplace.

JP2026028815APending Publication Date: 2026-02-20SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024131431
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2026-02-20

AI Technical Summary

Technical Problem

Employees often struggle to communicate their stress and mental health concerns effectively, leading to decreased productivity and increased resignation, and conventional mental health care systems lack personalized solutions and relaxation opportunities.

Method used

A system that utilizes voice analysis and natural language processing to generate personalized questions, evaluates stress levels, and provides tailored advice and soothing music to improve mental health management.

Benefits of technology

Enhances employee comfort in discussing mental health issues, reduces stress, and improves productivity by providing personalized advice and relaxation through voice analysis and natural language processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026028815000001_ABST
    Figure 2026028815000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system for managing an employee's mental health, the system comprising: a generation unit configured to generate a question related to mental health; an acquisition unit configured to acquire an employee's voice response to the question; an analysis unit configured to analyze the acquired voice response and evaluate an emotion or a stress level; an advice generation unit configured to generate specific advice or a countermeasure based on an evaluation result obtained by the analysis; and a presentation unit configured to present the advice or the countermeasure and reproduce healing music.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In today's corporate environment, employee mental health is a critical issue. In particular, a lack of stress management and mental health care can lead to decreased employee productivity and even employee resignation. However, it is often difficult for employees to fully communicate their stress and concerns through direct human interaction. Furthermore, conventional mental health care systems struggle to provide personalized solutions or appropriate relaxation opportunities. Therefore, the present invention aims to provide a system that allows employees to talk more easily and effectively care for their mental health. [Means for solving the problem]

[0005] The present invention provides a system for managing the mental health of employees, which includes a generation means for generating questions related to mental health, an acquisition means for acquiring employees' voice responses to the questions, an analysis means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generation means for generating specific advice and countermeasures based on the evaluation results obtained by the analysis, and a presentation means for presenting the advice and countermeasures and playing healing music.

[0006] This system uses voice analysis and natural language processing to accurately assess employees' stress levels and provide appropriate advice and countermeasures. It also plays soothing music to enhance employee relaxation. As a result, employees feel more comfortable talking about their mental health issues through dialogue with the robot, effectively reducing stress.

[0007] "Mental health" refers to the psychological and emotional well-being of employees and is a concept that encompasses stress, anxiety, depression, and other symptoms.

[0008] "Generator" refers to a device or program for automatically generating questions related to mental health.

[0009] "Acquisition means" refers to a device or program that collects employees' voice responses to questions and records them as data.

[0010] "Analysis means" refers to a device or program for analyzing acquired voice data and evaluating emotions and stress levels.

[0011] The "advice generation means" refers to a device or program that automatically generates specific advice and countermeasures based on the evaluation results obtained by the analysis means.

[0012] "Presentation means" refers to a device or program for presenting the generated advice and countermeasures to employees and simultaneously playing healing music.

[0013] "Voice analysis engine" refers to software or algorithms used to scientifically analyze voice data.

[0014] A "natural language processing (NLP) engine" refers to software or algorithms that analyze the content of voice data on a text-based basis and extract emotions and keywords.

[0015] "Healing music" refers to music used to enhance relaxation and is created by generative AI.

[0016] "Feedback" refers to opinions and impressions provided by users after use. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention is a system for managing employees' mental health and providing specific advice and countermeasures. This system includes functions for generating questions related to mental health, acquiring and analyzing audio data, generating advice, and playing healing music.

[0039] System Configuration

[0040] The system consists of the following main elements:

[0041] 1. Server: This is the central data processing device that generates questions, analyzes voice, evaluates stress, generates advice, and generates healing music.

[0042] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[0043] 3. User: An employee who uses the system.

[0044] System Operation

[0045] 1. Question generation and retrieval

[0046] The server generates questions about the user's mental health, using generative AI to create a personalized list of questions. For example:

[0047] "What kind of stress have you been experiencing at work recently?"

[0048] "Are you getting enough rest?"

[0049] The questions are displayed in order below.

[0050] The terminal presents the generated question to the user by voice or text.

[0051] The user responds to the question by voice, for example, "I felt a little stressed during the meeting."

[0052] The terminal records the user's voice response and transmits it to the server.

[0053] 2. Voice analysis and stress assessment

[0054] The server analyzes the received voice data using a speech analysis engine to analyze the intonation, speed, and volume of the voice, and a natural language processing (NLP) engine to analyze the text content, which then evaluates the user's emotional state and stress level.

[0055] Specifically, a high intonation and fast speech rate is judged to indicate a state of tension. Additionally, if the answer contains keywords such as "stress" or "anxiety," the stress level is also assessed as high.

[0056] 3. Advice Generation and Presentation

[0057] Based on the evaluation results, the server generates specific advice and countermeasures. The AI ​​generator makes optimal suggestions based on the user's situation. For example,

[0058] "Light exercise is good. Walking and stretching are recommended."

[0059] "Increasing communication with other employees is also effective."

[0060] The terminal displays the generated advice and countermeasures to the user.

[0061] 4. Play healing music

[0062] While the advice is displayed, the server uses a generation AI to create healing music and send it to the device.

[0063] The device plays soothing music along with a screen displaying advice and suggested solutions, allowing the user to receive the presented information in a relaxed state.

[0064] Specific examples

[0065] In one scenario, the following specific exchanges take place:

[0066] 1. Device: "Have you felt stressed at work recently?"

[0067] 2. User: "Yes, writing the report was very stressful."

[0068] 3. Server: Performs voice analysis and NLP analysis and determines that the stress level is high.

[0069] 4. Server: Generates the advice, "Take regular breaks while writing your report. It's also a good idea to do some light stretching."

[0070] 5. Device: Displays advice and simultaneously plays relaxing, healing music.

[0071] This allows users to receive specific advice and learn how to reduce stress in a relaxed environment.

[0072] By using this system, employees can improve their mental health through dialogue with the robot, creating a comfortable working environment.

[0073] The processing flow will be explained below.

[0074] Step 1:

[0075] The server starts up and makes the user database accessible, preparing the entire system for operation.

[0076] Step 2:

[0077] The terminal displays a login screen to the user, providing a screen with fields for entering a user ID and password.

[0078] Step 3:

[0079] The user enters login information (user ID and password) and presses the login button. The login information is sent to the system.

[0080] Step 4:

[0081] The terminal sends the user's login information to the server, which performs authentication and, if successful, proceeds to the next step.

[0082] Step 5:

[0083] The server uses generative AI to generate a personalized list of mental health questions for the user, based on the user's past data and general mental health questions.

[0084] Step 6:

[0085] The device presents the generated list of questions to the user via voice or text, for example, "How stressful are you feeling at work these days?"

[0086] Step 7:

[0087] The user answers the question by voice. For example, the user might say, "I felt a little stressed during the meeting."

[0088] Step 8:

[0089] The device records the user's voice response and sends it to the server, where the user's response arrives as recorded data.

[0090] Step 9:

[0091] The server uses a speech analysis engine to analyze the received voice data, extracting voice intonation, speed, and volume characteristics.

[0092] Step 10:

[0093] The server uses a natural language processing (NLP) engine to analyze the content of the voice data and extract keywords related to emotions and stress.

[0094] Step 11:

[0095] The server integrates the results of the voice and text analysis to evaluate the user's emotional state and stress level, and generates an evaluation result for the next step.

[0096] Step 12:

[0097] Based on the stress assessment results, the server uses AI to generate specific advice and countermeasures, such as "Light exercise would be good. Walking and stretching are recommended."

[0098] Step 13:

[0099] The server sends the generated advice and countermeasures to the device, which then prepares the countermeasures to be provided to the user.

[0100] Step 14:

[0101] The device displays advice and solutions to the user, while soothing music is played in the background.

[0102] Step 15:

[0103] The device plays soothing music while providing advice and solutions to the user, creating an environment that enhances relaxation.

[0104] Step 16:

[0105] The device displays a screen requesting the user for advice and feedback on the effects of the healing music.

[0106] Step 17:

[0107] The user provides feedback via voice or text, for example, "I found the music very relaxing."

[0108] Step 18:

[0109] The device sends the user's feedback to the server, which uses the feedback data to generate questions and advice for the next time.

[0110] Step 19:

[0111] The server records the feedback and uses the data to improve the system and user experience in the future.

[0112] Example 1

[0113] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0114] In today's workforce, managing employee mental health is becoming increasingly important. Many employees are exposed to stress and pressure on a daily basis, often resulting in decreased productivity and health problems. However, the means to provide effective mental health care remain limited. In particular, it is difficult to quickly provide personalized advice and measures tailored to each employee's situation. Therefore, there is a need for a comprehensive mental health management system that can accurately assess employees' stress levels, provide appropriate advice, and combine relaxation elements such as healing music.

[0115] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0116] In this invention, the server includes a generating means for generating questions related to mental health, an acquiring means for acquiring employees' voice responses to the questions, an analyzing means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generating means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, and a presenting means for presenting the advice and proposed measures and generating and playing healing music, thereby making it possible to provide personalized mental health care to each employee.

[0117] "Generation means" refers to a device or program for generating questions related to mental health.

[0118] The term "acquisition means" refers to a device or program for acquiring a user's voice response.

[0119] "Analysis means" refers to a device or program for analyzing the acquired voice response and evaluating emotions and stress levels.

[0120] The "advice generation means" refers to a device or program that generates specific advice or countermeasures based on the evaluation results obtained by analysis.

[0121] The "presentation means" refers to a device or program that presents the generated advice and countermeasures to the user and also generates and plays healing music.

[0122] "Emotions and stress levels" refers to the emotional state the user is feeling and the associated degree of stress.

[0123] "Healing music" refers to music that is intended to relax the user and reduce stress.

[0124] A "generative AI model" refers to an algorithm or program that uses artificial intelligence to generate questions, advice, music, etc.

[0125] "Voice analysis" refers to the process of analyzing acquired voice data in terms of intonation, speed, volume, etc.

[0126] A "natural language processing engine" refers to a program or algorithm that converts voice data into text and analyzes the content.

[0127] This invention is a system for managing employees' mental health and providing specific advice and countermeasures. This system includes functions for generating questions related to mental health, acquiring and analyzing audio data, generating advice, and playing healing music. The main elements of the system are a server, a terminal, and a user.

[0128] System Configuration

[0129] 1. Server

[0130] The server is the central data processing unit, generating questions using a generative AI model and analyzing the voice data. It also generates specific advice based on the analysis results and creates healing music. Software such as "PRAAT" and "SpaCy" are used for voice analysis.

[0131] 2. Terminal

[0132] The terminal is a device operated by the user, and has the function of presenting questions sent from the server, receiving voice responses from the user, displaying analysis results and advice, and playing healing music.

[0133] 3. Users

[0134] Users are employees who use the system, answer mental health questions through their terminals, and receive the advice provided.

[0135] System Operation

[0136] The server uses a generative AI model to generate mental health-related questions that are personalized based on the user's previous answers and general mental health care knowledge.

[0137] The device presents the generated question to the user by voice or text, for example, "What has been making you feel stressed recently?"

[0138] The user answers the questions by voice, for example, "I felt very stressed writing the report."

[0139] The terminal records the user's voice response and transmits it to the server.

[0140] The server uses the speech analysis engine "PRAAT" to analyze the intonation, speed, and volume of the voice, and the natural language processing engine "SpaCy" to analyze the text content, which then evaluates the user's emotional state and stress level.

[0141] Based on the evaluation results, the server uses a generative AI model to generate specific advice and countermeasures for the user. For example, it might suggest, "It would be good to do some light exercise. Walking and stretching are recommended."

[0142] Once the evaluation results and advice generation are complete, the server uses the generative AI model to create healing music and send it to the device.

[0143] The device displays the generated advice and countermeasures to the user while simultaneously playing soothing music, allowing the user to receive specific advice in a relaxed state.

[0144] Specific examples

[0145] For example, in a scenario on one day, the following exchange takes place:

[0146] 1. Device: "What has been stressing you out lately?"

[0147] 2. User: "Yes, writing the report was very stressful."

[0148] 3. The device records this response and sends it to the server.

[0149] 4. Server: Conduct voice analysis and NLP analysis to identify high stress levels associated with "report writing."

[0150] 5. Server: Generate the advice, "Take regular breaks while writing your report. It's also a good idea to do some light stretching."

[0151] 6. Device: Displays advice and simultaneously plays relaxing, healing music.

[0152] Example prompt sentence:

[0153] "Please write a natural language description of a system that manages employees' mental health and provides specific advice. The system includes functions for question generation, voice analysis, advice generation, and soothing music playback. Please describe the process in detail along the following lines."

[0154] By using this system, employees can manage their own mental health more effectively and create a comfortable working environment.

[0155] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0156] Step 1: Question Generation

[0157] The server uses a generative AI model to generate questions related to mental health. At this time, the server obtains the user's profile information and past answers and creates a personalized list of questions. The input data is the user's profile information and past answer data, and the output is the generated list of questions. For example, the server generates a question such as, "What has been making you feel stressed recently?"

[0158] Step 2: Question posing

[0159] The terminal presents the questions sent from the server to the user. The input data is the list of questions sent from the server, and the output is a display screen or audio presented to the user. For example, the terminal may read out the question "What has been making you feel stressed recently?" or display it on the screen.

[0160] Step 3: Get the answer

[0161] The user answers the presented question by voice. The input data is the question presented to the user, and the output is the user's voice response. For example, the user might respond, "Writing the report was very stressful."

[0162] Step 4: Acquire and transmit audio data

[0163] The terminal acquires and records the user's voice response. Then, it transmits this voice data to the server. The input data is the user's voice response, and the output is the voice data transmitted to the server. For example, the terminal transmits voice data such as "I felt very stressed while writing the report" to the server.

[0164] Step 5: Audio analysis

[0165] The server analyzes the received voice data. For the analysis, it uses the voice analysis engine "PRAAT" and the natural language processing engine "SpaCy." The input data is the user's voice data, and the output is the voice intonation, speed, volume, and text analysis results. Specifically, the server analyzes the voice data "I felt very stressed writing the report" for voice intonation, speed, and volume using "PRAAT," and converts it into text and analyzes it using "SpaCy."

[0166] Step 6: Stress Assessment

[0167] The server evaluates the user's stress level based on the results of voice and text analysis. The input data are the voice and text analysis results, and the output is the user's stress level assessment. For example, if the speech is high-pitched, fast-paced, and includes the word "stress," the server will assess the user's stress level as high.

[0168] Step 7: Advice Generation

[0169] The server uses the generative AI model to generate specific advice and countermeasures for the user based on the evaluation results. The input data is the stress level evaluation result, and the output is the generated advice and countermeasures. For example, it generates advice such as, "Take appropriate breaks while writing your report. It is also recommended that you do some light stretching."

[0170] Step 8: Providing advice

[0171] The terminal displays the advice and countermeasures sent from the server to the user. The input data is the advice and countermeasures sent from the server, and the output is the display screen and audio presented to the user. For example, advice such as "Take appropriate breaks while writing a report" can be displayed on the screen and read aloud.

[0172] Step 9: Generate and Play Healing Music

[0173] The server uses a generative AI model to create healing music and sends it to the device. The input data is the stress assessment results, and the output is the generated healing music. For example, a song with a relaxing effect is generated. The device receives this healing music and plays it along with the advice display. The user can receive advice in a relaxing environment.

[0174] (Application example 1)

[0175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0176] In today's work environment, many employees are increasingly experiencing stress and emotional strain, which increases the risk of decreased productivity and poor mental health. Furthermore, virtual store shopping experiences often leave users feeling overwhelmed by numerous options and information, leading to stress. There is a need for an effective mental health management system that can improve these situations and help employees and virtual store users feel refreshed, improving their productivity and satisfaction.

[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0178] In this invention, the server includes a generating means for generating questions related to mental health, an acquiring means for acquiring the user's voice responses to the questions, an analyzing means for analyzing the acquired voice responses and evaluating the user's emotions and stress level, an advice generating means for generating specific advice and proposed measures based on the evaluation results, a presenting means for presenting the advice and proposed measures and playing soothing music, and a refreshing means for presenting refreshing advice to improve the user's experience in the virtual space. This enables practical support to reduce stress felt by the user while shopping in a virtual store and to refresh the user.

[0179] The "means for generating mental health-related questions" is a device or method that dynamically generates appropriate questions to assess the mental health status of a user.

[0180] The "means for acquiring the employee's voice response to the question" refers to a device or method for acquiring voice data in which the user answers the question and transmitting that data to the system.

[0181] The "analysis means for analyzing the acquired voice response and evaluating the emotion and stress level" is a device or method for analyzing voice data and evaluating the user's emotional state and stress level.

[0182] The "advice generating means for generating specific advice and countermeasures" is a device or method for generating appropriate advice and countermeasures for the user based on the analysis results.

[0183] The "presentation means for presenting advice and measures and playing healing music" is a device or method for presenting the generated advice and measures to the user and playing healing music at the same time.

[0184] A "refreshment means for presenting refreshment advice to improve the user experience in a virtual space" is a device or method that presents specific advice to refresh a user when the user feels stressed in a virtual store.

[0185] MODE FOR CARRYING OUT THE INVENTION

[0186] This invention provides a system for managing employee mental health and improving user experience in a virtual store. The following configuration makes it possible to evaluate the user's emotional state, generate optimal advice, and provide soothing music.

[0187] System Configuration

[0188] The system consists of the following elements:

[0189] 1. Server: A central data processing device that generates questions, analyzes voice, evaluates stress, generates advice, and generates healing music.

[0190] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[0191] 3. User: A user who uses the system.

[0192] question generation

[0193] The server generates questions related to the user's mental health, personalized using a generative AI model. For example:

[0194] It generates questions such as, "Have you felt stressed while shopping online recently?"

[0195] Acquiring audio data

[0196] The device presents the generated question to the user by voice or text. The user answers the question by voice. For example, the user might say, "Yes, there are so many products, I don't know which one to choose." This voice data is captured by the device and sent to the server.

[0197] Voice analysis and stress assessment

[0198] The server analyzes the acquired voice data using a speech recognition engine and a natural language processing (NLP) engine. The server analyzes the intonation, speed, volume, etc. of the voice to evaluate the user's emotional state and stress level. For example, a high intonation and fast speed of the voice may be determined to indicate a state of tension.

[0199] Advice generation and presentation

[0200] The server generates specific advice and countermeasures based on the analysis results. The generated advice provides specific support for the user's actions. For example,

[0201] Advice such as "Take a deep breath to relax. We recommend drinking tea to refresh yourself" is generated.

[0202] Playing healing music

[0203] The device not only displays the generated advice but also plays soothing music at the same time, allowing the user to receive the advice in a relaxed state.

[0204] Specific examples

[0205] For example, if a user using a virtual store responds that they "felt stressed," the server analyzes the voice data and detects a high stress level. As a result, the server generates advice such as "Take a deep breath to relax. We recommend drinking tea to refresh yourself," and the device displays this advice to the user while playing soothing music.

[0206] Prompt Sentence Examples

[0207] "Please assess this user's mental health status and provide appropriate advice: Yes, there are so many options that it's hard to know which one to choose."

[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0209] Step 1:

[0210] The server generates questions related to the user's mental health, using a generative AI model to create a personalized list of questions, such as, "Have you felt stressed while shopping online recently?" These questions are customized based on the user's past feedback and shopping history.

[0211] Step 2:

[0212] The device presents the generated question to the user via voice or text. The user can receive the question through either visual or auditory means. For example, the device might ask the user aloud, "Have you felt stressed while shopping online recently?" The user listens to the question and understands its content.

[0213] Step 3:

[0214] The user answers the question by voice. For example, they might say, "Yes, there are so many products, I don't know which one to choose." The user's voice is picked up through the device's microphone.

[0215] Step 4:

[0216] The terminal records the user's voice responses and converts the data into a data format for transmission to the server. The terminal converts the raw voice data as input into an audio file and transmits it to the server over the network.

[0217] Step 5:

[0218] The server analyzes the acquired voice data and converts the voice into text using a voice recognition engine. This text data is then analyzed using an NLP engine to evaluate the emotional state and stress level. For example, the intonation and speed of the voice can be analyzed to determine the state of tension and stress level.

[0219] Step 6:

[0220] The server generates specific advice and countermeasures based on the evaluation results. Using a generative AI model, it creates advice tailored to the user's condition. For example, it generates advice such as "Take deep breaths to relax. We recommend drinking tea to refresh yourself." The input is analyzed emotion and stress data, and the output is specific written advice for the user.

[0221] Step 7:

[0222] The terminal presents the generated advice to the user in text or audio format. For example, the generated advice may be displayed on the terminal's display and simultaneously played back as audio, allowing the user to receive the advice both visually and audibly.

[0223] Step 8:

[0224] The device plays healing music while providing advice. The server sends the generated healing music data to the device, providing a relaxing environment for the user. By playing the healing music, the user can refresh their mind and body while receiving advice.

[0225] This allows users to receive practical support to reduce stress and feel refreshed while shopping in a virtual store.

[0226] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0227] This invention is a system for effectively managing employee mental health, and in particular, by combining an emotion engine, it is possible to provide more accurate stress assessments and personalized advice. This system integrates the following functions: generating mental health-related questions, acquiring and analyzing voice data, recognizing emotions, evaluating stress, generating advice, and playing healing music.

[0228] System Configuration

[0229] The system consists of the following main elements:

[0230] 1. Server: The central data processing center, which generates questions, analyzes voice, evaluates stress, recognizes emotions, generates advice, and generates soothing music.

[0231] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[0232] 3. User: An employee who uses the system.

[0233] 4. Emotion Engine: Recognizes emotions from the user's voice responses and helps assess stress levels.

[0234] System Operation

[0235] 1. Question generation and retrieval

[0236] The server generates questions about mental health. It uses generative AI to create a personalized list of questions, taking into account the user's past data and common mental health questions. For example, questions might include, "How stressed are you at work these days?" and "How do you spend your time outside of work?"

[0237] The terminal presents the generated question to the user by voice or text.

[0238] The user answers the questions by voice, for example, "I've had a lot of meetings this week and I'm feeling a bit stressed."

[0239] The terminal records the user's voice response and transmits it to the server.

[0240] 2. Voice Analysis and Emotion Recognition

[0241] The server analyzes the received voice data using a voice analysis engine, extracting voice characteristics such as intonation, speed, and volume, and then analyzes the text content using a natural language processing (NLP) engine.

[0242] The server then uses an emotion engine to recognize the user's emotions from the voice data, for example, recognizing emotional states such as "tension," "anxiety," or "relief" from the tone of voice and speaking style.

[0243] 3. Stress Assessment

[0244] The server evaluates the user's stress level based on emotional data recognized through voice analysis and the emotion engine. Specifically, if the emotion engine recognizes "tension" or "anxiety," the stress level is determined to be high.

[0245] 4. Advice Generation and Presentation

[0246] Based on the evaluation results, the server generates appropriate advice and countermeasures. The AI ​​generator provides personalized advice tailored to the user's situation. For example, specific advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension" is provided.

[0247] The device displays the generated advice and countermeasures to the user, and is set to play healing music at the same time.

[0248] 5. Play healing music

[0249] The server uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[0250] The device plays soothing music while displaying advice and suggested solutions on a screen, allowing the user to receive the presented information in a relaxed state.

[0251] Specific examples

[0252] A typical day scenario could involve the following exchange:

[0253] 1. Device: "Have you felt stressed at work recently?"

[0254] 2. User: "Yes, writing the report was very stressful."

[0255] 3. Server: Performs voice analysis and NLP analysis, and recognizes "tension" using the emotion engine.

[0256] 4. Server: Generates the advice, "Take regular breaks while writing your report. We also recommend doing some light stretching."

[0257] 5. Device: Displays advice and plays relaxing, soothing music.

[0258] This system allows employees to effectively relieve stress and improve their mental health at work. By incorporating an emotion engine, it provides more accurate assessments and specific advice tailored to each user's condition.

[0259] The processing flow will be explained below.

[0260] Step 1:

[0261] The server starts up and makes the user database accessible, preparing the entire system for operation.

[0262] Step 2:

[0263] The terminal displays a login screen to the user, providing a screen with fields for entering a user ID and password.

[0264] Step 3:

[0265] The user enters login information (user ID and password) and presses the login button. The login information is sent to the system.

[0266] Step 4:

[0267] The terminal sends the user's login information to the server, which performs authentication and, if successful, proceeds to the next step.

[0268] Step 5:

[0269] The server uses generative AI to generate a personalized list of mental health questions for the user, based on the user's past data and general mental health questions.

[0270] Step 6:

[0271] The device presents the generated list of questions to the user via voice or text, for example, "How stressful are you feeling at work these days?"

[0272] Step 7:

[0273] The user answers the question by voice. For example, the user might say, "I felt a little stressed during the meeting."

[0274] Step 8:

[0275] The device records the user's voice response and sends it to the server, where the user's response arrives as recorded data.

[0276] Step 9:

[0277] The server uses a speech analysis engine to analyze the received voice data, extracting voice intonation, speed, and volume characteristics.

[0278] Step 10:

[0279] The server uses a natural language processing (NLP) engine to analyze the content of the voice data and extract keywords related to emotions and stress.

[0280] Step 11:

[0281] The server uses an emotion engine to recognize the user's emotions from the voice data, for example, recognizing emotional states such as "tension," "anxiety," and "relief" from the tone of voice and speaking style.

[0282] Step 12:

[0283] The server integrates the results of the voice analysis and emotion recognition by the emotion engine to evaluate the user's stress level, and the evaluation result is generated for the next step.

[0284] Step 13:

[0285] Based on the stress assessment results, the server uses AI to generate specific advice and countermeasures, such as "Light exercise would be good. Walking and stretching are recommended."

[0286] Step 14:

[0287] The server sends the generated advice and countermeasures to the device, which then prepares the countermeasures to be provided to the user.

[0288] Step 15:

[0289] The device displays advice and suggested solutions to the user. The display is confirmed to be correct.

[0290] Step 16:

[0291] The device is designed to play healing music sent from the server, and will begin playing it while simultaneously displaying advice and suggestions for countermeasures.

[0292] Step 17:

[0293] The device displays a screen requesting the user for advice and feedback on the effects of the healing music.

[0294] Step 18:

[0295] The user provides feedback via voice or text, for example, "I found the music very relaxing."

[0296] Step 19:

[0297] The device sends the user's feedback to the server, which uses the feedback data to generate questions and advice for the next time.

[0298] Step 20:

[0299] The server records the feedback and uses the data to improve the system and user experience in the future.

[0300] Example 2

[0301] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0302] Conventional mental health management systems rely on general questions and self-reported assessments by users, making it difficult to accurately grasp emotions and stress levels. Furthermore, they lack the personalization required to provide appropriate advice and stress reduction methods, limiting their effectiveness in improving users' mental health. Therefore, there is a need for a system that provides more accurate and personalized emotion assessment and stress reduction.

[0303] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0304] In this invention, the server includes a generation means for generating questions related to mental health, an acquisition means for acquiring employees' voice responses to the questions, an analysis means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generation means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, a presentation means for presenting the advice and proposed measures and playing soothing music, a means for analyzing the intonation, speed, volume, etc. of the voice and analyzing the text using natural language processing, a means for recognizing the user's emotions from the voice data using an emotion engine, and a means for generating and playing soothing music optimized for the stress level. This makes it possible to accurately evaluate the user's emotions and stress level and provide personalized advice and stress relief methods.

[0305] A "generator" is a device or software function for generating questions related to a user's mental health.

[0306] The "acquisition means" is a device or software function for recording the user's voice response and inputting it into the system.

[0307] The "analysis means" is a device or software function for analyzing the acquired voice data and evaluating the user's emotions and stress level.

[0308] The "advice generation means" is a device or software function for creating specific advice or countermeasures for the user based on the evaluation results obtained by the analysis means.

[0309] The "presentation means" is a device or software function for showing the generated advice and countermeasures to the user and playing healing music.

[0310] "Means for analyzing voice intonation, speed, volume, etc." refers to the function of a device or software for analyzing voice characteristics such as intonation, speed, volume, etc. contained in voice data.

[0311] "Means for analyzing text using natural language processing" refers to a device or software function that uses natural language processing techniques to analyze text data converted from voice data.

[0312] The "means for recognizing a user's emotion from voice data using an emotion engine" is a function of the emotion engine used to recognize a user's emotion based on voice data.

[0313] The "means for generating and playing healing music optimized for a stress level" refers to a device or software function for creating and playing optimal healing music according to a user's stress level.

[0314] This invention is a system designed to effectively manage the mental health of employees. The system mainly consists of three entities: a server, a terminal, and a user, each of which plays a specific role.

[0315] Question generation and retrieval

[0316] The server generates questions related to employees' mental health. Specifically, it uses a generative AI model (e.g., OpenAI's GPT-4) to create a personalized list of questions based on the user's past data and common mental health questions. For example, questions include, "What kind of stress have you been feeling at work recently?" and "How do you spend your time outside of work?"

[0317] The device receives the questions sent from the server and presents them to the user in voice or text. For voice presentation, speech synthesis software (e.g., Google Cloud Text-to-Speech) is used.

[0318] The user answers the questions by voice, for example, "I've had a lot of meetings this week and I'm feeling a bit stressed." The device records this voice and sends it to the server. The voice data is first stored locally and then uploaded to the server using a secure communication protocol (e.g., HTTPS).

[0319] Voice analysis and emotion recognition

[0320] The server converts the received voice data into text data using a voice analysis engine (e.g., Google Cloud Speech-to-Text), which also extracts voice characteristics such as intonation, speed, and volume.

[0321] Next, the server uses a natural language processing (NLP) engine (e.g., OpenAI's GPT-4) to analyze the text content and understand what the user is saying. Using the analysis results and voice feature data, an emotion recognition engine (e.g., Microsoft's Azure Emotion API) recognizes the user's emotional state. Specifically, emotions such as "tension," "anxiety," and "relief" are detected from the tone of voice and speaking style.

[0322] Stress assessment

[0323] The server evaluates the user's stress level based on the results of voice analysis and emotion recognition. If the emotion recognition engine recognizes "tension" or "anxiety," the stress level is deemed high; conversely, if it recognizes "relief," it is deemed low. This evaluation result is recorded within the system and used for subsequent analysis and the development of countermeasures.

[0324] Advice generation and presentation

[0325] The server generates appropriate advice based on the stress assessment results. It uses a generative AI model (e.g., OpenAI's GPT-4) to create personalized advice based on the assessment results and past user data. For example, it provides specific advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension."

[0326] The device displays the advice sent from the server to the user and simultaneously plays healing music. Music playback is performed using a media player or streaming service (e.g., Spotify API).

[0327] Playing healing music

[0328] The server uses a generative AI model to generate soothing music optimized for the user's stress level and transmits it to the device, which then plays the soothing music in sync with the user's advice, providing a relaxing environment for the user.

[0329] Prompt Sentence Examples

[0330] In a one day scenario, the following prompts could be used:

[0331] "Tell me about your work environment recently. If you have had any stressful experiences, please tell me about them in detail."

[0332] This allows users to effectively manage stress and receive appropriate advice in a comfortable environment. In this way, the system is able to perform highly accurate emotion recognition and stress assessment according to the individual state of the user, and provide specific and personalized advice.

[0333] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0334] Step 1: Question Generation

[0335] The server uses a generative AI model to receive the user's past mental health data and general question data as input, and generates a personalized list of questions based on this. Specifically, questions such as "What kind of stress have you been feeling at work recently?" and "How do you spend your time outside of work?" are generated. The generated list of questions is then sent to the device in the next step.

[0336] Step 2: Posing the Question

[0337] The device receives a list of questions from the server as input and uses speech synthesis software to present the questions to the user either audibly or via text. For example, Google Cloud Text-to-Speech is used to present the questions to the user audibly, and the user can then respond verbally.

[0338] Step 3: Get an audio response

[0339] The user answers the questions by voice, for example, "I had a lot of meetings this week and I felt a bit stressed."

[0340] The device takes this voice response as input, stores it locally, and then transmits the voice data to the server using a secure communication protocol (e.g., HTTPS).

[0341] Step 4: Audio analysis

[0342] The server receives the voice data sent from the device as input and converts it into text using a voice analysis engine (for example, Google Cloud Speech-to-Text). During this process, it extracts voice characteristics such as intonation, speed, and volume, and passes them on to the next process along with the text data.

[0343] Step 5: Natural Language Processing (NLP)

[0344] The server receives the text data and speech feature data obtained through speech analysis as input, analyzes the text content using an NLP engine (for example, OpenAI's GPT-4), and outputs structured data that helps understand what the user said.

[0345] Step 6: Emotion Recognition

[0346] The server receives the analysis results of the NLP engine and the voice feature data as input, and uses an emotion recognition engine (such as Microsoft's Azure Emotion API) to recognize the user's emotions. For example, it detects emotions such as "tension," "anxiety," and "relief" from the tone of voice and speaking style, and generates emotion data based on the results.

[0347] Step 7: Stress Assessment

[0348] The server receives the emotion data output by the emotion recognition engine as input and evaluates the user's stress level. Specifically, if the emotion data is recognized as "tension" or "anxiety," it is judged to be high stress. The evaluation results are stored in an internal database.

[0349] Step 8: Advice Generation

[0350] The server receives the stress assessment results and past user data as input, and uses the generative AI model to generate specific advice and countermeasures. For example, advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension" is generated and sent to the device.

[0351] Step 9: Providing advice and playing healing music

[0352] The device receives advice sent from the server as input and displays it to the user in text. At the same time, it receives healing music optimized for the user's stress level from the server and plays it using a media player or streaming service (e.g., Spotify API). This allows the user to receive advice in a relaxed state.

[0353] (Application example 2)

[0354] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0355] Modern factory workers are under a lot of stress during their work, making mental health care in the workplace extremely important. However, traditional mental health management systems do not adequately consider the characteristics of the workplace or the emotional state of individual workers. Furthermore, to provide effective care, a method is needed to accurately assess workers' emotional states and provide individually tailored advice and measures. Therefore, there is a need to develop a new system that can reduce worker stress and improve their mental health.

[0356] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0357] In this invention, the server includes a generation means for generating questions related to mental health, an acquisition means for acquiring employees' voice responses to the questions, an analysis means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generation means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, a presentation means for presenting the advice and proposed measures and playing soothing music, and a factory robot that collects data from workers through a voice interface and provides personalized advice and soothing music for the mental health care of factory workers. This allows workers to continue working in a less stressful environment, and is expected to improve the overall working environment.

[0358] The "Employee Mental Health Management System" is a system that effectively manages employees' mental health, assesses their emotions and stress levels, and provides specific advice and healing music.

[0359] A "generation means" is a means that has the function of generating questions related to mental health.

[0360] The "acquisition means" is a means having a function of acquiring an employee's voice response to a question.

[0361] The "analysis means" is a means having a function of analyzing the acquired voice response and evaluating emotions and stress levels.

[0362] The "advice generation means" is a means having a function of generating specific advice and countermeasures based on the evaluation results obtained by the analysis.

[0363] The "presentation means" is a means that has the function of presenting advice and countermeasures as well as playing healing music.

[0364] The "Mental Health Care System for Factory Workers" is a system for managing the mental health of factory workers, assessing their emotional state and stress levels, and providing personalized advice and healing music.

[0365] A "voice interface" is an interface for collecting voice data from a user and conducting interaction.

[0366] A "factory robot" is a robot that supports workers in a factory and has mental health management functions.

[0367] This invention is a system for effectively managing the mental health of employees and factory workers, and is characterized by the combination of an emotion engine to realize mental care suited to the working environment. This system integrates the following functions: generating questions related to mental health, acquiring and analyzing voice data, recognizing emotions, evaluating stress, generating advice, and playing healing music.

[0368] System Configuration

[0369] The system consists of the following main components:

[0370] 1. Server: This is the central data processing center, and performs question generation, voice analysis, stress assessment, emotion recognition, advice generation, healing music generation, etc. The server can be, for example, a cloud-based server or an on-premise server.

[0371] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music. Terminals can be smartphones, tablets, desktop computers, etc.

[0372] 3. Users: Employees and workers who use the system.

[0373] 4. Emotion Engine: Recognizes emotions from the user's voice responses and helps assess stress levels. It can use Google Cloud Natural Language API and OpenAI's Emotion Recognition model.

[0374] System Operation

[0375] Question generation and retrieval

[0376] Server: Uses generative AI to create a personalized list of questions, taking into account the user's past data and common mental health-related questions.

[0377] Terminal: Presents the generated questions to the user by voice or text.

[0378] User: Answers questions verbally.

[0379] Terminal: Records the user's voice response and sends it to the server.

[0380] Voice analysis and emotion recognition

[0381] Server: The received voice data is analyzed using a voice analysis engine. Voice characteristics such as intonation, speed, and volume are extracted, and the text content is analyzed using a natural language processing (NLP) engine.

[0382] The server then uses an emotion engine to recognize the user's emotion from the voice data.

[0383] Stress assessment

[0384] Server: Evaluates the user's stress level based on emotional data recognized through voice analysis and the emotion engine. Specifically, if the emotion engine recognizes "tension" or "anxiety," the stress level is determined to be high.

[0385] Advice generation and presentation

[0386] Server: Based on the evaluation results, appropriate advice and countermeasures are generated. The generation AI provides personalized advice tailored to the user's situation.

[0387] Device: Displays the generated advice and countermeasures to the user. Set it to play healing music at the same time.

[0388] Playing healing music

[0389] Server: Uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[0390] Device: The device plays soothing music while displaying advice and suggested solutions, allowing the user to receive the information in a relaxed state.

[0391] Specific examples

[0392] A typical day scenario could involve the following exchange:

[0393] 1. Device: "Have you felt stressed at work recently?"

[0394] 2. User: "Yes, writing the report was very stressful."

[0395] 3. Server: Performs voice analysis and NLP analysis, and recognizes "tension" using the emotion engine.

[0396] 4. Server: Generates the advice, "Take regular breaks while writing your report. We also recommend doing some light stretching."

[0397] 5. Device: Displays advice and plays relaxing, soothing music.

[0398] This system will enable workers to continue working in a less stressful environment, which is expected to improve the overall working environment.

[0399] Examples of prompt statements

[0400] Examples of prompts the system might use when using a generative AI model include:

[0401] "If workers are feeling very stressed, generate advice to help them relieve it."

[0402] This prompt allows the system to provide appropriate advice for the worker.

[0403] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0404] Step 1:

[0405] The server uses a generative AI model to generate a personalized list of questions, taking into account the user's past data and common mental health-related questions.

[0406] Input: User history data, general question data

[0407] Data processing: Generate a list of questions using a generative AI model

[0408] Output: A personalized list of questions

[0409] Step 2:

[0410] The terminal presents the list of questions sent from the server to the user by voice or text.

[0411] Input: Personalized Question List

[0412] Output: Questions presented in audio or text format

[0413] What happens: When the device uses a speech synthesis engine to play a question aloud, it presents the user with a question such as "Have you felt stressed at work recently?"

[0414] Step 3:

[0415] The user answers the questions by voice, and the terminal records the voice response and transmits it to the server.

[0416] Input: User's spoken response

[0417] Output: Recorded audio data

[0418] Specific operation: When the user answers, for example, "Yes, writing the report was very stressful," the device records the voice and sends it to the server.

[0419] Step 4:

[0420] The server analyzes the received voice data using a voice analysis engine to extract voice characteristics such as intonation, speed, and volume. It also analyzes the text content using an NLP engine and recognizes the user's emotions using an emotion engine.

[0421] Input: Recorded audio data

[0422] Data processing: Analyze voice data with a voice analysis engine and extract features. Analyze text with an NLP engine. Recognize emotions with an emotion engine.

[0423] Output: Emotional state, stress level data

[0424] Specific operation: The server analyzes the voice data, and if the user says "I felt very stressed," it recognizes emotions such as "tension" and "anxiety."

[0425] Step 5:

[0426] The server evaluates the user's stress level based on emotion recognition data and generates appropriate advice and countermeasures.

[0427] Input: Emotional state, stress level data

[0428] Data processing: Based on the evaluation results, advice and countermeasures are created using an advice generation AI model.

[0429] Output: Specific advice and countermeasures

[0430] Specific actions: The server generates specific advice such as "Take appropriate breaks while writing the report. We also recommend doing some light stretching."

[0431] Step 6:

[0432] The terminal displays advice and countermeasures from the server to the user and simultaneously plays healing music.

[0433] Input: Specific advice, countermeasures, healing music data

[0434] Output: Displaying advice to the user, playing healing music

[0435] Specific operation: The device displays advice such as "Take regular breaks while writing your report" on the screen and plays relaxing music.

[0436] Step 7:

[0437] The server uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[0438] Input: User's stress level data

[0439] Data processing: Generating healing music using generative AI models

[0440] Output: Healing music data

[0441] Specific operation: The server generates relaxing music according to the user's stress level and sends it to the device.

[0442] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0443] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0444] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0445] [Second embodiment]

[0446] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0447] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0448] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0449] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0450] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0451] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0452] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0453] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0454] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0455] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0456] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0457] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0458] The present invention is a system for managing employees' mental health and providing specific advice and countermeasures. This system includes functions for generating questions related to mental health, acquiring and analyzing audio data, generating advice, and playing healing music.

[0459] System Configuration

[0460] The system consists of the following main elements:

[0461] 1. Server: This is the central data processing device that generates questions, analyzes voice, evaluates stress, generates advice, and generates healing music.

[0462] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[0463] 3. User: An employee who uses the system.

[0464] System Operation

[0465] 1. Question generation and retrieval

[0466] The server generates questions about the user's mental health, using generative AI to create a personalized list of questions. For example:

[0467] "What kind of stress have you been experiencing at work recently?"

[0468] "Are you getting enough rest?"

[0469] The questions are displayed in order below.

[0470] The terminal presents the generated question to the user by voice or text.

[0471] The user responds to the question by voice, for example, "I felt a little stressed during the meeting."

[0472] The terminal records the user's voice response and transmits it to the server.

[0473] 2. Voice analysis and stress assessment

[0474] The server analyzes the received voice data using a speech analysis engine to analyze the intonation, speed, and volume of the voice, and a natural language processing (NLP) engine to analyze the text content, which then evaluates the user's emotional state and stress level.

[0475] Specifically, a high intonation and fast speech rate is judged to indicate a state of tension. Additionally, if the answer contains keywords such as "stress" or "anxiety," the stress level is also assessed as high.

[0476] 3. Advice Generation and Presentation

[0477] Based on the evaluation results, the server generates specific advice and countermeasures. The AI ​​generator makes optimal suggestions based on the user's situation. For example,

[0478] "Light exercise is good. Walking and stretching are recommended."

[0479] "Increasing communication with other employees is also effective."

[0480] The terminal displays the generated advice and countermeasures to the user.

[0481] 4. Play healing music

[0482] While the advice is displayed, the server uses a generation AI to create healing music and send it to the device.

[0483] The device plays soothing music along with a screen displaying advice and suggested solutions, allowing the user to receive the presented information in a relaxed state.

[0484] Specific examples

[0485] In one scenario, the following specific exchanges take place:

[0486] 1. Device: "Have you felt stressed at work recently?"

[0487] 2. User: "Yes, writing the report was very stressful."

[0488] 3. Server: Performs voice analysis and NLP analysis and determines that the stress level is high.

[0489] 4. Server: Generates the advice, "Take regular breaks while writing your report. It's also a good idea to do some light stretching."

[0490] 5. Device: Displays advice and simultaneously plays relaxing, healing music.

[0491] This allows users to receive specific advice and learn how to reduce stress in a relaxed environment.

[0492] By using this system, employees can improve their mental health through dialogue with the robot, creating a comfortable working environment.

[0493] The processing flow will be explained below.

[0494] Step 1:

[0495] The server starts up and makes the user database accessible, preparing the entire system for operation.

[0496] Step 2:

[0497] The terminal displays a login screen to the user, providing a screen with fields for entering a user ID and password.

[0498] Step 3:

[0499] The user enters login information (user ID and password) and presses the login button. The login information is sent to the system.

[0500] Step 4:

[0501] The terminal sends the user's login information to the server, which performs authentication and, if successful, proceeds to the next step.

[0502] Step 5:

[0503] The server uses generative AI to generate a personalized list of mental health questions for the user, based on the user's past data and general mental health questions.

[0504] Step 6:

[0505] The device presents the generated list of questions to the user via voice or text, for example, "How stressful are you feeling at work these days?"

[0506] Step 7:

[0507] The user answers the question by voice. For example, the user might say, "I felt a little stressed during the meeting."

[0508] Step 8:

[0509] The device records the user's voice response and sends it to the server, where the user's response arrives as recorded data.

[0510] Step 9:

[0511] The server uses a speech analysis engine to analyze the received voice data, extracting voice intonation, speed, and volume characteristics.

[0512] Step 10:

[0513] The server uses a natural language processing (NLP) engine to analyze the content of the voice data and extract keywords related to emotions and stress.

[0514] Step 11:

[0515] The server integrates the results of the voice and text analysis to evaluate the user's emotional state and stress level, and generates an evaluation result for the next step.

[0516] Step 12:

[0517] Based on the stress assessment results, the server uses AI to generate specific advice and countermeasures, such as "Light exercise would be good. Walking and stretching are recommended."

[0518] Step 13:

[0519] The server sends the generated advice and countermeasures to the device, which then prepares the countermeasures to be provided to the user.

[0520] Step 14:

[0521] The device displays advice and solutions to the user, while soothing music is played in the background.

[0522] Step 15:

[0523] The device plays soothing music while providing advice and solutions to the user, creating an environment that enhances relaxation.

[0524] Step 16:

[0525] The device displays a screen requesting the user for advice and feedback on the effects of the healing music.

[0526] Step 17:

[0527] The user provides feedback via voice or text, for example, "I found the music very relaxing."

[0528] Step 18:

[0529] The device sends the user's feedback to the server, which uses the feedback data to generate questions and advice for the next time.

[0530] Step 19:

[0531] The server records the feedback and uses the data to improve the system and user experience in the future.

[0532] Example 1

[0533] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0534] In today's workforce, managing employee mental health is becoming increasingly important. Many employees are exposed to stress and pressure on a daily basis, often resulting in decreased productivity and health problems. However, the means to provide effective mental health care remain limited. In particular, it is difficult to quickly provide personalized advice and measures tailored to each employee's situation. Therefore, there is a need for a comprehensive mental health management system that can accurately assess employees' stress levels, provide appropriate advice, and combine relaxation elements such as healing music.

[0535] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0536] In this invention, the server includes a generating means for generating questions related to mental health, an acquiring means for acquiring employees' voice responses to the questions, an analyzing means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generating means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, and a presenting means for presenting the advice and proposed measures and generating and playing healing music, thereby making it possible to provide personalized mental health care to each employee.

[0537] "Generation means" refers to a device or program for generating questions related to mental health.

[0538] The term "acquisition means" refers to a device or program for acquiring a user's voice response.

[0539] "Analysis means" refers to a device or program for analyzing the acquired voice response and evaluating emotions and stress levels.

[0540] The "advice generation means" refers to a device or program that generates specific advice or countermeasures based on the evaluation results obtained by analysis.

[0541] The "presentation means" refers to a device or program that presents the generated advice and countermeasures to the user and also generates and plays healing music.

[0542] "Emotions and stress levels" refers to the emotional state the user is feeling and the associated degree of stress.

[0543] "Healing music" refers to music that is intended to relax the user and reduce stress.

[0544] A "generative AI model" refers to an algorithm or program that uses artificial intelligence to generate questions, advice, music, etc.

[0545] "Voice analysis" refers to the process of analyzing acquired voice data in terms of intonation, speed, volume, etc.

[0546] A "natural language processing engine" refers to a program or algorithm that converts voice data into text and analyzes the content.

[0547] This invention is a system for managing employees' mental health and providing specific advice and countermeasures. This system includes functions for generating questions related to mental health, acquiring and analyzing audio data, generating advice, and playing healing music. The main elements of the system are a server, a terminal, and a user.

[0548] System Configuration

[0549] 1. Server

[0550] The server is the central data processing unit, generating questions using a generative AI model and analyzing the voice data. It also generates specific advice based on the analysis results and creates healing music. Software such as "PRAAT" and "SpaCy" are used for voice analysis.

[0551] 2. Terminal

[0552] The terminal is a device operated by the user, and has the function of presenting questions sent from the server, receiving voice responses from the user, displaying analysis results and advice, and playing healing music.

[0553] 3. Users

[0554] Users are employees who use the system, answer mental health questions through their terminals, and receive the advice provided.

[0555] System Operation

[0556] The server uses a generative AI model to generate mental health-related questions that are personalized based on the user's previous answers and general mental health care knowledge.

[0557] The device presents the generated question to the user by voice or text, for example, "What has been making you feel stressed recently?"

[0558] The user answers the questions by voice, for example, "I felt very stressed writing the report."

[0559] The terminal records the user's voice response and transmits it to the server.

[0560] The server uses the speech analysis engine "PRAAT" to analyze the intonation, speed, and volume of the voice, and the natural language processing engine "SpaCy" to analyze the text content, which then evaluates the user's emotional state and stress level.

[0561] Based on the evaluation results, the server uses a generative AI model to generate specific advice and countermeasures for the user. For example, it might suggest, "It would be good to do some light exercise. Walking and stretching are recommended."

[0562] Once the evaluation results and advice generation are complete, the server uses the generative AI model to create healing music and send it to the device.

[0563] The device displays the generated advice and countermeasures to the user while simultaneously playing soothing music, allowing the user to receive specific advice in a relaxed state.

[0564] Specific examples

[0565] For example, in a scenario on one day, the following exchange takes place:

[0566] 1. Device: "What has been stressing you out lately?"

[0567] 2. User: "Yes, writing the report was very stressful."

[0568] 3. The device records this response and sends it to the server.

[0569] 4. Server: Conduct voice analysis and NLP analysis to identify high stress levels associated with "report writing."

[0570] 5. Server: Generate the advice, "Take regular breaks while writing your report. It's also a good idea to do some light stretching."

[0571] 6. Device: Displays advice and simultaneously plays relaxing, healing music.

[0572] Example prompt sentence:

[0573] "Please write a natural language description of a system that manages employees' mental health and provides specific advice. The system includes functions for question generation, voice analysis, advice generation, and soothing music playback. Please describe the process in detail along the following lines."

[0574] By using this system, employees can manage their own mental health more effectively and create a comfortable working environment.

[0575] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0576] Step 1: Question Generation

[0577] The server uses a generative AI model to generate questions related to mental health. At this time, the server obtains the user's profile information and past answers and creates a personalized list of questions. The input data is the user's profile information and past answer data, and the output is the generated list of questions. For example, the server generates a question such as, "What has been making you feel stressed recently?"

[0578] Step 2: Question posing

[0579] The terminal presents the questions sent from the server to the user. The input data is the list of questions sent from the server, and the output is a display screen or audio presented to the user. For example, the terminal may read out the question "What has been making you feel stressed recently?" or display it on the screen.

[0580] Step 3: Get the answer

[0581] The user answers the presented question by voice. The input data is the question presented to the user, and the output is the user's voice response. For example, the user might respond, "Writing the report was very stressful."

[0582] Step 4: Acquire and transmit audio data

[0583] The terminal acquires and records the user's voice response. Then, it transmits this voice data to the server. The input data is the user's voice response, and the output is the voice data transmitted to the server. For example, the terminal transmits voice data such as "I felt very stressed while writing the report" to the server.

[0584] Step 5: Audio analysis

[0585] The server analyzes the received voice data. For the analysis, it uses the voice analysis engine "PRAAT" and the natural language processing engine "SpaCy." The input data is the user's voice data, and the output is the voice intonation, speed, volume, and text analysis results. Specifically, the server analyzes the voice data "I felt very stressed writing the report" for voice intonation, speed, and volume using "PRAAT," and converts it into text and analyzes it using "SpaCy."

[0586] Step 6: Stress Assessment

[0587] The server evaluates the user's stress level based on the results of voice and text analysis. The input data are the voice and text analysis results, and the output is the user's stress level assessment. For example, if the speech is high-pitched, fast-paced, and includes the word "stress," the server will assess the user's stress level as high.

[0588] Step 7: Advice Generation

[0589] The server uses the generative AI model to generate specific advice and countermeasures for the user based on the evaluation results. The input data is the stress level evaluation result, and the output is the generated advice and countermeasures. For example, it generates advice such as, "Take appropriate breaks while writing your report. It is also recommended that you do some light stretching."

[0590] Step 8: Providing advice

[0591] The terminal displays the advice and countermeasures sent from the server to the user. The input data is the advice and countermeasures sent from the server, and the output is the display screen and audio presented to the user. For example, advice such as "Take appropriate breaks while writing a report" can be displayed on the screen and read aloud.

[0592] Step 9: Generate and Play Healing Music

[0593] The server uses a generative AI model to create healing music and sends it to the device. The input data is the stress assessment results, and the output is the generated healing music. For example, a song with a relaxing effect is generated. The device receives this healing music and plays it along with the advice display. The user can receive advice in a relaxing environment.

[0594] (Application example 1)

[0595] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0596] In today's work environment, many employees are increasingly experiencing stress and emotional strain, which increases the risk of decreased productivity and poor mental health. Furthermore, virtual store shopping experiences often leave users feeling overwhelmed by numerous options and information, leading to stress. There is a need for an effective mental health management system that can improve these situations and help employees and virtual store users feel refreshed, improving their productivity and satisfaction.

[0597] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0598] In this invention, the server includes a generating means for generating questions related to mental health, an acquiring means for acquiring the user's voice responses to the questions, an analyzing means for analyzing the acquired voice responses and evaluating the user's emotions and stress level, an advice generating means for generating specific advice and proposed measures based on the evaluation results, a presenting means for presenting the advice and proposed measures and playing soothing music, and a refreshing means for presenting refreshing advice to improve the user's experience in the virtual space. This enables practical support to reduce stress felt by the user while shopping in a virtual store and to refresh the user.

[0599] The "means for generating mental health-related questions" is a device or method that dynamically generates appropriate questions to assess the mental health status of a user.

[0600] The "means for acquiring the employee's voice response to the question" refers to a device or method for acquiring voice data in which the user answers the question and transmitting that data to the system.

[0601] The "analysis means for analyzing the acquired voice response and evaluating the emotion and stress level" is a device or method for analyzing voice data and evaluating the user's emotional state and stress level.

[0602] The "advice generating means for generating specific advice and countermeasures" is a device or method for generating appropriate advice and countermeasures for the user based on the analysis results.

[0603] The "presentation means for presenting advice and measures and playing healing music" is a device or method for presenting the generated advice and measures to the user and playing healing music at the same time.

[0604] A "refreshment means for presenting refreshment advice to improve the user experience in a virtual space" is a device or method that presents specific advice to refresh a user when the user feels stressed in a virtual store.

[0605] MODE FOR CARRYING OUT THE INVENTION

[0606] This invention provides a system for managing employee mental health and improving user experience in a virtual store. The following configuration makes it possible to evaluate the user's emotional state, generate optimal advice, and provide soothing music.

[0607] System Configuration

[0608] The system consists of the following elements:

[0609] 1. Server: A central data processing device that generates questions, analyzes voice, evaluates stress, generates advice, and generates healing music.

[0610] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[0611] 3. User: A user who uses the system.

[0612] question generation

[0613] The server generates questions related to the user's mental health, personalized using a generative AI model. For example:

[0614] It generates questions such as, "Have you felt stressed while shopping online recently?"

[0615] Acquiring audio data

[0616] The device presents the generated question to the user by voice or text. The user answers the question by voice. For example, the user might say, "Yes, there are so many products, I don't know which one to choose." This voice data is captured by the device and sent to the server.

[0617] Voice analysis and stress assessment

[0618] The server analyzes the acquired voice data using a speech recognition engine and a natural language processing (NLP) engine. The server analyzes the intonation, speed, volume, etc. of the voice to evaluate the user's emotional state and stress level. For example, a high intonation and fast speed of the voice may be determined to indicate a state of tension.

[0619] Advice generation and presentation

[0620] The server generates specific advice and countermeasures based on the analysis results. The generated advice provides specific support for the user's actions. For example,

[0621] Advice such as "Take a deep breath to relax. We recommend drinking tea to refresh yourself" is generated.

[0622] Playing healing music

[0623] The device not only displays the generated advice but also plays soothing music at the same time, allowing the user to receive the advice in a relaxed state.

[0624] Specific examples

[0625] For example, if a user using a virtual store responds that they "felt stressed," the server analyzes the voice data and detects a high stress level. As a result, the server generates advice such as "Take a deep breath to relax. We recommend drinking tea to refresh yourself," and the device displays this advice to the user while playing soothing music.

[0626] Prompt Sentence Examples

[0627] "Please assess this user's mental health status and provide appropriate advice: Yes, there are so many options that it's hard to know which one to choose."

[0628] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0629] Step 1:

[0630] The server generates questions related to the user's mental health, using a generative AI model to create a personalized list of questions, such as, "Have you felt stressed while shopping online recently?" These questions are customized based on the user's past feedback and shopping history.

[0631] Step 2:

[0632] The device presents the generated question to the user via voice or text. The user can receive the question through either visual or auditory means. For example, the device might ask the user aloud, "Have you felt stressed while shopping online recently?" The user listens to the question and understands its content.

[0633] Step 3:

[0634] The user answers the question by voice. For example, they might say, "Yes, there are so many products, I don't know which one to choose." The user's voice is picked up through the device's microphone.

[0635] Step 4:

[0636] The terminal records the user's voice responses and converts the data into a data format for transmission to the server. The terminal converts the raw voice data as input into an audio file and transmits it to the server over the network.

[0637] Step 5:

[0638] The server analyzes the acquired voice data and converts the voice into text using a voice recognition engine. This text data is then analyzed using an NLP engine to evaluate the emotional state and stress level. For example, the intonation and speed of the voice can be analyzed to determine the state of tension and stress level.

[0639] Step 6:

[0640] The server generates specific advice and countermeasures based on the evaluation results. Using a generative AI model, it creates advice tailored to the user's condition. For example, it generates advice such as "Take deep breaths to relax. We recommend drinking tea to refresh yourself." The input is analyzed emotion and stress data, and the output is specific written advice for the user.

[0641] Step 7:

[0642] The terminal presents the generated advice to the user in text or audio format. For example, the generated advice may be displayed on the terminal's display and simultaneously played back as audio, allowing the user to receive the advice both visually and audibly.

[0643] Step 8:

[0644] The device plays healing music while providing advice. The server sends the generated healing music data to the device, providing a relaxing environment for the user. By playing the healing music, the user can refresh their mind and body while receiving advice.

[0645] This allows users to receive practical support to reduce stress and feel refreshed while shopping in a virtual store.

[0646] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0647] This invention is a system for effectively managing employee mental health, and in particular, by combining an emotion engine, it is possible to provide more accurate stress assessments and personalized advice. This system integrates the following functions: generating mental health-related questions, acquiring and analyzing voice data, recognizing emotions, evaluating stress, generating advice, and playing healing music.

[0648] System Configuration

[0649] The system consists of the following main elements:

[0650] 1. Server: The central data processing center, which generates questions, analyzes voice, evaluates stress, recognizes emotions, generates advice, and generates soothing music.

[0651] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[0652] 3. User: An employee who uses the system.

[0653] 4. Emotion Engine: Recognizes emotions from the user's voice responses and helps assess stress levels.

[0654] System Operation

[0655] 1. Question generation and retrieval

[0656] The server generates questions about mental health. It uses generative AI to create a personalized list of questions, taking into account the user's past data and common mental health questions. For example, questions might include, "How stressed are you at work these days?" and "How do you spend your time outside of work?"

[0657] The terminal presents the generated question to the user by voice or text.

[0658] The user answers the questions by voice, for example, "I've had a lot of meetings this week and I'm feeling a bit stressed."

[0659] The terminal records the user's voice response and transmits it to the server.

[0660] 2. Voice Analysis and Emotion Recognition

[0661] The server analyzes the received voice data using a voice analysis engine, extracting voice characteristics such as intonation, speed, and volume, and then analyzes the text content using a natural language processing (NLP) engine.

[0662] The server then uses an emotion engine to recognize the user's emotions from the voice data, for example, recognizing emotional states such as "tension," "anxiety," or "relief" from the tone of voice and speaking style.

[0663] 3. Stress Assessment

[0664] The server evaluates the user's stress level based on emotional data recognized through voice analysis and the emotion engine. Specifically, if the emotion engine recognizes "tension" or "anxiety," the stress level is determined to be high.

[0665] 4. Advice Generation and Presentation

[0666] Based on the evaluation results, the server generates appropriate advice and countermeasures. The AI ​​generator provides personalized advice tailored to the user's situation. For example, specific advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension" is provided.

[0667] The device displays the generated advice and countermeasures to the user, and is set to play healing music at the same time.

[0668] 5. Play healing music

[0669] The server uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[0670] The device plays soothing music while displaying advice and suggested solutions on a screen, allowing the user to receive the presented information in a relaxed state.

[0671] Specific examples

[0672] A typical day scenario could involve the following exchange:

[0673] 1. Device: "Have you felt stressed at work recently?"

[0674] 2. User: "Yes, writing the report was very stressful."

[0675] 3. Server: Performs voice analysis and NLP analysis, and recognizes "tension" using the emotion engine.

[0676] 4. Server: Generates the advice, "Take regular breaks while writing your report. We also recommend doing some light stretching."

[0677] 5. Device: Displays advice and plays relaxing, soothing music.

[0678] This system allows employees to effectively relieve stress and improve their mental health at work. By incorporating an emotion engine, it provides more accurate assessments and specific advice tailored to each user's condition.

[0679] The processing flow will be explained below.

[0680] Step 1:

[0681] The server starts up and makes the user database accessible, preparing the entire system for operation.

[0682] Step 2:

[0683] The terminal displays a login screen to the user, providing a screen with fields for entering a user ID and password.

[0684] Step 3:

[0685] The user enters login information (user ID and password) and presses the login button. The login information is sent to the system.

[0686] Step 4:

[0687] The terminal sends the user's login information to the server, which performs authentication and, if successful, proceeds to the next step.

[0688] Step 5:

[0689] The server uses generative AI to generate a personalized list of mental health questions for the user, based on the user's past data and general mental health questions.

[0690] Step 6:

[0691] The device presents the generated list of questions to the user via voice or text, for example, "How stressful are you feeling at work these days?"

[0692] Step 7:

[0693] The user answers the question by voice. For example, the user might say, "I felt a little stressed during the meeting."

[0694] Step 8:

[0695] The device records the user's voice response and sends it to the server, where the user's response arrives as recorded data.

[0696] Step 9:

[0697] The server uses a speech analysis engine to analyze the received voice data, extracting voice intonation, speed, and volume characteristics.

[0698] Step 10:

[0699] The server uses a natural language processing (NLP) engine to analyze the content of the voice data and extract keywords related to emotions and stress.

[0700] Step 11:

[0701] The server uses an emotion engine to recognize the user's emotions from the voice data, for example, recognizing emotional states such as "tension," "anxiety," and "relief" from the tone of voice and speaking style.

[0702] Step 12:

[0703] The server integrates the results of the voice analysis and emotion recognition by the emotion engine to evaluate the user's stress level, and the evaluation result is generated for the next step.

[0704] Step 13:

[0705] Based on the stress assessment results, the server uses AI to generate specific advice and countermeasures, such as "Light exercise would be good. Walking and stretching are recommended."

[0706] Step 14:

[0707] The server sends the generated advice and countermeasures to the device, which then prepares the countermeasures to be provided to the user.

[0708] Step 15:

[0709] The device displays advice and suggested solutions to the user. The display is confirmed to be correct.

[0710] Step 16:

[0711] The device is designed to play healing music sent from the server, and will begin playing it while simultaneously displaying advice and suggestions for countermeasures.

[0712] Step 17:

[0713] The device displays a screen requesting the user for advice and feedback on the effects of the healing music.

[0714] Step 18:

[0715] The user provides feedback via voice or text, for example, "I found the music very relaxing."

[0716] Step 19:

[0717] The device sends the user's feedback to the server, which uses the feedback data to generate questions and advice for the next time.

[0718] Step 20:

[0719] The server records the feedback and uses the data to improve the system and user experience in the future.

[0720] Example 2

[0721] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0722] Conventional mental health management systems rely on general questions and self-reported assessments by users, making it difficult to accurately grasp emotions and stress levels. Furthermore, they lack the personalization required to provide appropriate advice and stress reduction methods, limiting their effectiveness in improving users' mental health. Therefore, there is a need for a system that provides more accurate and personalized emotion assessment and stress reduction.

[0723] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0724] In this invention, the server includes a generation means for generating questions related to mental health, an acquisition means for acquiring employees' voice responses to the questions, an analysis means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generation means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, a presentation means for presenting the advice and proposed measures and playing soothing music, a means for analyzing the intonation, speed, volume, etc. of the voice and analyzing the text using natural language processing, a means for recognizing the user's emotions from the voice data using an emotion engine, and a means for generating and playing soothing music optimized for the stress level. This makes it possible to accurately evaluate the user's emotions and stress level and provide personalized advice and stress relief methods.

[0725] A "generator" is a device or software function for generating questions related to a user's mental health.

[0726] The "acquisition means" is a device or software function for recording the user's voice response and inputting it into the system.

[0727] The "analysis means" is a device or software function for analyzing the acquired voice data and evaluating the user's emotions and stress level.

[0728] The "advice generation means" is a device or software function for creating specific advice or countermeasures for the user based on the evaluation results obtained by the analysis means.

[0729] The "presentation means" is a device or software function for showing the generated advice and countermeasures to the user and playing healing music.

[0730] "Means for analyzing voice intonation, speed, volume, etc." refers to the function of a device or software for analyzing voice characteristics such as intonation, speed, volume, etc. contained in voice data.

[0731] "Means for analyzing text using natural language processing" refers to a device or software function that uses natural language processing techniques to analyze text data converted from voice data.

[0732] The "means for recognizing a user's emotion from voice data using an emotion engine" is a function of the emotion engine used to recognize a user's emotion based on voice data.

[0733] The "means for generating and playing healing music optimized for a stress level" refers to a device or software function for creating and playing optimal healing music according to a user's stress level.

[0734] This invention is a system designed to effectively manage the mental health of employees. The system mainly consists of three entities: a server, a terminal, and a user, each of which plays a specific role.

[0735] Question generation and retrieval

[0736] The server generates questions related to employees' mental health. Specifically, it uses a generative AI model (e.g., OpenAI's GPT-4) to create a personalized list of questions based on the user's past data and common mental health questions. For example, questions include, "What kind of stress have you been feeling at work recently?" and "How do you spend your time outside of work?"

[0737] The device receives the questions sent from the server and presents them to the user in voice or text. For voice presentation, speech synthesis software (e.g., Google Cloud Text-to-Speech) is used.

[0738] The user answers the questions by voice, for example, "I've had a lot of meetings this week and I'm feeling a bit stressed." The device records this voice and sends it to the server. The voice data is first stored locally and then uploaded to the server using a secure communication protocol (e.g., HTTPS).

[0739] Voice analysis and emotion recognition

[0740] The server converts the received voice data into text data using a voice analysis engine (e.g., Google Cloud Speech-to-Text), which also extracts voice characteristics such as intonation, speed, and volume.

[0741] Next, the server uses a natural language processing (NLP) engine (e.g., OpenAI's GPT-4) to analyze the text content and understand what the user is saying. Using the analysis results and voice feature data, an emotion recognition engine (e.g., Microsoft's Azure Emotion API) recognizes the user's emotional state. Specifically, emotions such as "tension," "anxiety," and "relief" are detected from the tone of voice and speaking style.

[0742] Stress assessment

[0743] The server evaluates the user's stress level based on the results of voice analysis and emotion recognition. If the emotion recognition engine recognizes "tension" or "anxiety," the stress level is deemed high; conversely, if it recognizes "relief," it is deemed low. This evaluation result is recorded within the system and used for subsequent analysis and the development of countermeasures.

[0744] Advice generation and presentation

[0745] The server generates appropriate advice based on the stress assessment results. It uses a generative AI model (e.g., OpenAI's GPT-4) to create personalized advice based on the assessment results and past user data. For example, it provides specific advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension."

[0746] The device displays the advice sent from the server to the user and simultaneously plays healing music. Music playback is performed using a media player or streaming service (e.g., Spotify API).

[0747] Playing healing music

[0748] The server uses a generative AI model to generate soothing music optimized for the user's stress level and transmits it to the device, which then plays the soothing music in sync with the user's advice, providing a relaxing environment for the user.

[0749] Prompt Sentence Examples

[0750] In a one day scenario, the following prompts could be used:

[0751] "Tell me about your work environment recently. If you have had any stressful experiences, please tell me about them in detail."

[0752] This allows users to effectively manage stress and receive appropriate advice in a comfortable environment. In this way, the system is able to perform highly accurate emotion recognition and stress assessment according to the individual state of the user, and provide specific and personalized advice.

[0753] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0754] Step 1: Question Generation

[0755] The server uses a generative AI model to receive the user's past mental health data and general question data as input, and generates a personalized list of questions based on this. Specifically, questions such as "What kind of stress have you been feeling at work recently?" and "How do you spend your time outside of work?" are generated. The generated list of questions is then sent to the device in the next step.

[0756] Step 2: Posing the Question

[0757] The device receives a list of questions from the server as input and uses speech synthesis software to present the questions to the user either audibly or via text. For example, Google Cloud Text-to-Speech is used to present the questions to the user audibly, and the user can then respond verbally.

[0758] Step 3: Get an audio response

[0759] The user answers the questions by voice, for example, "I had a lot of meetings this week and I felt a bit stressed."

[0760] The device takes this voice response as input, stores it locally, and then transmits the voice data to the server using a secure communication protocol (e.g., HTTPS).

[0761] Step 4: Audio analysis

[0762] The server receives the voice data sent from the device as input and converts it into text using a voice analysis engine (for example, Google Cloud Speech-to-Text). During this process, it extracts voice characteristics such as intonation, speed, and volume, and passes them on to the next process along with the text data.

[0763] Step 5: Natural Language Processing (NLP)

[0764] The server receives the text data and speech feature data obtained through speech analysis as input, analyzes the text content using an NLP engine (for example, OpenAI's GPT-4), and outputs structured data that helps understand what the user said.

[0765] Step 6: Emotion Recognition

[0766] The server receives the analysis results of the NLP engine and the voice feature data as input, and uses an emotion recognition engine (such as Microsoft's Azure Emotion API) to recognize the user's emotions. For example, it detects emotions such as "tension," "anxiety," and "relief" from the tone of voice and speaking style, and generates emotion data based on the results.

[0767] Step 7: Stress Assessment

[0768] The server receives the emotion data output by the emotion recognition engine as input and evaluates the user's stress level. Specifically, if the emotion data is recognized as "tension" or "anxiety," it is judged to be high stress. The evaluation results are stored in an internal database.

[0769] Step 8: Advice Generation

[0770] The server receives the stress assessment results and past user data as input, and uses the generative AI model to generate specific advice and countermeasures. For example, advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension" is generated and sent to the device.

[0771] Step 9: Providing advice and playing healing music

[0772] The device receives advice sent from the server as input and displays it to the user in text. At the same time, it receives healing music optimized for the user's stress level from the server and plays it using a media player or streaming service (e.g., Spotify API). This allows the user to receive advice in a relaxed state.

[0773] (Application example 2)

[0774] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0775] Modern factory workers are under a lot of stress during their work, making mental health care in the workplace extremely important. However, traditional mental health management systems do not adequately consider the characteristics of the workplace or the emotional state of individual workers. Furthermore, to provide effective care, a method is needed to accurately assess workers' emotional states and provide individually tailored advice and measures. Therefore, there is a need to develop a new system that can reduce worker stress and improve their mental health.

[0776] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0777] In this invention, the server includes a generation means for generating questions related to mental health, an acquisition means for acquiring employees' voice responses to the questions, an analysis means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generation means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, a presentation means for presenting the advice and proposed measures and playing soothing music, and a factory robot that collects data from workers through a voice interface and provides personalized advice and soothing music for the mental health care of factory workers. This allows workers to continue working in a less stressful environment, and is expected to improve the overall working environment.

[0778] The "Employee Mental Health Management System" is a system that effectively manages employees' mental health, assesses their emotions and stress levels, and provides specific advice and healing music.

[0779] A "generation means" is a means that has the function of generating questions related to mental health.

[0780] The "acquisition means" is a means having a function of acquiring an employee's voice response to a question.

[0781] The "analysis means" is a means having a function of analyzing the acquired voice response and evaluating emotions and stress levels.

[0782] The "advice generation means" is a means having a function of generating specific advice and countermeasures based on the evaluation results obtained by the analysis.

[0783] The "presentation means" is a means that has the function of presenting advice and countermeasures as well as playing healing music.

[0784] The "Mental Health Care System for Factory Workers" is a system for managing the mental health of factory workers, assessing their emotional state and stress levels, and providing personalized advice and healing music.

[0785] A "voice interface" is an interface for collecting voice data from a user and conducting interaction.

[0786] A "factory robot" is a robot that supports workers in a factory and has mental health management functions.

[0787] This invention is a system for effectively managing the mental health of employees and factory workers, and is characterized by the combination of an emotion engine to realize mental care suited to the working environment. This system integrates the following functions: generating questions related to mental health, acquiring and analyzing voice data, recognizing emotions, evaluating stress, generating advice, and playing healing music.

[0788] System Configuration

[0789] The system consists of the following main components:

[0790] 1. Server: This is the central data processing center, and performs question generation, voice analysis, stress assessment, emotion recognition, advice generation, healing music generation, etc. The server can be, for example, a cloud-based server or an on-premise server.

[0791] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music. Terminals can be smartphones, tablets, desktop computers, etc.

[0792] 3. Users: Employees and workers who use the system.

[0793] 4. Emotion Engine: Recognizes emotions from the user's voice responses and helps assess stress levels. It can use Google Cloud Natural Language API and OpenAI's Emotion Recognition model.

[0794] System Operation

[0795] Question generation and retrieval

[0796] Server: Uses generative AI to create a personalized list of questions, taking into account the user's past data and common mental health-related questions.

[0797] Terminal: Presents the generated questions to the user by voice or text.

[0798] User: Answers questions verbally.

[0799] Terminal: Records the user's voice response and sends it to the server.

[0800] Voice analysis and emotion recognition

[0801] Server: The received voice data is analyzed using a voice analysis engine. Voice characteristics such as intonation, speed, and volume are extracted, and the text content is analyzed using a natural language processing (NLP) engine.

[0802] The server then uses an emotion engine to recognize the user's emotion from the voice data.

[0803] Stress assessment

[0804] Server: Evaluates the user's stress level based on emotional data recognized through voice analysis and the emotion engine. Specifically, if the emotion engine recognizes "tension" or "anxiety," the stress level is determined to be high.

[0805] Advice generation and presentation

[0806] Server: Based on the evaluation results, appropriate advice and countermeasures are generated. The generation AI provides personalized advice tailored to the user's situation.

[0807] Device: Displays the generated advice and countermeasures to the user. Set it to play healing music at the same time.

[0808] Playing healing music

[0809] Server: Uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[0810] Device: The device plays soothing music while displaying advice and suggested solutions, allowing the user to receive the information in a relaxed state.

[0811] Specific examples

[0812] A typical day scenario could involve the following exchange:

[0813] 1. Device: "Have you felt stressed at work recently?"

[0814] 2. User: "Yes, writing the report was very stressful."

[0815] 3. Server: Performs voice analysis and NLP analysis, and recognizes "tension" using the emotion engine.

[0816] 4. Server: Generates the advice, "Take regular breaks while writing your report. We also recommend doing some light stretching."

[0817] 5. Device: Displays advice and plays relaxing, soothing music.

[0818] This system will enable workers to continue working in a less stressful environment, which is expected to improve the overall working environment.

[0819] Examples of prompt statements

[0820] Examples of prompts the system might use when using a generative AI model include:

[0821] "If workers are feeling very stressed, generate advice to help them relieve it."

[0822] This prompt allows the system to provide appropriate advice for the worker.

[0823] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0824] Step 1:

[0825] The server uses a generative AI model to generate a personalized list of questions, taking into account the user's past data and common mental health-related questions.

[0826] Input: User history data, general question data

[0827] Data processing: Generate a list of questions using a generative AI model

[0828] Output: A personalized list of questions

[0829] Step 2:

[0830] The terminal presents the list of questions sent from the server to the user by voice or text.

[0831] Input: Personalized Question List

[0832] Output: Questions presented in audio or text format

[0833] What happens: When the device uses a speech synthesis engine to play a question aloud, it presents the user with a question such as "Have you felt stressed at work recently?"

[0834] Step 3:

[0835] The user answers the questions by voice, and the terminal records the voice response and transmits it to the server.

[0836] Input: User's spoken response

[0837] Output: Recorded audio data

[0838] Specific operation: When the user answers, for example, "Yes, writing the report was very stressful," the device records the voice and sends it to the server.

[0839] Step 4:

[0840] The server analyzes the received voice data using a voice analysis engine to extract voice characteristics such as intonation, speed, and volume. It also analyzes the text content using an NLP engine and recognizes the user's emotions using an emotion engine.

[0841] Input: Recorded audio data

[0842] Data processing: Analyze voice data with a voice analysis engine and extract features. Analyze text with an NLP engine. Recognize emotions with an emotion engine.

[0843] Output: Emotional state, stress level data

[0844] Specific operation: The server analyzes the voice data, and if the user says "I felt very stressed," it recognizes emotions such as "tension" and "anxiety."

[0845] Step 5:

[0846] The server evaluates the user's stress level based on emotion recognition data and generates appropriate advice and countermeasures.

[0847] Input: Emotional state, stress level data

[0848] Data processing: Based on the evaluation results, advice and countermeasures are created using an advice generation AI model.

[0849] Output: Specific advice and countermeasures

[0850] Specific actions: The server generates specific advice such as "Take appropriate breaks while writing the report. We also recommend doing some light stretching."

[0851] Step 6:

[0852] The terminal displays advice and countermeasures from the server to the user and simultaneously plays healing music.

[0853] Input: Specific advice, countermeasures, healing music data

[0854] Output: Displaying advice to the user, playing healing music

[0855] Specific operation: The device displays advice such as "Take regular breaks while writing your report" on the screen and plays relaxing music.

[0856] Step 7:

[0857] The server uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[0858] Input: User's stress level data

[0859] Data processing: Generating healing music using generative AI models

[0860] Output: Healing music data

[0861] Specific operation: The server generates relaxing music according to the user's stress level and sends it to the device.

[0862] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0863] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0864] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0865] [Third embodiment]

[0866] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0867] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0868] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0869] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0870] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0871] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0872] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0873] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0874] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0875] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0876] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0877] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0878] The present invention is a system for managing employees' mental health and providing specific advice and countermeasures. This system includes functions for generating questions related to mental health, acquiring and analyzing audio data, generating advice, and playing healing music.

[0879] System Configuration

[0880] The system consists of the following main elements:

[0881] 1. Server: This is the central data processing device that generates questions, analyzes voice, evaluates stress, generates advice, and generates healing music.

[0882] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[0883] 3. User: An employee who uses the system.

[0884] System Operation

[0885] 1. Question generation and retrieval

[0886] The server generates questions about the user's mental health, using generative AI to create a personalized list of questions. For example:

[0887] "What kind of stress have you been experiencing at work recently?"

[0888] "Are you getting enough rest?"

[0889] The questions are displayed in order below.

[0890] The terminal presents the generated question to the user by voice or text.

[0891] The user responds to the question by voice, for example, "I felt a little stressed during the meeting."

[0892] The terminal records the user's voice response and transmits it to the server.

[0893] 2. Voice analysis and stress assessment

[0894] The server analyzes the received voice data using a speech analysis engine to analyze the intonation, speed, and volume of the voice, and a natural language processing (NLP) engine to analyze the text content, which then evaluates the user's emotional state and stress level.

[0895] Specifically, a high intonation and fast speech rate is judged to indicate a state of tension. Additionally, if the answer contains keywords such as "stress" or "anxiety," the stress level is also assessed as high.

[0896] 3. Advice Generation and Presentation

[0897] Based on the evaluation results, the server generates specific advice and countermeasures. The AI ​​generator makes optimal suggestions based on the user's situation. For example,

[0898] "Light exercise is good. Walking and stretching are recommended."

[0899] "Increasing communication with other employees is also effective."

[0900] The terminal displays the generated advice and countermeasures to the user.

[0901] 4. Play healing music

[0902] While the advice is displayed, the server uses a generation AI to create healing music and send it to the device.

[0903] The device plays soothing music along with a screen displaying advice and suggested solutions, allowing the user to receive the presented information in a relaxed state.

[0904] Specific examples

[0905] In one scenario, the following specific exchanges take place:

[0906] 1. Device: "Have you felt stressed at work recently?"

[0907] 2. User: "Yes, writing the report was very stressful."

[0908] 3. Server: Performs voice analysis and NLP analysis and determines that the stress level is high.

[0909] 4. Server: Generates the advice, "Take regular breaks while writing your report. It's also a good idea to do some light stretching."

[0910] 5. Device: Displays advice and simultaneously plays relaxing, healing music.

[0911] This allows users to receive specific advice and learn how to reduce stress in a relaxed environment.

[0912] By using this system, employees can improve their mental health through dialogue with the robot, creating a comfortable working environment.

[0913] The processing flow will be explained below.

[0914] Step 1:

[0915] The server starts up and makes the user database accessible, preparing the entire system for operation.

[0916] Step 2:

[0917] The terminal displays a login screen to the user, providing a screen with fields for entering a user ID and password.

[0918] Step 3:

[0919] The user enters login information (user ID and password) and presses the login button. The login information is sent to the system.

[0920] Step 4:

[0921] The terminal sends the user's login information to the server, which performs authentication and, if successful, proceeds to the next step.

[0922] Step 5:

[0923] The server uses generative AI to generate a personalized list of mental health questions for the user, based on the user's past data and general mental health questions.

[0924] Step 6:

[0925] The device presents the generated list of questions to the user via voice or text, for example, "How stressful are you feeling at work these days?"

[0926] Step 7:

[0927] The user answers the question by voice. For example, the user might say, "I felt a little stressed during the meeting."

[0928] Step 8:

[0929] The device records the user's voice response and sends it to the server, where the user's response arrives as recorded data.

[0930] Step 9:

[0931] The server uses a speech analysis engine to analyze the received voice data, extracting voice intonation, speed, and volume characteristics.

[0932] Step 10:

[0933] The server uses a natural language processing (NLP) engine to analyze the content of the voice data and extract keywords related to emotions and stress.

[0934] Step 11:

[0935] The server integrates the results of the voice and text analysis to evaluate the user's emotional state and stress level, and generates an evaluation result for the next step.

[0936] Step 12:

[0937] Based on the stress assessment results, the server uses AI to generate specific advice and countermeasures, such as "Light exercise would be good. Walking and stretching are recommended."

[0938] Step 13:

[0939] The server sends the generated advice and countermeasures to the device, which then prepares the countermeasures to be provided to the user.

[0940] Step 14:

[0941] The device displays advice and solutions to the user, while soothing music is played in the background.

[0942] Step 15:

[0943] The device plays soothing music while providing advice and solutions to the user, creating an environment that enhances relaxation.

[0944] Step 16:

[0945] The device displays a screen requesting the user for advice and feedback on the effects of the healing music.

[0946] Step 17:

[0947] The user provides feedback via voice or text, for example, "I found the music very relaxing."

[0948] Step 18:

[0949] The device sends the user's feedback to the server, which uses the feedback data to generate questions and advice for the next time.

[0950] Step 19:

[0951] The server records the feedback and uses the data to improve the system and user experience in the future.

[0952] Example 1

[0953] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0954] In today's workforce, managing employee mental health is becoming increasingly important. Many employees are exposed to stress and pressure on a daily basis, often resulting in decreased productivity and health problems. However, the means to provide effective mental health care remain limited. In particular, it is difficult to quickly provide personalized advice and measures tailored to each employee's situation. Therefore, there is a need for a comprehensive mental health management system that can accurately assess employees' stress levels, provide appropriate advice, and combine relaxation elements such as healing music.

[0955] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0956] In this invention, the server includes a generating means for generating questions related to mental health, an acquiring means for acquiring employees' voice responses to the questions, an analyzing means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generating means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, and a presenting means for presenting the advice and proposed measures and generating and playing healing music, thereby making it possible to provide personalized mental health care to each employee.

[0957] "Generation means" refers to a device or program for generating questions related to mental health.

[0958] The term "acquisition means" refers to a device or program for acquiring a user's voice response.

[0959] "Analysis means" refers to a device or program for analyzing the acquired voice response and evaluating emotions and stress levels.

[0960] The "advice generation means" refers to a device or program that generates specific advice or countermeasures based on the evaluation results obtained by analysis.

[0961] The "presentation means" refers to a device or program that presents the generated advice and countermeasures to the user and also generates and plays healing music.

[0962] "Emotions and stress levels" refers to the emotional state the user is feeling and the associated degree of stress.

[0963] "Healing music" refers to music that is intended to relax the user and reduce stress.

[0964] A "generative AI model" refers to an algorithm or program that uses artificial intelligence to generate questions, advice, music, etc.

[0965] "Voice analysis" refers to the process of analyzing acquired voice data in terms of intonation, speed, volume, etc.

[0966] A "natural language processing engine" refers to a program or algorithm that converts voice data into text and analyzes the content.

[0967] This invention is a system for managing employees' mental health and providing specific advice and countermeasures. This system includes functions for generating questions related to mental health, acquiring and analyzing audio data, generating advice, and playing healing music. The main elements of the system are a server, a terminal, and a user.

[0968] System Configuration

[0969] 1. Server

[0970] The server is the central data processing unit, generating questions using a generative AI model and analyzing the voice data. It also generates specific advice based on the analysis results and creates healing music. Software such as "PRAAT" and "SpaCy" are used for voice analysis.

[0971] 2. Terminal

[0972] The terminal is a device operated by the user, and has the function of presenting questions sent from the server, receiving voice responses from the user, displaying analysis results and advice, and playing healing music.

[0973] 3. Users

[0974] Users are employees who use the system, answer mental health questions through their terminals, and receive the advice provided.

[0975] System Operation

[0976] The server uses a generative AI model to generate mental health-related questions that are personalized based on the user's previous answers and general mental health care knowledge.

[0977] The device presents the generated question to the user by voice or text, for example, "What has been making you feel stressed recently?"

[0978] The user answers the questions by voice, for example, "I felt very stressed writing the report."

[0979] The terminal records the user's voice response and transmits it to the server.

[0980] The server uses the speech analysis engine "PRAAT" to analyze the intonation, speed, and volume of the voice, and the natural language processing engine "SpaCy" to analyze the text content, which then evaluates the user's emotional state and stress level.

[0981] Based on the evaluation results, the server uses a generative AI model to generate specific advice and countermeasures for the user. For example, it might suggest, "It would be good to do some light exercise. Walking and stretching are recommended."

[0982] Once the evaluation results and advice generation are complete, the server uses the generative AI model to create healing music and send it to the device.

[0983] The device displays the generated advice and countermeasures to the user while simultaneously playing soothing music, allowing the user to receive specific advice in a relaxed state.

[0984] Specific examples

[0985] For example, in a scenario on one day, the following exchange takes place:

[0986] 1. Device: "What has been stressing you out lately?"

[0987] 2. User: "Yes, writing the report was very stressful."

[0988] 3. The device records this response and sends it to the server.

[0989] 4. Server: Conduct voice analysis and NLP analysis to identify high stress levels associated with "report writing."

[0990] 5. Server: Generate the advice, "Take regular breaks while writing your report. It's also a good idea to do some light stretching."

[0991] 6. Device: Displays advice and simultaneously plays relaxing, healing music.

[0992] Example prompt sentence:

[0993] "Please write a natural language description of a system that manages employees' mental health and provides specific advice. The system includes functions for question generation, voice analysis, advice generation, and soothing music playback. Please describe the process in detail along the following lines."

[0994] By using this system, employees can manage their own mental health more effectively and create a comfortable working environment.

[0995] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0996] Step 1: Question Generation

[0997] The server uses a generative AI model to generate questions related to mental health. At this time, the server obtains the user's profile information and past answers and creates a personalized list of questions. The input data is the user's profile information and past answer data, and the output is the generated list of questions. For example, the server generates a question such as, "What has been making you feel stressed recently?"

[0998] Step 2: Question posing

[0999] The terminal presents the questions sent from the server to the user. The input data is the list of questions sent from the server, and the output is a display screen or audio presented to the user. For example, the terminal may read out the question "What has been making you feel stressed recently?" or display it on the screen.

[1000] Step 3: Get the answer

[1001] The user answers the presented question by voice. The input data is the question presented to the user, and the output is the user's voice response. For example, the user might respond, "Writing the report was very stressful."

[1002] Step 4: Acquire and transmit audio data

[1003] The terminal acquires and records the user's voice response. Then, it transmits this voice data to the server. The input data is the user's voice response, and the output is the voice data transmitted to the server. For example, the terminal transmits voice data such as "I felt very stressed while writing the report" to the server.

[1004] Step 5: Audio analysis

[1005] The server analyzes the received voice data. For the analysis, it uses the voice analysis engine "PRAAT" and the natural language processing engine "SpaCy." The input data is the user's voice data, and the output is the voice intonation, speed, volume, and text analysis results. Specifically, the server analyzes the voice data "I felt very stressed writing the report" for voice intonation, speed, and volume using "PRAAT," and converts it into text and analyzes it using "SpaCy."

[1006] Step 6: Stress Assessment

[1007] The server evaluates the user's stress level based on the results of voice and text analysis. The input data are the voice and text analysis results, and the output is the user's stress level assessment. For example, if the speech is high-pitched, fast-paced, and includes the word "stress," the server will assess the user's stress level as high.

[1008] Step 7: Advice Generation

[1009] The server uses the generative AI model to generate specific advice and countermeasures for the user based on the evaluation results. The input data is the stress level evaluation result, and the output is the generated advice and countermeasures. For example, it generates advice such as, "Take appropriate breaks while writing your report. It is also recommended that you do some light stretching."

[1010] Step 8: Providing advice

[1011] The terminal displays the advice and countermeasures sent from the server to the user. The input data is the advice and countermeasures sent from the server, and the output is the display screen and audio presented to the user. For example, advice such as "Take appropriate breaks while writing a report" can be displayed on the screen and read aloud.

[1012] Step 9: Generate and Play Healing Music

[1013] The server uses a generative AI model to create healing music and sends it to the device. The input data is the stress assessment results, and the output is the generated healing music. For example, a song with a relaxing effect is generated. The device receives this healing music and plays it along with the advice display. The user can receive advice in a relaxing environment.

[1014] (Application example 1)

[1015] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1016] In today's work environment, many employees are increasingly experiencing stress and emotional strain, which increases the risk of decreased productivity and poor mental health. Furthermore, virtual store shopping experiences often leave users feeling overwhelmed by numerous options and information, leading to stress. There is a need for an effective mental health management system that can improve these situations and help employees and virtual store users feel refreshed, improving their productivity and satisfaction.

[1017] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1018] In this invention, the server includes a generating means for generating questions related to mental health, an acquiring means for acquiring the user's voice responses to the questions, an analyzing means for analyzing the acquired voice responses and evaluating the user's emotions and stress level, an advice generating means for generating specific advice and proposed measures based on the evaluation results, a presenting means for presenting the advice and proposed measures and playing soothing music, and a refreshing means for presenting refreshing advice to improve the user's experience in the virtual space. This enables practical support to reduce stress felt by the user while shopping in a virtual store and to refresh the user.

[1019] The "means for generating mental health-related questions" is a device or method that dynamically generates appropriate questions to assess the mental health status of a user.

[1020] The "means for acquiring the employee's voice response to the question" refers to a device or method for acquiring voice data in which the user answers the question and transmitting that data to the system.

[1021] The "analysis means for analyzing the acquired voice response and evaluating the emotion and stress level" is a device or method for analyzing voice data and evaluating the user's emotional state and stress level.

[1022] The "advice generating means for generating specific advice and countermeasures" is a device or method for generating appropriate advice and countermeasures for the user based on the analysis results.

[1023] The "presentation means for presenting advice and measures and playing healing music" is a device or method for presenting the generated advice and measures to the user and playing healing music at the same time.

[1024] A "refreshment means for presenting refreshment advice to improve the user experience in a virtual space" is a device or method that presents specific advice to refresh a user when the user feels stressed in a virtual store.

[1025] MODE FOR CARRYING OUT THE INVENTION

[1026] This invention provides a system for managing employee mental health and improving user experience in a virtual store. The following configuration makes it possible to evaluate the user's emotional state, generate optimal advice, and provide soothing music.

[1027] System Configuration

[1028] The system consists of the following elements:

[1029] 1. Server: A central data processing device that generates questions, analyzes voice, evaluates stress, generates advice, and generates healing music.

[1030] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[1031] 3. User: A user who uses the system.

[1032] question generation

[1033] The server generates questions related to the user's mental health, personalized using a generative AI model. For example:

[1034] It generates questions such as, "Have you felt stressed while shopping online recently?"

[1035] Acquiring audio data

[1036] The device presents the generated question to the user by voice or text. The user answers the question by voice. For example, the user might say, "Yes, there are so many products, I don't know which one to choose." This voice data is captured by the device and sent to the server.

[1037] Voice analysis and stress assessment

[1038] The server analyzes the acquired voice data using a speech recognition engine and a natural language processing (NLP) engine. The server analyzes the intonation, speed, volume, etc. of the voice to evaluate the user's emotional state and stress level. For example, a high intonation and fast speed of the voice may be determined to indicate a state of tension.

[1039] Advice generation and presentation

[1040] The server generates specific advice and countermeasures based on the analysis results. The generated advice provides specific support for the user's actions. For example,

[1041] Advice such as "Take a deep breath to relax. We recommend drinking tea to refresh yourself" is generated.

[1042] Playing healing music

[1043] The device not only displays the generated advice but also plays soothing music at the same time, allowing the user to receive the advice in a relaxed state.

[1044] Specific examples

[1045] For example, if a user using a virtual store responds that they "felt stressed," the server analyzes the voice data and detects a high stress level. As a result, the server generates advice such as "Take a deep breath to relax. We recommend drinking tea to refresh yourself," and the device displays this advice to the user while playing soothing music.

[1046] Prompt Sentence Examples

[1047] "Please assess this user's mental health status and provide appropriate advice: Yes, there are so many options that it's hard to know which one to choose."

[1048] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1049] Step 1:

[1050] The server generates questions related to the user's mental health, using a generative AI model to create a personalized list of questions, such as, "Have you felt stressed while shopping online recently?" These questions are customized based on the user's past feedback and shopping history.

[1051] Step 2:

[1052] The device presents the generated question to the user via voice or text. The user can receive the question through either visual or auditory means. For example, the device might ask the user aloud, "Have you felt stressed while shopping online recently?" The user listens to the question and understands its content.

[1053] Step 3:

[1054] The user answers the question by voice. For example, they might say, "Yes, there are so many products, I don't know which one to choose." The user's voice is picked up through the device's microphone.

[1055] Step 4:

[1056] The terminal records the user's voice responses and converts the data into a data format for transmission to the server. The terminal converts the raw voice data as input into an audio file and transmits it to the server over the network.

[1057] Step 5:

[1058] The server analyzes the acquired voice data and converts the voice into text using a voice recognition engine. This text data is then analyzed using an NLP engine to evaluate the emotional state and stress level. For example, the intonation and speed of the voice can be analyzed to determine the state of tension and stress level.

[1059] Step 6:

[1060] The server generates specific advice and countermeasures based on the evaluation results. Using a generative AI model, it creates advice tailored to the user's condition. For example, it generates advice such as "Take deep breaths to relax. We recommend drinking tea to refresh yourself." The input is analyzed emotion and stress data, and the output is specific written advice for the user.

[1061] Step 7:

[1062] The terminal presents the generated advice to the user in text or audio format. For example, the generated advice may be displayed on the terminal's display and simultaneously played back as audio, allowing the user to receive the advice both visually and audibly.

[1063] Step 8:

[1064] The device plays healing music while providing advice. The server sends the generated healing music data to the device, providing a relaxing environment for the user. By playing the healing music, the user can refresh their mind and body while receiving advice.

[1065] This allows users to receive practical support to reduce stress and feel refreshed while shopping in a virtual store.

[1066] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1067] This invention is a system for effectively managing employee mental health, and in particular, by combining an emotion engine, it is possible to provide more accurate stress assessments and personalized advice. This system integrates the following functions: generating mental health-related questions, acquiring and analyzing voice data, recognizing emotions, evaluating stress, generating advice, and playing healing music.

[1068] System Configuration

[1069] The system consists of the following main elements:

[1070] 1. Server: The central data processing center, which generates questions, analyzes voice, evaluates stress, recognizes emotions, generates advice, and generates soothing music.

[1071] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[1072] 3. User: An employee who uses the system.

[1073] 4. Emotion Engine: Recognizes emotions from the user's voice responses and helps assess stress levels.

[1074] System Operation

[1075] 1. Question generation and retrieval

[1076] The server generates questions about mental health. It uses generative AI to create a personalized list of questions, taking into account the user's past data and common mental health questions. For example, questions might include, "How stressed are you at work these days?" and "How do you spend your time outside of work?"

[1077] The terminal presents the generated question to the user by voice or text.

[1078] The user answers the questions by voice, for example, "I've had a lot of meetings this week and I'm feeling a bit stressed."

[1079] The terminal records the user's voice response and transmits it to the server.

[1080] 2. Voice Analysis and Emotion Recognition

[1081] The server analyzes the received voice data using a voice analysis engine, extracting voice characteristics such as intonation, speed, and volume, and then analyzes the text content using a natural language processing (NLP) engine.

[1082] The server then uses an emotion engine to recognize the user's emotions from the voice data, for example, recognizing emotional states such as "tension," "anxiety," or "relief" from the tone of voice and speaking style.

[1083] 3. Stress Assessment

[1084] The server evaluates the user's stress level based on emotional data recognized through voice analysis and the emotion engine. Specifically, if the emotion engine recognizes "tension" or "anxiety," the stress level is determined to be high.

[1085] 4. Advice Generation and Presentation

[1086] Based on the evaluation results, the server generates appropriate advice and countermeasures. The AI ​​generator provides personalized advice tailored to the user's situation. For example, specific advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension" is provided.

[1087] The device displays the generated advice and countermeasures to the user, and is set to play healing music at the same time.

[1088] 5. Play healing music

[1089] The server uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[1090] The device plays soothing music while displaying advice and suggested solutions on a screen, allowing the user to receive the presented information in a relaxed state.

[1091] Specific examples

[1092] A typical day scenario could involve the following exchange:

[1093] 1. Device: "Have you felt stressed at work recently?"

[1094] 2. User: "Yes, writing the report was very stressful."

[1095] 3. Server: Performs voice analysis and NLP analysis, and recognizes "tension" using the emotion engine.

[1096] 4. Server: Generates the advice, "Take regular breaks while writing your report. We also recommend doing some light stretching."

[1097] 5. Device: Displays advice and plays relaxing, soothing music.

[1098] This system allows employees to effectively relieve stress and improve their mental health at work. By incorporating an emotion engine, it provides more accurate assessments and specific advice tailored to each user's condition.

[1099] The processing flow will be explained below.

[1100] Step 1:

[1101] The server starts up and makes the user database accessible, preparing the entire system for operation.

[1102] Step 2:

[1103] The terminal displays a login screen to the user, providing a screen with fields for entering a user ID and password.

[1104] Step 3:

[1105] The user enters login information (user ID and password) and presses the login button. The login information is sent to the system.

[1106] Step 4:

[1107] The terminal sends the user's login information to the server, which performs authentication and, if successful, proceeds to the next step.

[1108] Step 5:

[1109] The server uses generative AI to generate a personalized list of mental health questions for the user, based on the user's past data and general mental health questions.

[1110] Step 6:

[1111] The device presents the generated list of questions to the user via voice or text, for example, "How stressful are you feeling at work these days?"

[1112] Step 7:

[1113] The user answers the question by voice. For example, the user might say, "I felt a little stressed during the meeting."

[1114] Step 8:

[1115] The device records the user's voice response and sends it to the server, where the user's response arrives as recorded data.

[1116] Step 9:

[1117] The server uses a speech analysis engine to analyze the received voice data, extracting voice intonation, speed, and volume characteristics.

[1118] Step 10:

[1119] The server uses a natural language processing (NLP) engine to analyze the content of the voice data and extract keywords related to emotions and stress.

[1120] Step 11:

[1121] The server uses an emotion engine to recognize the user's emotions from the voice data, for example, recognizing emotional states such as "tension," "anxiety," and "relief" from the tone of voice and speaking style.

[1122] Step 12:

[1123] The server integrates the results of the voice analysis and emotion recognition by the emotion engine to evaluate the user's stress level, and the evaluation result is generated for the next step.

[1124] Step 13:

[1125] Based on the stress assessment results, the server uses AI to generate specific advice and countermeasures, such as "Light exercise would be good. Walking and stretching are recommended."

[1126] Step 14:

[1127] The server sends the generated advice and countermeasures to the device, which then prepares the countermeasures to be provided to the user.

[1128] Step 15:

[1129] The device displays advice and suggested solutions to the user. The display is confirmed to be correct.

[1130] Step 16:

[1131] The device is designed to play healing music sent from the server, and will begin playing it while simultaneously displaying advice and suggestions for countermeasures.

[1132] Step 17:

[1133] The device displays a screen requesting the user for advice and feedback on the effects of the healing music.

[1134] Step 18:

[1135] The user provides feedback via voice or text, for example, "I found the music very relaxing."

[1136] Step 19:

[1137] The device sends the user's feedback to the server, which uses the feedback data to generate questions and advice for the next time.

[1138] Step 20:

[1139] The server records the feedback and uses the data to improve the system and user experience in the future.

[1140] Example 2

[1141] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1142] Conventional mental health management systems rely on general questions and self-reported assessments by users, making it difficult to accurately grasp emotions and stress levels. Furthermore, they lack the personalization required to provide appropriate advice and stress reduction methods, limiting their effectiveness in improving users' mental health. Therefore, there is a need for a system that provides more accurate and personalized emotion assessment and stress reduction.

[1143] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1144] In this invention, the server includes a generation means for generating questions related to mental health, an acquisition means for acquiring employees' voice responses to the questions, an analysis means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generation means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, a presentation means for presenting the advice and proposed measures and playing soothing music, a means for analyzing the intonation, speed, volume, etc. of the voice and analyzing the text using natural language processing, a means for recognizing the user's emotions from the voice data using an emotion engine, and a means for generating and playing soothing music optimized for the stress level. This makes it possible to accurately evaluate the user's emotions and stress level and provide personalized advice and stress relief methods.

[1145] A "generator" is a device or software function for generating questions related to a user's mental health.

[1146] The "acquisition means" is a device or software function for recording the user's voice response and inputting it into the system.

[1147] The "analysis means" is a device or software function for analyzing the acquired voice data and evaluating the user's emotions and stress level.

[1148] The "advice generation means" is a device or software function for creating specific advice or countermeasures for the user based on the evaluation results obtained by the analysis means.

[1149] The "presentation means" is a device or software function for showing the generated advice and countermeasures to the user and playing healing music.

[1150] "Means for analyzing voice intonation, speed, volume, etc." refers to the function of a device or software for analyzing voice characteristics such as intonation, speed, volume, etc. contained in voice data.

[1151] "Means for analyzing text using natural language processing" refers to a device or software function that uses natural language processing techniques to analyze text data converted from voice data.

[1152] The "means for recognizing a user's emotion from voice data using an emotion engine" is a function of the emotion engine used to recognize a user's emotion based on voice data.

[1153] The "means for generating and playing healing music optimized for a stress level" refers to a device or software function for creating and playing optimal healing music according to a user's stress level.

[1154] This invention is a system designed to effectively manage the mental health of employees. The system mainly consists of three entities: a server, a terminal, and a user, each of which plays a specific role.

[1155] Question generation and retrieval

[1156] The server generates questions related to employees' mental health. Specifically, it uses a generative AI model (e.g., OpenAI's GPT-4) to create a personalized list of questions based on the user's past data and common mental health questions. For example, questions include, "What kind of stress have you been feeling at work recently?" and "How do you spend your time outside of work?"

[1157] The device receives the questions sent from the server and presents them to the user in voice or text. For voice presentation, speech synthesis software (e.g., Google Cloud Text-to-Speech) is used.

[1158] The user answers the questions by voice, for example, "I've had a lot of meetings this week and I'm feeling a bit stressed." The device records this voice and sends it to the server. The voice data is first stored locally and then uploaded to the server using a secure communication protocol (e.g., HTTPS).

[1159] Voice analysis and emotion recognition

[1160] The server converts the received voice data into text data using a voice analysis engine (e.g., Google Cloud Speech-to-Text), which also extracts voice characteristics such as intonation, speed, and volume.

[1161] Next, the server uses a natural language processing (NLP) engine (e.g., OpenAI's GPT-4) to analyze the text content and understand what the user is saying. Using the analysis results and voice feature data, an emotion recognition engine (e.g., Microsoft's Azure Emotion API) recognizes the user's emotional state. Specifically, emotions such as "tension," "anxiety," and "relief" are detected from the tone of voice and speaking style.

[1162] Stress assessment

[1163] The server evaluates the user's stress level based on the results of voice analysis and emotion recognition. If the emotion recognition engine recognizes "tension" or "anxiety," the stress level is deemed high; conversely, if it recognizes "relief," it is deemed low. This evaluation result is recorded within the system and used for subsequent analysis and the development of countermeasures.

[1164] Advice generation and presentation

[1165] The server generates appropriate advice based on the stress assessment results. It uses a generative AI model (e.g., OpenAI's GPT-4) to create personalized advice based on the assessment results and past user data. For example, it provides specific advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension."

[1166] The device displays the advice sent from the server to the user and simultaneously plays healing music. Music playback is performed using a media player or streaming service (e.g., Spotify API).

[1167] Playing healing music

[1168] The server uses a generative AI model to generate soothing music optimized for the user's stress level and transmits it to the device, which then plays the soothing music in sync with the user's advice, providing a relaxing environment for the user.

[1169] Prompt Sentence Examples

[1170] In a one day scenario, the following prompts could be used:

[1171] "Tell me about your work environment recently. If you have had any stressful experiences, please tell me about them in detail."

[1172] This allows users to effectively manage stress and receive appropriate advice in a comfortable environment. In this way, the system is able to perform highly accurate emotion recognition and stress assessment according to the individual state of the user, and provide specific and personalized advice.

[1173] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1174] Step 1: Question Generation

[1175] The server uses a generative AI model to receive the user's past mental health data and general question data as input, and generates a personalized list of questions based on this. Specifically, questions such as "What kind of stress have you been feeling at work recently?" and "How do you spend your time outside of work?" are generated. The generated list of questions is then sent to the device in the next step.

[1176] Step 2: Posing the Question

[1177] The device receives a list of questions from the server as input and uses speech synthesis software to present the questions to the user either audibly or via text. For example, Google Cloud Text-to-Speech is used to present the questions to the user audibly, and the user can then respond verbally.

[1178] Step 3: Get an audio response

[1179] The user answers the questions by voice, for example, "I had a lot of meetings this week and I felt a bit stressed."

[1180] The device takes this voice response as input, stores it locally, and then transmits the voice data to the server using a secure communication protocol (e.g., HTTPS).

[1181] Step 4: Audio analysis

[1182] The server receives the voice data sent from the device as input and converts it into text using a voice analysis engine (for example, Google Cloud Speech-to-Text). During this process, it extracts voice characteristics such as intonation, speed, and volume, and passes them on to the next process along with the text data.

[1183] Step 5: Natural Language Processing (NLP)

[1184] The server receives the text data and speech feature data obtained through speech analysis as input, analyzes the text content using an NLP engine (for example, OpenAI's GPT-4), and outputs structured data that helps understand what the user said.

[1185] Step 6: Emotion Recognition

[1186] The server receives the analysis results of the NLP engine and the voice feature data as input, and uses an emotion recognition engine (such as Microsoft's Azure Emotion API) to recognize the user's emotions. For example, it detects emotions such as "tension," "anxiety," and "relief" from the tone of voice and speaking style, and generates emotion data based on the results.

[1187] Step 7: Stress Assessment

[1188] The server receives the emotion data output by the emotion recognition engine as input and evaluates the user's stress level. Specifically, if the emotion data is recognized as "tension" or "anxiety," it is judged to be high stress. The evaluation results are stored in an internal database.

[1189] Step 8: Advice Generation

[1190] The server receives the stress assessment results and past user data as input, and uses the generative AI model to generate specific advice and countermeasures. For example, advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension" is generated and sent to the device.

[1191] Step 9: Providing advice and playing healing music

[1192] The device receives advice sent from the server as input and displays it to the user in text. At the same time, it receives healing music optimized for the user's stress level from the server and plays it using a media player or streaming service (e.g., Spotify API). This allows the user to receive advice in a relaxed state.

[1193] (Application example 2)

[1194] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1195] Modern factory workers are under a lot of stress during their work, making mental health care in the workplace extremely important. However, traditional mental health management systems do not adequately consider the characteristics of the workplace or the emotional state of individual workers. Furthermore, to provide effective care, a method is needed to accurately assess workers' emotional states and provide individually tailored advice and measures. Therefore, there is a need to develop a new system that can reduce worker stress and improve their mental health.

[1196] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1197] In this invention, the server includes a generation means for generating questions related to mental health, an acquisition means for acquiring employees' voice responses to the questions, an analysis means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generation means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, a presentation means for presenting the advice and proposed measures and playing soothing music, and a factory robot that collects data from workers through a voice interface and provides personalized advice and soothing music for the mental health care of factory workers. This allows workers to continue working in a less stressful environment, and is expected to improve the overall working environment.

[1198] The "Employee Mental Health Management System" is a system that effectively manages employees' mental health, assesses their emotions and stress levels, and provides specific advice and healing music.

[1199] A "generation means" is a means that has the function of generating questions related to mental health.

[1200] The "acquisition means" is a means having a function of acquiring an employee's voice response to a question.

[1201] The "analysis means" is a means having a function of analyzing the acquired voice response and evaluating emotions and stress levels.

[1202] The "advice generation means" is a means having a function of generating specific advice and countermeasures based on the evaluation results obtained by the analysis.

[1203] The "presentation means" is a means that has the function of presenting advice and countermeasures as well as playing healing music.

[1204] The "Mental Health Care System for Factory Workers" is a system for managing the mental health of factory workers, assessing their emotional state and stress levels, and providing personalized advice and healing music.

[1205] A "voice interface" is an interface for collecting voice data from a user and conducting interaction.

[1206] A "factory robot" is a robot that supports workers in a factory and has mental health management functions.

[1207] This invention is a system for effectively managing the mental health of employees and factory workers, and is characterized by the combination of an emotion engine to realize mental care suited to the working environment. This system integrates the following functions: generating questions related to mental health, acquiring and analyzing voice data, recognizing emotions, evaluating stress, generating advice, and playing healing music.

[1208] System Configuration

[1209] The system consists of the following main components:

[1210] 1. Server: This is the central data processing center, and performs question generation, voice analysis, stress assessment, emotion recognition, advice generation, healing music generation, etc. The server can be, for example, a cloud-based server or an on-premise server.

[1211] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music. Terminals can be smartphones, tablets, desktop computers, etc.

[1212] 3. Users: Employees and workers who use the system.

[1213] 4. Emotion Engine: Recognizes emotions from the user's voice responses and helps assess stress levels. It can use Google Cloud Natural Language API and OpenAI's Emotion Recognition model.

[1214] System Operation

[1215] Question generation and retrieval

[1216] Server: Uses generative AI to create a personalized list of questions, taking into account the user's past data and common mental health-related questions.

[1217] Terminal: Presents the generated questions to the user by voice or text.

[1218] User: Answers questions verbally.

[1219] Terminal: Records the user's voice response and sends it to the server.

[1220] Voice analysis and emotion recognition

[1221] Server: The received voice data is analyzed using a voice analysis engine. Voice characteristics such as intonation, speed, and volume are extracted, and the text content is analyzed using a natural language processing (NLP) engine.

[1222] The server then uses an emotion engine to recognize the user's emotion from the voice data.

[1223] Stress assessment

[1224] Server: Evaluates the user's stress level based on emotional data recognized through voice analysis and the emotion engine. Specifically, if the emotion engine recognizes "tension" or "anxiety," the stress level is determined to be high.

[1225] Advice generation and presentation

[1226] Server: Based on the evaluation results, appropriate advice and countermeasures are generated. The generation AI provides personalized advice tailored to the user's situation.

[1227] Device: Displays the generated advice and countermeasures to the user. Set it to play healing music at the same time.

[1228] Playing healing music

[1229] Server: Uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[1230] Device: The device plays soothing music while displaying advice and suggested solutions, allowing the user to receive the information in a relaxed state.

[1231] Specific examples

[1232] A typical day scenario could involve the following exchange:

[1233] 1. Device: "Have you felt stressed at work recently?"

[1234] 2. User: "Yes, writing the report was very stressful."

[1235] 3. Server: Performs voice analysis and NLP analysis, and recognizes "tension" using the emotion engine.

[1236] 4. Server: Generates the advice, "Take regular breaks while writing your report. We also recommend doing some light stretching."

[1237] 5. Device: Displays advice and plays relaxing, soothing music.

[1238] This system will enable workers to continue working in a less stressful environment, which is expected to improve the overall working environment.

[1239] Examples of prompt statements

[1240] Examples of prompts the system might use when using a generative AI model include:

[1241] "If workers are feeling very stressed, generate advice to help them relieve it."

[1242] This prompt allows the system to provide appropriate advice for the worker.

[1243] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1244] Step 1:

[1245] The server uses a generative AI model to generate a personalized list of questions, taking into account the user's past data and common mental health-related questions.

[1246] Input: User history data, general question data

[1247] Data processing: Generate a list of questions using a generative AI model

[1248] Output: A personalized list of questions

[1249] Step 2:

[1250] The terminal presents the list of questions sent from the server to the user by voice or text.

[1251] Input: Personalized Question List

[1252] Output: Questions presented in audio or text format

[1253] What happens: When the device uses a speech synthesis engine to play a question aloud, it presents the user with a question such as "Have you felt stressed at work recently?"

[1254] Step 3:

[1255] The user answers the questions by voice, and the terminal records the voice response and transmits it to the server.

[1256] Input: User's spoken response

[1257] Output: Recorded audio data

[1258] Specific operation: When the user answers, for example, "Yes, writing the report was very stressful," the device records the voice and sends it to the server.

[1259] Step 4:

[1260] The server analyzes the received voice data using a voice analysis engine to extract voice characteristics such as intonation, speed, and volume. It also analyzes the text content using an NLP engine and recognizes the user's emotions using an emotion engine.

[1261] Input: Recorded audio data

[1262] Data processing: Analyze voice data with a voice analysis engine and extract features. Analyze text with an NLP engine. Recognize emotions with an emotion engine.

[1263] Output: Emotional state, stress level data

[1264] Specific operation: The server analyzes the voice data, and if the user says "I felt very stressed," it recognizes emotions such as "tension" and "anxiety."

[1265] Step 5:

[1266] The server evaluates the user's stress level based on emotion recognition data and generates appropriate advice and countermeasures.

[1267] Input: Emotional state, stress level data

[1268] Data processing: Based on the evaluation results, advice and countermeasures are created using an advice generation AI model.

[1269] Output: Specific advice and countermeasures

[1270] Specific actions: The server generates specific advice such as "Take appropriate breaks while writing the report. We also recommend doing some light stretching."

[1271] Step 6:

[1272] The terminal displays advice and countermeasures from the server to the user and simultaneously plays healing music.

[1273] Input: Specific advice, countermeasures, healing music data

[1274] Output: Displaying advice to the user, playing healing music

[1275] Specific operation: The device displays advice such as "Take regular breaks while writing your report" on the screen and plays relaxing music.

[1276] Step 7:

[1277] The server uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[1278] Input: User's stress level data

[1279] Data processing: Generating healing music using generative AI models

[1280] Output: Healing music data

[1281] Specific operation: The server generates relaxing music according to the user's stress level and sends it to the device.

[1282] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1283] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1284] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1285] [Fourth embodiment]

[1286] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1287] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1288] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1289] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1290] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1291] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1292] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1293] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1294] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1295] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1296] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1297] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1298] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1299] The present invention is a system for managing employees' mental health and providing specific advice and countermeasures. This system includes functions for generating questions related to mental health, acquiring and analyzing audio data, generating advice, and playing healing music.

[1300] System Configuration

[1301] The system consists of the following main elements:

[1302] 1. Server: This is the central data processing device that generates questions, analyzes voice, evaluates stress, generates advice, and generates healing music.

[1303] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[1304] 3. User: An employee who uses the system.

[1305] System Operation

[1306] 1. Question generation and retrieval

[1307] The server generates questions about the user's mental health, using generative AI to create a personalized list of questions. For example:

[1308] "What kind of stress have you been experiencing at work recently?"

[1309] "Are you getting enough rest?"

[1310] The questions are displayed in order below.

[1311] The terminal presents the generated question to the user by voice or text.

[1312] The user responds to the question by voice, for example, "I felt a little stressed during the meeting."

[1313] The terminal records the user's voice response and transmits it to the server.

[1314] 2. Voice analysis and stress assessment

[1315] The server analyzes the received voice data using a speech analysis engine to analyze the intonation, speed, and volume of the voice, and a natural language processing (NLP) engine to analyze the text content, which then evaluates the user's emotional state and stress level.

[1316] Specifically, a high intonation and fast speech rate is judged to indicate a state of tension. Additionally, if the answer contains keywords such as "stress" or "anxiety," the stress level is also assessed as high.

[1317] 3. Advice Generation and Presentation

[1318] Based on the evaluation results, the server generates specific advice and countermeasures. The AI ​​generator makes optimal suggestions based on the user's situation. For example,

[1319] "Light exercise is good. Walking and stretching are recommended."

[1320] "Increasing communication with other employees is also effective."

[1321] The terminal displays the generated advice and countermeasures to the user.

[1322] 4. Play healing music

[1323] While the advice is displayed, the server uses a generation AI to create healing music and send it to the device.

[1324] The device plays soothing music along with a screen displaying advice and suggested solutions, allowing the user to receive the presented information in a relaxed state.

[1325] Specific examples

[1326] In one scenario, the following specific exchanges take place:

[1327] 1. Device: "Have you felt stressed at work recently?"

[1328] 2. User: "Yes, writing the report was very stressful."

[1329] 3. Server: Performs voice analysis and NLP analysis and determines that the stress level is high.

[1330] 4. Server: Generates the advice, "Take regular breaks while writing your report. It's also a good idea to do some light stretching."

[1331] 5. Device: Displays advice and simultaneously plays relaxing, healing music.

[1332] This allows users to receive specific advice and learn how to reduce stress in a relaxed environment.

[1333] By using this system, employees can improve their mental health through dialogue with the robot, creating a comfortable working environment.

[1334] The processing flow will be explained below.

[1335] Step 1:

[1336] The server starts up and makes the user database accessible, preparing the entire system for operation.

[1337] Step 2:

[1338] The terminal displays a login screen to the user, providing a screen with fields for entering a user ID and password.

[1339] Step 3:

[1340] The user enters login information (user ID and password) and presses the login button. The login information is sent to the system.

[1341] Step 4:

[1342] The terminal sends the user's login information to the server, which performs authentication and, if successful, proceeds to the next step.

[1343] Step 5:

[1344] The server uses generative AI to generate a personalized list of mental health questions for the user, based on the user's past data and general mental health questions.

[1345] Step 6:

[1346] The device presents the generated list of questions to the user via voice or text, for example, "How stressful are you feeling at work these days?"

[1347] Step 7:

[1348] The user answers the question by voice. For example, the user might say, "I felt a little stressed during the meeting."

[1349] Step 8:

[1350] The device records the user's voice response and sends it to the server, where the user's response arrives as recorded data.

[1351] Step 9:

[1352] The server uses a speech analysis engine to analyze the received voice data, extracting voice intonation, speed, and volume characteristics.

[1353] Step 10:

[1354] The server uses a natural language processing (NLP) engine to analyze the content of the voice data and extract keywords related to emotions and stress.

[1355] Step 11:

[1356] The server integrates the results of the voice and text analysis to evaluate the user's emotional state and stress level, and generates an evaluation result for the next step.

[1357] Step 12:

[1358] Based on the stress assessment results, the server uses AI to generate specific advice and countermeasures, such as "Light exercise would be good. Walking and stretching are recommended."

[1359] Step 13:

[1360] The server sends the generated advice and countermeasures to the device, which then prepares the countermeasures to be provided to the user.

[1361] Step 14:

[1362] The device displays advice and solutions to the user, while soothing music is played in the background.

[1363] Step 15:

[1364] The device plays soothing music while providing advice and solutions to the user, creating an environment that enhances relaxation.

[1365] Step 16:

[1366] The device displays a screen requesting the user for advice and feedback on the effects of the healing music.

[1367] Step 17:

[1368] The user provides feedback via voice or text, for example, "I found the music very relaxing."

[1369] Step 18:

[1370] The device sends the user's feedback to the server, which uses the feedback data to generate questions and advice for the next time.

[1371] Step 19:

[1372] The server records the feedback and uses the data to improve the system and user experience in the future.

[1373] Example 1

[1374] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1375] In today's workforce, managing employee mental health is becoming increasingly important. Many employees are exposed to stress and pressure on a daily basis, often resulting in decreased productivity and health problems. However, the means to provide effective mental health care remain limited. In particular, it is difficult to quickly provide personalized advice and measures tailored to each employee's situation. Therefore, there is a need for a comprehensive mental health management system that can accurately assess employees' stress levels, provide appropriate advice, and combine relaxation elements such as healing music.

[1376] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1377] In this invention, the server includes a generating means for generating questions related to mental health, an acquiring means for acquiring employees' voice responses to the questions, an analyzing means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generating means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, and a presenting means for presenting the advice and proposed measures and generating and playing healing music, thereby making it possible to provide personalized mental health care to each employee.

[1378] "Generation means" refers to a device or program for generating questions related to mental health.

[1379] The term "acquisition means" refers to a device or program for acquiring a user's voice response.

[1380] "Analysis means" refers to a device or program for analyzing the acquired voice response and evaluating emotions and stress levels.

[1381] The "advice generation means" refers to a device or program that generates specific advice or countermeasures based on the evaluation results obtained by analysis.

[1382] The "presentation means" refers to a device or program that presents the generated advice and countermeasures to the user and also generates and plays healing music.

[1383] "Emotions and stress levels" refers to the emotional state the user is feeling and the associated degree of stress.

[1384] "Healing music" refers to music that is intended to relax the user and reduce stress.

[1385] A "generative AI model" refers to an algorithm or program that uses artificial intelligence to generate questions, advice, music, etc.

[1386] "Voice analysis" refers to the process of analyzing acquired voice data in terms of intonation, speed, volume, etc.

[1387] A "natural language processing engine" refers to a program or algorithm that converts voice data into text and analyzes the content.

[1388] This invention is a system for managing employees' mental health and providing specific advice and countermeasures. This system includes functions for generating questions related to mental health, acquiring and analyzing audio data, generating advice, and playing healing music. The main elements of the system are a server, a terminal, and a user.

[1389] System Configuration

[1390] 1. Server

[1391] The server is the central data processing unit, generating questions using a generative AI model and analyzing the voice data. It also generates specific advice based on the analysis results and creates healing music. Software such as "PRAAT" and "SpaCy" are used for voice analysis.

[1392] 2. Terminal

[1393] The terminal is a device operated by the user, and has the function of presenting questions sent from the server, receiving voice responses from the user, displaying analysis results and advice, and playing healing music.

[1394] 3. Users

[1395] Users are employees who use the system, answer mental health questions through their terminals, and receive the advice provided.

[1396] System Operation

[1397] The server uses a generative AI model to generate mental health-related questions that are personalized based on the user's previous answers and general mental health care knowledge.

[1398] The device presents the generated question to the user by voice or text, for example, "What has been making you feel stressed recently?"

[1399] The user answers the questions by voice, for example, "I felt very stressed writing the report."

[1400] The terminal records the user's voice response and transmits it to the server.

[1401] The server uses the speech analysis engine "PRAAT" to analyze the intonation, speed, and volume of the voice, and the natural language processing engine "SpaCy" to analyze the text content, which then evaluates the user's emotional state and stress level.

[1402] Based on the evaluation results, the server uses a generative AI model to generate specific advice and countermeasures for the user. For example, it might suggest, "It would be good to do some light exercise. Walking and stretching are recommended."

[1403] Once the evaluation results and advice generation are complete, the server uses the generative AI model to create healing music and send it to the device.

[1404] The device displays the generated advice and countermeasures to the user while simultaneously playing soothing music, allowing the user to receive specific advice in a relaxed state.

[1405] Specific examples

[1406] For example, in a scenario on one day, the following exchange takes place:

[1407] 1. Device: "What has been stressing you out lately?"

[1408] 2. User: "Yes, writing the report was very stressful."

[1409] 3. The device records this response and sends it to the server.

[1410] 4. Server: Conduct voice analysis and NLP analysis to identify high stress levels associated with "report writing."

[1411] 5. Server: Generate the advice, "Take regular breaks while writing your report. It's also a good idea to do some light stretching."

[1412] 6. Device: Displays advice and simultaneously plays relaxing, healing music.

[1413] Example prompt sentence:

[1414] "Please write a natural language description of a system that manages employees' mental health and provides specific advice. The system includes functions for question generation, voice analysis, advice generation, and soothing music playback. Please describe the process in detail along the following lines."

[1415] By using this system, employees can manage their own mental health more effectively and create a comfortable working environment.

[1416] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1417] Step 1: Question Generation

[1418] The server uses a generative AI model to generate questions related to mental health. At this time, the server obtains the user's profile information and past answers and creates a personalized list of questions. The input data is the user's profile information and past answer data, and the output is the generated list of questions. For example, the server generates a question such as, "What has been making you feel stressed recently?"

[1419] Step 2: Question posing

[1420] The terminal presents the questions sent from the server to the user. The input data is the list of questions sent from the server, and the output is a display screen or audio presented to the user. For example, the terminal may read out the question "What has been making you feel stressed recently?" or display it on the screen.

[1421] Step 3: Get the answer

[1422] The user answers the presented question by voice. The input data is the question presented to the user, and the output is the user's voice response. For example, the user might respond, "Writing the report was very stressful."

[1423] Step 4: Acquire and transmit audio data

[1424] The terminal acquires and records the user's voice response. Then, it transmits this voice data to the server. The input data is the user's voice response, and the output is the voice data transmitted to the server. For example, the terminal transmits voice data such as "I felt very stressed while writing the report" to the server.

[1425] Step 5: Audio analysis

[1426] The server analyzes the received voice data. For the analysis, it uses the voice analysis engine "PRAAT" and the natural language processing engine "SpaCy." The input data is the user's voice data, and the output is the voice intonation, speed, volume, and text analysis results. Specifically, the server analyzes the voice data "I felt very stressed writing the report" for voice intonation, speed, and volume using "PRAAT," and converts it into text and analyzes it using "SpaCy."

[1427] Step 6: Stress Assessment

[1428] The server evaluates the user's stress level based on the results of voice and text analysis. The input data are the voice and text analysis results, and the output is the user's stress level assessment. For example, if the speech is high-pitched, fast-paced, and includes the word "stress," the server will assess the user's stress level as high.

[1429] Step 7: Advice Generation

[1430] The server uses the generative AI model to generate specific advice and countermeasures for the user based on the evaluation results. The input data is the stress level evaluation result, and the output is the generated advice and countermeasures. For example, it generates advice such as, "Take appropriate breaks while writing your report. It is also recommended that you do some light stretching."

[1431] Step 8: Providing advice

[1432] The terminal displays the advice and countermeasures sent from the server to the user. The input data is the advice and countermeasures sent from the server, and the output is the display screen and audio presented to the user. For example, advice such as "Take appropriate breaks while writing a report" can be displayed on the screen and read aloud.

[1433] Step 9: Generate and Play Healing Music

[1434] The server uses a generative AI model to create healing music and sends it to the device. The input data is the stress assessment results, and the output is the generated healing music. For example, a song with a relaxing effect is generated. The device receives this healing music and plays it along with the advice display. The user can receive advice in a relaxing environment.

[1435] (Application example 1)

[1436] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1437] In today's work environment, many employees are increasingly experiencing stress and emotional strain, which increases the risk of decreased productivity and poor mental health. Furthermore, virtual store shopping experiences often leave users feeling overwhelmed by numerous options and information, leading to stress. There is a need for an effective mental health management system that can improve these situations and help employees and virtual store users feel refreshed, improving their productivity and satisfaction.

[1438] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1439] In this invention, the server includes a generating means for generating questions related to mental health, an acquiring means for acquiring the user's voice responses to the questions, an analyzing means for analyzing the acquired voice responses and evaluating the user's emotions and stress level, an advice generating means for generating specific advice and proposed measures based on the evaluation results, a presenting means for presenting the advice and proposed measures and playing soothing music, and a refreshing means for presenting refreshing advice to improve the user's experience in the virtual space. This enables practical support to reduce stress felt by the user while shopping in a virtual store and to refresh the user.

[1440] The "means for generating mental health-related questions" is a device or method that dynamically generates appropriate questions to assess the mental health status of a user.

[1441] The "means for acquiring the employee's voice response to the question" refers to a device or method for acquiring voice data in which the user answers the question and transmitting that data to the system.

[1442] The "analysis means for analyzing the acquired voice response and evaluating the emotion and stress level" is a device or method for analyzing voice data and evaluating the user's emotional state and stress level.

[1443] The "advice generating means for generating specific advice and countermeasures" is a device or method for generating appropriate advice and countermeasures for the user based on the analysis results.

[1444] The "presentation means for presenting advice and measures and playing healing music" is a device or method for presenting the generated advice and measures to the user and playing healing music at the same time.

[1445] A "refreshment means for presenting refreshment advice to improve the user experience in a virtual space" is a device or method that presents specific advice to refresh a user when the user feels stressed in a virtual store.

[1446] MODE FOR CARRYING OUT THE INVENTION

[1447] This invention provides a system for managing employee mental health and improving user experience in a virtual store. The following configuration makes it possible to evaluate the user's emotional state, generate optimal advice, and provide soothing music.

[1448] System Configuration

[1449] The system consists of the following elements:

[1450] 1. Server: A central data processing device that generates questions, analyzes voice, evaluates stress, generates advice, and generates healing music.

[1451] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[1452] 3. User: A user who uses the system.

[1453] question generation

[1454] The server generates questions related to the user's mental health, personalized using a generative AI model. For example:

[1455] It generates questions such as, "Have you felt stressed while shopping online recently?"

[1456] Acquiring audio data

[1457] The device presents the generated question to the user by voice or text. The user answers the question by voice. For example, the user might say, "Yes, there are so many products, I don't know which one to choose." This voice data is captured by the device and sent to the server.

[1458] Voice analysis and stress assessment

[1459] The server analyzes the acquired voice data using a speech recognition engine and a natural language processing (NLP) engine. The server analyzes the intonation, speed, volume, etc. of the voice to evaluate the user's emotional state and stress level. For example, a high intonation and fast speed of the voice may be determined to indicate a state of tension.

[1460] Advice generation and presentation

[1461] The server generates specific advice and countermeasures based on the analysis results. The generated advice provides specific support for the user's actions. For example,

[1462] Advice such as "Take a deep breath to relax. We recommend drinking tea to refresh yourself" is generated.

[1463] Playing healing music

[1464] The device not only displays the generated advice but also plays soothing music at the same time, allowing the user to receive the advice in a relaxed state.

[1465] Specific examples

[1466] For example, if a user using a virtual store responds that they "felt stressed," the server analyzes the voice data and detects a high stress level. As a result, the server generates advice such as "Take a deep breath to relax. We recommend drinking tea to refresh yourself," and the device displays this advice to the user while playing soothing music.

[1467] Prompt Sentence Examples

[1468] "Please assess this user's mental health status and provide appropriate advice: Yes, there are so many options that it's hard to know which one to choose."

[1469] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1470] Step 1:

[1471] The server generates questions related to the user's mental health, using a generative AI model to create a personalized list of questions, such as, "Have you felt stressed while shopping online recently?" These questions are customized based on the user's past feedback and shopping history.

[1472] Step 2:

[1473] The device presents the generated question to the user via voice or text. The user can receive the question through either visual or auditory means. For example, the device might ask the user aloud, "Have you felt stressed while shopping online recently?" The user listens to the question and understands its content.

[1474] Step 3:

[1475] The user answers the question by voice. For example, they might say, "Yes, there are so many products, I don't know which one to choose." The user's voice is picked up through the device's microphone.

[1476] Step 4:

[1477] The terminal records the user's voice responses and converts the data into a data format for transmission to the server. The terminal converts the raw voice data as input into an audio file and transmits it to the server over the network.

[1478] Step 5:

[1479] The server analyzes the acquired voice data and converts the voice into text using a voice recognition engine. This text data is then analyzed using an NLP engine to evaluate the emotional state and stress level. For example, the intonation and speed of the voice can be analyzed to determine the state of tension and stress level.

[1480] Step 6:

[1481] The server generates specific advice and countermeasures based on the evaluation results. Using a generative AI model, it creates advice tailored to the user's condition. For example, it generates advice such as "Take deep breaths to relax. We recommend drinking tea to refresh yourself." The input is analyzed emotion and stress data, and the output is specific written advice for the user.

[1482] Step 7:

[1483] The terminal presents the generated advice to the user in text or audio format. For example, the generated advice may be displayed on the terminal's display and simultaneously played back as audio, allowing the user to receive the advice both visually and audibly.

[1484] Step 8:

[1485] The device plays healing music while providing advice. The server sends the generated healing music data to the device, providing a relaxing environment for the user. By playing the healing music, the user can refresh their mind and body while receiving advice.

[1486] This allows users to receive practical support to reduce stress and feel refreshed while shopping in a virtual store.

[1487] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1488] This invention is a system for effectively managing employee mental health, and in particular, by combining an emotion engine, it is possible to provide more accurate stress assessments and personalized advice. This system integrates the following functions: generating mental health-related questions, acquiring and analyzing voice data, recognizing emotions, evaluating stress, generating advice, and playing healing music.

[1489] System Configuration

[1490] The system consists of the following main elements:

[1491] 1. Server: The central data processing center, which generates questions, analyzes voice, evaluates stress, recognizes emotions, generates advice, and generates soothing music.

[1492] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music.

[1493] 3. User: An employee who uses the system.

[1494] 4. Emotion Engine: Recognizes emotions from the user's voice responses and helps assess stress levels.

[1495] System Operation

[1496] 1. Question generation and retrieval

[1497] The server generates questions about mental health. It uses generative AI to create a personalized list of questions, taking into account the user's past data and common mental health questions. For example, questions might include, "How stressed are you at work these days?" and "How do you spend your time outside of work?"

[1498] The terminal presents the generated question to the user by voice or text.

[1499] The user answers the questions by voice, for example, "I've had a lot of meetings this week and I'm feeling a bit stressed."

[1500] The terminal records the user's voice response and transmits it to the server.

[1501] 2. Voice Analysis and Emotion Recognition

[1502] The server analyzes the received voice data using a voice analysis engine, extracting voice characteristics such as intonation, speed, and volume, and then analyzes the text content using a natural language processing (NLP) engine.

[1503] The server then uses an emotion engine to recognize the user's emotions from the voice data, for example, recognizing emotional states such as "tension," "anxiety," or "relief" from the tone of voice and speaking style.

[1504] 3. Stress Assessment

[1505] The server evaluates the user's stress level based on emotional data recognized through voice analysis and the emotion engine. Specifically, if the emotion engine recognizes "tension" or "anxiety," the stress level is determined to be high.

[1506] 4. Advice Generation and Presentation

[1507] Based on the evaluation results, the server generates appropriate advice and countermeasures. The AI ​​generator provides personalized advice tailored to the user's situation. For example, specific advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension" is provided.

[1508] The device displays the generated advice and countermeasures to the user, and is set to play healing music at the same time.

[1509] 5. Play healing music

[1510] The server uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[1511] The device plays soothing music while displaying advice and suggested solutions on a screen, allowing the user to receive the presented information in a relaxed state.

[1512] Specific examples

[1513] A typical day scenario could involve the following exchange:

[1514] 1. Device: "Have you felt stressed at work recently?"

[1515] 2. User: "Yes, writing the report was very stressful."

[1516] 3. Server: Performs voice analysis and NLP analysis, and recognizes "tension" using the emotion engine.

[1517] 4. Server: Generates the advice, "Take regular breaks while writing your report. We also recommend doing some light stretching."

[1518] 5. Device: Displays advice and plays relaxing, soothing music.

[1519] This system allows employees to effectively relieve stress and improve their mental health at work. By incorporating an emotion engine, it provides more accurate assessments and specific advice tailored to each user's condition.

[1520] The processing flow will be explained below.

[1521] Step 1:

[1522] The server starts up and makes the user database accessible, preparing the entire system for operation.

[1523] Step 2:

[1524] The terminal displays a login screen to the user, providing a screen with fields for entering a user ID and password.

[1525] Step 3:

[1526] The user enters login information (user ID and password) and presses the login button. The login information is sent to the system.

[1527] Step 4:

[1528] The terminal sends the user's login information to the server, which performs authentication and, if successful, proceeds to the next step.

[1529] Step 5:

[1530] The server uses generative AI to generate a personalized list of mental health questions for the user, based on the user's past data and general mental health questions.

[1531] Step 6:

[1532] The device presents the generated list of questions to the user via voice or text, for example, "How stressful are you feeling at work these days?"

[1533] Step 7:

[1534] The user answers the question by voice. For example, the user might say, "I felt a little stressed during the meeting."

[1535] Step 8:

[1536] The device records the user's voice response and sends it to the server, where the user's response arrives as recorded data.

[1537] Step 9:

[1538] The server uses a speech analysis engine to analyze the received voice data, extracting voice intonation, speed, and volume characteristics.

[1539] Step 10:

[1540] The server uses a natural language processing (NLP) engine to analyze the content of the voice data and extract keywords related to emotions and stress.

[1541] Step 11:

[1542] The server uses an emotion engine to recognize the user's emotions from the voice data, for example, recognizing emotional states such as "tension," "anxiety," and "relief" from the tone of voice and speaking style.

[1543] Step 12:

[1544] The server integrates the results of the voice analysis and emotion recognition by the emotion engine to evaluate the user's stress level, and the evaluation result is generated for the next step.

[1545] Step 13:

[1546] Based on the stress assessment results, the server uses AI to generate specific advice and countermeasures, such as "Light exercise would be good. Walking and stretching are recommended."

[1547] Step 14:

[1548] The server sends the generated advice and countermeasures to the device, which then prepares the countermeasures to be provided to the user.

[1549] Step 15:

[1550] The device displays advice and suggested solutions to the user. The display is confirmed to be correct.

[1551] Step 16:

[1552] The device is designed to play healing music sent from the server, and will begin playing it while simultaneously displaying advice and suggestions for countermeasures.

[1553] Step 17:

[1554] The device displays a screen requesting the user for advice and feedback on the effects of the healing music.

[1555] Step 18:

[1556] The user provides feedback via voice or text, for example, "I found the music very relaxing."

[1557] Step 19:

[1558] The device sends the user's feedback to the server, which uses the feedback data to generate questions and advice for the next time.

[1559] Step 20:

[1560] The server records the feedback and uses the data to improve the system and user experience in the future.

[1561] Example 2

[1562] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1563] Conventional mental health management systems rely on general questions and self-reported assessments by users, making it difficult to accurately grasp emotions and stress levels. Furthermore, they lack the personalization required to provide appropriate advice and stress reduction methods, limiting their effectiveness in improving users' mental health. Therefore, there is a need for a system that provides more accurate and personalized emotion assessment and stress reduction.

[1564] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1565] In this invention, the server includes a generation means for generating questions related to mental health, an acquisition means for acquiring employees' voice responses to the questions, an analysis means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generation means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, a presentation means for presenting the advice and proposed measures and playing soothing music, a means for analyzing the intonation, speed, volume, etc. of the voice and analyzing the text using natural language processing, a means for recognizing the user's emotions from the voice data using an emotion engine, and a means for generating and playing soothing music optimized for the stress level. This makes it possible to accurately evaluate the user's emotions and stress level and provide personalized advice and stress relief methods.

[1566] A "generator" is a device or software function for generating questions related to a user's mental health.

[1567] The "acquisition means" is a device or software function for recording the user's voice response and inputting it into the system.

[1568] The "analysis means" is a device or software function for analyzing the acquired voice data and evaluating the user's emotions and stress level.

[1569] The "advice generation means" is a device or software function for creating specific advice or countermeasures for the user based on the evaluation results obtained by the analysis means.

[1570] The "presentation means" is a device or software function for showing the generated advice and countermeasures to the user and playing healing music.

[1571] "Means for analyzing voice intonation, speed, volume, etc." refers to the function of a device or software for analyzing voice characteristics such as intonation, speed, volume, etc. contained in voice data.

[1572] "Means for analyzing text using natural language processing" refers to a device or software function that uses natural language processing techniques to analyze text data converted from voice data.

[1573] The "means for recognizing a user's emotion from voice data using an emotion engine" is a function of the emotion engine used to recognize a user's emotion based on voice data.

[1574] The "means for generating and playing healing music optimized for a stress level" refers to a device or software function for creating and playing optimal healing music according to a user's stress level.

[1575] This invention is a system designed to effectively manage the mental health of employees. The system mainly consists of three entities: a server, a terminal, and a user, each of which plays a specific role.

[1576] Question generation and retrieval

[1577] The server generates questions related to employees' mental health. Specifically, it uses a generative AI model (e.g., OpenAI's GPT-4) to create a personalized list of questions based on the user's past data and common mental health questions. For example, questions include, "What kind of stress have you been feeling at work recently?" and "How do you spend your time outside of work?"

[1578] The device receives the questions sent from the server and presents them to the user in voice or text. For voice presentation, speech synthesis software (e.g., Google Cloud Text-to-Speech) is used.

[1579] The user answers the questions by voice, for example, "I've had a lot of meetings this week and I'm feeling a bit stressed." The device records this voice and sends it to the server. The voice data is first stored locally and then uploaded to the server using a secure communication protocol (e.g., HTTPS).

[1580] Voice analysis and emotion recognition

[1581] The server converts the received voice data into text data using a voice analysis engine (e.g., Google Cloud Speech-to-Text), which also extracts voice characteristics such as intonation, speed, and volume.

[1582] Next, the server uses a natural language processing (NLP) engine (e.g., OpenAI's GPT-4) to analyze the text content and understand what the user is saying. Using the analysis results and voice feature data, an emotion recognition engine (e.g., Microsoft's Azure Emotion API) recognizes the user's emotional state. Specifically, emotions such as "tension," "anxiety," and "relief" are detected from the tone of voice and speaking style.

[1583] Stress assessment

[1584] The server evaluates the user's stress level based on the results of voice analysis and emotion recognition. If the emotion recognition engine recognizes "tension" or "anxiety," the stress level is deemed high; conversely, if it recognizes "relief," it is deemed low. This evaluation result is recorded within the system and used for subsequent analysis and the development of countermeasures.

[1585] Advice generation and presentation

[1586] The server generates appropriate advice based on the stress assessment results. It uses a generative AI model (e.g., OpenAI's GPT-4) to create personalized advice based on the assessment results and past user data. For example, it provides specific advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension."

[1587] The device displays the advice sent from the server to the user and simultaneously plays healing music. Music playback is performed using a media player or streaming service (e.g., Spotify API).

[1588] Playing healing music

[1589] The server uses a generative AI model to generate soothing music optimized for the user's stress level and transmits it to the device, which then plays the soothing music in sync with the user's advice, providing a relaxing environment for the user.

[1590] Prompt Sentence Examples

[1591] In a one day scenario, the following prompts could be used:

[1592] "Tell me about your work environment recently. If you have had any stressful experiences, please tell me about them in detail."

[1593] This allows users to effectively manage stress and receive appropriate advice in a comfortable environment. In this way, the system is able to perform highly accurate emotion recognition and stress assessment according to the individual state of the user, and provide specific and personalized advice.

[1594] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1595] Step 1: Question Generation

[1596] The server uses a generative AI model to receive the user's past mental health data and general question data as input, and generates a personalized list of questions based on this. Specifically, questions such as "What kind of stress have you been feeling at work recently?" and "How do you spend your time outside of work?" are generated. The generated list of questions is then sent to the device in the next step.

[1597] Step 2: Posing the Question

[1598] The device receives a list of questions from the server as input and uses speech synthesis software to present the questions to the user either audibly or via text. For example, Google Cloud Text-to-Speech is used to present the questions to the user audibly, and the user can then respond verbally.

[1599] Step 3: Get an audio response

[1600] The user answers the questions by voice, for example, "I had a lot of meetings this week and I felt a bit stressed."

[1601] The device takes this voice response as input, stores it locally, and then transmits the voice data to the server using a secure communication protocol (e.g., HTTPS).

[1602] Step 4: Audio analysis

[1603] The server receives the voice data sent from the device as input and converts it into text using a voice analysis engine (for example, Google Cloud Speech-to-Text). During this process, it extracts voice characteristics such as intonation, speed, and volume, and passes them on to the next process along with the text data.

[1604] Step 5: Natural Language Processing (NLP)

[1605] The server receives the text data and speech feature data obtained through speech analysis as input, analyzes the text content using an NLP engine (for example, OpenAI's GPT-4), and outputs structured data that helps understand what the user said.

[1606] Step 6: Emotion Recognition

[1607] The server receives the analysis results of the NLP engine and the voice feature data as input, and uses an emotion recognition engine (such as Microsoft's Azure Emotion API) to recognize the user's emotions. For example, it detects emotions such as "tension," "anxiety," and "relief" from the tone of voice and speaking style, and generates emotion data based on the results.

[1608] Step 7: Stress Assessment

[1609] The server receives the emotion data output by the emotion recognition engine as input and evaluates the user's stress level. Specifically, if the emotion data is recognized as "tension" or "anxiety," it is judged to be high stress. The evaluation results are stored in an internal database.

[1610] Step 8: Advice Generation

[1611] The server receives the stress assessment results and past user data as input, and uses the generative AI model to generate specific advice and countermeasures. For example, advice such as "It's a good idea to do some light exercise to relax after work" or "Take a few deep breaths to relieve tension" is generated and sent to the device.

[1612] Step 9: Providing advice and playing healing music

[1613] The device receives advice sent from the server as input and displays it to the user in text. At the same time, it receives healing music optimized for the user's stress level from the server and plays it using a media player or streaming service (e.g., Spotify API). This allows the user to receive advice in a relaxed state.

[1614] (Application example 2)

[1615] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1616] Modern factory workers are under a lot of stress during their work, making mental health care in the workplace extremely important. However, traditional mental health management systems do not adequately consider the characteristics of the workplace or the emotional state of individual workers. Furthermore, to provide effective care, a method is needed to accurately assess workers' emotional states and provide individually tailored advice and measures. Therefore, there is a need to develop a new system that can reduce worker stress and improve their mental health.

[1617] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1618] In this invention, the server includes a generation means for generating questions related to mental health, an acquisition means for acquiring employees' voice responses to the questions, an analysis means for analyzing the acquired voice responses and evaluating their emotions and stress levels, an advice generation means for generating specific advice and proposed measures based on the evaluation results obtained by the analysis, a presentation means for presenting the advice and proposed measures and playing soothing music, and a factory robot that collects data from workers through a voice interface and provides personalized advice and soothing music for the mental health care of factory workers. This allows workers to continue working in a less stressful environment, and is expected to improve the overall working environment.

[1619] The "Employee Mental Health Management System" is a system that effectively manages employees' mental health, assesses their emotions and stress levels, and provides specific advice and healing music.

[1620] A "generation means" is a means that has the function of generating questions related to mental health.

[1621] The "acquisition means" is a means having a function of acquiring an employee's voice response to a question.

[1622] The "analysis means" is a means having a function of analyzing the acquired voice response and evaluating emotions and stress levels.

[1623] The "advice generation means" is a means having a function of generating specific advice and countermeasures based on the evaluation results obtained by the analysis.

[1624] The "presentation means" is a means that has the function of presenting advice and countermeasures as well as playing healing music.

[1625] The "Mental Health Care System for Factory Workers" is a system for managing the mental health of factory workers, assessing their emotional state and stress levels, and providing personalized advice and healing music.

[1626] A "voice interface" is an interface for collecting voice data from a user and conducting interaction.

[1627] A "factory robot" is a robot that supports workers in a factory and has mental health management functions.

[1628] This invention is a system for effectively managing the mental health of employees and factory workers, and is characterized by the combination of an emotion engine to realize mental care suited to the working environment. This system integrates the following functions: generating questions related to mental health, acquiring and analyzing voice data, recognizing emotions, evaluating stress, generating advice, and playing healing music.

[1629] System Configuration

[1630] The system consists of the following main components:

[1631] 1. Server: This is the central data processing center, and performs question generation, voice analysis, stress assessment, emotion recognition, advice generation, healing music generation, etc. The server can be, for example, a cloud-based server or an on-premise server.

[1632] 2. Terminal: A device operated by the user that presents questions, acquires audio data, displays analysis results and advice, and plays healing music. Terminals can be smartphones, tablets, desktop computers, etc.

[1633] 3. Users: Employees and workers who use the system.

[1634] 4. Emotion Engine: Recognizes emotions from the user's voice responses and helps assess stress levels. It can use Google Cloud Natural Language API and OpenAI's Emotion Recognition model.

[1635] System Operation

[1636] Question generation and retrieval

[1637] Server: Uses generative AI to create a personalized list of questions, taking into account the user's past data and common mental health-related questions.

[1638] Terminal: Presents the generated questions to the user by voice or text.

[1639] User: Answers questions verbally.

[1640] Terminal: Records the user's voice response and sends it to the server.

[1641] Voice analysis and emotion recognition

[1642] Server: The received voice data is analyzed using a voice analysis engine. Voice characteristics such as intonation, speed, and volume are extracted, and the text content is analyzed using a natural language processing (NLP) engine.

[1643] The server then uses an emotion engine to recognize the user's emotion from the voice data.

[1644] Stress assessment

[1645] Server: Evaluates the user's stress level based on emotional data recognized through voice analysis and the emotion engine. Specifically, if the emotion engine recognizes "tension" or "anxiety," the stress level is determined to be high.

[1646] Advice generation and presentation

[1647] Server: Based on the evaluation results, appropriate advice and countermeasures are generated. The generation AI provides personalized advice tailored to the user's situation.

[1648] Device: Displays the generated advice and countermeasures to the user. Set it to play healing music at the same time.

[1649] Playing healing music

[1650] Server: Uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[1651] Device: The device plays soothing music while displaying advice and suggested solutions, allowing the user to receive the information in a relaxed state.

[1652] Specific examples

[1653] A typical day scenario could involve the following exchange:

[1654] 1. Device: "Have you felt stressed at work recently?"

[1655] 2. User: "Yes, writing the report was very stressful."

[1656] 3. Server: Performs voice analysis and NLP analysis, and recognizes "tension" using the emotion engine.

[1657] 4. Server: Generates the advice, "Take regular breaks while writing your report. We also recommend doing some light stretching."

[1658] 5. Device: Displays advice and plays relaxing, soothing music.

[1659] This system will enable workers to continue working in a less stressful environment, which is expected to improve the overall working environment.

[1660] Examples of prompt statements

[1661] Examples of prompts the system might use when using a generative AI model include:

[1662] "If workers are feeling very stressed, generate advice to help them relieve it."

[1663] This prompt allows the system to provide appropriate advice for the worker.

[1664] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1665] Step 1:

[1666] The server uses a generative AI model to generate a personalized list of questions, taking into account the user's past data and common mental health-related questions.

[1667] Input: User history data, general question data

[1668] Data processing: Generate a list of questions using a generative AI model

[1669] Output: A personalized list of questions

[1670] Step 2:

[1671] The terminal presents the list of questions sent from the server to the user by voice or text.

[1672] Input: Personalized Question List

[1673] Output: Questions presented in audio or text format

[1674] What happens: When the device uses a speech synthesis engine to play a question aloud, it presents the user with a question such as "Have you felt stressed at work recently?"

[1675] Step 3:

[1676] The user answers the questions by voice, and the terminal records the voice response and transmits it to the server.

[1677] Input: User's spoken response

[1678] Output: Recorded audio data

[1679] Specific operation: When the user answers, for example, "Yes, writing the report was very stressful," the device records the voice and sends it to the server.

[1680] Step 4:

[1681] The server analyzes the received voice data using a voice analysis engine to extract voice characteristics such as intonation, speed, and volume. It also analyzes the text content using an NLP engine and recognizes the user's emotions using an emotion engine.

[1682] Input: Recorded audio data

[1683] Data processing: Analyze voice data with a voice analysis engine and extract features. Analyze text with an NLP engine. Recognize emotions with an emotion engine.

[1684] Output: Emotional state, stress level data

[1685] Specific operation: The server analyzes the voice data, and if the user says "I felt very stressed," it recognizes emotions such as "tension" and "anxiety."

[1686] Step 5:

[1687] The server evaluates the user's stress level based on emotion recognition data and generates appropriate advice and countermeasures.

[1688] Input: Emotional state, stress level data

[1689] Data processing: Based on the evaluation results, advice and countermeasures are created using an advice generation AI model.

[1690] Output: Specific advice and countermeasures

[1691] Specific actions: The server generates specific advice such as "Take appropriate breaks while writing the report. We also recommend doing some light stretching."

[1692] Step 6:

[1693] The terminal displays advice and countermeasures from the server to the user and simultaneously plays healing music.

[1694] Input: Specific advice, countermeasures, healing music data

[1695] Output: Displaying advice to the user, playing healing music

[1696] Specific operation: The device displays advice such as "Take regular breaks while writing your report" on the screen and plays relaxing music.

[1697] Step 7:

[1698] The server uses generative AI to create soothing music tailored to the user's stress level and sends it to the device.

[1699] Input: User's stress level data

[1700] Data processing: Generating healing music using generative AI models

[1701] Output: Healing music data

[1702] Specific operation: The server generates relaxing music according to the user's stress level and sends it to the device.

[1703] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1704] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1705] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1706] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1707] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1708] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1709] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1710] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1711] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1712] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1713] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1714] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1715] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1716] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1717] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1718] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1719] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1720] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1721] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1722] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1723] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1724] The following is further disclosed regarding the above embodiment.

[1725] (Claim 1)

[1726] A system for managing employee mental health,

[1727] a generating means for generating questions related to mental health;

[1728] an acquisition means for acquiring an employee's voice response to a question;

[1729] an analysis means for analyzing the acquired voice response and evaluating emotions and stress levels;

[1730] an advice generation means for generating specific advice and countermeasures based on the evaluation results obtained by the analysis;

[1731] A presentation method that provides advice and countermeasures, as well as playing healing music,

[1732] A system including:

[1733] (Claim 2)

[1734] 10. The system of claim 1, wherein the system performs speech intonation, rate, volume, and text analysis to assess emotion and stress level.

[1735] (Claim 3)

[1736] 10. The system of claim 1, wherein the system takes into account past employee feedback when generating advice and proposed actions.

[1737] "Example 1"

[1738] (Claim 1)

[1739] A system for managing employee mental health,

[1740] a generating means for generating questions related to mental health;

[1741] an acquisition means for acquiring an employee's voice response to a question;

[1742] an analysis means for analyzing the acquired voice response and evaluating emotions and stress levels;

[1743] an advice generation means for generating specific advice and countermeasures based on the evaluation results obtained by the analysis;

[1744] A presentation means for generating and playing healing music while providing advice and countermeasures;

[1745] A system including:

[1746] (Claim 2)

[1747] 10. The system of claim 1, wherein the system performs speech intonation, rate, volume, and text analysis to assess emotion and stress level.

[1748] (Claim 3)

[1749] 10. The system of claim 1, wherein the system takes into account past employee feedback when generating advice and proposed actions.

[1750] "Application Example 1"

[1751] (Claim 1)

[1752] A system for managing employee mental health,

[1753] a generating means for generating questions related to mental health;

[1754] an acquisition means for acquiring an employee's voice response to a question;

[1755] an analysis means for analyzing the acquired voice response and evaluating emotions and stress levels;

[1756] an advice generation means for generating specific advice and countermeasures based on the evaluation results obtained by the analysis;

[1757] A presentation method that provides advice and countermeasures, as well as playing healing music,

[1758] a refreshment means for presenting refreshment advice for improving the user experience in the virtual space;

[1759] A system including:

[1760] (Claim 2)

[1761] 10. The system of claim 1, wherein the system performs speech intonation, rate, volume, and text analysis to assess emotion and stress level.

[1762] (Claim 3)

[1763] 10. The system of claim 1, wherein the system takes into account past employee feedback when generating advice and proposed actions.

[1764] "Example 2: Combining Emotion Engines"

[1765] (Claim 1)

[1766] A system for managing employee mental health,

[1767] a generating means for generating questions related to mental health;

[1768] an acquisition means for acquiring an employee's voice response to a question;

[1769] an analysis means for analyzing the acquired voice response and evaluating emotions and stress levels;

[1770] an advice generation means for generating specific advice and countermeasures based on the evaluation results obtained by the analysis;

[1771] A presentation method that provides advice and countermeasures, as well as playing healing music,

[1772] A means of analyzing voice intonation, speed, volume, etc. and analyzing text using natural language processing;

[1773] means for recognizing a user's emotion from the speech data using an emotion engine;

[1774] A means for generating and playing healing music optimized for stress levels;

[1775] A system including:

[1776] (Claim 2)

[1777] 10. The system of claim 1, wherein the system performs speech intonation, rate, volume, and text analysis to assess emotion and stress level.

[1778] (Claim 3)

[1779] 10. The system of claim 1, wherein the system takes into account past employee feedback when generating advice and proposed actions.

[1780] "Application example 2 when combining emotion engines"

[1781] (Claim 1)

[1782] A system for managing employee mental health,

[1783] a generating means for generating questions related to mental health;

[1784] an acquisition means for acquiring an employee's voice response to a question;

[1785] an analysis means for analyzing the acquired voice response and evaluating emotions and stress levels;

[1786] an advice generation means for generating specific advice and countermeasures based on the evaluation results obtained by the analysis;

[1787] A presentation method that provides advice and countermeasures, as well as playing healing music,

[1788] A system for mental health care for factory workers that includes a factory robot that collects data from workers through a voice interface and provides personalized advice and soothing music.

[1789] (Claim 2)

[1790] 10. The system of claim 1, wherein the system performs speech intonation, rate, volume, and text analysis to assess emotion and stress level.

[1791] (Claim 3)

[1792] 10. The system of claim 1, wherein the system takes into account past employee feedback when generating advice and proposed actions. [Explanation of symbols]

[1793] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A system for managing employee mental health, a generating means for generating questions related to mental health; an acquisition means for acquiring an employee's voice response to a question; an analysis means for analyzing the acquired voice response and evaluating emotions and stress levels; an advice generation means for generating specific advice and countermeasures based on the evaluation results obtained by the analysis; A presentation method that provides advice and countermeasures, as well as playing healing music, A system including:

2. 10. The system of claim 1, wherein the system performs speech intonation, rate, volume, and text analysis to assess emotion and stress level.

3. 10. The system of claim 1, wherein the system takes into account past employee feedback when generating advice and countermeasures.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A