System

A voice-activated system with emotion analysis and synthesis provides 24/7 consultation services, addressing the need for rapid support in mental health crises by converting user voice into text, analyzing emotions, and generating timely responses.

JP2026019024APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024120432
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Suicides in Japan are often triggered by mental distress and stress, with consultation services being slow at night, and the COVID-19 pandemic exacerbating the risk, necessitating a 24/7 system for rapid and effective support.

Method used

A system utilizing voice recognition, emotion analysis, and speech synthesis to provide 24/7 consultation services by converting user voice input into text, analyzing emotions and intentions, generating appropriate responses, and transmitting them via a user terminal, with the option to transfer to experts when needed.

Benefits of technology

Enables users to receive immediate and appropriate support 24/7, reducing the risk of suicide by providing rapid and effective consultation services, including emergency responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019024000001_ABST
    Figure 2026019024000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for receiving voice data; voice recognition means for converting the voice data into text data; means for analyzing a user's emotion and intention from the text data; generation means for generating an answer based on the analyzed emotion and intention; means for converting the generated answer into voice data; and means for transmitting the voice data to a user terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Many suicides occur in Japan every year, many of which are caused by mental distress and stress. Consultation services are particularly slow at night, making it difficult to receive appropriate support in emergencies. Furthermore, economic hardship and social anxiety caused by the COVID-19 pandemic have increased the risk of suicide, creating a need for rapid and effective countermeasures. The present invention aims to provide a system that allows users at risk of suicide to receive consultation 24 hours a day, 365 days a year in such emergencies. [Means for solving the problem]

[0005] The present invention solves the above problems by the following means. First, the system includes a means for receiving voice data. Next, the system includes a voice recognition means for converting the voice data into text data. Furthermore, the system includes a generation means for generating an answer based on the analyzed emotions and intentions using a means for analyzing the user's emotions and intentions from the text data. The system also includes a means for converting the generated answer into voice data and finally a means for transmitting the voice data to a user terminal. Furthermore, the present invention can include a means for receiving a user's voice input and forwarding it to an expert based on the analyzed emotions and intentions. Furthermore, the system includes a means for realizing 24 / 7 operation to enhance consultation services even at night. This provides an environment where users at risk of suicide can seek consultation at any time, and strengthens emergency response.

[0006] "Audio data" is data that represents an audio signal in digital form.

[0007] The term "means" refers to a method or apparatus used to achieve a particular function or process.

[0008] "Speech recognition means" refers to technology or devices for converting voice data into text data.

[0009] "Text data" is character string data converted from voice data.

[0010] "Means for analyzing emotions and intentions" refers to technology or devices for determining a user's emotions and intentions based on text data.

[0011] A "generator" is a technology or device that creates answers based on the analyzed emotions and intentions.

[0012] A "speech synthesis engine" is a technology or device for converting text data into voice data.

[0013] A "user terminal" is a device used by a user, such as a computer or smartphone.

[0014] "24 hours a day, 365 days a year" refers to a state in which a system is operational at all times of the day and all year round.

[0015] The "means for transferring to an expert" refers to a technology or device for connecting to an expert when further detailed consultation is required based on the analysis results. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention is a system that provides a 24 / 7 consultation service by combining voice recognition technology and an emotion analysis engine. It is specifically designed to allow users at risk of suicide to talk about their worries by voice.

[0038] Server Roles

[0039] Voice input acceptance

[0040] The server receives the voice data sent by the user in real time. When the user calls the helpline, the voice is transmitted directly to the server.

[0041] Analysis of audio data

[0042] When the server receives the voice data, it first converts it into text data using a speech recognition library, using various speech recognition engines (e.g., speech recognition APIs) to achieve high accuracy.

[0043] Emotion and Intention Analysis

[0044] The text data is then fed into a sentiment and intent analysis engine, which uses natural language processing techniques to extract the user's emotions (e.g., anxiety, sadness, urgency, etc.) and intent (e.g., help-seeking, advice-seeking, etc.) from the text data.

[0045] Answer generation

[0046] The server generates answers based on the analyzed emotions and intent. To generate answers, it uses an AI model that has previously learned the knowledge of doctors and counselors. This model creates appropriate responses based on the information and emotions the user is seeking.

[0047] Conversion to audio

[0048] The generated text response is converted into audio data using a speech synthesis engine, which uses speech synthesis technology (e.g., speech synthesis API) to provide a human-like voice response to the user.

[0049] Sending audio data

[0050] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[0051] Device Role

[0052] Voice input

[0053] When a user calls to talk about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0054] Audio Output

[0055] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0056] User Roles

[0057] System Usage

[0058] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0059] Additional consultation

[0060] If the user wants to continue the consultation, they can repeat the same process as many times as they like. If professional help is required, the server will automatically transfer the user to an expert.

[0061] Specific examples

[0062] For example, if a user is suffering from work-related stress, they might say, "Work hasn't been going well lately, and I don't know what to do." This speech is sent to the server, where it is recognized. From the converted text, an emotion analysis engine extracts the emotions of "stress" and "hopelessness." The AI ​​model then generates "advice for overcoming work-related problems," which is then converted into speech by a speech synthesis engine. Specific advice such as "The key to overcoming work-related problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[0063] Role Alignment

[0064] These components work together to provide an environment where users at high risk of suicide can seek advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[0065] In this way, the present invention realizes a consultation system that can respond quickly and effectively at any time, including at night, thereby contributing to reducing the risk of suicide.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] The user dials the phone number of the IVR system using a smartphone or landline, which initiates voice input by the user.

[0069] Step 2:

[0070] The device records the user's voice in real time and transmits the voice data to the server using a secure protocol (e.g., HTTPS).

[0071] Step 3:

[0072] The server converts the received voice data into text data using a speech recognition library (e.g., speech recognition API). This conversion process turns the voice signal into a string of characters.

[0073] Step 4:

[0074] The converted text data is sent to an emotion and intent analysis engine, which extracts the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data.

[0075] Step 5:

[0076] Based on the extracted emotions and intent, the server uses an AI model to generate appropriate answers. This model has previously learned the knowledge of doctors and counselors to create specific and appropriate responses.

[0077] Step 6:

[0078] The generated response text data is converted into voice data using a speech synthesis engine (e.g., speech synthesis API), resulting in a human-like voice response.

[0079] Step 7:

[0080] The server transmits the generated voice data to the user's terminal, which receives the voice data and plays it back to the user.

[0081] Step 8:

[0082] The user receives the voice response provided through the terminal and, if necessary, can request further consultation, in which case the same process (steps 1 to 7) can be repeated.

[0083] Step 9:

[0084] If the user's consultation requires professional help, the server transfers the user to a specialist based on the results of emotion and intent analysis. To do this, it searches for the specialist's contact information and relays the connection using a conference call system.

[0085] The processing steps of this system provide an environment where users can receive consultation 24 hours a day, 365 days a year.

[0086] Example 1

[0087] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0088] In modern society, many people suffer from stress in their daily lives and work, feelings of loneliness, and a variety of other worries. One particularly serious problem is consultations that involve the risk of suicide. Current counseling services and consultation centers often have limited response times and are unable to provide prompt and appropriate responses in emergencies. Furthermore, professional assistance is often unavailable at night or on holidays, which can leave users without the support they need and put them in danger. Given this current situation, there is a need for a system that is available 24 hours a day, 365 days a year, can accurately analyze users' emotions and intentions, and provide appropriate responses in real time.

[0089] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0090] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, means for analyzing the user's emotions and intentions from the converted text data, generation means using a generative AI model to generate answers based on the analyzed emotions and intentions, speech synthesis means for converting the generated answers into voice data, and means for transmitting the voice data to a user terminal. This enables real-time analysis of voice input from the user and immediate provision of appropriate answers. Furthermore, by transferring the call to a specialist as needed, more specialized support can be received, realizing consultation services 24 hours a day, 365 days a year, including nights and holidays. This provides a safe and reliable system that can handle serious consultations, including those regarding suicide risk.

[0091] A "server" is a centralized control device for receiving and processing data sent by users.

[0092] "Voice data" refers to data in which voice information uttered by a user is recorded as a digital signal.

[0093] "Speech recognition means" refers to a technology or device that analyzes voice data and converts it into text data.

[0094] "Text data" is character information converted from voice data by a voice recognition means.

[0095] The "analysis means" is a technique or device for identifying a user's emotions and intentions from text data.

[0096] A "generative AI model" is an artificial intelligence model that has learned a huge amount of data in advance and generates appropriate answers based on the user's requests and emotions.

[0097] "Generation means" refers to a technology or device that uses a generative AI model to create an answer based on information obtained from the analysis means.

[0098] "Speech synthesis means" refers to a technique or device that generates voice data based on generated text data.

[0099] A "user terminal" is a device used by a user to communicate with a server. This includes smartphones and landlines.

[0100] "Transfer means" refers to a technology or device that transmits data to an expert based on the user's voice input and the analysis results.

[0101] This invention is a system that provides a 24 / 7 consultation service by combining voice recognition technology and an emotion analysis engine. It is specifically designed to allow users at risk of suicide to talk about their worries by voice.

[0102] Server Roles

[0103] Voice input acceptance

[0104] The server receives the voice data sent by the user in real time. When the user calls the consultation service, the voice is transmitted directly to the server. The voice data is transmitted to the server using a secure communication protocol (e.g., HTTPS).

[0105] Analysis of audio data

[0106] When the server receives the voice data, it first converts it into text using a speech recognition library, using a speech recognition engine such as the Google Cloud Speech-to-Text API to achieve high accuracy.

[0107] Emotion and Intention Analysis

[0108] The text data is then sent to a sentiment and intent analysis engine, which uses natural language processing techniques to extract the user's emotions (e.g., anxiety, sadness, urgency, etc.) and intent (e.g., help-seeking, advice-seeking) from the text data. This process uses tools such as the Microsoft Azure Text Analytics API.

[0109] Answer generation

[0110] Based on the analyzed emotions and intent, the server generates an answer using a generative AI model (e.g., OpenAI GPT-4) that has previously learned the knowledge of doctors and counselors. This model creates an appropriate response based on the information and emotions the user is seeking.

[0111] Conversion to audio

[0112] The generated text response is converted into audio data using a speech synthesis engine, using voice synthesis technologies such as Amazon Polly to provide the user with a human-like voice.

[0113] Sending audio data

[0114] Finally, the generated voice data is sent to the user's terminal, where the user can receive a voice response to their inquiry. This voice data is also sent using a secure communication protocol.

[0115] Device Role

[0116] Voice input

[0117] When a user calls to talk about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0118] Audio Output

[0119] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0120] User Roles

[0121] System Usage

[0122] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0123] Additional consultation

[0124] If the user wants to continue the consultation, they can repeat the same process as many times as they like. If professional help is required, the server will automatically transfer the user to an expert.

[0125] Specific examples

[0126] For example, if a user is suffering from work-related stress, they might say, "Work hasn't been going well lately, and I don't know what to do." This speech is sent to the server, where it is recognized. The converted text is then processed by an emotion analysis engine, which extracts the emotions of "stress" and "hopelessness." The generative AI model then generates "advice for overcoming work-related problems," which is then converted into speech by a speech synthesis engine. Specific advice such as "The key to overcoming work-related problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is then provided to the user via voice.

[0127] Examples of prompt statements

[0128] Below is an example of a prompt that can be fed into a generative AI model:

[0129] "A user says, 'Work hasn't been going well lately, and I don't know what to do.' Please provide appropriate advice to address the stress and despair the user is feeling."

[0130] This system provides an environment where users at high risk of suicide can receive advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[0131] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0132] Step 1:

[0133] User voice input

[0134] When a user has a problem, they call the helpline using their smartphone or landline. They then speak to the helpline, describing their specific problem, such as, "Recently, things haven't been going well at work, and I don't know what to do." This voice data becomes the input.

[0135] Step 2:

[0136] Sending audio data by the device

[0137] The device (smartphone or landline phone) records the user's voice in real time and sends it to the server. This voice data is sent to the server using a secure communication protocol (e.g., HTTPS). The input is the user's voice data, and the output is the voice data sent to the server.

[0138] Step 3:

[0139] Receiving audio data by the server

[0140] The server receives the voice data sent from the device in real time. The received voice data is temporarily stored in the server's memory. The input is the voice data sent from the device, and the output is the voice data stored in the server.

[0141] Step 4:

[0142] Server-based speech recognition and text conversion

[0143] The server converts the received voice data into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text API). The voice data is analyzed and the corresponding text data is generated. The input is the voice data stored on the server, and the output is the converted text data.

[0144] Step 5:

[0145] Server-based emotion and intent analysis

[0146] The server sends the converted text data to an emotion and intent analysis engine (e.g., Microsoft Azure Text Analytics API), which extracts the user's emotion (e.g., stress, despair) and intent (e.g., asking for help) from the text data. The input is the text data, and the output is the emotion and intent analysis results.

[0147] Step 6:

[0148] Server-generated answers

[0149] Based on the results of the emotion and intent analysis, the server sends a prompt to a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate answer. An example of a prompt sentence is, "The user said, 'Recently, things haven't been going well at work, and I don't know what to do anymore.' Please provide appropriate advice to help the user with the stress and despair they are feeling." The input is the results of the emotion and intent analysis, and the output is the generated answer in text format.

[0150] Step 7:

[0151] Server-based speech synthesis

[0152] The generated text response is converted back into audio data by a speech synthesis engine (e.g., Amazon Polly). This process uses techniques to express natural pronunciation and emotion in the voice. The input is the generated text answer, and the output is audio data.

[0153] Step 8:

[0154] Server sends audio data

[0155] The server then transmits the generated voice data back to the device, again using a secure communication protocol. The input is the generated voice data, and the output is the voice data sent to the device.

[0156] Step 9:

[0157] Receiving and playing audio data by a terminal

[0158] The terminal receives the voice data sent from the server and plays it back to the user. The user can receive a specific answer to their question by listening to the played back voice. The input is the voice data sent from the server, and the output is the voice played back to the user.

[0159] (Application example 1)

[0160] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0161] In modern society, the number of people suffering from stress and mental health problems is increasing, making immediate and appropriate support essential, especially for those at high risk of suicide. However, existing systems have difficulty analyzing emotions and generating responses in real time, and are therefore inadequate, especially at night. Furthermore, there is a need for a method that allows users to effectively utilize smart devices and receive a more realistic experience.

[0162] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0163] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, means for analyzing a user's emotions and intentions from the text data, means for generating a response based on the analyzed emotions and intentions, means for converting the generated response into voice data, means for transmitting the voice data to a user terminal, and means for receiving a user's voice input in real time and providing a response to the user through smart glasses or a head-mounted display. This allows for 24-hour, 365-day service, and enables users to receive immediate and effective support by providing emotion analysis and appropriate responses in real time through the smart device.

[0164] "Voice data" refers to data that digitally represents a user's speech.

[0165] "Text data" refers to data expressed as character information converted by a voice recognition means.

[0166] "Speech recognition means" refers to a device or software for converting voice data into text data.

[0167] "Means for analyzing emotions and intentions" refers to a device or software for analyzing a user's emotions and intentions from text data.

[0168] The "means for generating an answer" refers to a device or software that generates an answer to be provided to a user based on the analyzed emotions and intentions.

[0169] "Means for converting into voice data" refers to a device or software for converting the generated response into voice data.

[0170] A "user terminal" refers to a device used by a user, including a smartphone, smart glasses, a head-mounted display, etc.

[0171] The "means for transmitting voice data to a user terminal" refers to a device or software for transmitting the generated voice data to a user terminal.

[0172] "Smart glasses" are glasses-type devices equipped with functions such as augmented reality (AR).

[0173] A "head-mounted display" is a display device that is worn on the user's head and provides images and sounds.

[0174] "Real-time" means that data processing occurs almost immediately, allowing users to receive results without waiting.

[0175] This invention combines voice recognition technology with an emotion analysis engine to provide a 24 / 7 consultation service system for users at risk of suicide. In particular, it utilizes smart glasses and head-mounted displays to provide appropriate support to users in real time. The main components of this system include a voice recognition unit, an emotion and intention analysis unit, a response generation unit, a voice synthesis unit, and a voice data transmission unit.

[0176] Handling voice input and reception

[0177] The server receives the user's voice input in real time. It uses the Python speech_recognition library to convert the user's speech into text data. This voice data is acquired through the microphone in the smart glasses or head-mounted display.

[0178] Emotion and Intention Analysis

[0179] When the server receives the text data, it uses Hugging Face's sentiment-analysis model to analyze the user's emotions (e.g., anxiety, sadness, urgency). It also determines the user's intent based on the analyzed emotions and the text content. The results of this analysis are used to accurately identify situations in which advice or support is needed.

[0180] Generate answers

[0181] The server uses Hugging Face's text-generation (GPT-3) model to generate appropriate answers based on the analyzed emotions and intent, providing high-quality responses in real time that mimic expert advice and support.

[0182] Speech synthesis and output

[0183] The generated text response is converted into audio data using the gTTS (Google Text-to-Speech) library, and this audio data is provided to the user through the speakers of smart glasses or head-mounted displays, allowing the user to listen to the audio response directly using the device in front of them.

[0184] Specific examples

[0185] For example, if a user says, "I've been having trouble sleeping lately, please help me," this speech is sent to the server and converted into text using speech recognition. Then, emotion and intent analysis extracts the emotions "anxiety" and "seeking help." The AI ​​model then generates an answer using the following prompt:

[0186] Example prompt sentence:

[0187] User: I've been having trouble sleeping lately, please help.

[0188] System: Please provide advice for insomnia.

[0189] Ultimately, the voice will provide advice such as, "The best way to combat your insomnia is to relax before bed. Try meditating or doing some gentle stretching." This allows users to receive the support they need in real time.

[0190] This invention allows users to receive support 24 hours a day, 365 days a year with peace of mind, and provides an environment where immediate response is possible, especially for people at high risk of suicide.

[0191] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0192] Step 1:

[0193] The user inputs voice through smart glasses or a head-mounted display. The device records this voice in real time and transmits it to the server as voice data. The transmitted voice data is the input.

[0194] Step 2:

[0195] The server receives the transmitted voice data and converts it to text using the Python speech_recognition library. This converted text data is the output. This process is performed by the speech recognition means.

[0196] Step 3:

[0197] The server inputs the text data obtained from the speech recognition means into Hugging Face's sentiment-analysis model, which analyzes the user's emotions from the text data. This analysis generates emotional information such as anxiety, sadness, and urgency as output. The emotion and intent analysis means performs this process.

[0198] Step 4:

[0199] The server generates an appropriate answer for the user based on the results of the emotion analysis (emotion information) by inputting it into Hugging Face's text-generation (GPT-3) model. This generated text response is the output. This process is carried out by the generation means for generating the answer.

[0200] Step 5:

[0201] The server inputs the generated text response into the gTTS (Google Text-to-Speech) library and converts it to speech. This converted speech response is the output. The speech conversion method performs this process.

[0202] Step 6:

[0203] The server converts the response into voice data and sends it to the terminal. The terminal plays back the received voice data and provides it to the user. The voice response is played back through the user's smart glasses or head-mounted display. The provision of this voice response is the final output. This process is performed by a means for transmitting voice data to the user terminal.

[0204] Step 7:

[0205] The user listens to the voice response provided by the device and can re-voice the patient if necessary, and if necessary, repeat the consultation or request a transfer to a specialist. This interactive process continues.

[0206] Through this specific processing step, the system analyzes the user's emotions and intentions in real time and provides appropriate support.

[0207] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0208] This invention is a system that combines voice recognition technology and an emotion analysis engine to provide a 24 / 7 consultation service. It is designed to enable users at risk of suicide to discuss their concerns via voice. The invention adds an emotion engine with the ability to recognize the user's emotions and adaptively adjust responses.

[0209] Server Roles

[0210] Voice input acceptance

[0211] The server receives the voice data sent by the user in real time. When the user calls the helpline, the voice is transmitted directly to the server.

[0212] Analysis of audio data

[0213] When the server receives the voice data, it first converts it into text using a voice recognition library, and uses a variety of voice recognition engines to achieve high accuracy.

[0214] Emotion and Intention Analysis

[0215] The converted text data is sent to a sentiment analysis engine, which the server uses to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data. Furthermore, the sentiment analysis engine has the ability to adaptively adjust responses.

[0216] Answer generation

[0217] The server generates answers based on the analyzed emotions and intent. To generate answers, it uses an AI model that has previously learned the knowledge of doctors and counselors. This model creates appropriate responses based on the information and emotions the user is seeking.

[0218] Adaptive response adjustment

[0219] The emotion engine adaptively adjusts the generated response based on the user's perceived emotion, a process that provides a response that best suits the user's emotional state.

[0220] Conversion to audio

[0221] The generated response text is converted into voice data using a speech synthesis engine, which uses speech synthesis technology to provide answers to users in a human-like voice.

[0222] Sending audio data

[0223] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[0224] Device Role

[0225] Voice input

[0226] When a user calls and talks about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0227] Audio Output

[0228] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0229] User Roles

[0230] System Usage

[0231] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0232] Additional consultation

[0233] If the user wants to continue the consultation, they can repeat the same process (steps 1 to 7). Also, if professional help is required, the server can automatically transfer the user to a specialist.

[0234] Specific examples

[0235] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent to a server for speech recognition. The converted text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. The AI ​​model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "The key to overcoming work problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[0236] Role Alignment

[0237] These components work together to provide an environment where users at high risk of suicide can seek advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[0238] In this way, the present invention realizes a consultation system that can respond quickly and effectively at any time, including at night, thereby contributing to reducing the risk of suicide.

[0239] The processing flow will be explained below.

[0240] Step 1:

[0241] A user dials the phone number of the IVR system using a smartphone or landline, which initiates voice input.

[0242] Step 2:

[0243] The device records the user's voice in real time and transmits the voice data to the server. The data is transferred using a secure communication protocol (e.g., HTTPS).

[0244] Step 3:

[0245] The server converts the received voice data into text data using a voice recognition library (e.g., a voice recognition API). In this step, the voice signal is represented as a string.

[0246] Step 4:

[0247] The converted text data is sent to a sentiment analysis engine, which the server uses to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for information, etc.) from the text data.

[0248] Step 5:

[0249] The sentiment analysis engine identifies the user's current emotional state based on the perceived emotions and intent, which is important for subsequent response generation.

[0250] Step 6:

[0251] The server uses an AI model to generate responses based on the extracted emotions and intent, which is based on knowledge previously learned from experts (e.g., doctors, counselors, etc.).

[0252] Step 7:

[0253] The generated response text data is adaptively adjusted through an emotion engine, which selects the most appropriate expression and tone for the user's current emotional state.

[0254] Step 8:

[0255] The adjusted response text data is converted into audio data using a speech synthesis engine (e.g., speech synthesis API). This process produces a natural, human-like voice.

[0256] Step 9:

[0257] The server transmits the generated voice data to the user's terminal, which receives the voice data and plays it back to the user.

[0258] Step 10:

[0259] The user receives the voice response provided through the terminal and, if necessary, can request further consultation, in which case the same process (steps 1 to 9) can be repeated.

[0260] Step 11:

[0261] If the user's consultation needs professional help, the server will transfer the user to a specialist based on the results of emotion and intent analysis. To do this, it will search for the specialist's contact information and establish a relay connection using a telephone conference system.

[0262] In this way, the system of the present invention recognizes the user's emotional state and adaptively adjusts responses based on that state, thereby providing effective mental health counseling. The processes in steps 1 through 11 work together to provide an environment where users at high risk of suicide can seek counseling safely 24 hours a day, 365 days a year.

[0263] Example 2

[0264] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0265] In modern society, the number of people suffering from mental health problems and stress is increasing, and there is an urgent need to provide prompt and appropriate support, especially for those at high risk of suicide. However, conventional consultation response systems have difficulty providing support 24 hours a day, 365 days a year, and it is difficult to accurately analyze emotions and intentions and generate appropriate responses. For this reason, there is a growing need for a system that can flexibly respond to changes in emotions and provide effective and prompt support to users at high risk of suicide.

[0266] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0267] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, emotion analysis means for analyzing a user's emotion and intention from the text data, generation means for generating an answer based on the analyzed emotion and intention, means for adaptively adjusting the generated answer based on the user's emotional state, means for converting the adjusted answer into voice data, and means for transmitting the voice data to a user terminal, thereby making it possible to provide an appropriate answer according to the user's emotional state in real time, 24 hours a day, 365 days a year.

[0268] The "means for receiving voice data" refers to a device or software that has the function of acquiring the voice uttered by the user as digital data, converting it into a format that can be processed within the system, and receiving it.

[0269] The "voice recognition means for converting voice data into text data" refers to a device or software that has the function of analyzing received voice data and converting it into corresponding text data.

[0270] "Emotion analysis means for analyzing user emotions and intentions from text data" refers to a device or software that has the function of analyzing a user's emotional state (e.g., anxiety, sadness, despair, etc.) and intentions (e.g., asking for help, asking for advice, etc.) based on converted text data.

[0271] The "generation means for generating a response based on the analyzed emotion and intention" is a device or software that has the function of generating an appropriate response based on data from the emotion analysis means.

[0272] "Means for adaptively adjusting the generated response based on the user's emotional state" refers to a device or software that has the function of fine-tuning the generated response in accordance with the user's emotional state and providing it in an optimal form.

[0273] The "means for converting the adjusted response into voice data" refers to a device or software that has the function of converting adaptively adjusted text data into voice data and providing the voice data to the user as a natural voice.

[0274] The "means for transmitting audio data to a user terminal" refers to a device or software that has the function of transmitting the generated and converted audio data to a user's device and playing it on that device.

[0275] This invention combines voice recognition technology with an emotion analysis engine to provide a consultation service system that is available 24 hours a day, 365 days a year. It is designed specifically to enable users at risk of suicide to discuss their concerns via voice. The system is equipped with an emotion engine that recognizes the user's emotions and adaptively adjusts responses.

[0276] Server Roles

[0277] Voice input acceptance

[0278] The server receives voice data sent from the user in real time. For example, when a user calls a helpline, the voice is transmitted directly to the server.

[0279] Analysis of audio data

[0280] When the server receives the voice data, it first converts it into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text, Amazon Transcribe), thereby obtaining text information about what the user said.

[0281] Emotion and Intention Analysis

[0282] The converted text data is sent to a sentiment analysis engine (e.g., IBM Watson Tone Analyzer, Microsoft Azure Text Analytics), which uses the engine to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data.

[0283] Answer generation

[0284] Based on the analyzed emotions and intent, the server generates answers using a generative AI model (e.g., OpenAI GPT-3), which has previously learned the knowledge of doctors and counselors to create appropriate responses to the information and emotions the user is seeking.

[0285] Adaptive response adjustment

[0286] The emotion engine adaptively adjusts the generated response based on the user's perceived emotion, a process that provides a response that best suits the user's emotional state.

[0287] Conversion to audio

[0288] The generated response text is converted into audio data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech, Amazon Polly), which uses speech synthesis technology to provide answers to users in a human-like voice.

[0289] Sending audio data

[0290] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[0291] Device Role

[0292] Voice input

[0293] When a user calls and talks about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0294] Audio Output

[0295] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0296] User Roles

[0297] System Usage

[0298] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0299] Additional consultation

[0300] If the user wants to continue with the consultation, they can repeat the same process, and if professional help is needed, the server can automatically transfer them to a specialist.

[0301] Specific examples

[0302] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent to a server for speech recognition. The converted text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. A generative AI model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "The key to overcoming work problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[0303] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0304] Step 1: Accept voice input

[0305] Input: User's voice

[0306] Process: The user calls the system and tells it about their concerns. The user's voice is recorded by the device (smartphone, landline, etc.).

[0307] Output: Recorded audio data (digital format)

[0308] Specific operation: When a user speaks about a problem such as "Work hasn't been going well lately," the device captures the voice in real time and sends it to the server as audio data.

[0309] Step 2: Analyzing the audio data

[0310] Input: Recorded audio data

[0311] Processing: The server sends the audio data to a speech recognition library, which converts it into text. Examples of speech recognition libraries used here include Google Cloud Speech-to-Text and Amazon Transcribe.

[0312] Output: Converted text data

[0313] Specific operation: The speech recognition library analyzes the voice data "Work is not going well" received by the server, and generates the corresponding text "Work is not going well."

[0314] Step 3: Emotion and Intent Analysis

[0315] Input: Converted text data

[0316] Processing: The server sends the text data to a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., help-seeking, advice-seeking, etc.).

[0317] Output: Sentiment and intent analysis results

[0318] Specific operation: The server inputs the text data "My work is not going well" into the emotion analysis engine, and the engine determines that the user is feeling "stress" or "despair."

[0319] Step 4: Answer Generation

[0320] Input: Sentiment and intent analysis results

[0321] Processing: The server uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate answer based on the analysis results. This model has been trained with the knowledge of doctors and counselors.

[0322] Output: Generated answer text

[0323] Specific behavior: If emotion analysis identifies a high level of "despair," the generative AI model will create specific advice such as "First, stay calm. It's important to understand your own feelings."

[0324] Step 5: Adaptively adjust the response

[0325] Input: Generated answer text

[0326] Processing: The server re-analyzes the generated answer text with the emotion engine and fine-tunes the response to best suit the user's emotional state.

[0327] Output: Adaptively adjusted answer text

[0328] What it does: The emotion engine translates the response created by the generative AI model, "It's important to understand your own feelings," into a "tone that conveys a calm and welcoming atmosphere."

[0329] Step 6: Convert to audio

[0330] Input: Adaptively adjusted answer text

[0331] Processing: The server converts the tailored response text into audio data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech, Amazon Polly).

[0332] Output: Generated audio data

[0333] What happens: The adjusted text response is sent to a speech synthesis engine, which converts it into speech data, generating a human-like voice saying, "It's important to understand your own feelings."

[0334] Step 7: Sending audio data

[0335] Input: Generated audio data

[0336] Processing: The server sends the generated audio data to the user's device.

[0337] Output: The audio response the user receives

[0338] Specific operation: The server sends the voice data to the user's device, which then plays it back. The user receives the advice, "It's important to understand your own feelings."

[0339] (Application example 2)

[0340] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0341] In recent years, the number of individuals at risk of suicide due to psychological stress and social anxiety has been increasing. It is often difficult to find an appropriate counselor or support, especially during late night or early morning hours. Furthermore, existing counseling systems lack the ability to properly analyze emotions and provide optimal responses to users. Furthermore, if the counseling process is not carried out quickly and effectively, there is a risk that the user's psychological burden will increase. Therefore, a system that operates 24 hours a day, 365 days a year and provides appropriate advice based on emotions is needed.

[0342] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice data, voice recognition means for converting the voice data into text data, emotion analysis means for analyzing the user's emotions and intentions from the text data, generation means for generating an answer based on the analyzed emotions and intentions, adjustment means for adaptively adjusting the generated answer, voice synthesis means for converting the adjusted answer into voice data, means for transmitting the voice data to the user terminal, and means for receiving the user's voice input and providing voice and visual information through a head-mounted display. This makes it possible to provide appropriate advice based on the user's emotions in the form of voice and visual information 24 hours a day, 365 days a year.

[0343] "Voice data" refers to voice information input by a user.

[0344] "Means" refers to a combination of devices and software for realizing a specific function.

[0345] "Speech recognition means" refers to a device or algorithm for converting input voice data into text data.

[0346] "Text data" refers to character string information converted by speech recognition means.

[0347] "Emotion analysis means" refers to devices and algorithms for analyzing user emotions and intentions from text data.

[0348] "Generator" refers to the system or algorithm that generates answers based on analyzed sentiment and intent.

[0349] "Adjustment means" refers to a system or algorithm that adaptively adjusts the generated answers based on the user's emotions.

[0350] "Speech synthesis means" refers to a device or algorithm for converting conditioned responses into speech data.

[0351] A "user terminal" is a device used by a user, and refers to an apparatus for inputting and outputting voice data.

[0352] A "head-mounted display" is a display device that provides information to the wearer's field of vision and has the function of simultaneously providing audio and visual information.

[0353] "Adaptive" refers to the ability to change adaptively according to the situation or conditions.

[0354] "Visual information" refers to visual data, icons, and messages presented to the user.

[0355] This invention is a system that uses a head-mounted display (HMD) to provide a voice-based consultation service available 24 hours a day, 365 days a year. This system combines voice recognition technology and an emotion analysis engine to provide appropriate advice to users and aim to reduce the risk of suicide.

[0356] Server Roles

[0357] The server processes the voice data and generates a response using the following means:

[0358] Voice input acceptance

[0359] The server receives the voice data sent by the user in real time. When the user speaks through the microphone built into the HMD, the voice data is sent to the server.

[0360] Analysis of audio data

[0361] When the server receives the voice data, it converts the voice data into text data using a voice recognition library (e.g., Python's speech_recognition).

[0362] Emotion and Intention Analysis

[0363] The converted text data is sent to an emotion analysis engine (e.g., Hugging Face's transformers library), which the server uses to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent from the text data.

[0364] Answer generation

[0365] Based on the analyzed sentiment and intent, the server utilizes a generative AI model (e.g., GPT-3) to generate appropriate answers, using prompts based on pre-trained datasets.

[0366] Adaptive response adjustment

[0367] The generated responses are adaptively adjusted in tone and content based on the user's emotions as analyzed by the emotion engine.

[0368] Conversion to audio

[0369] The tailored answers are then converted into audio data using a speech synthesis engine, allowing the answers to be provided to the user in a human-like voice.

[0370] Sending audio data

[0371] Finally, the generated audio data is sent to the user's HMD and played back to the user through the HMD's speakers.

[0372] Device Role

[0373] Voice input

[0374] When the user talks about their concerns through the HMD's microphone, the device records the audio and sends it to the server in real time.

[0375] Audio Output

[0376] The voice data sent from the server is played back to the user through the HMD speaker, allowing the user to receive appropriate support in real time.

[0377] User Roles

[0378] System Usage

[0379] When a user has a problem, they can wear the HMD and talk about it through a microphone. By talking about specific problems, the system can provide appropriate support.

[0380] Additional consultation

[0381] If the user wishes to continue the consultation, they can repeat the same process. There is also a function to transfer the call to a specialist if necessary, ensuring appropriate treatment in emergencies or serious cases.

[0382] Specific examples

[0383] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent from the HMD to the server. The server performs speech recognition, and the generated text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. The AI ​​model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "To overcome work problems, you need to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via the HMD via voice.

[0384] Prompt Sentence Examples

[0385] For users who are in an anxious emotional state, provide appropriate advice on the following: Things have been going badly at work lately, and I don't know what to do.

[0386] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0387] Step 1:

[0388] The user voice-inputs their concerns through the HMD microphone. The voice data is recorded by the HMD and sent to the server in real time. The consultation process begins when the user's voice input is transferred to the server as input data.

[0389] Step 2:

[0390] Once the server receives the audio data, it uses a speech recognition library (e.g., Python's speech_recognition) to convert the audio data into text. This step takes the audio data as input and generates highly accurate text data. Once the audio data is converted to text, it is ready for the next processing step.

[0391] Step 3:

[0392] The converted text data is sent to an emotion analysis engine (e.g., Hugging Face's transformers library). The server uses this engine to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent from the text data. As output, the user's emotions and intent are extracted, providing the basis for the next step.

[0393] Step 4:

[0394] Based on the analyzed sentiment and intent, the server utilizes a generative AI model (e.g., GPT-3) to generate an appropriate answer. A prompt sentence is used to instruct the model to generate an answer text appropriate to the user's situation. Sentiment and intent are used as input, and a generated answer is obtained as output.

[0395] Step 5:

[0396] The generated answer is adaptively adjusted in tone and content based on the user's emotions analyzed by the emotion engine. This step takes the generated answer text as input and optimizes the response with the appropriate tone and nuance depending on the user's emotional state. The final adjusted answer text is output.

[0397] Step 6:

[0398] The adjusted answer text is converted into voice data using a speech synthesis engine. The server uses this engine to generate natural-sounding speech and outputs it as voice data to be conveyed to the user. This conversion process uses the adjusted text data as input and generates human-like voice data.

[0399] Step 7:

[0400] The final generated voice data is sent to the user's HMD and played back to the user through the HMD's speaker. The server sends the voice data to the HMD, and the user receives a response in real time. The voice data is delivered to the HMD as output, thus completing the user's consultation process.

[0401] Specific examples of operation

[0402] For example, if a user says, "Things haven't been going well at work lately, and I don't know what to do anymore," then:

[0403] Step 1: The voice input "Work hasn't been going well lately, and I don't know what to do anymore" is sent from the HMD to the server.

[0404] Step 2: The speech recognition library converts this speech into text data, generating the text "Work hasn't been going well lately and I don't know what to do."

[0405] Step 3: The sentiment analysis engine analyzes the text and recognizes the emotions "anxiety" and "despair."

[0406] Step 4: The generative AI model generates an appropriate response for the user based on this emotion and intent. Example prompt: "For a user who is in an anxious emotional state, please provide appropriate advice on the following: Things haven't been going well at work recently, and I don't know what to do anymore."

[0407] Step 5: The generated answer, "The key to overcoming work-related worries is to understand your own feelings," is adaptively adjusted to optimize the tone to best fit the user's current emotions.

[0408] Step 6: The adjusted answers are converted into voice data by a speech synthesis engine.

[0409] Step 7: The final voice data is sent to the HMD, and the user receives the answer in real time through the speaker.

[0410] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0411] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0412] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0413] [Second embodiment]

[0414] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0415] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0416] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0417] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0418] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0420] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0421] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0422] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0423] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0424] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0425] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0426] This invention is a system that provides a 24 / 7 consultation service by combining voice recognition technology and an emotion analysis engine. It is specifically designed to allow users at risk of suicide to talk about their worries by voice.

[0427] Server Roles

[0428] Voice input acceptance

[0429] The server receives the voice data sent by the user in real time. When the user calls the helpline, the voice is transmitted directly to the server.

[0430] Analysis of audio data

[0431] When the server receives the voice data, it first converts it into text data using a speech recognition library, using various speech recognition engines (e.g., speech recognition APIs) to achieve high accuracy.

[0432] Emotion and Intention Analysis

[0433] The text data is then fed into a sentiment and intent analysis engine, which uses natural language processing techniques to extract the user's emotions (e.g., anxiety, sadness, urgency, etc.) and intent (e.g., help-seeking, advice-seeking, etc.) from the text data.

[0434] Answer generation

[0435] The server generates answers based on the analyzed emotions and intent. To generate answers, it uses an AI model that has previously learned the knowledge of doctors and counselors. This model creates appropriate responses based on the information and emotions the user is seeking.

[0436] Conversion to audio

[0437] The generated text response is converted into audio data using a speech synthesis engine, which uses speech synthesis technology (e.g., speech synthesis API) to provide a human-like voice response to the user.

[0438] Sending audio data

[0439] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[0440] Device Role

[0441] Voice input

[0442] When a user calls to talk about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0443] Audio Output

[0444] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0445] User Roles

[0446] System Usage

[0447] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0448] Additional consultation

[0449] If the user wants to continue the consultation, they can repeat the same process as many times as they like. If professional help is required, the server will automatically transfer the user to an expert.

[0450] Specific examples

[0451] For example, if a user is suffering from work-related stress, they might say, "Work hasn't been going well lately, and I don't know what to do." This speech is sent to the server, where it is recognized. From the converted text, an emotion analysis engine extracts the emotions of "stress" and "hopelessness." The AI ​​model then generates "advice for overcoming work-related problems," which is then converted into speech by a speech synthesis engine. Specific advice such as "The key to overcoming work-related problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[0452] Role Alignment

[0453] These components work together to provide an environment where users at high risk of suicide can seek advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[0454] In this way, the present invention realizes a consultation system that can respond quickly and effectively at any time, including at night, thereby contributing to reducing the risk of suicide.

[0455] The processing flow will be explained below.

[0456] Step 1:

[0457] The user dials the phone number of the IVR system using a smartphone or landline, which initiates voice input by the user.

[0458] Step 2:

[0459] The device records the user's voice in real time and transmits the voice data to the server using a secure protocol (e.g., HTTPS).

[0460] Step 3:

[0461] The server converts the received voice data into text data using a speech recognition library (e.g., speech recognition API). This conversion process turns the voice signal into a string of characters.

[0462] Step 4:

[0463] The converted text data is sent to an emotion and intent analysis engine, which extracts the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data.

[0464] Step 5:

[0465] Based on the extracted emotions and intent, the server uses an AI model to generate appropriate answers. This model has previously learned the knowledge of doctors and counselors to create specific and appropriate responses.

[0466] Step 6:

[0467] The generated response text data is converted into voice data using a speech synthesis engine (e.g., speech synthesis API), resulting in a human-like voice response.

[0468] Step 7:

[0469] The server transmits the generated voice data to the user's terminal, which receives the voice data and plays it back to the user.

[0470] Step 8:

[0471] The user receives the voice response provided through the terminal and, if necessary, can request further consultation, in which case the same process (steps 1 to 7) can be repeated.

[0472] Step 9:

[0473] If the user's consultation requires professional help, the server transfers the user to a specialist based on the results of emotion and intent analysis. To do this, it searches for the specialist's contact information and relays the connection using a conference call system.

[0474] The processing steps of this system provide an environment where users can receive consultation 24 hours a day, 365 days a year.

[0475] Example 1

[0476] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0477] In modern society, many people suffer from stress in their daily lives and work, feelings of loneliness, and a variety of other worries. One particularly serious problem is consultations that involve the risk of suicide. Current counseling services and consultation centers often have limited response times and are unable to provide prompt and appropriate responses in emergencies. Furthermore, professional assistance is often unavailable at night or on holidays, which can leave users without the support they need and put them in danger. Given this current situation, there is a need for a system that is available 24 hours a day, 365 days a year, can accurately analyze users' emotions and intentions, and provide appropriate responses in real time.

[0478] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0479] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, means for analyzing the user's emotions and intentions from the converted text data, generation means using a generative AI model to generate answers based on the analyzed emotions and intentions, speech synthesis means for converting the generated answers into voice data, and means for transmitting the voice data to a user terminal. This enables real-time analysis of voice input from the user and immediate provision of appropriate answers. Furthermore, by transferring the call to a specialist as needed, more specialized support can be received, realizing consultation services 24 hours a day, 365 days a year, including nights and holidays. This provides a safe and reliable system that can handle serious consultations, including those regarding suicide risk.

[0480] A "server" is a centralized control device for receiving and processing data sent by users.

[0481] "Voice data" refers to data in which voice information uttered by a user is recorded as a digital signal.

[0482] "Speech recognition means" refers to a technology or device that analyzes voice data and converts it into text data.

[0483] "Text data" is character information converted from voice data by a voice recognition means.

[0484] The "analysis means" is a technique or device for identifying a user's emotions and intentions from text data.

[0485] A "generative AI model" is an artificial intelligence model that has learned a huge amount of data in advance and generates appropriate answers based on the user's requests and emotions.

[0486] "Generation means" refers to a technology or device that uses a generative AI model to create an answer based on information obtained from the analysis means.

[0487] "Speech synthesis means" refers to a technique or device that generates voice data based on generated text data.

[0488] A "user terminal" is a device used by a user to communicate with a server. This includes smartphones and landlines.

[0489] "Transfer means" refers to a technology or device that transmits data to an expert based on the user's voice input and the analysis results.

[0490] This invention is a system that provides a 24 / 7 consultation service by combining voice recognition technology and an emotion analysis engine. It is specifically designed to allow users at risk of suicide to talk about their worries by voice.

[0491] Server Roles

[0492] Voice input acceptance

[0493] The server receives the voice data sent by the user in real time. When the user calls the consultation service, the voice is transmitted directly to the server. The voice data is transmitted to the server using a secure communication protocol (e.g., HTTPS).

[0494] Analysis of audio data

[0495] When the server receives the voice data, it first converts it into text using a speech recognition library, using a speech recognition engine such as the Google Cloud Speech-to-Text API to achieve high accuracy.

[0496] Emotion and Intention Analysis

[0497] The text data is then sent to a sentiment and intent analysis engine, which uses natural language processing techniques to extract the user's emotions (e.g., anxiety, sadness, urgency, etc.) and intent (e.g., help-seeking, advice-seeking) from the text data. This process uses tools such as the Microsoft Azure Text Analytics API.

[0498] Answer generation

[0499] Based on the analyzed emotions and intent, the server generates an answer using a generative AI model (e.g., OpenAI GPT-4) that has previously learned the knowledge of doctors and counselors. This model creates an appropriate response based on the information and emotions the user is seeking.

[0500] Conversion to audio

[0501] The generated text response is converted into audio data using a speech synthesis engine, using voice synthesis technologies such as Amazon Polly to provide the user with a human-like voice.

[0502] Sending audio data

[0503] Finally, the generated voice data is sent to the user's terminal, where the user can receive a voice response to their inquiry. This voice data is also sent using a secure communication protocol.

[0504] Device Role

[0505] Voice input

[0506] When a user calls to talk about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0507] Audio Output

[0508] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0509] User Roles

[0510] System Usage

[0511] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0512] Additional consultation

[0513] If the user wants to continue the consultation, they can repeat the same process as many times as they like. If professional help is required, the server will automatically transfer the user to an expert.

[0514] Specific examples

[0515] For example, if a user is suffering from work-related stress, they might say, "Work hasn't been going well lately, and I don't know what to do." This speech is sent to the server, where it is recognized. The converted text is then processed by an emotion analysis engine, which extracts the emotions of "stress" and "hopelessness." The generative AI model then generates "advice for overcoming work-related problems," which is then converted into speech by a speech synthesis engine. Specific advice such as "The key to overcoming work-related problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is then provided to the user via voice.

[0516] Examples of prompt statements

[0517] Below is an example of a prompt that can be fed into a generative AI model:

[0518] "A user says, 'Work hasn't been going well lately, and I don't know what to do.' Please provide appropriate advice to address the stress and despair the user is feeling."

[0519] This system provides an environment where users at high risk of suicide can receive advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[0520] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0521] Step 1:

[0522] User voice input

[0523] When a user has a problem, they call the helpline using their smartphone or landline. They then speak to the helpline, describing their specific problem, such as, "Recently, things haven't been going well at work, and I don't know what to do." This voice data becomes the input.

[0524] Step 2:

[0525] Sending audio data by the device

[0526] The device (smartphone or landline phone) records the user's voice in real time and sends it to the server. This voice data is sent to the server using a secure communication protocol (e.g., HTTPS). The input is the user's voice data, and the output is the voice data sent to the server.

[0527] Step 3:

[0528] Receiving audio data by the server

[0529] The server receives the voice data sent from the device in real time. The received voice data is temporarily stored in the server's memory. The input is the voice data sent from the device, and the output is the voice data stored in the server.

[0530] Step 4:

[0531] Server-based speech recognition and text conversion

[0532] The server converts the received voice data into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text API). The voice data is analyzed and the corresponding text data is generated. The input is the voice data stored on the server, and the output is the converted text data.

[0533] Step 5:

[0534] Server-based emotion and intent analysis

[0535] The server sends the converted text data to an emotion and intent analysis engine (e.g., Microsoft Azure Text Analytics API), which extracts the user's emotion (e.g., stress, despair) and intent (e.g., asking for help) from the text data. The input is the text data, and the output is the emotion and intent analysis results.

[0536] Step 6:

[0537] Server-generated answers

[0538] Based on the results of the emotion and intent analysis, the server sends a prompt to a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate answer. An example of a prompt sentence is, "The user said, 'Recently, things haven't been going well at work, and I don't know what to do anymore.' Please provide appropriate advice to help the user with the stress and despair they are feeling." The input is the results of the emotion and intent analysis, and the output is the generated answer in text format.

[0539] Step 7:

[0540] Server-based speech synthesis

[0541] The generated text response is converted back into audio data by a speech synthesis engine (e.g., Amazon Polly). This process uses techniques to express natural pronunciation and emotion in the voice. The input is the generated text answer, and the output is audio data.

[0542] Step 8:

[0543] Server sends audio data

[0544] The server then transmits the generated voice data back to the device, again using a secure communication protocol. The input is the generated voice data, and the output is the voice data sent to the device.

[0545] Step 9:

[0546] Receiving and playing audio data by a terminal

[0547] The terminal receives the voice data sent from the server and plays it back to the user. The user can receive a specific answer to their question by listening to the played back voice. The input is the voice data sent from the server, and the output is the voice played back to the user.

[0548] (Application example 1)

[0549] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0550] In modern society, the number of people suffering from stress and mental health problems is increasing, making immediate and appropriate support essential, especially for those at high risk of suicide. However, existing systems have difficulty analyzing emotions and generating responses in real time, and are therefore inadequate, especially at night. Furthermore, there is a need for a method that allows users to effectively utilize smart devices and receive a more realistic experience.

[0551] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0552] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, means for analyzing a user's emotions and intentions from the text data, means for generating a response based on the analyzed emotions and intentions, means for converting the generated response into voice data, means for transmitting the voice data to a user terminal, and means for receiving a user's voice input in real time and providing a response to the user through smart glasses or a head-mounted display. This allows for 24-hour, 365-day service, and enables users to receive immediate and effective support by providing emotion analysis and appropriate responses in real time through the smart device.

[0553] "Voice data" refers to data that digitally represents a user's speech.

[0554] "Text data" refers to data expressed as character information converted by a voice recognition means.

[0555] "Speech recognition means" refers to a device or software for converting voice data into text data.

[0556] "Means for analyzing emotions and intentions" refers to a device or software for analyzing a user's emotions and intentions from text data.

[0557] The "means for generating an answer" refers to a device or software that generates an answer to be provided to a user based on the analyzed emotions and intentions.

[0558] "Means for converting into voice data" refers to a device or software for converting the generated response into voice data.

[0559] A "user terminal" refers to a device used by a user, including a smartphone, smart glasses, a head-mounted display, etc.

[0560] The "means for transmitting voice data to a user terminal" refers to a device or software for transmitting the generated voice data to a user terminal.

[0561] "Smart glasses" are glasses-type devices equipped with functions such as augmented reality (AR).

[0562] A "head-mounted display" is a display device that is worn on the user's head and provides images and sounds.

[0563] "Real-time" means that data processing occurs almost immediately, allowing users to receive results without waiting.

[0564] This invention combines voice recognition technology with an emotion analysis engine to provide a 24 / 7 consultation service system for users at risk of suicide. In particular, it utilizes smart glasses and head-mounted displays to provide appropriate support to users in real time. The main components of this system include a voice recognition unit, an emotion and intention analysis unit, a response generation unit, a voice synthesis unit, and a voice data transmission unit.

[0565] Handling voice input and reception

[0566] The server receives the user's voice input in real time. It uses the Python speech_recognition library to convert the user's speech into text data. This voice data is acquired through the microphone in the smart glasses or head-mounted display.

[0567] Emotion and Intention Analysis

[0568] When the server receives the text data, it uses Hugging Face's sentiment-analysis model to analyze the user's emotions (e.g., anxiety, sadness, urgency). It also determines the user's intent based on the analyzed emotions and the text content. The results of this analysis are used to accurately identify situations in which advice or support is needed.

[0569] Generate answers

[0570] The server uses Hugging Face's text-generation (GPT-3) model to generate appropriate answers based on the analyzed emotions and intent, providing high-quality responses in real time that mimic expert advice and support.

[0571] Speech synthesis and output

[0572] The generated text response is converted into audio data using the gTTS (Google Text-to-Speech) library, and this audio data is provided to the user through the speakers of smart glasses or head-mounted displays, allowing the user to listen to the audio response directly using the device in front of them.

[0573] Specific examples

[0574] For example, if a user says, "I've been having trouble sleeping lately, please help me," this speech is sent to the server and converted into text using speech recognition. Then, emotion and intent analysis extracts the emotions "anxiety" and "seeking help." The AI ​​model then generates an answer using the following prompt:

[0575] Example prompt sentence:

[0576] User: I've been having trouble sleeping lately, please help.

[0577] System: Please provide advice for insomnia.

[0578] Ultimately, the voice will provide advice such as, "The best way to combat your insomnia is to relax before bed. Try meditating or doing some gentle stretching." This allows users to receive the support they need in real time.

[0579] This invention allows users to receive support 24 hours a day, 365 days a year with peace of mind, and provides an environment where immediate response is possible, especially for people at high risk of suicide.

[0580] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0581] Step 1:

[0582] The user inputs voice through smart glasses or a head-mounted display. The device records this voice in real time and transmits it to the server as voice data. The transmitted voice data is the input.

[0583] Step 2:

[0584] The server receives the transmitted voice data and converts it to text using the Python speech_recognition library. This converted text data is the output. This process is performed by the speech recognition means.

[0585] Step 3:

[0586] The server inputs the text data obtained from the speech recognition means into Hugging Face's sentiment-analysis model, which analyzes the user's emotions from the text data. This analysis generates emotional information such as anxiety, sadness, and urgency as output. The emotion and intent analysis means performs this process.

[0587] Step 4:

[0588] The server generates an appropriate answer for the user based on the results of the emotion analysis (emotion information) by inputting it into Hugging Face's text-generation (GPT-3) model. This generated text response is the output. This process is carried out by the generation means for generating the answer.

[0589] Step 5:

[0590] The server inputs the generated text response into the gTTS (Google Text-to-Speech) library and converts it to speech. This converted speech response is the output. The speech conversion method performs this process.

[0591] Step 6:

[0592] The server converts the response into voice data and sends it to the terminal. The terminal plays back the received voice data and provides it to the user. The voice response is played back through the user's smart glasses or head-mounted display. The provision of this voice response is the final output. This process is performed by a means for transmitting voice data to the user terminal.

[0593] Step 7:

[0594] The user listens to the voice response provided by the device and can re-voice the patient if necessary, and if necessary, repeat the consultation or request a transfer to a specialist. This interactive process continues.

[0595] Through this specific processing step, the system analyzes the user's emotions and intentions in real time and provides appropriate support.

[0596] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0597] This invention is a system that combines voice recognition technology and an emotion analysis engine to provide a 24 / 7 consultation service. It is designed to enable users at risk of suicide to discuss their concerns via voice. The invention adds an emotion engine with the ability to recognize the user's emotions and adaptively adjust responses.

[0598] Server Roles

[0599] Voice input acceptance

[0600] The server receives the voice data sent by the user in real time. When the user calls the helpline, the voice is transmitted directly to the server.

[0601] Analysis of audio data

[0602] When the server receives the voice data, it first converts it into text using a voice recognition library, and uses a variety of voice recognition engines to achieve high accuracy.

[0603] Emotion and Intention Analysis

[0604] The converted text data is sent to a sentiment analysis engine, which the server uses to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data. Furthermore, the sentiment analysis engine has the ability to adaptively adjust responses.

[0605] Answer generation

[0606] The server generates answers based on the analyzed emotions and intent. To generate answers, it uses an AI model that has previously learned the knowledge of doctors and counselors. This model creates appropriate responses based on the information and emotions the user is seeking.

[0607] Adaptive response adjustment

[0608] The emotion engine adaptively adjusts the generated response based on the user's perceived emotion, a process that provides a response that best suits the user's emotional state.

[0609] Conversion to audio

[0610] The generated response text is converted into voice data using a speech synthesis engine, which uses speech synthesis technology to provide answers to users in a human-like voice.

[0611] Sending audio data

[0612] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[0613] Device Role

[0614] Voice input

[0615] When a user calls and talks about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0616] Audio Output

[0617] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0618] User Roles

[0619] System Usage

[0620] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0621] Additional consultation

[0622] If the user wants to continue the consultation, they can repeat the same process (steps 1 to 7). Also, if professional help is required, the server can automatically transfer the user to a specialist.

[0623] Specific examples

[0624] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent to a server for speech recognition. The converted text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. The AI ​​model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "The key to overcoming work problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[0625] Role Alignment

[0626] These components work together to provide an environment where users at high risk of suicide can seek advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[0627] In this way, the present invention realizes a consultation system that can respond quickly and effectively at any time, including at night, thereby contributing to reducing the risk of suicide.

[0628] The processing flow will be explained below.

[0629] Step 1:

[0630] A user dials the phone number of the IVR system using a smartphone or landline, which initiates voice input.

[0631] Step 2:

[0632] The device records the user's voice in real time and transmits the voice data to the server. The data is transferred using a secure communication protocol (e.g., HTTPS).

[0633] Step 3:

[0634] The server converts the received voice data into text data using a voice recognition library (e.g., a voice recognition API). In this step, the voice signal is represented as a string.

[0635] Step 4:

[0636] The converted text data is sent to a sentiment analysis engine, which the server uses to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for information, etc.) from the text data.

[0637] Step 5:

[0638] The sentiment analysis engine identifies the user's current emotional state based on the perceived emotions and intent, which is important for subsequent response generation.

[0639] Step 6:

[0640] The server uses an AI model to generate responses based on the extracted emotions and intent, which is based on knowledge previously learned from experts (e.g., doctors, counselors, etc.).

[0641] Step 7:

[0642] The generated response text data is adaptively adjusted through an emotion engine, which selects the most appropriate expression and tone for the user's current emotional state.

[0643] Step 8:

[0644] The adjusted response text data is converted into audio data using a speech synthesis engine (e.g., speech synthesis API). This process produces a natural, human-like voice.

[0645] Step 9:

[0646] The server transmits the generated voice data to the user's terminal, which receives the voice data and plays it back to the user.

[0647] Step 10:

[0648] The user receives the voice response provided through the terminal and, if necessary, can request further consultation, in which case the same process (steps 1 to 9) can be repeated.

[0649] Step 11:

[0650] If the user's consultation needs professional help, the server will transfer the user to a specialist based on the results of emotion and intent analysis. To do this, it will search for the specialist's contact information and establish a relay connection using a telephone conference system.

[0651] In this way, the system of the present invention recognizes the user's emotional state and adaptively adjusts responses based on that state, thereby providing effective mental health counseling. The processes in steps 1 through 11 work together to provide an environment where users at high risk of suicide can seek counseling safely 24 hours a day, 365 days a year.

[0652] Example 2

[0653] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0654] In modern society, the number of people suffering from mental health problems and stress is increasing, and there is an urgent need to provide prompt and appropriate support, especially for those at high risk of suicide. However, conventional consultation response systems have difficulty providing support 24 hours a day, 365 days a year, and it is difficult to accurately analyze emotions and intentions and generate appropriate responses. For this reason, there is a growing need for a system that can flexibly respond to changes in emotions and provide effective and prompt support to users at high risk of suicide.

[0655] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0656] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, emotion analysis means for analyzing a user's emotion and intention from the text data, generation means for generating an answer based on the analyzed emotion and intention, means for adaptively adjusting the generated answer based on the user's emotional state, means for converting the adjusted answer into voice data, and means for transmitting the voice data to a user terminal, thereby making it possible to provide an appropriate answer according to the user's emotional state in real time, 24 hours a day, 365 days a year.

[0657] The "means for receiving voice data" refers to a device or software that has the function of acquiring the voice uttered by the user as digital data, converting it into a format that can be processed within the system, and receiving it.

[0658] The "voice recognition means for converting voice data into text data" refers to a device or software that has the function of analyzing received voice data and converting it into corresponding text data.

[0659] "Emotion analysis means for analyzing user emotions and intentions from text data" refers to a device or software that has the function of analyzing a user's emotional state (e.g., anxiety, sadness, despair, etc.) and intentions (e.g., asking for help, asking for advice, etc.) based on converted text data.

[0660] The "generation means for generating a response based on the analyzed emotion and intention" is a device or software that has the function of generating an appropriate response based on data from the emotion analysis means.

[0661] "Means for adaptively adjusting the generated response based on the user's emotional state" refers to a device or software that has the function of fine-tuning the generated response in accordance with the user's emotional state and providing it in an optimal form.

[0662] The "means for converting the adjusted response into voice data" refers to a device or software that has the function of converting adaptively adjusted text data into voice data and providing the voice data to the user as a natural voice.

[0663] The "means for transmitting audio data to a user terminal" refers to a device or software that has the function of transmitting the generated and converted audio data to a user's device and playing it on that device.

[0664] This invention combines voice recognition technology with an emotion analysis engine to provide a consultation service system that is available 24 hours a day, 365 days a year. It is designed specifically to enable users at risk of suicide to discuss their concerns via voice. The system is equipped with an emotion engine that recognizes the user's emotions and adaptively adjusts responses.

[0665] Server Roles

[0666] Voice input acceptance

[0667] The server receives voice data sent from the user in real time. For example, when a user calls a helpline, the voice is transmitted directly to the server.

[0668] Analysis of audio data

[0669] When the server receives the voice data, it first converts it into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text, Amazon Transcribe), thereby obtaining text information about what the user said.

[0670] Emotion and Intention Analysis

[0671] The converted text data is sent to a sentiment analysis engine (e.g., IBM Watson Tone Analyzer, Microsoft Azure Text Analytics), which uses the engine to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data.

[0672] Answer generation

[0673] Based on the analyzed emotions and intent, the server generates answers using a generative AI model (e.g., OpenAI GPT-3), which has previously learned the knowledge of doctors and counselors to create appropriate responses to the information and emotions the user is seeking.

[0674] Adaptive response adjustment

[0675] The emotion engine adaptively adjusts the generated response based on the user's perceived emotion, a process that provides a response that best suits the user's emotional state.

[0676] Conversion to audio

[0677] The generated response text is converted into audio data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech, Amazon Polly), which uses speech synthesis technology to provide answers to users in a human-like voice.

[0678] Sending audio data

[0679] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[0680] Device Role

[0681] Voice input

[0682] When a user calls and talks about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0683] Audio Output

[0684] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0685] User Roles

[0686] System Usage

[0687] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0688] Additional consultation

[0689] If the user wants to continue with the consultation, they can repeat the same process, and if professional help is needed, the server can automatically transfer them to a specialist.

[0690] Specific examples

[0691] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent to a server for speech recognition. The converted text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. A generative AI model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "The key to overcoming work problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[0692] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0693] Step 1: Accept voice input

[0694] Input: User's voice

[0695] Process: The user calls the system and tells it about their concerns. The user's voice is recorded by the device (smartphone, landline, etc.).

[0696] Output: Recorded audio data (digital format)

[0697] Specific operation: When a user speaks about a problem such as "Work hasn't been going well lately," the device captures the voice in real time and sends it to the server as audio data.

[0698] Step 2: Analyzing the audio data

[0699] Input: Recorded audio data

[0700] Processing: The server sends the audio data to a speech recognition library, which converts it into text. Examples of speech recognition libraries used here include Google Cloud Speech-to-Text and Amazon Transcribe.

[0701] Output: Converted text data

[0702] Specific operation: The speech recognition library analyzes the voice data "Work is not going well" received by the server, and generates the corresponding text "Work is not going well."

[0703] Step 3: Emotion and Intent Analysis

[0704] Input: Converted text data

[0705] Processing: The server sends the text data to a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., help-seeking, advice-seeking, etc.).

[0706] Output: Sentiment and intent analysis results

[0707] Specific operation: The server inputs the text data "My work is not going well" into the emotion analysis engine, and the engine determines that the user is feeling "stress" or "despair."

[0708] Step 4: Answer Generation

[0709] Input: Sentiment and intent analysis results

[0710] Processing: The server uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate answer based on the analysis results. This model has been trained with the knowledge of doctors and counselors.

[0711] Output: Generated answer text

[0712] Specific behavior: If emotion analysis identifies a high level of "despair," the generative AI model will create specific advice such as "First, stay calm. It's important to understand your own feelings."

[0713] Step 5: Adaptively adjust the response

[0714] Input: Generated answer text

[0715] Processing: The server re-analyzes the generated answer text with the emotion engine and fine-tunes the response to best suit the user's emotional state.

[0716] Output: Adaptively adjusted answer text

[0717] What it does: The emotion engine translates the response created by the generative AI model, "It's important to understand your own feelings," into a "tone that conveys a calm and welcoming atmosphere."

[0718] Step 6: Convert to audio

[0719] Input: Adaptively adjusted answer text

[0720] Processing: The server converts the tailored response text into audio data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech, Amazon Polly).

[0721] Output: Generated audio data

[0722] What happens: The adjusted text response is sent to a speech synthesis engine, which converts it into speech data, generating a human-like voice saying, "It's important to understand your own feelings."

[0723] Step 7: Sending audio data

[0724] Input: Generated audio data

[0725] Processing: The server sends the generated audio data to the user's device.

[0726] Output: The audio response the user receives

[0727] Specific operation: The server sends the voice data to the user's device, which then plays it back. The user receives the advice, "It's important to understand your own feelings."

[0728] (Application example 2)

[0729] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0730] In recent years, the number of individuals at risk of suicide due to psychological stress and social anxiety has been increasing. It is often difficult to find an appropriate counselor or support, especially during late night or early morning hours. Furthermore, existing counseling systems lack the ability to properly analyze emotions and provide optimal responses to users. Furthermore, if the counseling process is not carried out quickly and effectively, there is a risk that the user's psychological burden will increase. Therefore, a system that operates 24 hours a day, 365 days a year and provides appropriate advice based on emotions is needed.

[0731] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice data, voice recognition means for converting the voice data into text data, emotion analysis means for analyzing the user's emotions and intentions from the text data, generation means for generating an answer based on the analyzed emotions and intentions, adjustment means for adaptively adjusting the generated answer, voice synthesis means for converting the adjusted answer into voice data, means for transmitting the voice data to the user terminal, and means for receiving the user's voice input and providing voice and visual information through a head-mounted display. This makes it possible to provide appropriate advice based on the user's emotions in the form of voice and visual information 24 hours a day, 365 days a year.

[0732] "Voice data" refers to voice information input by a user.

[0733] "Means" refers to a combination of devices and software for realizing a specific function.

[0734] "Speech recognition means" refers to a device or algorithm for converting input voice data into text data.

[0735] "Text data" refers to character string information converted by speech recognition means.

[0736] "Emotion analysis means" refers to devices and algorithms for analyzing user emotions and intentions from text data.

[0737] "Generator" refers to the system or algorithm that generates answers based on analyzed sentiment and intent.

[0738] "Adjustment means" refers to a system or algorithm that adaptively adjusts the generated answers based on the user's emotions.

[0739] "Speech synthesis means" refers to a device or algorithm for converting conditioned responses into speech data.

[0740] A "user terminal" is a device used by a user, and refers to an apparatus for inputting and outputting voice data.

[0741] A "head-mounted display" is a display device that provides information to the wearer's field of vision and has the function of simultaneously providing audio and visual information.

[0742] "Adaptive" refers to the ability to change adaptively according to the situation or conditions.

[0743] "Visual information" refers to visual data, icons, and messages presented to the user.

[0744] This invention is a system that uses a head-mounted display (HMD) to provide a voice-based consultation service available 24 hours a day, 365 days a year. This system combines voice recognition technology and an emotion analysis engine to provide appropriate advice to users and aim to reduce the risk of suicide.

[0745] Server Roles

[0746] The server processes the voice data and generates a response using the following means:

[0747] Voice input acceptance

[0748] The server receives the voice data sent by the user in real time. When the user speaks through the microphone built into the HMD, the voice data is sent to the server.

[0749] Analysis of audio data

[0750] When the server receives the voice data, it converts the voice data into text data using a voice recognition library (e.g., Python's speech_recognition).

[0751] Emotion and Intention Analysis

[0752] The converted text data is sent to an emotion analysis engine (e.g., Hugging Face's transformers library), which the server uses to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent from the text data.

[0753] Answer generation

[0754] Based on the analyzed sentiment and intent, the server utilizes a generative AI model (e.g., GPT-3) to generate appropriate answers, using prompts based on pre-trained datasets.

[0755] Adaptive response adjustment

[0756] The generated responses are adaptively adjusted in tone and content based on the user's emotions as analyzed by the emotion engine.

[0757] Conversion to audio

[0758] The tailored answers are then converted into audio data using a speech synthesis engine, allowing the answers to be provided to the user in a human-like voice.

[0759] Sending audio data

[0760] Finally, the generated audio data is sent to the user's HMD and played back to the user through the HMD's speakers.

[0761] Device Role

[0762] Voice input

[0763] When the user talks about their concerns through the HMD's microphone, the device records the audio and sends it to the server in real time.

[0764] Audio Output

[0765] The voice data sent from the server is played back to the user through the HMD speaker, allowing the user to receive appropriate support in real time.

[0766] User Roles

[0767] System Usage

[0768] When a user has a problem, they can wear the HMD and talk about it through a microphone. By talking about specific problems, the system can provide appropriate support.

[0769] Additional consultation

[0770] If the user wishes to continue the consultation, they can repeat the same process. There is also a function to transfer the call to a specialist if necessary, ensuring appropriate treatment in emergencies or serious cases.

[0771] Specific examples

[0772] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent from the HMD to the server. The server performs speech recognition, and the generated text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. The AI ​​model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "To overcome work problems, you need to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via the HMD via voice.

[0773] Prompt Sentence Examples

[0774] For users who are in an anxious emotional state, provide appropriate advice on the following: Things have been going badly at work lately, and I don't know what to do.

[0775] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0776] Step 1:

[0777] The user voice-inputs their concerns through the HMD microphone. The voice data is recorded by the HMD and sent to the server in real time. The consultation process begins when the user's voice input is transferred to the server as input data.

[0778] Step 2:

[0779] Once the server receives the audio data, it uses a speech recognition library (e.g., Python's speech_recognition) to convert the audio data into text. This step takes the audio data as input and generates highly accurate text data. Once the audio data is converted to text, it is ready for the next processing step.

[0780] Step 3:

[0781] The converted text data is sent to an emotion analysis engine (e.g., Hugging Face's transformers library). The server uses this engine to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent from the text data. As output, the user's emotions and intent are extracted, providing the basis for the next step.

[0782] Step 4:

[0783] Based on the analyzed sentiment and intent, the server utilizes a generative AI model (e.g., GPT-3) to generate an appropriate answer. A prompt sentence is used to instruct the model to generate an answer text appropriate to the user's situation. Sentiment and intent are used as input, and a generated answer is obtained as output.

[0784] Step 5:

[0785] The generated answer is adaptively adjusted in tone and content based on the user's emotions analyzed by the emotion engine. This step takes the generated answer text as input and optimizes the response with the appropriate tone and nuance depending on the user's emotional state. The final adjusted answer text is output.

[0786] Step 6:

[0787] The adjusted answer text is converted into voice data using a speech synthesis engine. The server uses this engine to generate natural-sounding speech and outputs it as voice data to be conveyed to the user. This conversion process uses the adjusted text data as input and generates human-like voice data.

[0788] Step 7:

[0789] The final generated voice data is sent to the user's HMD and played back to the user through the HMD's speaker. The server sends the voice data to the HMD, and the user receives a response in real time. The voice data is delivered to the HMD as output, thus completing the user's consultation process.

[0790] Specific examples of operation

[0791] For example, if a user says, "Things haven't been going well at work lately, and I don't know what to do anymore," then:

[0792] Step 1: The voice input "Work hasn't been going well lately, and I don't know what to do anymore" is sent from the HMD to the server.

[0793] Step 2: The speech recognition library converts this speech into text data, generating the text "Work hasn't been going well lately and I don't know what to do."

[0794] Step 3: The sentiment analysis engine analyzes the text and recognizes the emotions "anxiety" and "despair."

[0795] Step 4: The generative AI model generates an appropriate response for the user based on this emotion and intent. Example prompt: "For a user who is in an anxious emotional state, please provide appropriate advice on the following: Things haven't been going well at work recently, and I don't know what to do anymore."

[0796] Step 5: The generated answer, "The key to overcoming work-related worries is to understand your own feelings," is adaptively adjusted to optimize the tone to best fit the user's current emotions.

[0797] Step 6: The adjusted answers are converted into voice data by a speech synthesis engine.

[0798] Step 7: The final voice data is sent to the HMD, and the user receives the answer in real time through the speaker.

[0799] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0800] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0801] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0802] [Third embodiment]

[0803] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0804] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0805] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0806] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0807] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0808] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0809] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0810] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0811] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0812] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0813] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0814] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0815] This invention is a system that provides a 24 / 7 consultation service by combining voice recognition technology and an emotion analysis engine. It is specifically designed to allow users at risk of suicide to talk about their worries by voice.

[0816] Server Roles

[0817] Voice input acceptance

[0818] The server receives the voice data sent by the user in real time. When the user calls the helpline, the voice is transmitted directly to the server.

[0819] Analysis of audio data

[0820] When the server receives the voice data, it first converts it into text data using a speech recognition library, using various speech recognition engines (e.g., speech recognition APIs) to achieve high accuracy.

[0821] Emotion and Intention Analysis

[0822] The text data is then fed into a sentiment and intent analysis engine, which uses natural language processing techniques to extract the user's emotions (e.g., anxiety, sadness, urgency, etc.) and intent (e.g., help-seeking, advice-seeking, etc.) from the text data.

[0823] Answer generation

[0824] The server generates answers based on the analyzed emotions and intent. To generate answers, it uses an AI model that has previously learned the knowledge of doctors and counselors. This model creates appropriate responses based on the information and emotions the user is seeking.

[0825] Conversion to audio

[0826] The generated text response is converted into audio data using a speech synthesis engine, which uses speech synthesis technology (e.g., speech synthesis API) to provide a human-like voice response to the user.

[0827] Sending audio data

[0828] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[0829] Device Role

[0830] Voice input

[0831] When a user calls to talk about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0832] Audio Output

[0833] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0834] User Roles

[0835] System Usage

[0836] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0837] Additional consultation

[0838] If the user wants to continue the consultation, they can repeat the same process as many times as they like. If professional help is required, the server will automatically transfer the user to an expert.

[0839] Specific examples

[0840] For example, if a user is suffering from work-related stress, they might say, "Work hasn't been going well lately, and I don't know what to do." This speech is sent to the server, where it is recognized. From the converted text, an emotion analysis engine extracts the emotions of "stress" and "hopelessness." The AI ​​model then generates "advice for overcoming work-related problems," which is then converted into speech by a speech synthesis engine. Specific advice such as "The key to overcoming work-related problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[0841] Role Alignment

[0842] These components work together to provide an environment where users at high risk of suicide can seek advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[0843] In this way, the present invention realizes a consultation system that can respond quickly and effectively at any time, including at night, thereby contributing to reducing the risk of suicide.

[0844] The processing flow will be explained below.

[0845] Step 1:

[0846] The user dials the phone number of the IVR system using a smartphone or landline, which initiates voice input by the user.

[0847] Step 2:

[0848] The device records the user's voice in real time and transmits the voice data to the server using a secure protocol (e.g., HTTPS).

[0849] Step 3:

[0850] The server converts the received voice data into text data using a speech recognition library (e.g., speech recognition API). This conversion process turns the voice signal into a string of characters.

[0851] Step 4:

[0852] The converted text data is sent to an emotion and intent analysis engine, which extracts the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data.

[0853] Step 5:

[0854] Based on the extracted emotions and intent, the server uses an AI model to generate appropriate answers. This model has previously learned the knowledge of doctors and counselors to create specific and appropriate responses.

[0855] Step 6:

[0856] The generated response text data is converted into voice data using a speech synthesis engine (e.g., speech synthesis API), resulting in a human-like voice response.

[0857] Step 7:

[0858] The server transmits the generated voice data to the user's terminal, which receives the voice data and plays it back to the user.

[0859] Step 8:

[0860] The user receives the voice response provided through the terminal and, if necessary, can request further consultation, in which case the same process (steps 1 to 7) can be repeated.

[0861] Step 9:

[0862] If the user's consultation requires professional help, the server transfers the user to a specialist based on the results of emotion and intent analysis. To do this, it searches for the specialist's contact information and relays the connection using a conference call system.

[0863] The processing steps of this system provide an environment where users can receive consultation 24 hours a day, 365 days a year.

[0864] Example 1

[0865] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0866] In modern society, many people suffer from stress in their daily lives and work, feelings of loneliness, and a variety of other worries. One particularly serious problem is consultations that involve the risk of suicide. Current counseling services and consultation centers often have limited response times and are unable to provide prompt and appropriate responses in emergencies. Furthermore, professional assistance is often unavailable at night or on holidays, which can leave users without the support they need and put them in danger. Given this current situation, there is a need for a system that is available 24 hours a day, 365 days a year, can accurately analyze users' emotions and intentions, and provide appropriate responses in real time.

[0867] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0868] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, means for analyzing the user's emotions and intentions from the converted text data, generation means using a generative AI model to generate answers based on the analyzed emotions and intentions, speech synthesis means for converting the generated answers into voice data, and means for transmitting the voice data to a user terminal. This enables real-time analysis of voice input from the user and immediate provision of appropriate answers. Furthermore, by transferring the call to a specialist as needed, more specialized support can be received, realizing consultation services 24 hours a day, 365 days a year, including nights and holidays. This provides a safe and reliable system that can handle serious consultations, including those regarding suicide risk.

[0869] A "server" is a centralized control device for receiving and processing data sent by users.

[0870] "Voice data" refers to data in which voice information uttered by a user is recorded as a digital signal.

[0871] "Speech recognition means" refers to a technology or device that analyzes voice data and converts it into text data.

[0872] "Text data" is character information converted from voice data by a voice recognition means.

[0873] The "analysis means" is a technique or device for identifying a user's emotions and intentions from text data.

[0874] A "generative AI model" is an artificial intelligence model that has learned a huge amount of data in advance and generates appropriate answers based on the user's requests and emotions.

[0875] "Generation means" refers to a technology or device that uses a generative AI model to create an answer based on information obtained from the analysis means.

[0876] "Speech synthesis means" refers to a technique or device that generates voice data based on generated text data.

[0877] A "user terminal" is a device used by a user to communicate with a server. This includes smartphones and landlines.

[0878] "Transfer means" refers to a technology or device that transmits data to an expert based on the user's voice input and the analysis results.

[0879] This invention is a system that provides a 24 / 7 consultation service by combining voice recognition technology and an emotion analysis engine. It is specifically designed to allow users at risk of suicide to talk about their worries by voice.

[0880] Server Roles

[0881] Voice input acceptance

[0882] The server receives the voice data sent by the user in real time. When the user calls the consultation service, the voice is transmitted directly to the server. The voice data is transmitted to the server using a secure communication protocol (e.g., HTTPS).

[0883] Analysis of audio data

[0884] When the server receives the voice data, it first converts it into text using a speech recognition library, using a speech recognition engine such as the Google Cloud Speech-to-Text API to achieve high accuracy.

[0885] Emotion and Intention Analysis

[0886] The text data is then sent to a sentiment and intent analysis engine, which uses natural language processing techniques to extract the user's emotions (e.g., anxiety, sadness, urgency, etc.) and intent (e.g., help-seeking, advice-seeking) from the text data. This process uses tools such as the Microsoft Azure Text Analytics API.

[0887] Answer generation

[0888] Based on the analyzed emotions and intent, the server generates an answer using a generative AI model (e.g., OpenAI GPT-4) that has previously learned the knowledge of doctors and counselors. This model creates an appropriate response based on the information and emotions the user is seeking.

[0889] Conversion to audio

[0890] The generated text response is converted into audio data using a speech synthesis engine, using voice synthesis technologies such as Amazon Polly to provide the user with a human-like voice.

[0891] Sending audio data

[0892] Finally, the generated voice data is sent to the user's terminal, where the user can receive a voice response to their inquiry. This voice data is also sent using a secure communication protocol.

[0893] Device Role

[0894] Voice input

[0895] When a user calls to talk about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[0896] Audio Output

[0897] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[0898] User Roles

[0899] System Usage

[0900] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[0901] Additional consultation

[0902] If the user wants to continue the consultation, they can repeat the same process as many times as they like. If professional help is required, the server will automatically transfer the user to an expert.

[0903] Specific examples

[0904] For example, if a user is suffering from work-related stress, they might say, "Work hasn't been going well lately, and I don't know what to do." This speech is sent to the server, where it is recognized. The converted text is then processed by an emotion analysis engine, which extracts the emotions of "stress" and "hopelessness." The generative AI model then generates "advice for overcoming work-related problems," which is then converted into speech by a speech synthesis engine. Specific advice such as "The key to overcoming work-related problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is then provided to the user via voice.

[0905] Examples of prompt statements

[0906] Below is an example of a prompt that can be fed into a generative AI model:

[0907] "A user says, 'Work hasn't been going well lately, and I don't know what to do.' Please provide appropriate advice to address the stress and despair the user is feeling."

[0908] This system provides an environment where users at high risk of suicide can receive advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[0909] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0910] Step 1:

[0911] User voice input

[0912] When a user has a problem, they call the helpline using their smartphone or landline. They then speak to the helpline, describing their specific problem, such as, "Recently, things haven't been going well at work, and I don't know what to do." This voice data becomes the input.

[0913] Step 2:

[0914] Sending audio data by the device

[0915] The device (smartphone or landline phone) records the user's voice in real time and sends it to the server. This voice data is sent to the server using a secure communication protocol (e.g., HTTPS). The input is the user's voice data, and the output is the voice data sent to the server.

[0916] Step 3:

[0917] Receiving audio data by the server

[0918] The server receives the voice data sent from the device in real time. The received voice data is temporarily stored in the server's memory. The input is the voice data sent from the device, and the output is the voice data stored in the server.

[0919] Step 4:

[0920] Server-based speech recognition and text conversion

[0921] The server converts the received voice data into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text API). The voice data is analyzed and the corresponding text data is generated. The input is the voice data stored on the server, and the output is the converted text data.

[0922] Step 5:

[0923] Server-based emotion and intent analysis

[0924] The server sends the converted text data to an emotion and intent analysis engine (e.g., Microsoft Azure Text Analytics API), which extracts the user's emotion (e.g., stress, despair) and intent (e.g., asking for help) from the text data. The input is the text data, and the output is the emotion and intent analysis results.

[0925] Step 6:

[0926] Server-generated answers

[0927] Based on the results of the emotion and intent analysis, the server sends a prompt to a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate answer. An example of a prompt sentence is, "The user said, 'Recently, things haven't been going well at work, and I don't know what to do anymore.' Please provide appropriate advice to help the user with the stress and despair they are feeling." The input is the results of the emotion and intent analysis, and the output is the generated answer in text format.

[0928] Step 7:

[0929] Server-based speech synthesis

[0930] The generated text response is converted back into audio data by a speech synthesis engine (e.g., Amazon Polly). This process uses techniques to express natural pronunciation and emotion in the voice. The input is the generated text answer, and the output is audio data.

[0931] Step 8:

[0932] Server sends audio data

[0933] The server then transmits the generated voice data back to the device, again using a secure communication protocol. The input is the generated voice data, and the output is the voice data sent to the device.

[0934] Step 9:

[0935] Receiving and playing audio data by a terminal

[0936] The terminal receives the voice data sent from the server and plays it back to the user. The user can receive a specific answer to their question by listening to the played back voice. The input is the voice data sent from the server, and the output is the voice played back to the user.

[0937] (Application example 1)

[0938] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0939] In modern society, the number of people suffering from stress and mental health problems is increasing, making immediate and appropriate support essential, especially for those at high risk of suicide. However, existing systems have difficulty analyzing emotions and generating responses in real time, and are therefore inadequate, especially at night. Furthermore, there is a need for a method that allows users to effectively utilize smart devices and receive a more realistic experience.

[0940] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0941] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, means for analyzing a user's emotions and intentions from the text data, means for generating a response based on the analyzed emotions and intentions, means for converting the generated response into voice data, means for transmitting the voice data to a user terminal, and means for receiving a user's voice input in real time and providing a response to the user through smart glasses or a head-mounted display. This allows for 24-hour, 365-day service, and enables users to receive immediate and effective support by providing emotion analysis and appropriate responses in real time through the smart device.

[0942] "Voice data" refers to data that digitally represents a user's speech.

[0943] "Text data" refers to data expressed as character information converted by a voice recognition means.

[0944] "Speech recognition means" refers to a device or software for converting voice data into text data.

[0945] "Means for analyzing emotions and intentions" refers to a device or software for analyzing a user's emotions and intentions from text data.

[0946] The "means for generating an answer" refers to a device or software that generates an answer to be provided to a user based on the analyzed emotions and intentions.

[0947] "Means for converting into voice data" refers to a device or software for converting the generated response into voice data.

[0948] A "user terminal" refers to a device used by a user, including a smartphone, smart glasses, a head-mounted display, etc.

[0949] The "means for transmitting voice data to a user terminal" refers to a device or software for transmitting the generated voice data to a user terminal.

[0950] "Smart glasses" are glasses-type devices equipped with functions such as augmented reality (AR).

[0951] A "head-mounted display" is a display device that is worn on the user's head and provides images and sounds.

[0952] "Real-time" means that data processing occurs almost immediately, allowing users to receive results without waiting.

[0953] This invention combines voice recognition technology with an emotion analysis engine to provide a 24 / 7 consultation service system for users at risk of suicide. In particular, it utilizes smart glasses and head-mounted displays to provide appropriate support to users in real time. The main components of this system include a voice recognition unit, an emotion and intention analysis unit, a response generation unit, a voice synthesis unit, and a voice data transmission unit.

[0954] Handling voice input and reception

[0955] The server receives the user's voice input in real time. It uses the Python speech_recognition library to convert the user's speech into text data. This voice data is acquired through the microphone in the smart glasses or head-mounted display.

[0956] Emotion and Intention Analysis

[0957] When the server receives the text data, it uses Hugging Face's sentiment-analysis model to analyze the user's emotions (e.g., anxiety, sadness, urgency). It also determines the user's intent based on the analyzed emotions and the text content. The results of this analysis are used to accurately identify situations in which advice or support is needed.

[0958] Generate answers

[0959] The server uses Hugging Face's text-generation (GPT-3) model to generate appropriate answers based on the analyzed emotions and intent, providing high-quality responses in real time that mimic expert advice and support.

[0960] Speech synthesis and output

[0961] The generated text response is converted into audio data using the gTTS (Google Text-to-Speech) library, and this audio data is provided to the user through the speakers of smart glasses or head-mounted displays, allowing the user to listen to the audio response directly using the device in front of them.

[0962] Specific examples

[0963] For example, if a user says, "I've been having trouble sleeping lately, please help me," this speech is sent to the server and converted into text using speech recognition. Then, emotion and intent analysis extracts the emotions "anxiety" and "seeking help." The AI ​​model then generates an answer using the following prompt:

[0964] Example prompt sentence:

[0965] User: I've been having trouble sleeping lately, please help.

[0966] System: Please provide advice for insomnia.

[0967] Ultimately, the voice will provide advice such as, "The best way to combat your insomnia is to relax before bed. Try meditating or doing some gentle stretching." This allows users to receive the support they need in real time.

[0968] This invention allows users to receive support 24 hours a day, 365 days a year with peace of mind, and provides an environment where immediate response is possible, especially for people at high risk of suicide.

[0969] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0970] Step 1:

[0971] The user inputs voice through smart glasses or a head-mounted display. The device records this voice in real time and transmits it to the server as voice data. The transmitted voice data is the input.

[0972] Step 2:

[0973] The server receives the transmitted voice data and converts it to text using the Python speech_recognition library. This converted text data is the output. This process is performed by the speech recognition means.

[0974] Step 3:

[0975] The server inputs the text data obtained from the speech recognition means into Hugging Face's sentiment-analysis model, which analyzes the user's emotions from the text data. This analysis generates emotional information such as anxiety, sadness, and urgency as output. The emotion and intent analysis means performs this process.

[0976] Step 4:

[0977] The server generates an appropriate answer for the user based on the results of the emotion analysis (emotion information) by inputting it into Hugging Face's text-generation (GPT-3) model. This generated text response is the output. This process is carried out by the generation means for generating the answer.

[0978] Step 5:

[0979] The server inputs the generated text response into the gTTS (Google Text-to-Speech) library and converts it to speech. This converted speech response is the output. The speech conversion method performs this process.

[0980] Step 6:

[0981] The server converts the response into voice data and sends it to the terminal. The terminal plays back the received voice data and provides it to the user. The voice response is played back through the user's smart glasses or head-mounted display. The provision of this voice response is the final output. This process is performed by a means for transmitting voice data to the user terminal.

[0982] Step 7:

[0983] The user listens to the voice response provided by the device and can re-voice the patient if necessary, and if necessary, repeat the consultation or request a transfer to a specialist. This interactive process continues.

[0984] Through this specific processing step, the system analyzes the user's emotions and intentions in real time and provides appropriate support.

[0985] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0986] This invention is a system that combines voice recognition technology and an emotion analysis engine to provide a 24 / 7 consultation service. It is designed to enable users at risk of suicide to discuss their concerns via voice. The invention adds an emotion engine with the ability to recognize the user's emotions and adaptively adjust responses.

[0987] Server Roles

[0988] Voice input acceptance

[0989] The server receives the voice data sent by the user in real time. When the user calls the helpline, the voice is transmitted directly to the server.

[0990] Analysis of audio data

[0991] When the server receives the voice data, it first converts it into text using a voice recognition library, and uses a variety of voice recognition engines to achieve high accuracy.

[0992] Emotion and Intention Analysis

[0993] The converted text data is sent to a sentiment analysis engine, which the server uses to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data. Furthermore, the sentiment analysis engine has the ability to adaptively adjust responses.

[0994] Answer generation

[0995] The server generates answers based on the analyzed emotions and intent. To generate answers, it uses an AI model that has previously learned the knowledge of doctors and counselors. This model creates appropriate responses based on the information and emotions the user is seeking.

[0996] Adaptive response adjustment

[0997] The emotion engine adaptively adjusts the generated response based on the user's perceived emotion, a process that provides a response that best suits the user's emotional state.

[0998] Conversion to audio

[0999] The generated response text is converted into voice data using a speech synthesis engine, which uses speech synthesis technology to provide answers to users in a human-like voice.

[1000] Sending audio data

[1001] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[1002] Device Role

[1003] Voice input

[1004] When a user calls and talks about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[1005] Audio Output

[1006] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[1007] User Roles

[1008] System Usage

[1009] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[1010] Additional consultation

[1011] If the user wants to continue the consultation, they can repeat the same process (steps 1 to 7). Also, if professional help is required, the server can automatically transfer the user to a specialist.

[1012] Specific examples

[1013] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent to a server for speech recognition. The converted text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. The AI ​​model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "The key to overcoming work problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[1014] Role Alignment

[1015] These components work together to provide an environment where users at high risk of suicide can seek advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[1016] In this way, the present invention realizes a consultation system that can respond quickly and effectively at any time, including at night, thereby contributing to reducing the risk of suicide.

[1017] The processing flow will be explained below.

[1018] Step 1:

[1019] A user dials the phone number of the IVR system using a smartphone or landline, which initiates voice input.

[1020] Step 2:

[1021] The device records the user's voice in real time and transmits the voice data to the server. The data is transferred using a secure communication protocol (e.g., HTTPS).

[1022] Step 3:

[1023] The server converts the received voice data into text data using a voice recognition library (e.g., a voice recognition API). In this step, the voice signal is represented as a string.

[1024] Step 4:

[1025] The converted text data is sent to a sentiment analysis engine, which the server uses to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for information, etc.) from the text data.

[1026] Step 5:

[1027] The sentiment analysis engine identifies the user's current emotional state based on the perceived emotions and intent, which is important for subsequent response generation.

[1028] Step 6:

[1029] The server uses an AI model to generate responses based on the extracted emotions and intent, which is based on knowledge previously learned from experts (e.g., doctors, counselors, etc.).

[1030] Step 7:

[1031] The generated response text data is adaptively adjusted through an emotion engine, which selects the most appropriate expression and tone for the user's current emotional state.

[1032] Step 8:

[1033] The adjusted response text data is converted into audio data using a speech synthesis engine (e.g., speech synthesis API). This process produces a natural, human-like voice.

[1034] Step 9:

[1035] The server transmits the generated voice data to the user's terminal, which receives the voice data and plays it back to the user.

[1036] Step 10:

[1037] The user receives the voice response provided through the terminal and, if necessary, can request further consultation, in which case the same process (steps 1 to 9) can be repeated.

[1038] Step 11:

[1039] If the user's consultation needs professional help, the server will transfer the user to a specialist based on the results of emotion and intent analysis. To do this, it will search for the specialist's contact information and establish a relay connection using a telephone conference system.

[1040] In this way, the system of the present invention recognizes the user's emotional state and adaptively adjusts responses based on that state, thereby providing effective mental health counseling. The processes in steps 1 through 11 work together to provide an environment where users at high risk of suicide can seek counseling safely 24 hours a day, 365 days a year.

[1041] Example 2

[1042] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1043] In modern society, the number of people suffering from mental health problems and stress is increasing, and there is an urgent need to provide prompt and appropriate support, especially for those at high risk of suicide. However, conventional consultation response systems have difficulty providing support 24 hours a day, 365 days a year, and it is difficult to accurately analyze emotions and intentions and generate appropriate responses. For this reason, there is a growing need for a system that can flexibly respond to changes in emotions and provide effective and prompt support to users at high risk of suicide.

[1044] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1045] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, emotion analysis means for analyzing a user's emotion and intention from the text data, generation means for generating an answer based on the analyzed emotion and intention, means for adaptively adjusting the generated answer based on the user's emotional state, means for converting the adjusted answer into voice data, and means for transmitting the voice data to a user terminal, thereby making it possible to provide an appropriate answer according to the user's emotional state in real time, 24 hours a day, 365 days a year.

[1046] The "means for receiving voice data" refers to a device or software that has the function of acquiring the voice uttered by the user as digital data, converting it into a format that can be processed within the system, and receiving it.

[1047] The "voice recognition means for converting voice data into text data" refers to a device or software that has the function of analyzing received voice data and converting it into corresponding text data.

[1048] "Emotion analysis means for analyzing user emotions and intentions from text data" refers to a device or software that has the function of analyzing a user's emotional state (e.g., anxiety, sadness, despair, etc.) and intentions (e.g., asking for help, asking for advice, etc.) based on converted text data.

[1049] The "generation means for generating a response based on the analyzed emotion and intention" is a device or software that has the function of generating an appropriate response based on data from the emotion analysis means.

[1050] "Means for adaptively adjusting the generated response based on the user's emotional state" refers to a device or software that has the function of fine-tuning the generated response in accordance with the user's emotional state and providing it in an optimal form.

[1051] The "means for converting the adjusted response into voice data" refers to a device or software that has the function of converting adaptively adjusted text data into voice data and providing the voice data to the user as a natural voice.

[1052] The "means for transmitting audio data to a user terminal" refers to a device or software that has the function of transmitting the generated and converted audio data to a user's device and playing it on that device.

[1053] This invention combines voice recognition technology with an emotion analysis engine to provide a consultation service system that is available 24 hours a day, 365 days a year. It is designed specifically to enable users at risk of suicide to discuss their concerns via voice. The system is equipped with an emotion engine that recognizes the user's emotions and adaptively adjusts responses.

[1054] Server Roles

[1055] Voice input acceptance

[1056] The server receives voice data sent from the user in real time. For example, when a user calls a helpline, the voice is transmitted directly to the server.

[1057] Analysis of audio data

[1058] When the server receives the voice data, it first converts it into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text, Amazon Transcribe), thereby obtaining text information about what the user said.

[1059] Emotion and Intention Analysis

[1060] The converted text data is sent to a sentiment analysis engine (e.g., IBM Watson Tone Analyzer, Microsoft Azure Text Analytics), which uses the engine to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data.

[1061] Answer generation

[1062] Based on the analyzed emotions and intent, the server generates answers using a generative AI model (e.g., OpenAI GPT-3), which has previously learned the knowledge of doctors and counselors to create appropriate responses to the information and emotions the user is seeking.

[1063] Adaptive response adjustment

[1064] The emotion engine adaptively adjusts the generated response based on the user's perceived emotion, a process that provides a response that best suits the user's emotional state.

[1065] Conversion to audio

[1066] The generated response text is converted into audio data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech, Amazon Polly), which uses speech synthesis technology to provide answers to users in a human-like voice.

[1067] Sending audio data

[1068] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[1069] Device Role

[1070] Voice input

[1071] When a user calls and talks about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[1072] Audio Output

[1073] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[1074] User Roles

[1075] System Usage

[1076] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[1077] Additional consultation

[1078] If the user wants to continue with the consultation, they can repeat the same process, and if professional help is needed, the server can automatically transfer them to a specialist.

[1079] Specific examples

[1080] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent to a server for speech recognition. The converted text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. A generative AI model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "The key to overcoming work problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[1081] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1082] Step 1: Accept voice input

[1083] Input: User's voice

[1084] Process: The user calls the system and tells it about their concerns. The user's voice is recorded by the device (smartphone, landline, etc.).

[1085] Output: Recorded audio data (digital format)

[1086] Specific operation: When a user speaks about a problem such as "Work hasn't been going well lately," the device captures the voice in real time and sends it to the server as audio data.

[1087] Step 2: Analyzing the audio data

[1088] Input: Recorded audio data

[1089] Processing: The server sends the audio data to a speech recognition library, which converts it into text. Examples of speech recognition libraries used here include Google Cloud Speech-to-Text and Amazon Transcribe.

[1090] Output: Converted text data

[1091] Specific operation: The speech recognition library analyzes the voice data "Work is not going well" received by the server, and generates the corresponding text "Work is not going well."

[1092] Step 3: Emotion and Intent Analysis

[1093] Input: Converted text data

[1094] Processing: The server sends the text data to a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., help-seeking, advice-seeking, etc.).

[1095] Output: Sentiment and intent analysis results

[1096] Specific operation: The server inputs the text data "My work is not going well" into the emotion analysis engine, and the engine determines that the user is feeling "stress" or "despair."

[1097] Step 4: Answer Generation

[1098] Input: Sentiment and intent analysis results

[1099] Processing: The server uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate answer based on the analysis results. This model has been trained with the knowledge of doctors and counselors.

[1100] Output: Generated answer text

[1101] Specific behavior: If emotion analysis identifies a high level of "despair," the generative AI model will create specific advice such as "First, stay calm. It's important to understand your own feelings."

[1102] Step 5: Adaptively adjust the response

[1103] Input: Generated answer text

[1104] Processing: The server re-analyzes the generated answer text with the emotion engine and fine-tunes the response to best suit the user's emotional state.

[1105] Output: Adaptively adjusted answer text

[1106] What it does: The emotion engine translates the response created by the generative AI model, "It's important to understand your own feelings," into a "tone that conveys a calm and welcoming atmosphere."

[1107] Step 6: Convert to audio

[1108] Input: Adaptively adjusted answer text

[1109] Processing: The server converts the tailored response text into audio data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech, Amazon Polly).

[1110] Output: Generated audio data

[1111] What happens: The adjusted text response is sent to a speech synthesis engine, which converts it into speech data, generating a human-like voice saying, "It's important to understand your own feelings."

[1112] Step 7: Sending audio data

[1113] Input: Generated audio data

[1114] Processing: The server sends the generated audio data to the user's device.

[1115] Output: The audio response the user receives

[1116] Specific operation: The server sends the voice data to the user's device, which then plays it back. The user receives the advice, "It's important to understand your own feelings."

[1117] (Application example 2)

[1118] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1119] In recent years, the number of individuals at risk of suicide due to psychological stress and social anxiety has been increasing. It is often difficult to find an appropriate counselor or support, especially during late night or early morning hours. Furthermore, existing counseling systems lack the ability to properly analyze emotions and provide optimal responses to users. Furthermore, if the counseling process is not carried out quickly and effectively, there is a risk that the user's psychological burden will increase. Therefore, a system that operates 24 hours a day, 365 days a year and provides appropriate advice based on emotions is needed.

[1120] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice data, voice recognition means for converting the voice data into text data, emotion analysis means for analyzing the user's emotions and intentions from the text data, generation means for generating an answer based on the analyzed emotions and intentions, adjustment means for adaptively adjusting the generated answer, voice synthesis means for converting the adjusted answer into voice data, means for transmitting the voice data to the user terminal, and means for receiving the user's voice input and providing voice and visual information through a head-mounted display. This makes it possible to provide appropriate advice based on the user's emotions in the form of voice and visual information 24 hours a day, 365 days a year.

[1121] "Voice data" refers to voice information input by a user.

[1122] "Means" refers to a combination of devices and software for realizing a specific function.

[1123] "Speech recognition means" refers to a device or algorithm for converting input voice data into text data.

[1124] "Text data" refers to character string information converted by speech recognition means.

[1125] "Emotion analysis means" refers to devices and algorithms for analyzing user emotions and intentions from text data.

[1126] "Generator" refers to the system or algorithm that generates answers based on analyzed sentiment and intent.

[1127] "Adjustment means" refers to a system or algorithm that adaptively adjusts the generated answers based on the user's emotions.

[1128] "Speech synthesis means" refers to a device or algorithm for converting conditioned responses into speech data.

[1129] A "user terminal" is a device used by a user, and refers to an apparatus for inputting and outputting voice data.

[1130] A "head-mounted display" is a display device that provides information to the wearer's field of vision and has the function of simultaneously providing audio and visual information.

[1131] "Adaptive" refers to the ability to change adaptively according to the situation or conditions.

[1132] "Visual information" refers to visual data, icons, and messages presented to the user.

[1133] This invention is a system that uses a head-mounted display (HMD) to provide a voice-based consultation service available 24 hours a day, 365 days a year. This system combines voice recognition technology and an emotion analysis engine to provide appropriate advice to users and aim to reduce the risk of suicide.

[1134] Server Roles

[1135] The server processes the voice data and generates a response using the following means:

[1136] Voice input acceptance

[1137] The server receives the voice data sent by the user in real time. When the user speaks through the microphone built into the HMD, the voice data is sent to the server.

[1138] Analysis of audio data

[1139] When the server receives the voice data, it converts the voice data into text data using a voice recognition library (e.g., Python's speech_recognition).

[1140] Emotion and Intention Analysis

[1141] The converted text data is sent to an emotion analysis engine (e.g., Hugging Face's transformers library), which the server uses to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent from the text data.

[1142] Answer generation

[1143] Based on the analyzed sentiment and intent, the server utilizes a generative AI model (e.g., GPT-3) to generate appropriate answers, using prompts based on pre-trained datasets.

[1144] Adaptive response adjustment

[1145] The generated responses are adaptively adjusted in tone and content based on the user's emotions as analyzed by the emotion engine.

[1146] Conversion to audio

[1147] The tailored answers are then converted into audio data using a speech synthesis engine, allowing the answers to be provided to the user in a human-like voice.

[1148] Sending audio data

[1149] Finally, the generated audio data is sent to the user's HMD and played back to the user through the HMD's speakers.

[1150] Device Role

[1151] Voice input

[1152] When the user talks about their concerns through the HMD's microphone, the device records the audio and sends it to the server in real time.

[1153] Audio Output

[1154] The voice data sent from the server is played back to the user through the HMD speaker, allowing the user to receive appropriate support in real time.

[1155] User Roles

[1156] System Usage

[1157] When a user has a problem, they can wear the HMD and talk about it through a microphone. By talking about specific problems, the system can provide appropriate support.

[1158] Additional consultation

[1159] If the user wishes to continue the consultation, they can repeat the same process. There is also a function to transfer the call to a specialist if necessary, ensuring appropriate treatment in emergencies or serious cases.

[1160] Specific examples

[1161] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent from the HMD to the server. The server performs speech recognition, and the generated text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. The AI ​​model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "To overcome work problems, you need to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via the HMD via voice.

[1162] Prompt Sentence Examples

[1163] For users who are in an anxious emotional state, provide appropriate advice on the following: Things have been going badly at work lately, and I don't know what to do.

[1164] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1165] Step 1:

[1166] The user voice-inputs their concerns through the HMD microphone. The voice data is recorded by the HMD and sent to the server in real time. The consultation process begins when the user's voice input is transferred to the server as input data.

[1167] Step 2:

[1168] Once the server receives the audio data, it uses a speech recognition library (e.g., Python's speech_recognition) to convert the audio data into text. This step takes the audio data as input and generates highly accurate text data. Once the audio data is converted to text, it is ready for the next processing step.

[1169] Step 3:

[1170] The converted text data is sent to an emotion analysis engine (e.g., Hugging Face's transformers library). The server uses this engine to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent from the text data. As output, the user's emotions and intent are extracted, providing the basis for the next step.

[1171] Step 4:

[1172] Based on the analyzed sentiment and intent, the server utilizes a generative AI model (e.g., GPT-3) to generate an appropriate answer. A prompt sentence is used to instruct the model to generate an answer text appropriate to the user's situation. Sentiment and intent are used as input, and a generated answer is obtained as output.

[1173] Step 5:

[1174] The generated answer is adaptively adjusted in tone and content based on the user's emotions analyzed by the emotion engine. This step takes the generated answer text as input and optimizes the response with the appropriate tone and nuance depending on the user's emotional state. The final adjusted answer text is output.

[1175] Step 6:

[1176] The adjusted answer text is converted into voice data using a speech synthesis engine. The server uses this engine to generate natural-sounding speech and outputs it as voice data to be conveyed to the user. This conversion process uses the adjusted text data as input and generates human-like voice data.

[1177] Step 7:

[1178] The final generated voice data is sent to the user's HMD and played back to the user through the HMD's speaker. The server sends the voice data to the HMD, and the user receives a response in real time. The voice data is delivered to the HMD as output, thus completing the user's consultation process.

[1179] Specific examples of operation

[1180] For example, if a user says, "Things haven't been going well at work lately, and I don't know what to do anymore," then:

[1181] Step 1: The voice input "Work hasn't been going well lately, and I don't know what to do anymore" is sent from the HMD to the server.

[1182] Step 2: The speech recognition library converts this speech into text data, generating the text "Work hasn't been going well lately and I don't know what to do."

[1183] Step 3: The sentiment analysis engine analyzes the text and recognizes the emotions "anxiety" and "despair."

[1184] Step 4: The generative AI model generates an appropriate response for the user based on this emotion and intent. Example prompt: "For a user who is in an anxious emotional state, please provide appropriate advice on the following: Things haven't been going well at work recently, and I don't know what to do anymore."

[1185] Step 5: The generated answer, "The key to overcoming work-related worries is to understand your own feelings," is adaptively adjusted to optimize the tone to best fit the user's current emotions.

[1186] Step 6: The adjusted answers are converted into voice data by a speech synthesis engine.

[1187] Step 7: The final voice data is sent to the HMD, and the user receives the answer in real time through the speaker.

[1188] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1189] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1190] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1191] [Fourth embodiment]

[1192] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1193] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1194] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1195] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1196] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1197] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1198] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1199] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1200] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1201] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1202] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1203] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1204] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1205] This invention is a system that provides a 24 / 7 consultation service by combining voice recognition technology and an emotion analysis engine. It is specifically designed to allow users at risk of suicide to talk about their worries by voice.

[1206] Server Roles

[1207] Voice input acceptance

[1208] The server receives the voice data sent by the user in real time. When the user calls the helpline, the voice is transmitted directly to the server.

[1209] Analysis of audio data

[1210] When the server receives the voice data, it first converts it into text data using a speech recognition library, using various speech recognition engines (e.g., speech recognition APIs) to achieve high accuracy.

[1211] Emotion and Intention Analysis

[1212] The text data is then fed into a sentiment and intent analysis engine, which uses natural language processing techniques to extract the user's emotions (e.g., anxiety, sadness, urgency, etc.) and intent (e.g., help-seeking, advice-seeking, etc.) from the text data.

[1213] Answer generation

[1214] The server generates answers based on the analyzed emotions and intent. To generate answers, it uses an AI model that has previously learned the knowledge of doctors and counselors. This model creates appropriate responses based on the information and emotions the user is seeking.

[1215] Conversion to audio

[1216] The generated text response is converted into audio data using a speech synthesis engine, which uses speech synthesis technology (e.g., speech synthesis API) to provide a human-like voice response to the user.

[1217] Sending audio data

[1218] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[1219] Device Role

[1220] Voice input

[1221] When a user calls to talk about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[1222] Audio Output

[1223] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[1224] User Roles

[1225] System Usage

[1226] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[1227] Additional consultation

[1228] If the user wants to continue the consultation, they can repeat the same process as many times as they like. If professional help is required, the server will automatically transfer the user to an expert.

[1229] Specific examples

[1230] For example, if a user is suffering from work-related stress, they might say, "Work hasn't been going well lately, and I don't know what to do." This speech is sent to the server, where it is recognized. From the converted text, an emotion analysis engine extracts the emotions of "stress" and "hopelessness." The AI ​​model then generates "advice for overcoming work-related problems," which is then converted into speech by a speech synthesis engine. Specific advice such as "The key to overcoming work-related problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[1231] Role Alignment

[1232] These components work together to provide an environment where users at high risk of suicide can seek advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[1233] In this way, the present invention realizes a consultation system that can respond quickly and effectively at any time, including at night, thereby contributing to reducing the risk of suicide.

[1234] The processing flow will be explained below.

[1235] Step 1:

[1236] The user dials the phone number of the IVR system using a smartphone or landline, which initiates voice input by the user.

[1237] Step 2:

[1238] The device records the user's voice in real time and transmits the voice data to the server using a secure protocol (e.g., HTTPS).

[1239] Step 3:

[1240] The server converts the received voice data into text data using a speech recognition library (e.g., speech recognition API). This conversion process turns the voice signal into a string of characters.

[1241] Step 4:

[1242] The converted text data is sent to an emotion and intent analysis engine, which extracts the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data.

[1243] Step 5:

[1244] Based on the extracted emotions and intent, the server uses an AI model to generate appropriate answers. This model has previously learned the knowledge of doctors and counselors to create specific and appropriate responses.

[1245] Step 6:

[1246] The generated response text data is converted into voice data using a speech synthesis engine (e.g., speech synthesis API), resulting in a human-like voice response.

[1247] Step 7:

[1248] The server transmits the generated voice data to the user's terminal, which receives the voice data and plays it back to the user.

[1249] Step 8:

[1250] The user receives the voice response provided through the terminal and, if necessary, can request further consultation, in which case the same process (steps 1 to 7) can be repeated.

[1251] Step 9:

[1252] If the user's consultation requires professional help, the server transfers the user to a specialist based on the results of emotion and intent analysis. To do this, it searches for the specialist's contact information and relays the connection using a conference call system.

[1253] The processing steps of this system provide an environment where users can receive consultation 24 hours a day, 365 days a year.

[1254] Example 1

[1255] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1256] In modern society, many people suffer from stress in their daily lives and work, feelings of loneliness, and a variety of other worries. One particularly serious problem is consultations that involve the risk of suicide. Current counseling services and consultation centers often have limited response times and are unable to provide prompt and appropriate responses in emergencies. Furthermore, professional assistance is often unavailable at night or on holidays, which can leave users without the support they need and put them in danger. Given this current situation, there is a need for a system that is available 24 hours a day, 365 days a year, can accurately analyze users' emotions and intentions, and provide appropriate responses in real time.

[1257] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1258] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, means for analyzing the user's emotions and intentions from the converted text data, generation means using a generative AI model to generate answers based on the analyzed emotions and intentions, speech synthesis means for converting the generated answers into voice data, and means for transmitting the voice data to a user terminal. This enables real-time analysis of voice input from the user and immediate provision of appropriate answers. Furthermore, by transferring the call to a specialist as needed, more specialized support can be received, realizing consultation services 24 hours a day, 365 days a year, including nights and holidays. This provides a safe and reliable system that can handle serious consultations, including those regarding suicide risk.

[1259] A "server" is a centralized control device for receiving and processing data sent by users.

[1260] "Voice data" refers to data in which voice information uttered by a user is recorded as a digital signal.

[1261] "Speech recognition means" refers to a technology or device that analyzes voice data and converts it into text data.

[1262] "Text data" is character information converted from voice data by a voice recognition means.

[1263] The "analysis means" is a technique or device for identifying a user's emotions and intentions from text data.

[1264] A "generative AI model" is an artificial intelligence model that has learned a huge amount of data in advance and generates appropriate answers based on the user's requests and emotions.

[1265] "Generation means" refers to a technology or device that uses a generative AI model to create an answer based on information obtained from the analysis means.

[1266] "Speech synthesis means" refers to a technique or device that generates voice data based on generated text data.

[1267] A "user terminal" is a device used by a user to communicate with a server. This includes smartphones and landlines.

[1268] "Transfer means" refers to a technology or device that transmits data to an expert based on the user's voice input and the analysis results.

[1269] This invention is a system that provides a 24 / 7 consultation service by combining voice recognition technology and an emotion analysis engine. It is specifically designed to allow users at risk of suicide to talk about their worries by voice.

[1270] Server Roles

[1271] Voice input acceptance

[1272] The server receives the voice data sent by the user in real time. When the user calls the consultation service, the voice is transmitted directly to the server. The voice data is transmitted to the server using a secure communication protocol (e.g., HTTPS).

[1273] Analysis of audio data

[1274] When the server receives the voice data, it first converts it into text using a speech recognition library, using a speech recognition engine such as the Google Cloud Speech-to-Text API to achieve high accuracy.

[1275] Emotion and Intention Analysis

[1276] The text data is then sent to a sentiment and intent analysis engine, which uses natural language processing techniques to extract the user's emotions (e.g., anxiety, sadness, urgency, etc.) and intent (e.g., help-seeking, advice-seeking) from the text data. This process uses tools such as the Microsoft Azure Text Analytics API.

[1277] Answer generation

[1278] Based on the analyzed emotions and intent, the server generates an answer using a generative AI model (e.g., OpenAI GPT-4) that has previously learned the knowledge of doctors and counselors. This model creates an appropriate response based on the information and emotions the user is seeking.

[1279] Conversion to audio

[1280] The generated text response is converted into audio data using a speech synthesis engine, using voice synthesis technologies such as Amazon Polly to provide the user with a human-like voice.

[1281] Sending audio data

[1282] Finally, the generated voice data is sent to the user's terminal, where the user can receive a voice response to their inquiry. This voice data is also sent using a secure communication protocol.

[1283] Device Role

[1284] Voice input

[1285] When a user calls to talk about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[1286] Audio Output

[1287] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[1288] User Roles

[1289] System Usage

[1290] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[1291] Additional consultation

[1292] If the user wants to continue the consultation, they can repeat the same process as many times as they like. If professional help is required, the server will automatically transfer the user to an expert.

[1293] Specific examples

[1294] For example, if a user is suffering from work-related stress, they might say, "Work hasn't been going well lately, and I don't know what to do." This speech is sent to the server, where it is recognized. The converted text is then processed by an emotion analysis engine, which extracts the emotions of "stress" and "hopelessness." The generative AI model then generates "advice for overcoming work-related problems," which is then converted into speech by a speech synthesis engine. Specific advice such as "The key to overcoming work-related problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is then provided to the user via voice.

[1295] Examples of prompt statements

[1296] Below is an example of a prompt that can be fed into a generative AI model:

[1297] "A user says, 'Work hasn't been going well lately, and I don't know what to do.' Please provide appropriate advice to address the stress and despair the user is feeling."

[1298] This system provides an environment where users at high risk of suicide can receive advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[1299] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1300] Step 1:

[1301] User voice input

[1302] When a user has a problem, they call the helpline using their smartphone or landline. They then speak to the helpline, describing their specific problem, such as, "Recently, things haven't been going well at work, and I don't know what to do." This voice data becomes the input.

[1303] Step 2:

[1304] Sending audio data by the device

[1305] The device (smartphone or landline phone) records the user's voice in real time and sends it to the server. This voice data is sent to the server using a secure communication protocol (e.g., HTTPS). The input is the user's voice data, and the output is the voice data sent to the server.

[1306] Step 3:

[1307] Receiving audio data by the server

[1308] The server receives the voice data sent from the device in real time. The received voice data is temporarily stored in the server's memory. The input is the voice data sent from the device, and the output is the voice data stored in the server.

[1309] Step 4:

[1310] Server-based speech recognition and text conversion

[1311] The server converts the received voice data into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text API). The voice data is analyzed and the corresponding text data is generated. The input is the voice data stored on the server, and the output is the converted text data.

[1312] Step 5:

[1313] Server-based emotion and intent analysis

[1314] The server sends the converted text data to an emotion and intent analysis engine (e.g., Microsoft Azure Text Analytics API), which extracts the user's emotion (e.g., stress, despair) and intent (e.g., asking for help) from the text data. The input is the text data, and the output is the emotion and intent analysis results.

[1315] Step 6:

[1316] Server-generated answers

[1317] Based on the results of the emotion and intent analysis, the server sends a prompt to a generative AI model (e.g., OpenAI GPT-4) to generate an appropriate answer. An example of a prompt sentence is, "The user said, 'Recently, things haven't been going well at work, and I don't know what to do anymore.' Please provide appropriate advice to help the user with the stress and despair they are feeling." The input is the results of the emotion and intent analysis, and the output is the generated answer in text format.

[1318] Step 7:

[1319] Server-based speech synthesis

[1320] The generated text response is converted back into audio data by a speech synthesis engine (e.g., Amazon Polly). This process uses techniques to express natural pronunciation and emotion in the voice. The input is the generated text answer, and the output is audio data.

[1321] Step 8:

[1322] Server sends audio data

[1323] The server then transmits the generated voice data back to the device, again using a secure communication protocol. The input is the generated voice data, and the output is the voice data sent to the device.

[1324] Step 9:

[1325] Receiving and playing audio data by a terminal

[1326] The terminal receives the voice data sent from the server and plays it back to the user. The user can receive a specific answer to their question by listening to the played back voice. The input is the voice data sent from the server, and the output is the voice played back to the user.

[1327] (Application example 1)

[1328] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1329] In modern society, the number of people suffering from stress and mental health problems is increasing, making immediate and appropriate support essential, especially for those at high risk of suicide. However, existing systems have difficulty analyzing emotions and generating responses in real time, and are therefore inadequate, especially at night. Furthermore, there is a need for a method that allows users to effectively utilize smart devices and receive a more realistic experience.

[1330] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1331] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, means for analyzing a user's emotions and intentions from the text data, means for generating a response based on the analyzed emotions and intentions, means for converting the generated response into voice data, means for transmitting the voice data to a user terminal, and means for receiving a user's voice input in real time and providing a response to the user through smart glasses or a head-mounted display. This allows for 24-hour, 365-day service, and enables users to receive immediate and effective support by providing emotion analysis and appropriate responses in real time through the smart device.

[1332] "Voice data" refers to data that digitally represents a user's speech.

[1333] "Text data" refers to data expressed as character information converted by a voice recognition means.

[1334] "Speech recognition means" refers to a device or software for converting voice data into text data.

[1335] "Means for analyzing emotions and intentions" refers to a device or software for analyzing a user's emotions and intentions from text data.

[1336] The "means for generating an answer" refers to a device or software that generates an answer to be provided to a user based on the analyzed emotions and intentions.

[1337] "Means for converting into voice data" refers to a device or software for converting the generated response into voice data.

[1338] A "user terminal" refers to a device used by a user, including a smartphone, smart glasses, a head-mounted display, etc.

[1339] The "means for transmitting voice data to a user terminal" refers to a device or software for transmitting the generated voice data to a user terminal.

[1340] "Smart glasses" are glasses-type devices equipped with functions such as augmented reality (AR).

[1341] A "head-mounted display" is a display device that is worn on the user's head and provides images and sounds.

[1342] "Real-time" means that data processing occurs almost immediately, allowing users to receive results without waiting.

[1343] This invention combines voice recognition technology with an emotion analysis engine to provide a 24 / 7 consultation service system for users at risk of suicide. In particular, it utilizes smart glasses and head-mounted displays to provide appropriate support to users in real time. The main components of this system include a voice recognition unit, an emotion and intention analysis unit, a response generation unit, a voice synthesis unit, and a voice data transmission unit.

[1344] Handling voice input and reception

[1345] The server receives the user's voice input in real time. It uses the Python speech_recognition library to convert the user's speech into text data. This voice data is acquired through the microphone in the smart glasses or head-mounted display.

[1346] Emotion and Intention Analysis

[1347] When the server receives the text data, it uses Hugging Face's sentiment-analysis model to analyze the user's emotions (e.g., anxiety, sadness, urgency). It also determines the user's intent based on the analyzed emotions and the text content. The results of this analysis are used to accurately identify situations in which advice or support is needed.

[1348] Generate answers

[1349] The server uses Hugging Face's text-generation (GPT-3) model to generate appropriate answers based on the analyzed emotions and intent, providing high-quality responses in real time that mimic expert advice and support.

[1350] Speech synthesis and output

[1351] The generated text response is converted into audio data using the gTTS (Google Text-to-Speech) library, and this audio data is provided to the user through the speakers of smart glasses or head-mounted displays, allowing the user to listen to the audio response directly using the device in front of them.

[1352] Specific examples

[1353] For example, if a user says, "I've been having trouble sleeping lately, please help me," this speech is sent to the server and converted into text using speech recognition. Then, emotion and intent analysis extracts the emotions "anxiety" and "seeking help." The AI ​​model then generates an answer using the following prompt:

[1354] Example prompt sentence:

[1355] User: I've been having trouble sleeping lately, please help.

[1356] System: Please provide advice for insomnia.

[1357] Ultimately, the voice will provide advice such as, "The best way to combat your insomnia is to relax before bed. Try meditating or doing some gentle stretching." This allows users to receive the support they need in real time.

[1358] This invention allows users to receive support 24 hours a day, 365 days a year with peace of mind, and provides an environment where immediate response is possible, especially for people at high risk of suicide.

[1359] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1360] Step 1:

[1361] The user inputs voice through smart glasses or a head-mounted display. The device records this voice in real time and transmits it to the server as voice data. The transmitted voice data is the input.

[1362] Step 2:

[1363] The server receives the transmitted voice data and converts it to text using the Python speech_recognition library. This converted text data is the output. This process is performed by the speech recognition means.

[1364] Step 3:

[1365] The server inputs the text data obtained from the speech recognition means into Hugging Face's sentiment-analysis model, which analyzes the user's emotions from the text data. This analysis generates emotional information such as anxiety, sadness, and urgency as output. The emotion and intent analysis means performs this process.

[1366] Step 4:

[1367] The server generates an appropriate answer for the user based on the results of the emotion analysis (emotion information) by inputting it into Hugging Face's text-generation (GPT-3) model. This generated text response is the output. This process is carried out by the generation means for generating the answer.

[1368] Step 5:

[1369] The server inputs the generated text response into the gTTS (Google Text-to-Speech) library and converts it to speech. This converted speech response is the output. The speech conversion method performs this process.

[1370] Step 6:

[1371] The server converts the response into voice data and sends it to the terminal. The terminal plays back the received voice data and provides it to the user. The voice response is played back through the user's smart glasses or head-mounted display. The provision of this voice response is the final output. This process is performed by a means for transmitting voice data to the user terminal.

[1372] Step 7:

[1373] The user listens to the voice response provided by the device and can re-voice the patient if necessary, and if necessary, repeat the consultation or request a transfer to a specialist. This interactive process continues.

[1374] Through this specific processing step, the system analyzes the user's emotions and intentions in real time and provides appropriate support.

[1375] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1376] This invention is a system that combines voice recognition technology and an emotion analysis engine to provide a 24 / 7 consultation service. It is designed to enable users at risk of suicide to discuss their concerns via voice. The invention adds an emotion engine with the ability to recognize the user's emotions and adaptively adjust responses.

[1377] Server Roles

[1378] Voice input acceptance

[1379] The server receives the voice data sent by the user in real time. When the user calls the helpline, the voice is transmitted directly to the server.

[1380] Analysis of audio data

[1381] When the server receives the voice data, it first converts it into text using a voice recognition library, and uses a variety of voice recognition engines to achieve high accuracy.

[1382] Emotion and Intention Analysis

[1383] The converted text data is sent to a sentiment analysis engine, which the server uses to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data. Furthermore, the sentiment analysis engine has the ability to adaptively adjust responses.

[1384] Answer generation

[1385] The server generates answers based on the analyzed emotions and intent. To generate answers, it uses an AI model that has previously learned the knowledge of doctors and counselors. This model creates appropriate responses based on the information and emotions the user is seeking.

[1386] Adaptive response adjustment

[1387] The emotion engine adaptively adjusts the generated response based on the user's perceived emotion, a process that provides a response that best suits the user's emotional state.

[1388] Conversion to audio

[1389] The generated response text is converted into voice data using a speech synthesis engine, which uses speech synthesis technology to provide answers to users in a human-like voice.

[1390] Sending audio data

[1391] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[1392] Device Role

[1393] Voice input

[1394] When a user calls and talks about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[1395] Audio Output

[1396] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[1397] User Roles

[1398] System Usage

[1399] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[1400] Additional consultation

[1401] If the user wants to continue the consultation, they can repeat the same process (steps 1 to 7). Also, if professional help is required, the server can automatically transfer the user to a specialist.

[1402] Specific examples

[1403] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent to a server for speech recognition. The converted text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. The AI ​​model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "The key to overcoming work problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[1404] Role Alignment

[1405] These components work together to provide an environment where users at high risk of suicide can seek advice safely 24 hours a day, 365 days a year. It also includes a function to transfer calls to specialists as needed, allowing for appropriate response in emergencies and serious cases.

[1406] In this way, the present invention realizes a consultation system that can respond quickly and effectively at any time, including at night, thereby contributing to reducing the risk of suicide.

[1407] The processing flow will be explained below.

[1408] Step 1:

[1409] A user dials the phone number of the IVR system using a smartphone or landline, which initiates voice input.

[1410] Step 2:

[1411] The device records the user's voice in real time and transmits the voice data to the server. The data is transferred using a secure communication protocol (e.g., HTTPS).

[1412] Step 3:

[1413] The server converts the received voice data into text data using a voice recognition library (e.g., a voice recognition API). In this step, the voice signal is represented as a string.

[1414] Step 4:

[1415] The converted text data is sent to a sentiment analysis engine, which the server uses to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for information, etc.) from the text data.

[1416] Step 5:

[1417] The sentiment analysis engine identifies the user's current emotional state based on the perceived emotions and intent, which is important for subsequent response generation.

[1418] Step 6:

[1419] The server uses an AI model to generate responses based on the extracted emotions and intent, which is based on knowledge previously learned from experts (e.g., doctors, counselors, etc.).

[1420] Step 7:

[1421] The generated response text data is adaptively adjusted through an emotion engine, which selects the most appropriate expression and tone for the user's current emotional state.

[1422] Step 8:

[1423] The adjusted response text data is converted into audio data using a speech synthesis engine (e.g., speech synthesis API). This process produces a natural, human-like voice.

[1424] Step 9:

[1425] The server transmits the generated voice data to the user's terminal, which receives the voice data and plays it back to the user.

[1426] Step 10:

[1427] The user receives the voice response provided through the terminal and, if necessary, can request further consultation, in which case the same process (steps 1 to 9) can be repeated.

[1428] Step 11:

[1429] If the user's consultation needs professional help, the server will transfer the user to a specialist based on the results of emotion and intent analysis. To do this, it will search for the specialist's contact information and establish a relay connection using a telephone conference system.

[1430] In this way, the system of the present invention recognizes the user's emotional state and adaptively adjusts responses based on that state, thereby providing effective mental health counseling. The processes in steps 1 through 11 work together to provide an environment where users at high risk of suicide can seek counseling safely 24 hours a day, 365 days a year.

[1431] Example 2

[1432] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1433] In modern society, the number of people suffering from mental health problems and stress is increasing, and there is an urgent need to provide prompt and appropriate support, especially for those at high risk of suicide. However, conventional consultation response systems have difficulty providing support 24 hours a day, 365 days a year, and it is difficult to accurately analyze emotions and intentions and generate appropriate responses. For this reason, there is a growing need for a system that can flexibly respond to changes in emotions and provide effective and prompt support to users at high risk of suicide.

[1434] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1435] In this invention, the server includes means for receiving voice data, speech recognition means for converting the voice data into text data, emotion analysis means for analyzing a user's emotion and intention from the text data, generation means for generating an answer based on the analyzed emotion and intention, means for adaptively adjusting the generated answer based on the user's emotional state, means for converting the adjusted answer into voice data, and means for transmitting the voice data to a user terminal, thereby making it possible to provide an appropriate answer according to the user's emotional state in real time, 24 hours a day, 365 days a year.

[1436] The "means for receiving voice data" refers to a device or software that has the function of acquiring the voice uttered by the user as digital data, converting it into a format that can be processed within the system, and receiving it.

[1437] The "voice recognition means for converting voice data into text data" refers to a device or software that has the function of analyzing received voice data and converting it into corresponding text data.

[1438] "Emotion analysis means for analyzing user emotions and intentions from text data" refers to a device or software that has the function of analyzing a user's emotional state (e.g., anxiety, sadness, despair, etc.) and intentions (e.g., asking for help, asking for advice, etc.) based on converted text data.

[1439] The "generation means for generating a response based on the analyzed emotion and intention" is a device or software that has the function of generating an appropriate response based on data from the emotion analysis means.

[1440] "Means for adaptively adjusting the generated response based on the user's emotional state" refers to a device or software that has the function of fine-tuning the generated response in accordance with the user's emotional state and providing it in an optimal form.

[1441] The "means for converting the adjusted response into voice data" refers to a device or software that has the function of converting adaptively adjusted text data into voice data and providing the voice data to the user as a natural voice.

[1442] The "means for transmitting audio data to a user terminal" refers to a device or software that has the function of transmitting the generated and converted audio data to a user's device and playing it on that device.

[1443] This invention combines voice recognition technology with an emotion analysis engine to provide a consultation service system that is available 24 hours a day, 365 days a year. It is designed specifically to enable users at risk of suicide to discuss their concerns via voice. The system is equipped with an emotion engine that recognizes the user's emotions and adaptively adjusts responses.

[1444] Server Roles

[1445] Voice input acceptance

[1446] The server receives voice data sent from the user in real time. For example, when a user calls a helpline, the voice is transmitted directly to the server.

[1447] Analysis of audio data

[1448] When the server receives the voice data, it first converts it into text data using a speech recognition library (e.g., Google Cloud Speech-to-Text, Amazon Transcribe), thereby obtaining text information about what the user said.

[1449] Emotion and Intention Analysis

[1450] The converted text data is sent to a sentiment analysis engine (e.g., IBM Watson Tone Analyzer, Microsoft Azure Text Analytics), which uses the engine to extract the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., asking for help, asking for advice, etc.) from the text data.

[1451] Answer generation

[1452] Based on the analyzed emotions and intent, the server generates answers using a generative AI model (e.g., OpenAI GPT-3), which has previously learned the knowledge of doctors and counselors to create appropriate responses to the information and emotions the user is seeking.

[1453] Adaptive response adjustment

[1454] The emotion engine adaptively adjusts the generated response based on the user's perceived emotion, a process that provides a response that best suits the user's emotional state.

[1455] Conversion to audio

[1456] The generated response text is converted into audio data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech, Amazon Polly), which uses speech synthesis technology to provide answers to users in a human-like voice.

[1457] Sending audio data

[1458] Finally, the generated voice data is transmitted to the user's terminal, and the user can receive a voice response to his or her inquiry.

[1459] Device Role

[1460] Voice input

[1461] When a user calls and talks about their concerns, the device records the voice and sends it to the server in real time. This function is available to users using devices such as smartphones and landlines.

[1462] Audio Output

[1463] The device that receives the voice data generated by the server plays the voice and provides a response to the user, allowing the user to receive appropriate support in real time.

[1464] User Roles

[1465] System Usage

[1466] When users have a problem, they can access the system by calling in. By following the voice guidance and describing the specific problem, they can receive appropriate support.

[1467] Additional consultation

[1468] If the user wants to continue with the consultation, they can repeat the same process, and if professional help is needed, the server can automatically transfer them to a specialist.

[1469] Specific examples

[1470] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent to a server for speech recognition. The converted text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. A generative AI model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "The key to overcoming work problems is to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via voice.

[1471] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1472] Step 1: Accept voice input

[1473] Input: User's voice

[1474] Process: The user calls the system and tells it about their concerns. The user's voice is recorded by the device (smartphone, landline, etc.).

[1475] Output: Recorded audio data (digital format)

[1476] Specific operation: When a user speaks about a problem such as "Work hasn't been going well lately," the device captures the voice in real time and sends it to the server as audio data.

[1477] Step 2: Analyzing the audio data

[1478] Input: Recorded audio data

[1479] Processing: The server sends the audio data to a speech recognition library, which converts it into text. Examples of speech recognition libraries used here include Google Cloud Speech-to-Text and Amazon Transcribe.

[1480] Output: Converted text data

[1481] Specific operation: The speech recognition library analyzes the voice data "Work is not going well" received by the server, and generates the corresponding text "Work is not going well."

[1482] Step 3: Emotion and Intent Analysis

[1483] Input: Converted text data

[1484] Processing: The server sends the text data to a sentiment analysis engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent (e.g., help-seeking, advice-seeking, etc.).

[1485] Output: Sentiment and intent analysis results

[1486] Specific operation: The server inputs the text data "My work is not going well" into the emotion analysis engine, and the engine determines that the user is feeling "stress" or "despair."

[1487] Step 4: Answer Generation

[1488] Input: Sentiment and intent analysis results

[1489] Processing: The server uses a generative AI model (e.g., OpenAI GPT-3) to generate an appropriate answer based on the analysis results. This model has been trained with the knowledge of doctors and counselors.

[1490] Output: Generated answer text

[1491] Specific behavior: If emotion analysis identifies a high level of "despair," the generative AI model will create specific advice such as "First, stay calm. It's important to understand your own feelings."

[1492] Step 5: Adaptively adjust the response

[1493] Input: Generated answer text

[1494] Processing: The server re-analyzes the generated answer text with the emotion engine and fine-tunes the response to best suit the user's emotional state.

[1495] Output: Adaptively adjusted answer text

[1496] What it does: The emotion engine translates the response created by the generative AI model, "It's important to understand your own feelings," into a "tone that conveys a calm and welcoming atmosphere."

[1497] Step 6: Convert to audio

[1498] Input: Adaptively adjusted answer text

[1499] Processing: The server converts the tailored response text into audio data using a speech synthesis engine (e.g., Google Cloud Text-to-Speech, Amazon Polly).

[1500] Output: Generated audio data

[1501] What happens: The adjusted text response is sent to a speech synthesis engine, which converts it into speech data, generating a human-like voice saying, "It's important to understand your own feelings."

[1502] Step 7: Sending audio data

[1503] Input: Generated audio data

[1504] Processing: The server sends the generated audio data to the user's device.

[1505] Output: The audio response the user receives

[1506] Specific operation: The server sends the voice data to the user's device, which then plays it back. The user receives the advice, "It's important to understand your own feelings."

[1507] (Application example 2)

[1508] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1509] In recent years, the number of individuals at risk of suicide due to psychological stress and social anxiety has been increasing. It is often difficult to find an appropriate counselor or support, especially during late night or early morning hours. Furthermore, existing counseling systems lack the ability to properly analyze emotions and provide optimal responses to users. Furthermore, if the counseling process is not carried out quickly and effectively, there is a risk that the user's psychological burden will increase. Therefore, a system that operates 24 hours a day, 365 days a year and provides appropriate advice based on emotions is needed.

[1510] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving voice data, voice recognition means for converting the voice data into text data, emotion analysis means for analyzing the user's emotions and intentions from the text data, generation means for generating an answer based on the analyzed emotions and intentions, adjustment means for adaptively adjusting the generated answer, voice synthesis means for converting the adjusted answer into voice data, means for transmitting the voice data to the user terminal, and means for receiving the user's voice input and providing voice and visual information through a head-mounted display. This makes it possible to provide appropriate advice based on the user's emotions in the form of voice and visual information 24 hours a day, 365 days a year.

[1511] "Voice data" refers to voice information input by a user.

[1512] "Means" refers to a combination of devices and software for realizing a specific function.

[1513] "Speech recognition means" refers to a device or algorithm for converting input voice data into text data.

[1514] "Text data" refers to character string information converted by speech recognition means.

[1515] "Emotion analysis means" refers to devices and algorithms for analyzing user emotions and intentions from text data.

[1516] "Generator" refers to the system or algorithm that generates answers based on analyzed sentiment and intent.

[1517] "Adjustment means" refers to a system or algorithm that adaptively adjusts the generated answers based on the user's emotions.

[1518] "Speech synthesis means" refers to a device or algorithm for converting conditioned responses into speech data.

[1519] A "user terminal" is a device used by a user, and refers to an apparatus for inputting and outputting voice data.

[1520] A "head-mounted display" is a display device that provides information to the wearer's field of vision and has the function of simultaneously providing audio and visual information.

[1521] "Adaptive" refers to the ability to change adaptively according to the situation or conditions.

[1522] "Visual information" refers to visual data, icons, and messages presented to the user.

[1523] This invention is a system that uses a head-mounted display (HMD) to provide a voice-based consultation service available 24 hours a day, 365 days a year. This system combines voice recognition technology and an emotion analysis engine to provide appropriate advice to users and aim to reduce the risk of suicide.

[1524] Server Roles

[1525] The server processes the voice data and generates a response using the following means:

[1526] Voice input acceptance

[1527] The server receives the voice data sent by the user in real time. When the user speaks through the microphone built into the HMD, the voice data is sent to the server.

[1528] Analysis of audio data

[1529] When the server receives the voice data, it converts the voice data into text data using a voice recognition library (e.g., Python's speech_recognition).

[1530] Emotion and Intention Analysis

[1531] The converted text data is sent to an emotion analysis engine (e.g., Hugging Face's transformers library), which the server uses to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent from the text data.

[1532] Answer generation

[1533] Based on the analyzed sentiment and intent, the server utilizes a generative AI model (e.g., GPT-3) to generate appropriate answers, using prompts based on pre-trained datasets.

[1534] Adaptive response adjustment

[1535] The generated responses are adaptively adjusted in tone and content based on the user's emotions as analyzed by the emotion engine.

[1536] Conversion to audio

[1537] The tailored answers are then converted into audio data using a speech synthesis engine, allowing the answers to be provided to the user in a human-like voice.

[1538] Sending audio data

[1539] Finally, the generated audio data is sent to the user's HMD and played back to the user through the HMD's speakers.

[1540] Device Role

[1541] Voice input

[1542] When the user talks about their concerns through the HMD's microphone, the device records the audio and sends it to the server in real time.

[1543] Audio Output

[1544] The voice data sent from the server is played back to the user through the HMD speaker, allowing the user to receive appropriate support in real time.

[1545] User Roles

[1546] System Usage

[1547] When a user has a problem, they can wear the HMD and talk about it through a microphone. By talking about specific problems, the system can provide appropriate support.

[1548] Additional consultation

[1549] If the user wishes to continue the consultation, they can repeat the same process. There is also a function to transfer the call to a specialist if necessary, ensuring appropriate treatment in emergencies or serious cases.

[1550] Specific examples

[1551] For example, if a user says, "Work hasn't been going well lately, and I don't know what to do," this speech is sent from the HMD to the server. The server performs speech recognition, and the generated text data is sent to an emotion analysis engine, which recognizes feelings of stress and despair. The AI ​​model then generates "advice for overcoming work problems," and the emotion engine fine-tunes the tone and content of the response based on the user's current emotions. Specific advice such as "To overcome work problems, you need to fully understand your own feelings. It's also effective to find someone you can trust to share your worries with" is provided to the user via the HMD via voice.

[1552] Prompt Sentence Examples

[1553] For users who are in an anxious emotional state, provide appropriate advice on the following: Things have been going badly at work lately, and I don't know what to do.

[1554] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1555] Step 1:

[1556] The user voice-inputs their concerns through the HMD microphone. The voice data is recorded by the HMD and sent to the server in real time. The consultation process begins when the user's voice input is transferred to the server as input data.

[1557] Step 2:

[1558] Once the server receives the audio data, it uses a speech recognition library (e.g., Python's speech_recognition) to convert the audio data into text. This step takes the audio data as input and generates highly accurate text data. Once the audio data is converted to text, it is ready for the next processing step.

[1559] Step 3:

[1560] The converted text data is sent to an emotion analysis engine (e.g., Hugging Face's transformers library). The server uses this engine to analyze the user's emotions (e.g., anxiety, sadness, despair, etc.) and intent from the text data. As output, the user's emotions and intent are extracted, providing the basis for the next step.

[1561] Step 4:

[1562] Based on the analyzed sentiment and intent, the server utilizes a generative AI model (e.g., GPT-3) to generate an appropriate answer. A prompt sentence is used to instruct the model to generate an answer text appropriate to the user's situation. Sentiment and intent are used as input, and a generated answer is obtained as output.

[1563] Step 5:

[1564] The generated answer is adaptively adjusted in tone and content based on the user's emotions analyzed by the emotion engine. This step takes the generated answer text as input and optimizes the response with the appropriate tone and nuance depending on the user's emotional state. The final adjusted answer text is output.

[1565] Step 6:

[1566] The adjusted answer text is converted into voice data using a speech synthesis engine. The server uses this engine to generate natural-sounding speech and outputs it as voice data to be conveyed to the user. This conversion process uses the adjusted text data as input and generates human-like voice data.

[1567] Step 7:

[1568] The final generated voice data is sent to the user's HMD and played back to the user through the HMD's speaker. The server sends the voice data to the HMD, and the user receives a response in real time. The voice data is delivered to the HMD as output, thus completing the user's consultation process.

[1569] Specific examples of operation

[1570] For example, if a user says, "Things haven't been going well at work lately, and I don't know what to do anymore," then:

[1571] Step 1: The voice input "Work hasn't been going well lately, and I don't know what to do anymore" is sent from the HMD to the server.

[1572] Step 2: The speech recognition library converts this speech into text data, generating the text "Work hasn't been going well lately and I don't know what to do."

[1573] Step 3: The sentiment analysis engine analyzes the text and recognizes the emotions "anxiety" and "despair."

[1574] Step 4: The generative AI model generates an appropriate response for the user based on this emotion and intent. Example prompt: "For a user who is in an anxious emotional state, please provide appropriate advice on the following: Things haven't been going well at work recently, and I don't know what to do anymore."

[1575] Step 5: The generated answer, "The key to overcoming work-related worries is to understand your own feelings," is adaptively adjusted to optimize the tone to best fit the user's current emotions.

[1576] Step 6: The adjusted answers are converted into voice data by a speech synthesis engine.

[1577] Step 7: The final voice data is sent to the HMD, and the user receives the answer in real time through the speaker.

[1578] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1579] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1580] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1581] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1582] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1583] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1584] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1585] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1586] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1587] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1588] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1589] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1590] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1591] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1592] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1593] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1594] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1595] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1596] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1597] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1598] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1599] The following is further disclosed regarding the above embodiment.

[1600] (Claim 1)

[1601] means for receiving audio data;

[1602] a speech recognition means for converting speech data into text data;

[1603] means for analyzing user emotions and intentions from text data;

[1604] generating means for generating an answer based on the analyzed sentiment and intent;

[1605] means for converting the generated answers into audio data;

[1606] means for transmitting voice data to a user terminal;

[1607] A system including:

[1608] (Claim 2)

[1609] 10. The system of claim 1, further comprising means for receiving a user's voice input and routing to an expert based on the analyzed sentiment and intent.

[1610] (Claim 3)

[1611] 2. The system according to claim 1, further comprising means for realizing 24-hour, 365-day operation in order to enhance consultation response during the night.

[1612] "Example 1"

[1613] (Claim 1)

[1614] means for receiving audio data;

[1615] a speech recognition means for converting speech data into text data;

[1616] means for analyzing the user's emotions and intentions from the converted text data;

[1617] a generating means using a generative AI model to generate answers based on the analyzed emotions and intentions;

[1618] a speech synthesis means for converting the generated answer into speech data;

[1619] means for transmitting voice data to a user terminal;

[1620] A system including:

[1621] (Claim 2)

[1622] 10. The system of claim 1, further comprising a forwarding means for receiving a user's voice input and forwarding to an expert based on the analyzed sentiment and intent.

[1623] (Claim 3)

[1624] 2. The system according to claim 1, further comprising means for realizing 24-hour, 365-day operation in order to enhance consultation response during the night.

[1625] "Application Example 1"

[1626] (Claim 1)

[1627] means for receiving audio data;

[1628] a speech recognition means for converting speech data into text data;

[1629] means for analyzing user emotions and intentions from text data;

[1630] generating means for generating an answer based on the analyzed sentiment and intent;

[1631] means for converting the generated answers into audio data;

[1632] means for transmitting voice data to a user terminal;

[1633] means for receiving a user's voice input in real time and providing a response to the user through the smart glasses or head mounted display;

[1634] A system including:

[1635] (Claim 2)

[1636] 10. The system of claim 1, further comprising means for receiving a user's voice input and routing to an expert based on the analyzed sentiment and intent.

[1637] (Claim 3)

[1638] 2. The system according to claim 1, further comprising means for realizing 24-hour, 365-day operation in order to enhance consultation response during the night.

[1639] "Example 2: Combining Emotion Engines"

[1640] (Claim 1)

[1641] means for receiving audio data;

[1642] a speech recognition means for converting speech data into text data;

[1643] emotion analysis means for analyzing user emotions and intentions from text data;

[1644] generating means for generating an answer based on the analyzed sentiment and intent;

[1645] means for adaptively adjusting the generated answers based on the user's emotional state;

[1646] means for converting the adjusted response into audio data;

[1647] means for transmitting voice data to a user terminal;

[1648] A system including:

[1649] (Claim 2)

[1650] 10. The system of claim 1, further comprising means for receiving a user's voice input and routing to an expert based on the analyzed sentiment and intent.

[1651] (Claim 3)

[1652] 2. The system according to claim 1, further comprising means for realizing 24-hour, 365-day operation in order to enhance consultation response during the night.

[1653] "Application example 2 when combining emotion engines"

[1654] (Claim 1)

[1655] means for receiving audio data;

[1656] a speech recognition means for converting speech data into text data;

[1657] emotion analysis means for analyzing user emotions and intentions from text data;

[1658] generating means for generating an answer based on the analyzed sentiment and intent;

[1659] an adjustment means for adaptively adjusting the generated answers;

[1660] a speech synthesis means for converting the adjusted response into speech data;

[1661] means for transmitting voice data to a user terminal;

[1662] means for receiving user voice input and providing audio and visual information through the head mounted display;

[1663] A system including:

[1664] (Claim 2)

[1665] 10. The system of claim 1, further comprising means for receiving a user's voice input and routing to an expert based on the analyzed sentiment and intent.

[1666] (Claim 3)

[1667] 2. The system according to claim 1, further comprising means for realizing 24-hour, 365-day operation in order to enhance consultation response during the night. [Explanation of symbols]

[1668] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving audio data; a speech recognition means for converting speech data into text data; means for analyzing user emotions and intentions from text data; generating means for generating an answer based on the analyzed sentiment and intent; means for converting the generated answers into audio data; means for transmitting voice data to a user terminal; A system including:

2. 10. The system of claim 1, further comprising means for receiving a user's voice input and routing to an expert based on the analyzed sentiment and intent.

3. 2. The system according to claim 1, further comprising means for realizing 24-hour, 365-day operation in order to enhance consultation response during the night.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A