system

The system uses generative AI to analyze vital signs and generate speech for effective communication, addressing the challenge of expressing feelings and thoughts in individuals with speech impairments or language barriers.

JP7858008B2Active Publication Date: 2026-05-13SOFTBANK GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-09-19
Publication Date
2026-05-13

AI Technical Summary

Technical Problem

Existing systems fail to adequately understand and express feelings and thoughts based on vital information, limiting effective communication, especially for individuals with speech impairments or language barriers.

Method used

A system comprising a collection unit, analysis unit, and speech unit that utilizes generative AI to analyze vital information such as heart rate and electroencephalogram data to generate and speak appropriate words and sentences, facilitating communication through wearable devices and speech-generating AI.

Benefits of technology

Enables accurate expression of emotions and thoughts, overcoming speech impairments and language barriers, allowing individuals to convey their intentions and emotions effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007858008000001
    Figure 0007858008000001
  • Figure 0007858008000002
    Figure 0007858008000002
  • Figure 0007858008000003
    Figure 0007858008000003
Patent Text Reader

Abstract

To provide a system according to an embodiment which understands feelings or emotions on the basis of vital information and generates proper sentences or languages to produce a speech.SOLUTION: The system according to an embodiment includes a collection unit, an analysis unit, a generation unit, and a speech unit. The collection unit collects vital information. The analysis unit analyzes information collected by the collection unit and understands emotions or feelings. The generation unit generates languages or sentences on the basis of the result of analysis obtained by the analysis unit. The speech unit produces speeches from the languages or sentences generated by the generation unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, there is a problem that understanding feelings and thoughts based on vital information and generating and uttering appropriate words and sentences have not been sufficiently performed.

[0005] The system according to the embodiment aims to understand feelings and thoughts based on vital information and generate and utter appropriate words and sentences.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a collection unit, an analysis unit, a generation unit, and a speech unit. The collection unit collects vital information. The analysis unit analyzes the information collected by the collection unit and understands emotions or thoughts. The generation unit generates words or sentences based on the analysis results obtained by the analysis unit. The speech unit speaks the words or sentences generated by the generation unit. [Effects of the Invention]

[0007] The system according to this embodiment can understand emotions and thoughts based on vital information, and generate and speak appropriate words and sentences. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F manages communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The communication system according to an embodiment of the present invention is a system that collects vital information, understands a person's emotions and thoughts, generates words and sentences using a generative AI, and has the person speak using a speech-generating AI. This system analyzes the collected vital information and generates words and sentences using a generative AI. Furthermore, by having the person speak using a speech-generating AI, communication can be facilitated. The intended users are ALS patients, people who are intubated and unable to speak, and it is expected that communication will expand between people with different native languages. For example, an ALS patient wears an electroencephalogram (EEG) sensor to have their emotions and thoughts analyzed. The generative AI generates words such as "I want to drink water," and the speech-generating AI speaks those words. This allows the patient to convey their intentions to others. Also, when people with different native languages ​​communicate, the generative AI generates appropriate words and sentences, and the speech-generating AI speaks them, enabling communication that transcends language barriers. In this way, the communication system can convey the user's emotions and thoughts to others.

[0029] The communication system according to this embodiment comprises a collection unit, an analysis unit, a generation unit, and a speech unit. The collection unit collects vital information. Vital information includes, but is not limited to, heart rate, electroencephalogram (EEG), and body temperature. The collection unit collects vital information using, for example, an EEG sensor and a heart rate sensor. Examples of EEG sensors include EEG sensors and fNIRS sensors. Examples of heart rate sensors include photoelectric heart rate sensors and electrical heart rate sensors. The analysis unit uses a generation AI to analyze the vital information collected by the collection unit and understand emotions and thoughts. Examples of generation AI include deep learning models and natural language processing models. The generation unit uses a generation AI to generate words and sentences based on the analysis results obtained by the analysis unit. Examples of generation AI include text generation AI (e.g., LLM) and multimodal generation AI. The speech unit uses a speech generation AI to speak the words and sentences generated by the generation unit. The speech generation AI utilizes, for example, speech synthesis models and text-to-speech conversion technologies. This allows the communication system according to the embodiment to convey the user's emotions and thoughts to others.

[0030] The data collection unit collects vital information. This vital information includes, but is not limited to, heart rate, electroencephalogram (EEG), and body temperature. The data collection unit collects vital information using, for example, an EEG sensor or a heart rate sensor. Examples of EEG sensors include EEG sensors and fNIRS sensors. EEG sensors measure the electrical activity of the brain by being attached to the scalp and acquire EEG data in real time. fNIRS sensors measure changes in blood flow to the brain using near-infrared light to understand the state of brain activity. Examples of heart rate sensors include photoelectric heart rate sensors and electrocardiographic heart rate sensors. Photoelectric heart rate sensors detect heart rate by irradiating light onto the skin and measuring the reflected light. Electrocardiographic heart rate sensors detect heart rate by attaching electrodes to the skin and measuring the electrical activity of the heart. These sensors can be incorporated into wearable devices and medical equipment to continuously monitor the user's vital information. The collected vital information is transmitted to a central database using wireless communication technology and provided to the analysis unit in real time. This allows the data collection unit to accurately understand the user's health and emotional state and provide the data necessary for the next analysis step. Furthermore, the data collection unit can integrate data from multiple sensors and perform noise reduction and outlier filtering to improve data accuracy and reliability. As a result, the data collection unit can provide more accurate vital information and improve the overall system performance.

[0031] The analysis unit uses generative AI to analyze vital information collected by the data collection unit and understand emotions and thoughts. Generative AI may include deep learning models or natural language processing models. Specifically, deep learning models classify the user's emotional state using collected electroencephalogram (EEG) and heart rate data as input. For example, EEG data is analyzed to identify alpha and beta wave patterns, determining whether the user is relaxed or focused. Heart rate data is analyzed to assess stress levels and excitement levels. Natural language processing models express the user's emotions and thoughts in text format based on these analysis results. For example, it may output a specific emotional state such as "The user is currently relaxed." Furthermore, the analysis unit can also analyze changes and trends in emotions by utilizing past data and user history information. This allows for a grasp of the user's long-term emotional patterns and provides more accurate analysis results. The analysis unit can also use anomaly detection algorithms to detect unusual emotional states or abnormal vital data, issuing early warnings. This allows the analysis unit to handle not only real-time sentiment analysis but also long-term sentiment management and anomaly detection, thereby improving the reliability and security of the entire system.

[0032] The generation unit uses a generation AI to generate words and sentences based on the analysis results obtained by the analysis unit. Examples of generation AIs include text generation AI (e.g., LLM) and multimodal generation AI. Text generation AI takes emotional states and thoughts provided by the analysis unit as input and generates natural-sounding words and sentences. For example, it generates specific sentences such as, "The user is currently relaxed and feeling calm." Multimodal generation AI can integrate not only text but also other modalities such as images and audio to generate richer expressions. For example, it can generate images and audio messages corresponding to the user's emotional state, conveying emotions through sight and sound. The generation unit outputs the generated words and sentences in an appropriate format and provides them to the speech unit. Furthermore, the generation unit can collect user feedback and continuously improve the generation AI model. For example, the user can evaluate the generated sentences as "accurate" or "inaccurate" to improve the accuracy of the generation AI. The generation unit can also generate customized sentences according to specific situations and contexts. This allows the generation unit to accurately and naturally express the user's emotions and thoughts, and provide information for conveying them to others.

[0033] The speech unit uses speech generation AI to speak the words and sentences generated by the generation unit. Speech generation AI can utilize technologies such as speech synthesis models or text-to-speech conversion. Speech synthesis models use text provided by the generation unit as input to generate natural-sounding speech. For example, they can generate speech with tone and intonation that matches the user's emotional state, conveying emotions more effectively. Text-to-speech conversion technology adjusts pronunciation and accent when converting text to speech, producing easy-to-understand speech. The speech unit plays the generated speech through speakers or headsets, conveying the user's emotions and thoughts to others. Furthermore, the speech unit can collect user feedback and continuously improve the speech generation AI model. For example, users can evaluate the generated speech as "easy to understand" or "difficult to understand," improving the accuracy of the speech generation AI. The speech unit can also generate customized speech tailored to specific situations and contexts. This allows the speech unit to accurately and naturally express the user's emotions and thoughts through speech, providing information for communication with others.

[0034] The data acquisition unit may include an electroencephalogram (EEG) sensor or a heart rate sensor. Examples of EEG sensors include EEG sensors and fNIRS sensors. An EEG sensor is a sensor that electrically measures brain waves, and an fNIRS sensor is a sensor that measures blood flow in the brain using near-infrared light. Examples of heart rate sensors include photoelectric heart rate sensors and electrocardiographic heart rate sensors. A photoelectric heart rate sensor is a sensor that measures blood flow using light, and an electrocardiographic heart rate sensor is a sensor that electrically measures heart rate. As a result, the data acquisition unit can collect more accurate vital information by using EEG sensors and heart rate sensors. Some or all of the above processing in the data acquisition unit may be performed using AI, for example, or without AI.

[0035] The analysis unit can analyze emotions and thoughts using generative AI. Generative AI includes, for example, deep learning models and natural language processing models. Deep learning models are models that learn from large amounts of data and have advanced analytical capabilities. Natural language processing models are models that analyze text data and understand its meaning. As a result, the analysis unit can improve the accuracy of its analysis of emotions and thoughts by using generative AI. Some or all of the above-mentioned processes in the analysis unit are performed using generative AI.

[0036] The generation unit can generate words and sentences using generative AI. Generative AI includes, for example, text generation AI (e.g., LLM) and multimodal generation AI. Text generation AI is a model that learns from large amounts of text data and has advanced natural language processing capabilities. Multimodal generation AI is a model that can handle multiple modals, such as images and audio, in addition to text. As a result, the generation unit can improve the accuracy of word and sentence generation by using generative AI. Some or all of the above-mentioned processes in the generation unit are performed using generative AI.

[0037] The speech generation unit can produce words and sentences using speech generation AI. Speech generation AI includes, for example, speech synthesis models and text-to-speech conversion technologies. A speech synthesis model is a model that converts text data into speech, and text-to-speech conversion technology is a technology that converts text data into speech. As a result, the speech generation unit improves the accuracy of its speech by using speech generation AI. Some or all of the above-mentioned processes in the speech generation unit are performed using speech generation AI.

[0038] The data acquisition unit can collect brain waves using an electroencephalogram (EEG) sensor. Examples of EEG sensors include EEG sensors and fNIRS sensors. An EEG sensor electrically measures brain waves, while an fNIRS sensor measures cerebral blood flow using near-infrared light. This allows the data acquisition unit to collect brain waves using an EEG sensor. Some or all of the above-described processing in the data acquisition unit may be performed using AI, or without AI.

[0039] The data collection unit can collect heart rate data using a heart rate sensor. Examples of heart rate sensors include photoelectric heart rate sensors and electrical heart rate sensors. A photoelectric heart rate sensor measures blood flow using light, while an electrical heart rate sensor measures heart rate electrically. This allows the data collection unit to collect heart rate data using a heart rate sensor. Some or all of the above-described processing in the data collection unit may be performed using AI, or without AI.

[0040] The data collection unit can analyze the user's past vital information and select an appropriate collection method. For example, the data collection unit can select the most stable collection timing based on the user's past vital information. The data collection unit can also concentrate data collection during specific time periods based on the user's past vital information. Furthermore, the data collection unit can analyze the user's past vital information and determine the optimal sensor placement. This allows for the selection of the optimal collection method by analyzing past vital information. Some or all of the above processing in the data collection unit may be performed using AI or not.

[0041] The data collection unit can filter vital information based on the user's current health status and activity level. For example, if the user is exercising, the data collection unit will prioritize collecting vital information related to exercise. It can also collect vital information related to relaxation if the user is resting. Furthermore, if the user is ill, the data collection unit can collect detailed vital information related to their health status. This allows for the collection of more relevant vital information by filtering based on health status and activity level. Some or all of the above processing in the data collection unit may be performed using AI or not.

[0042] The data collection unit can prioritize the collection of highly relevant information by considering the user's geographical location when collecting vital information. For example, if the user is at high altitude, the data collection unit will prioritize the collection of oxygen saturation and respiratory rate. Furthermore, if the user is in an urban area, the data collection unit can prioritize the collection of stress-related vital information. Additionally, if the user is in a natural environment, the data collection unit can prioritize the collection of relaxation-related vital information. This allows for the collection of highly relevant vital information by considering geographical location. Some or all of the processing described above in the data collection unit may be performed using AI, or it may be performed without AI.

[0043] The data collection unit can analyze the user's social media activity and collect relevant information when collecting vital information. For example, if the user is experiencing stress on social media, the data collection unit will prioritize collecting stress-related vital information. It can also prioritize collecting relaxation-related vital information if the user is relaxed on social media. Furthermore, if the user is excited on social media, the data collection unit can prioritize collecting excitement-related vital information. This allows for the collection of relevant vital information by analyzing social media activity. Some or all of the processing described above in the data collection unit may be performed using AI or not.

[0044] The analysis unit can adjust the level of detail of the analysis based on the importance of the vital information during the analysis. For example, the analysis unit can perform a detailed analysis of vital information of high importance. It can also perform a simplified analysis of vital information of low importance. Furthermore, it can perform an analysis of vital information of moderate importance with an appropriate level of detail. By adjusting the level of detail based on importance, efficient analysis becomes possible. Some or all of the above processing in the analysis unit may be performed using AI, or it may be performed without using AI.

[0045] The analysis unit can apply different analysis algorithms depending on the category of vital information during analysis. For example, the analysis unit can apply a heart rate analysis algorithm to vital information related to heart rate. It can also apply an electroencephalogram (EEG) analysis algorithm to vital information related to electroencephalograms. Furthermore, it can apply a respiratory analysis algorithm to vital information related to respiratory rate. By applying an analysis algorithm according to the category, the accuracy of the analysis is improved. Some or all of the above processing in the analysis unit may be performed using AI, or it may be performed without using AI.

[0046] The analysis unit can determine the priority of analysis based on the timing of vital information collection during the analysis. For example, the analysis unit may prioritize the analysis of the most recent vital information. The analysis unit can also analyze current vital information while referring to past vital information. Furthermore, the analysis unit can prioritize the analysis of vital information collected during a specific time period. This allows for the prioritization of the most recent information by determining the priority based on the collection timing. Some or all of the above processing in the analysis unit may be performed using AI, or it may be performed without using AI.

[0047] The analysis unit can adjust the order of analysis based on the relevance of vital information during the analysis. For example, the analysis unit can prioritize the analysis of highly relevant vital information. It can also postpone the analysis of less relevant vital information. Furthermore, the analysis unit can analyze vital information of moderate relevance in an appropriate order. By adjusting the order based on relevance, efficient analysis becomes possible. Some or all of the above processing in the analysis unit may be performed using AI, or it may be performed without using AI.

[0048] The generation unit can adjust the level of detail in the generation of words and sentences based on the intensity of emotion. For example, if the intensity of emotion is high, the generation unit will generate detailed words and sentences. Conversely, if the intensity of emotion is low, the generation unit can also generate concise words and sentences. Furthermore, if the intensity of emotion is moderate, the generation unit can generate words and sentences with an appropriate level of detail. In this way, by adjusting the level of detail based on the intensity of emotion, it is possible to generate words and sentences with an appropriate level of detail. Some or all of the above processing in the generation unit may be performed using a generation AI, or it may be performed without using a generation AI.

[0049] The generation unit can apply different generation algorithms depending on the type of emotion when generating words and sentences. For example, the generation unit can apply a positive expression generation algorithm to the emotion of joy. It can also apply an expression generation algorithm that soothes the emotion of sadness. Furthermore, it can apply an expression generation algorithm that maintains composure to the emotion of anger. In this way, by applying a generation algorithm according to the type of emotion, it is possible to generate words and sentences with appropriate expressions. Some or all of the above processing in the generation unit may be performed using a generation AI, or it may be performed without using a generation AI.

[0050] The generation unit can determine the generation priority of words and sentences based on when emotions arise. For example, the generation unit can prioritize generating words and sentences based on the most recent emotion. The generation unit can also generate words and sentences based on the current emotion, while referring to past emotions. Furthermore, the generation unit can generate words and sentences based on emotions that occurred during a specific time period. By determining priorities based on the timing of occurrence, it is possible to generate words and sentences based on the most recent emotion. Some or all of the above processing in the generation unit may be performed using a generation AI, or it may be performed without using a generation AI.

[0051] The generation unit can adjust the generation order based on emotional relevance when generating words and sentences. For example, the generation unit can prioritize generating words and sentences based on highly relevant emotions. It can also postpone the generation of words and sentences based on less relevant emotions. Furthermore, it can generate words and sentences in an appropriate order based on moderately relevant emotions. This allows for efficient generation by adjusting the order based on relevance. Some or all of the above processing in the generation unit may be performed using a generation AI, or it may be performed without a generation AI.

[0052] The speech unit can adjust the level of detail in its speech based on the importance of the generated words and sentences. For example, it will pronounce highly important words and sentences in detail. It can also pronounce less important words and sentences concisely. Furthermore, it can pronounce words and sentences of moderate importance with an appropriate level of detail. This allows for efficient speech by adjusting the level of detail based on importance. Some or all of the above processing in the speech unit may be performed using AI, or it may be performed without AI.

[0053] The speech unit can apply different speech algorithms depending on the category of the generated words or sentences when it speaks. For example, the speech unit can apply a questioning speech algorithm to words or sentences related to questions. It can also apply a clear and directive speech algorithm to words or sentences related to instructions. Furthermore, it can apply a speech algorithm that conveys gratitude to words or sentences related to gratitude. By applying a speech algorithm according to the category, appropriate speech is possible. Some or all of the above processing in the speech unit may be performed using AI or not.

[0054] The speech unit can determine the priority of utterances based on the timing of the generated words and sentences. For example, the speech unit prioritizes the most recent words and sentences. It can also utter current words and sentences while referring to past words and sentences. Furthermore, the speech unit can prioritize the utterance of words and sentences generated within a specific time period. This allows for the prioritization of the most recent words and sentences by determining priorities based on the timing of their generation. Some or all of the above processing in the speech unit may be performed using AI, or it may be performed without AI.

[0055] The speech unit can adjust the order of utterances based on the relevance of the generated words and sentences during speech production. For example, the speech unit can prioritize the utterance of highly relevant words and sentences. It can also postpone the utterance of less relevant words and sentences. Furthermore, it can utter words and sentences of moderate relevance in an appropriate order. This allows for efficient speech production by adjusting the order based on relevance. Some or all of the above processing in the speech unit may be performed using AI, or it may be performed without AI.

[0056] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0057] The communication system can also include a history analysis unit that analyzes the user's past communication history and generates appropriate words and sentences. For example, it can learn phrases and expressions that the user has frequently used in the past and generate natural conversations based on that. The history analysis unit can also consider the user's past emotional state and provide appropriate expressions at the appropriate time. Furthermore, the history analysis unit can analyze the user's past communication patterns and prepare predicted questions and responses in advance. This makes it possible to have more natural and effective communication by utilizing the user's past communication history.

[0058] The generation unit can further generate words and sentences while considering the user's preferences and interests. For example, if a user is interested in a particular topic, it will generate words and sentences related to that topic. The generation unit can also learn the user's preferred expressions and styles based on their past statements and actions, and generate words and sentences that reflect these. Furthermore, the generation unit can consider the user's current situation and context to generate appropriate words and sentences. This enables communication that reflects the user's preferences and interests.

[0059] The speech unit can further learn the characteristics of the user's voice and generate personalized speech. For example, it can learn the pitch, tone, and rhythm of the user's voice and generate natural-sounding speech based on that. The speech unit can also provide speech that closely resembles the user's own voice by generating speech that reflects the characteristics of the user's voice. Furthermore, the speech unit can generate speech corresponding to different emotional states based on the characteristics of the user's voice. This enables personalized speech generation that reflects the characteristics of the user's voice.

[0060] The generation unit can further consider the user's cultural background when generating words and sentences. For example, if the user belongs to a specific cultural sphere, it will generate expressions and words appropriate to that culture. The generation unit can also consider the user's religion and customs when generating appropriate words and sentences. Furthermore, it can learn regional expressions and slang specific to the user's area and generate words and sentences that reflect them. This enables communication that takes cultural background into account.

[0061] The speech generator can further consider the user's auditory characteristics when generating speech. For example, if the user has difficulty hearing high frequencies, it can generate speech that emphasizes low frequencies. Furthermore, if the user has an easy time hearing a specific frequency range, the speech generator can also generate speech that emphasizes that frequency range. In addition, the speech generator can adjust the speed and rhythm of the speech according to the user's auditory characteristics. This enables speech generation that takes auditory characteristics into account.

[0062] The following briefly describes the processing flow for example form 1.

[0063] Step 1: The data acquisition unit collects vital information. This vital information includes heart rate, electroencephalogram (EEG), and body temperature. The data acquisition unit collects vital information using an EEG sensor and a heart rate sensor. EEG sensors and fNIRS sensors are used as EEG sensors, and photoelectric heart rate sensors and electrical heart rate sensors are used as heart rate sensors. Step 2: The analysis unit uses generative AI to analyze the vital information collected by the collection unit and understand emotions and thoughts. Deep learning models and natural language processing models are used as generative AI. Step 3: The generation unit uses a generation AI to generate words and sentences based on the analysis results obtained by the analysis unit. The generation AI used may include text generation AI (e.g., LLM) or multimodal generation AI. Step 4: The speech unit uses speech generation AI to pronounce the words and sentences generated by the generation unit. Speech synthesis models and text-to-speech technologies are used as speech generation AI.

[0064] (Example of form 2) The communication system according to an embodiment of the present invention is a system that collects vital information, understands a person's emotions and thoughts, generates words and sentences using a generative AI, and has the person speak using a speech-generating AI. This system analyzes the collected vital information and generates words and sentences using a generative AI. Furthermore, by having the person speak using a speech-generating AI, communication can be facilitated. The intended users are ALS patients, people who are intubated and unable to speak, and it is expected that communication will expand between people with different native languages. For example, an ALS patient wears an electroencephalogram (EEG) sensor to have their emotions and thoughts analyzed. The generative AI generates words such as "I want to drink water," and the speech-generating AI speaks those words. This allows the patient to convey their intentions to others. Also, when people with different native languages ​​communicate, the generative AI generates appropriate words and sentences, and the speech-generating AI speaks them, enabling communication that transcends language barriers. In this way, the communication system can convey the user's emotions and thoughts to others.

[0065] The communication system according to this embodiment comprises a collection unit, an analysis unit, a generation unit, and a speech unit. The collection unit collects vital information. Vital information includes, but is not limited to, heart rate, electroencephalogram (EEG), and body temperature. The collection unit collects vital information using, for example, an EEG sensor and a heart rate sensor. Examples of EEG sensors include EEG sensors and fNIRS sensors. Examples of heart rate sensors include photoelectric heart rate sensors and electrical heart rate sensors. The analysis unit uses a generation AI to analyze the vital information collected by the collection unit and understand emotions and thoughts. Examples of generation AI include deep learning models and natural language processing models. The generation unit uses a generation AI to generate words and sentences based on the analysis results obtained by the analysis unit. Examples of generation AI include text generation AI (e.g., LLM) and multimodal generation AI. The speech unit uses a speech generation AI to speak the words and sentences generated by the generation unit. The speech generation AI utilizes, for example, speech synthesis models and text-to-speech conversion technologies. This allows the communication system according to the embodiment to convey the user's emotions and thoughts to others.

[0066] The data collection unit collects vital information. This vital information includes, but is not limited to, heart rate, electroencephalogram (EEG), and body temperature. The data collection unit collects vital information using, for example, an EEG sensor or a heart rate sensor. Examples of EEG sensors include EEG sensors and fNIRS sensors. EEG sensors measure the electrical activity of the brain by being attached to the scalp and acquire EEG data in real time. fNIRS sensors measure changes in blood flow to the brain using near-infrared light to understand the state of brain activity. Examples of heart rate sensors include photoelectric heart rate sensors and electrocardiographic heart rate sensors. Photoelectric heart rate sensors detect heart rate by irradiating light onto the skin and measuring the reflected light. Electrocardiographic heart rate sensors detect heart rate by attaching electrodes to the skin and measuring the electrical activity of the heart. These sensors can be incorporated into wearable devices and medical equipment to continuously monitor the user's vital information. The collected vital information is transmitted to a central database using wireless communication technology and provided to the analysis unit in real time. This allows the data collection unit to accurately understand the user's health and emotional state and provide the data necessary for the next analysis step. Furthermore, the data collection unit can integrate data from multiple sensors and perform noise reduction and outlier filtering to improve data accuracy and reliability. As a result, the data collection unit can provide more accurate vital information and improve the overall system performance.

[0067] The analysis unit uses generative AI to analyze vital information collected by the data collection unit and understand emotions and thoughts. Generative AI may include deep learning models or natural language processing models. Specifically, deep learning models classify the user's emotional state using collected electroencephalogram (EEG) and heart rate data as input. For example, EEG data is analyzed to identify alpha and beta wave patterns, determining whether the user is relaxed or focused. Heart rate data is analyzed to assess stress levels and excitement levels. Natural language processing models express the user's emotions and thoughts in text format based on these analysis results. For example, it may output a specific emotional state such as "The user is currently relaxed." Furthermore, the analysis unit can also analyze changes and trends in emotions by utilizing past data and user history information. This allows for a grasp of the user's long-term emotional patterns and provides more accurate analysis results. The analysis unit can also use anomaly detection algorithms to detect unusual emotional states or abnormal vital data, issuing early warnings. This allows the analysis unit to handle not only real-time sentiment analysis but also long-term sentiment management and anomaly detection, thereby improving the reliability and security of the entire system.

[0068] The generation unit uses a generation AI to generate words and sentences based on the analysis results obtained by the analysis unit. Examples of generation AIs include text generation AI (e.g., LLM) and multimodal generation AI. Text generation AI takes emotional states and thoughts provided by the analysis unit as input and generates natural-sounding words and sentences. For example, it generates specific sentences such as, "The user is currently relaxed and feeling calm." Multimodal generation AI can integrate not only text but also other modalities such as images and audio to generate richer expressions. For example, it can generate images and audio messages corresponding to the user's emotional state, conveying emotions through sight and sound. The generation unit outputs the generated words and sentences in an appropriate format and provides them to the speech unit. Furthermore, the generation unit can collect user feedback and continuously improve the generation AI model. For example, the user can evaluate the generated sentences as "accurate" or "inaccurate" to improve the accuracy of the generation AI. The generation unit can also generate customized sentences according to specific situations and contexts. This allows the generation unit to accurately and naturally express the user's emotions and thoughts, and provide information for conveying them to others.

[0069] The speech unit uses speech generation AI to speak the words and sentences generated by the generation unit. Speech generation AI can utilize technologies such as speech synthesis models or text-to-speech conversion. Speech synthesis models use text provided by the generation unit as input to generate natural-sounding speech. For example, they can generate speech with tone and intonation that matches the user's emotional state, conveying emotions more effectively. Text-to-speech conversion technology adjusts pronunciation and accent when converting text to speech, producing easy-to-understand speech. The speech unit plays the generated speech through speakers or headsets, conveying the user's emotions and thoughts to others. Furthermore, the speech unit can collect user feedback and continuously improve the speech generation AI model. For example, users can evaluate the generated speech as "easy to understand" or "difficult to understand," improving the accuracy of the speech generation AI. The speech unit can also generate customized speech tailored to specific situations and contexts. This allows the speech unit to accurately and naturally express the user's emotions and thoughts through speech, providing information for communication with others.

[0070] The data acquisition unit may include an electroencephalogram (EEG) sensor or a heart rate sensor. Examples of EEG sensors include EEG sensors and fNIRS sensors. An EEG sensor is a sensor that electrically measures brain waves, and an fNIRS sensor is a sensor that measures blood flow in the brain using near-infrared light. Examples of heart rate sensors include photoelectric heart rate sensors and electrocardiographic heart rate sensors. A photoelectric heart rate sensor is a sensor that measures blood flow using light, and an electrocardiographic heart rate sensor is a sensor that electrically measures heart rate. As a result, the data acquisition unit can collect more accurate vital information by using EEG sensors and heart rate sensors. Some or all of the above processing in the data acquisition unit may be performed using AI, for example, or without AI.

[0071] The analysis unit can analyze emotions and thoughts using generative AI. Generative AI includes, for example, deep learning models and natural language processing models. Deep learning models are models that learn from large amounts of data and have advanced analytical capabilities. Natural language processing models are models that analyze text data and understand its meaning. As a result, the analysis unit can improve the accuracy of its analysis of emotions and thoughts by using generative AI. Some or all of the above-mentioned processes in the analysis unit are performed using generative AI.

[0072] The generation unit can generate words and sentences using generative AI. Generative AI includes, for example, text generation AI (e.g., LLM) and multimodal generation AI. Text generation AI is a model that learns from large amounts of text data and has advanced natural language processing capabilities. Multimodal generation AI is a model that can handle multiple modals, such as images and audio, in addition to text. As a result, the generation unit can improve the accuracy of word and sentence generation by using generative AI. Some or all of the above-mentioned processes in the generation unit are performed using generative AI.

[0073] The speech generation unit can produce words and sentences using speech generation AI. Speech generation AI includes, for example, speech synthesis models and text-to-speech conversion technologies. A speech synthesis model is a model that converts text data into speech, and text-to-speech conversion technology is a technology that converts text data into speech. As a result, the speech generation unit improves the accuracy of its speech by using speech generation AI. Some or all of the above-mentioned processes in the speech generation unit are performed using speech generation AI.

[0074] The data acquisition unit can collect brain waves using an electroencephalogram (EEG) sensor. Examples of EEG sensors include EEG sensors and fNIRS sensors. An EEG sensor electrically measures brain waves, while an fNIRS sensor measures cerebral blood flow using near-infrared light. This allows the data acquisition unit to collect brain waves using an EEG sensor. Some or all of the above-described processing in the data acquisition unit may be performed using AI, or without AI.

[0075] The data collection unit can collect heart rate data using a heart rate sensor. Examples of heart rate sensors include photoelectric heart rate sensors and electrical heart rate sensors. A photoelectric heart rate sensor measures blood flow using light, while an electrical heart rate sensor measures heart rate electrically. This allows the data collection unit to collect heart rate data using a heart rate sensor. Some or all of the above-described processing in the data collection unit may be performed using AI, or without AI.

[0076] The data collection unit can estimate the user's emotions and adjust the timing of vital information collection based on the estimated emotions. For example, if the user is relaxed, the data collection unit can set a low collection frequency and collect only the minimum necessary data. Conversely, if the user is stressed, the data collection unit can set a high collection frequency and collect detailed data. Furthermore, if the user is excited, the data collection unit can bring the collection timing closer to real-time to collect immediate data. This allows for the collection of more appropriate vital information by adjusting the collection timing based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the data collection unit may be performed using AI or not.

[0077] The data collection unit can analyze the user's past vital information and select an appropriate collection method. For example, the data collection unit can select the most stable collection timing based on the user's past vital information. The data collection unit can also concentrate data collection during specific time periods based on the user's past vital information. Furthermore, the data collection unit can analyze the user's past vital information and determine the optimal sensor placement. This allows for the selection of the optimal collection method by analyzing past vital information. Some or all of the above processing in the data collection unit may be performed using AI or not.

[0078] The data collection unit can filter vital information based on the user's current health status and activity level. For example, if the user is exercising, the data collection unit will prioritize collecting vital information related to exercise. It can also collect vital information related to relaxation if the user is resting. Furthermore, if the user is ill, the data collection unit can collect detailed vital information related to their health status. This allows for the collection of more relevant vital information by filtering based on health status and activity level. Some or all of the above processing in the data collection unit may be performed using AI or not.

[0079] The data collection unit can estimate the user's emotions and determine the priority of vital information to collect based on the estimated emotions. For example, if the user is tense, the data collection unit will prioritize collecting stress-related vital information such as heart rate and blood pressure. Similarly, if the user is relaxed, the data collection unit can prioritize collecting relaxation-related vital information such as brain waves and respiratory rate. Furthermore, if the user is excited, the data collection unit can prioritize collecting excitement-related vital information such as brain waves and heart rate. This allows for the priority collection of important vital information based on emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. Generative AI may include, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processing in the data collection unit may be performed using AI or not.

[0080] The data collection unit can prioritize the collection of highly relevant information by considering the user's geographical location when collecting vital information. For example, if the user is at high altitude, the data collection unit will prioritize the collection of oxygen saturation and respiratory rate. Furthermore, if the user is in an urban area, the data collection unit can prioritize the collection of stress-related vital information. Additionally, if the user is in a natural environment, the data collection unit can prioritize the collection of relaxation-related vital information. This allows for the collection of highly relevant vital information by considering geographical location. Some or all of the processing described above in the data collection unit may be performed using AI, or it may be performed without AI.

[0081] The data collection unit can analyze the user's social media activity and collect relevant information when collecting vital information. For example, if the user is experiencing stress on social media, the data collection unit will prioritize collecting stress-related vital information. It can also prioritize collecting relaxation-related vital information if the user is relaxed on social media. Furthermore, if the user is excited on social media, the data collection unit can prioritize collecting excitement-related vital information. This allows for the collection of relevant vital information by analyzing social media activity. Some or all of the processing described above in the data collection unit may be performed using AI or not.

[0082] The analysis unit can estimate the user's emotions and adjust the presentation of the analysis based on the estimated emotions. For example, if the user is tense, the analysis unit can provide simple and easy-to-understand analysis results. If the user is relaxed, the analysis unit can also provide detailed analysis results. Furthermore, if the user is excited, the analysis unit can provide visually stimulating analysis results. In this way, by adjusting the presentation of the analysis based on emotions, it is possible to provide analysis results that are easy for the user to understand. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI or not using AI.

[0083] The analysis unit can adjust the level of detail of the analysis based on the importance of the vital information during the analysis. For example, the analysis unit can perform a detailed analysis of vital information of high importance. It can also perform a simplified analysis of vital information of low importance. Furthermore, it can perform an analysis of vital information of moderate importance with an appropriate level of detail. By adjusting the level of detail based on importance, efficient analysis becomes possible. Some or all of the above processing in the analysis unit may be performed using AI, or it may be performed without using AI.

[0084] The analysis unit can apply different analysis algorithms depending on the category of vital information during analysis. For example, the analysis unit can apply a heart rate analysis algorithm to vital information related to heart rate. It can also apply an electroencephalogram (EEG) analysis algorithm to vital information related to electroencephalograms. Furthermore, it can apply a respiratory analysis algorithm to vital information related to respiratory rate. By applying an analysis algorithm according to the category, the accuracy of the analysis is improved. Some or all of the above processing in the analysis unit may be performed using AI, or it may be performed without using AI.

[0085] The analysis unit can estimate the user's emotions and adjust the length of the analysis based on the estimated emotions. For example, if the user is in a hurry, the analysis unit can provide a short, concise analysis. If the user is relaxed, the analysis unit can also provide a detailed analysis. Furthermore, if the user is excited, the analysis unit can provide a visually stimulating analysis. By adjusting the length of the analysis based on emotions, the analysis unit can provide the user with an analysis of an appropriate length. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI may be, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the analysis unit may be performed using AI or not.

[0086] The analysis unit can determine the priority of analysis based on the timing of vital information collection during the analysis. For example, the analysis unit may prioritize the analysis of the most recent vital information. The analysis unit can also analyze current vital information while referring to past vital information. Furthermore, the analysis unit can prioritize the analysis of vital information collected during a specific time period. This allows for the prioritization of the most recent information by determining the priority based on the collection timing. Some or all of the above processing in the analysis unit may be performed using AI, or it may be performed without using AI.

[0087] The analysis unit can adjust the order of analysis based on the relevance of vital information during the analysis. For example, the analysis unit can prioritize the analysis of highly relevant vital information. It can also postpone the analysis of less relevant vital information. Furthermore, the analysis unit can analyze vital information of moderate relevance in an appropriate order. By adjusting the order based on relevance, efficient analysis becomes possible. Some or all of the above processing in the analysis unit may be performed using AI, or it may be performed without using AI.

[0088] The generation unit can estimate the user's emotions and adjust the expression of the words and sentences it generates based on the estimated emotions. For example, if the user is relaxed, the generation unit will generate words and sentences with calm expressions. It can also generate words and sentences with concise and clear expressions if the user is tense. Furthermore, if the user is excited, the generation unit can generate words and sentences with emphasized emotions. This allows for the generation of words and sentences that are appropriate for the user by adjusting the expression based on their emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or a generation AI. The generation AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the generation unit are performed using a generation AI.

[0089] The generation unit can adjust the level of detail in the generation of words and sentences based on the intensity of emotion. For example, if the intensity of emotion is high, the generation unit will generate detailed words and sentences. Conversely, if the intensity of emotion is low, the generation unit can also generate concise words and sentences. Furthermore, if the intensity of emotion is moderate, the generation unit can generate words and sentences with an appropriate level of detail. In this way, by adjusting the level of detail based on the intensity of emotion, it is possible to generate words and sentences with an appropriate level of detail. Some or all of the above processing in the generation unit may be performed using a generation AI, or it may be performed without using a generation AI.

[0090] The generation unit can apply different generation algorithms depending on the type of emotion when generating words and sentences. For example, the generation unit can apply a positive expression generation algorithm to the emotion of joy. It can also apply an expression generation algorithm that soothes the emotion of sadness. Furthermore, it can apply an expression generation algorithm that maintains composure to the emotion of anger. In this way, by applying a generation algorithm according to the type of emotion, it is possible to generate words and sentences with appropriate expressions. Some or all of the above processing in the generation unit may be performed using a generation AI, or it may be performed without using a generation AI.

[0091] The generation unit can estimate the user's emotions and adjust the length of the words and sentences it generates based on the estimated emotions. For example, if the user is in a hurry, the generation unit can generate short, concise words and sentences. If the user is relaxed, the generation unit can also generate longer words and sentences that include detailed explanations. Furthermore, if the user is excited, the generation unit can generate words and sentences of an appropriate length that emphasize the emotion. In this way, by adjusting the length based on emotions, it is possible to generate words and sentences of an appropriate length for the user. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generation AI. The generation AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples. Some or all of the above processing in the generation unit is performed using a generation AI.

[0092] The generation unit can determine the generation priority of words and sentences based on when emotions arise. For example, the generation unit can prioritize generating words and sentences based on the most recent emotion. The generation unit can also generate words and sentences based on the current emotion, while referring to past emotions. Furthermore, the generation unit can generate words and sentences based on emotions that occurred during a specific time period. By determining priorities based on the timing of occurrence, it is possible to generate words and sentences based on the most recent emotion. Some or all of the above processing in the generation unit may be performed using a generation AI, or it may be performed without using a generation AI.

[0093] The generation unit can adjust the generation order based on emotional relevance when generating words and sentences. For example, the generation unit can prioritize generating words and sentences based on highly relevant emotions. It can also postpone the generation of words and sentences based on less relevant emotions. Furthermore, it can generate words and sentences in an appropriate order based on moderately relevant emotions. This allows for efficient generation by adjusting the order based on relevance. Some or all of the above processing in the generation unit may be performed using a generation AI, or it may be performed without a generation AI.

[0094] The speech unit can estimate the user's emotions and adjust its speech expression based on the estimated emotions. For example, if the user is relaxed, the speech unit will speak in a calm voice. If the user is tense, the speech unit can also speak in a clear and calm voice. Furthermore, if the user is excited, the speech unit can speak in an emotionally emphasized voice. In this way, by adjusting the expression based on emotions, it is possible to deliver speech that is appropriate for the user. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the speech unit may be performed using AI or not.

[0095] The speech unit can adjust the level of detail in its speech based on the importance of the generated words and sentences. For example, it will pronounce highly important words and sentences in detail. It can also pronounce less important words and sentences concisely. Furthermore, it can pronounce words and sentences of moderate importance with an appropriate level of detail. This allows for efficient speech by adjusting the level of detail based on importance. Some or all of the above processing in the speech unit may be performed using AI, or it may be performed without AI.

[0096] The speech unit can apply different speech algorithms depending on the category of the generated words or sentences when it speaks. For example, the speech unit can apply a questioning speech algorithm to words or sentences related to questions. It can also apply a clear and directive speech algorithm to words or sentences related to instructions. Furthermore, it can apply a speech algorithm that conveys gratitude to words or sentences related to gratitude. By applying a speech algorithm according to the category, appropriate speech is possible. Some or all of the above processing in the speech unit may be performed using AI or not.

[0097] The speech unit can estimate the user's emotions and adjust the length of its utterances based on the estimated emotions. For example, if the user is in a hurry, the speech unit will make short, concise utterances. If the user is relaxed, the speech unit can also make longer utterances that include detailed explanations. Furthermore, if the user is excited, the speech unit can make utterances of an appropriate length that emphasize the emotion. In this way, by adjusting the length based on emotions, the speech unit can make utterances of an appropriate length for the user. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the speech unit may be performed using AI or not.

[0098] The speech unit can determine the priority of utterances based on the timing of the generated words and sentences. For example, the speech unit prioritizes the most recent words and sentences. It can also utter current words and sentences while referring to past words and sentences. Furthermore, the speech unit can prioritize the utterance of words and sentences generated within a specific time period. This allows for the prioritization of the most recent words and sentences by determining priorities based on the timing of their generation. Some or all of the above processing in the speech unit may be performed using AI, or it may be performed without AI.

[0099] The speech unit can adjust the order of utterances based on the relevance of the generated words and sentences during speech production. For example, the speech unit can prioritize the utterance of highly relevant words and sentences. It can also postpone the utterance of less relevant words and sentences. Furthermore, it can utter words and sentences of moderate relevance in an appropriate order. This allows for efficient speech production by adjusting the order based on relevance. Some or all of the above processing in the speech unit may be performed using AI, or it may be performed without AI.

[0100] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0101] The communication system can also include a history analysis unit that analyzes the user's past communication history and generates appropriate words and sentences. For example, it can learn phrases and expressions that the user has frequently used in the past and generate natural conversations based on that. The history analysis unit can also consider the user's past emotional state and provide appropriate expressions at the appropriate time. Furthermore, the history analysis unit can analyze the user's past communication patterns and prepare predicted questions and responses in advance. This makes it possible to have more natural and effective communication by utilizing the user's past communication history.

[0102] The data collection unit can further collect ambient sounds from the user, and the analysis unit can analyze those sounds. For example, if the user is in a noisy environment, the data collection unit will collect those ambient sounds, and the analysis unit will perform noise filtering. Similarly, if the user is in a quiet environment, the data collection unit can collect those ambient sounds, and the analysis unit can estimate the user's relaxed state. Furthermore, if the user is listening to a specific sound (e.g., music or nature sounds), the data collection unit can collect that sound, and the analysis unit can estimate the user's emotional state. This allows for a more accurate analysis of vital information by considering ambient sounds.

[0103] The analysis unit can further include a facial expression analysis unit that analyzes the user's facial expressions. For example, the user's face can be photographed using a camera, and the facial expression analysis unit can analyze those expressions. The facial expression analysis unit can also detect subtle changes in the user's facial expressions and estimate their emotional state. Furthermore, the facial expression analysis unit can analyze the user's eye and mouth movements to estimate their emotional state in more detail. As a result, facial expression analysis can provide a more accurate understanding of the user's emotional state.

[0104] The generation unit can further generate words and sentences while considering the user's preferences and interests. For example, if a user is interested in a particular topic, it will generate words and sentences related to that topic. The generation unit can also learn the user's preferred expressions and styles based on their past statements and actions, and generate words and sentences that reflect these. Furthermore, the generation unit can consider the user's current situation and context to generate appropriate words and sentences. This enables communication that reflects the user's preferences and interests.

[0105] The speech unit can further learn the characteristics of the user's voice and generate personalized speech. For example, it can learn the pitch, tone, and rhythm of the user's voice and generate natural-sounding speech based on that. The speech unit can also provide speech that closely resembles the user's own voice by generating speech that reflects the characteristics of the user's voice. Furthermore, the speech unit can generate speech corresponding to different emotional states based on the characteristics of the user's voice. This enables personalized speech generation that reflects the characteristics of the user's voice.

[0106] The data collection unit can further detect the user's physical movements and collect vital information based on those movements. For example, if the user is exercising, it can detect those movements and collect vital information related to the exercise. The data collection unit can also detect if the user is relaxed and collect vital information related to that relaxed state. Furthermore, if the user is performing a specific action (e.g., raising an arm, walking), the data collection unit can detect that action and collect appropriate vital information. This allows for more accurate collection of vital information by taking physical movements into consideration.

[0107] The analysis unit can further analyze the user's voice and estimate their emotional state from that voice. For example, it can analyze the tone, rhythm, and speed of the user's voice to estimate their emotional state. The analysis unit can also analyze the volume and intonation of the user's voice to estimate their emotional state. Furthermore, the analysis unit can analyze changes in the user's voice in real time and estimate their emotional state immediately. As a result, voice analysis can more accurately grasp the user's emotional state.

[0108] The generation unit can further consider the user's cultural background when generating words and sentences. For example, if the user belongs to a specific cultural sphere, it will generate expressions and words appropriate to that culture. The generation unit can also consider the user's religion and customs when generating appropriate words and sentences. Furthermore, it can learn regional expressions and slang specific to the user's area and generate words and sentences that reflect them. This enables communication that takes cultural background into account.

[0109] The speech generator can further consider the user's auditory characteristics when generating speech. For example, if the user has difficulty hearing high frequencies, it can generate speech that emphasizes low frequencies. Furthermore, if the user has an easy time hearing a specific frequency range, the speech generator can also generate speech that emphasizes that frequency range. In addition, the speech generator can adjust the speed and rhythm of the speech according to the user's auditory characteristics. This enables speech generation that takes auditory characteristics into account.

[0110] The data collection unit can further measure the user's skin conductance, and the analysis unit can analyze that data. For example, if the user is tense, the unit can detect an increase in skin conductance and analyze that data. The data collection unit can also detect a decrease in skin conductance when the user is relaxed and analyze that data. Furthermore, if the user is excited, the data collection unit can detect changes in skin conductance in real time and immediately analyze them in the analysis unit. This allows for a more accurate understanding of the user's emotional state by measuring skin conductance.

[0111] The following briefly describes the processing flow for example form 2.

[0112] Step 1: The data acquisition unit collects vital information. This vital information includes heart rate, electroencephalogram (EEG), and body temperature. The data acquisition unit collects vital information using an EEG sensor and a heart rate sensor. EEG sensors and fNIRS sensors are used as EEG sensors, and photoelectric heart rate sensors and electrical heart rate sensors are used as heart rate sensors. Step 2: The analysis unit uses generative AI to analyze the vital information collected by the collection unit and understand emotions and thoughts. Deep learning models and natural language processing models are used as generative AI. Step 3: The generation unit uses a generation AI to generate words and sentences based on the analysis results obtained by the analysis unit. The generation AI used may include text generation AI (e.g., LLM) or multimodal generation AI. Step 4: The speech unit uses speech generation AI to pronounce the words and sentences generated by the generation unit. Speech synthesis models and text-to-speech technologies are used as speech generation AI.

[0113] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0114] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0115] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0116] Each of the multiple elements described above, including the collection unit, analysis unit, generation unit, and speech unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the collection unit collects vital information using the electroencephalogram (EEG) sensor and heart rate sensor of the smart device 14. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and analyzes the collected vital information to understand emotions and thoughts. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and generates words and sentences based on the analysis results. The speech unit speaks the words and sentences generated using the voice generation AI of the smart device 14, for example. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0117] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0118] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0119] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0120] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0121] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0122] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0123] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0124] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0125] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0126] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0127] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0128] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0129] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0130] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0131] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0132] Each of the multiple elements described above, including the collection unit, analysis unit, generation unit, and speech unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the collection unit collects vital information using the electroencephalogram (EEG) sensor and heart rate sensor of the smart glasses 214. The analysis unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, and analyzes the collected vital information to understand emotions and thoughts. The generation unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, and generates words and sentences based on the analysis results. The speech unit speaks the words and sentences generated using, for example, the voice generation AI of the smart glasses 214. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0133] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0134] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0135] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0136] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0137] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0138] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0139] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0140] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0141] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0142] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0143] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0144] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0145] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0146] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0147] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0148] Each of the multiple elements described above, including the collection unit, analysis unit, generation unit, and speech unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the collection unit collects vital information using the electroencephalogram (EEG) sensor and heart rate sensor of the headset terminal 314. The analysis unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and analyzes the collected vital information to understand emotions and thoughts. The generation unit is implemented in the specific processing unit 290 of the data processing unit 12, for example, and generates words and sentences based on the analysis results. The speech unit speaks the words and sentences generated using the voice generation AI of the headset terminal 314, for example. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0149] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0150] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0151] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0152] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0153] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0154] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0155] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0156] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0157] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0158] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0159] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0160] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0161] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0162] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0163] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0164] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0165] Each of the multiple elements described above, including the collection unit, analysis unit, generation unit, and speech unit, is implemented in, for example, at least one of the robot 414 and the data processing unit 12. For example, the collection unit collects vital information using the electroencephalogram (EEG) sensor and heart rate sensor of the robot 414. The analysis unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, and analyzes the collected vital information to understand emotions and thoughts. The generation unit is implemented, for example, by the specific processing unit 290 of the data processing unit 12, and generates words and sentences based on the analysis results. The speech unit speaks the words and sentences generated using, for example, the voice generation AI of the robot 414. The correspondence between each unit and the device or control unit is not limited to the example described above, and various modifications are possible.

[0166] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0167] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0168] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0169] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0170] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0171] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0172] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0173] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0174] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0175] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0176] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0177] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0178] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0179] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0180] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0181] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0182] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0183] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0184] (Note 1) A collection unit that collects vital information, An analysis unit analyzes the information collected by the aforementioned collection unit to understand emotions or thoughts, A generation unit that generates words or sentences based on the analysis results obtained by the analysis unit, The system comprises a speech unit that speaks words and sentences generated by the generation unit. A system characterized by the following features. (Note 2) The aforementioned collection unit is Includes an electroencephalogram (EEG) sensor or a heart rate sensor. The system described in Appendix 1, characterized by the features described herein. (Note 3) The aforementioned analysis unit, Using generative AI to analyze emotions and thoughts. The system described in Appendix 1, characterized by the features described herein. (Note 4) The generating unit is Generative AI is used to generate words and sentences. The system described in Appendix 1, characterized by the features described herein. (Note 5) The aforementioned speech unit is, Use speech generation AI to produce words and sentences. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned collection unit is Brainwaves are collected using an electroencephalogram (EEG) sensor. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned collection unit is Collect heart rate data using a heart rate sensor. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned collection unit is The system estimates the user's emotions and adjusts the timing of vital information collection based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned collection unit is Analyze the user's past vital information and select the appropriate data collection method. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned collection unit is When collecting vital information, filtering is performed based on the user's current health status and activity level. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned collection unit is It estimates the user's emotions and determines the priority of vital information to collect based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned collection unit is When collecting vital information, the system prioritizes collecting highly relevant information by considering the user's geographical location. The system described in Appendix 1, characterized by the features described herein. (Note 13) The aforementioned collection unit is When collecting vital information, we analyze the user's social media activity and collect relevant information. The system described in Appendix 1, characterized by the features described herein. (Note 14) The aforementioned analysis unit, The system estimates the user's emotions and adjusts the representation of the analysis based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 15) The aforementioned analysis unit, During analysis, the level of detail of the analysis is adjusted based on the importance of the vital information. The system described in Appendix 1, characterized by the features described herein. (Note 16) The aforementioned analysis unit, During analysis, different analysis algorithms are applied depending on the category of vital information. The system described in Appendix 1, characterized by the features described herein. (Note 17) The aforementioned analysis unit, It estimates the user's emotions and adjusts the length of the analysis based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 18) The aforementioned analysis unit, During analysis, the priority of the analysis is determined based on when vital information was collected. The system described in Appendix 1, characterized by the features described herein. (Note 19) The aforementioned analysis unit, During analysis, the order of analysis is adjusted based on the relevance of vital information. The system described in Appendix 1, characterized by the features described herein. (Note 20) The generating unit is It estimates the user's emotions and adjusts the way words and sentences are expressed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 21) The generating unit is When generating words and sentences, adjust the level of detail based on the intensity of emotion. The system described in Appendix 1, characterized by the features described herein. (Note 22) The generating unit is When generating words and sentences, different generation algorithms are applied depending on the type of emotion. The system described in Appendix 1, characterized by the features described herein. (Note 23) The generating unit is It estimates the user's emotions and adjusts the length of the words and sentences generated based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 24) The generating unit is When generating words or sentences, the priority of generation is determined based on when emotions arise. The system described in Appendix 1, characterized by the features described herein. (Note 25) The generating unit is When generating words and sentences, the order of generation is adjusted based on emotional relevance. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned speech unit is, It estimates the user's emotions and adjusts the way the speech is expressed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned speech unit is, When uttering, the level of detail of the utterance is adjusted based on the importance of the generated words and sentences. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned speech unit is, When speech is produced, different speech algorithms are applied depending on the category of the generated words or sentences. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned speech unit is, It estimates the user's emotions and adjusts the length of utterances based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned speech unit is, When uttering, the priority of utterance is determined based on the timing of the generation of the words or sentences. The system described in Appendix 1, characterized by the features described herein. (Note 31) The aforementioned speech unit is, When uttering words, the order of utterances is adjusted based on the relationships between the generated words and sentences. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0185] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. A collection unit that collects multiple types of vital information, including heart rate, electroencephalogram, blood pressure and respiratory rate, The vital information collected by the collection unit is analyzed using generating AI, and the analysis unit understands the user's emotions or thoughts. A generation unit that generates words or sentences based on the analysis results obtained by the analysis unit, The system comprises a speech unit that speaks the words and sentences generated by the generation unit, The aforementioned collection unit is The analysis unit determines that if the user's emotions are in a relaxed state, the frequency of collecting vital information is set to a low level, and if the user's emotions are in a stressed state, the frequency of collecting vital information is set to a high level. Furthermore, when the user is in a state of tension, stress-related vital information, including heart rate and blood pressure, is preferentially collected; when the user is in a state of relaxation, relaxation-related vital information, including brain waves and respiratory rate, is preferentially collected; and when the user is in a state of excitement, excitement-related vital information, including brain waves and heart rate, is preferentially collected. A system characterized by the following features.

2. The generating unit is Generating words and sentences using generative AI. The system according to feature 1.

3. The aforementioned speech unit is, Use generative AI to produce words and sentences. The system according to feature 1.

4. The aforementioned collection unit is Brainwaves are collected using an electroencephalogram (EEG) sensor. The system according to feature 1.

5. The aforementioned collection unit is Collect heart rate data using a heart rate sensor. The system according to feature 1.