system

The system addresses real-time conversation understanding and ambient sound handling by using generative AI for transcription, summarization, emotion recognition, and alerting, improving communication clarity and safety.

JP2026072528APending Publication Date: 2026-05-01SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Conventional technologies fail to adequately support real-time conversation understanding and handling of ambient environmental sounds.

Method used

A system comprising a transcription unit, summarization unit, emotion recognition unit, and alert unit, utilizing generative AI for real-time conversation transcription, summarization, emotion analysis, and ambient sound monitoring, respectively.

Benefits of technology

Enables real-time understanding of conversations, summarizes important points, recognizes emotions, and alerts users to important sounds, enhancing communication clarity and safety in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026072528000001_ABST
    Figure 2026072528000001_ABST
Patent Text Reader

Abstract

The system according to this embodiment aims to understand conversations in real time while supporting hearing and responding to ambient noise. [Solution] The system according to the embodiment comprises a transcription unit, a summarization unit, an emotion recognition unit, a feedback unit, and an alert unit. The transcription unit transcribes the conversation in real time. The summarization unit summarizes the conversation transcribed by the transcription unit. The emotion recognition unit analyzes the emotions of the conversation partner based on the information summarized by the summarization unit. The feedback unit provides feedback based on the emotions analyzed by the emotion recognition unit. The alert unit monitors ambient sounds and notifies important warning sounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional technology, there is a problem that real-time conversation understanding while supporting hearing and coping with ambient environmental sounds have not been sufficiently achieved.

[0005] The system according to the embodiment aims to understand conversations in real time while supporting hearing and also cope with ambient environmental sounds.

Means for Solving the Problems

[0006] The system according to this embodiment comprises a transcription unit, a summarization unit, an emotion recognition unit, a feedback unit, and an alert unit. The transcription unit transcribes the conversation in real time. The summarization unit summarizes the conversation transcribed by the transcription unit. The emotion recognition unit analyzes the emotions of the conversation partner based on the information summarized by the summarization unit. The feedback unit provides feedback based on the emotions analyzed by the emotion recognition unit. The alert unit monitors ambient sounds and notifies important warning sounds. [Effects of the Invention]

[0007] The system according to this embodiment can understand conversations in real time while supporting hearing and can also respond to ambient sounds. [Brief explanation of the drawing]

[0008] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Modes for carrying out the invention]

[0009] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0010] First, let's explain the terminology used in the following explanation.

[0011] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), or TPU (Tensor Processing Unit).

[0012] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0013] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0014] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F manages communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0016] [First Embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0017] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0019] The smart device 14 comprises a computer 36, a receiving device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The receiving device 38, output device 40, and camera 42 are also connected to the bus 52.

[0020] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, and accepts user input. The touch panel 38A accepts user input via touch by detecting contact with an object (e.g., a pen or finger). The microphone 38B accepts user input via voice by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 (see Figure 2) acquires the data indicating the user input.

[0021] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user by outputting the data in a form perceptible to the user (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0022] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0023] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0024] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0025] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0026] In the smart device 14, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart device 14 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0027] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device having the data generation model 58. The data processing device 12 may also be a server device or a terminal device owned by a user (e.g., a mobile phone, robot, home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.

[0028] (Example of form 1) The Kiko-Yell System, according to an embodiment of the present invention, is a system that supports daily life while assisting hearing by utilizing hearing aids, smartphones, and generative AI. The Kiko-Yell System provides various functions to enable users to enjoy communication comfortably and with peace of mind. The Kiko-Yell System primarily solves communication problems for people with hearing difficulties (such as the elderly, the hearing impaired, and workers in noisy environments). For example, the Kiko-Yell System provides a function to analyze long conversations in everyday life and business settings in real time and summarize the main points concisely. This allows for efficient capture of important information that might otherwise be missed. Furthermore, the Kiko-Yell System improves the quality of communication by using emotion recognition to infer the emotions of the person being spoken to and providing feedback based on the analysis of voice and facial expressions. In addition, the Kiko-Yell System makes important sounds easier to hear even in noisy environments through advanced voice filtering using generative AI. It also provides an alert function that quickly detects abnormal sounds and warning sounds in emergencies and notifies the user. The Kiko-Yell System provides the following functions: 1. Conversation transcription function: Transcribes conversations in real time and displays them on the smartphone. 1. This allows users to review parts they missed or had difficulty hearing in text form. 2. Conversation summarization function: The generating AI summarizes long conversations, displaying only the important points concisely. This allows for efficient information acquisition. 3. Emotion recognition function: The generating AI analyzes the tone of voice and facial expressions of the person being spoken to, conveying their emotions and intentions to the user. 4. Emergency alert function: The generating AI monitors ambient noise and quickly notifies the user of important warning sounds and sirens. These functions enable the KikoYell system to achieve not only "hearing" but also "clearly understood" communication. For example, using the conversation transcription function, users can review parts they missed during a meeting on their smartphone. The emotion recognition function makes it easier to understand the emotions of the person being spoken to, enabling better communication. The emergency alert function ensures that important warning sounds are not missed even in noisy environments. In this way, the KikoYell system supports the user's hearing and makes daily life more comfortable.

[0029] The Kiko-Yell system according to this embodiment comprises a transcription unit, a summarization unit, an emotion recognition unit, a feedback unit, and an alert unit. The transcription unit transcribes conversations in real time. The transcription unit converts speech to text using, for example, a generative AI. The transcription unit can accurately transcribe the content of conversations using speech recognition technology with the generative AI. The transcription unit can also transcribe by analyzing the characteristics of speech and identifying the voices of the speakers using the generative AI. The summarization unit summarizes the conversations transcribed by the transcription unit. The summarization unit can summarize long conversations using, for example, a generative AI and concisely display only the important points. The summarization unit can extract and summarize the main points of a conversation using natural language processing technology with the generative AI. The summarization unit can also prioritize summarizing important information by understanding the context of the conversation with the generative AI. The emotion recognition unit analyzes the emotions of the conversation partners based on the information summarized by the summarization unit. The emotion recognition unit, for example, uses generative AI to analyze the tone of voice and facial expressions of the person it is talking to, and conveys emotions and intentions to the user. The emotion recognition unit can detect changes in emotion by having the generative AI analyze the pitch and volume of the voice. The emotion recognition unit can also estimate emotions from facial expressions by having the generative AI analyze the movement of facial feature points. The feedback unit provides feedback based on the emotions analyzed by the emotion recognition unit. The feedback unit, for example, uses generative AI to provide appropriate feedback to the user. The feedback unit can provide encouraging messages and advice according to the user's emotional state using the generative AI. The feedback unit can also provide feedback in real time using the generative AI and respond immediately to changes in the user's emotions. The alert unit monitors ambient sounds and notifies important warning sounds. The alert unit, for example, uses generative AI to analyze ambient sounds and quickly notifies important warning sounds and sirens. The alert unit can detect emergency warning sounds by having the generative AI analyze the type and volume of sounds. The alert unit can also use a generating AI to monitor ambient sounds in real time and detect abnormal sounds. As a result, the Kiko-Yell system according to this embodiment can transcribe conversations, summarize them, recognize emotions, provide feedback, and send alert notifications.

[0030] The transcription unit transcribes conversations in real time. For example, it uses generative AI to convert speech to text. Specifically, the generative AI utilizes advanced speech recognition technology to analyze the audio signal and accurately transcribe the spoken content. The generative AI analyzes the frequency spectrum of the speech to identify phonemes and words. Furthermore, the generative AI learns the characteristics of each speaker's voice, allowing it to identify and transcribe individual utterances even when multiple speakers are present. For example, in meetings or interviews, each speaker's utterance is recorded as a separate text, making later review and analysis easier. The generative AI also incorporates noise reduction technology, effectively removing background noise and other unwanted sounds to obtain clear audio data. This improves transcription accuracy and reduces the risk of misrecognition. Additionally, the generative AI considers context during the speech recognition process and automatically corrects homonyms and grammatical errors. This results in natural and easy-to-read generated text. Furthermore, the transcription unit provides an interface that displays the generated text in real time, allowing users to immediately review the content. This allows you to grasp important points without missing anything, even as the conversation is in progress.

[0031] The summarization unit summarizes conversations transcribed by the transcription unit. For example, the summarization unit uses generative AI to summarize long conversations, displaying only the most important points concisely. Specifically, the generative AI uses natural language processing technology to extract important information from text data and generate a summary. The generative AI understands the context of the conversation and identifies key keywords and phrases. For example, when summarizing meeting minutes, the generative AI prioritizes extracting agenda items, decisions, and action items to create a concise summary. The generative AI analyzes the text structure and omits redundant parts and repetitive expressions to improve the accuracy of the summary. Furthermore, the generative AI can adjust the level of detail in the summary according to the user's needs. For example, if a detailed summary is required, it generates a summary containing more information; if a concise summary is needed, it extracts only the most important points. In addition, the summarization unit provides a visual interface to display the generated summary, making it easy for users to understand. This allows for efficient comprehension of long conversations and large amounts of text data, and quick identification of important information.

[0032] The emotion recognition unit analyzes the emotions of the conversation partner based on the information summarized by the summarization unit. For example, the emotion recognition unit uses generative AI to analyze the tone of voice and facial expressions of the conversation partner and conveys their emotions and intentions to the user. Specifically, the generative AI analyzes features such as pitch, volume, and rhythm of the voice to detect changes in emotion. For example, a higher tone of voice may indicate excitement or joy, while a lower tone may indicate calmness or sadness. The generative AI learns these voice features and identifies patterns of emotion. The generative AI also analyzes the movement of facial feature points and estimates emotions from facial expressions. For example, it analyzes eyebrow movements and the degree to which the corners of the mouth are turned up to identify emotions such as joy, anger, and sadness. The generative AI integrates this information and can analyze the emotional state of the conversation partner in real time. Furthermore, the emotion recognition unit provides an interface that visually displays the analysis results, allowing the user to intuitively understand the emotions of the conversation partner. This enables the user to grasp the other person's emotions during a conversation and take appropriate action.

[0033] The feedback unit provides feedback based on emotions analyzed by the emotion recognition unit. For example, the feedback unit uses generative AI to provide appropriate feedback to the user. Specifically, the generative AI provides encouraging messages and advice according to the user's emotional state. For instance, if the user is stressed, the generative AI provides advice and words of encouragement to help them relax. If the user is feeling joy or excitement, the generative AI shares that emotion and provides positive feedback. The generative AI can monitor changes in the user's emotions in real time and respond immediately. Furthermore, the feedback unit collects user feedback and uses it as training data for the generative AI. This allows the generative AI to learn the user's preferences and tendencies, enabling it to provide more personalized feedback. The feedback unit combines visual and auditory feedback to communicate effectively with the user. This allows the user to receive appropriate support tailored to their emotional state.

[0034] The alert unit monitors ambient sounds and notifies users of important warning sounds. For example, it uses a generative AI to analyze ambient sounds and quickly notify users of important warning sounds and sirens. Specifically, the generative AI analyzes the type and volume of sounds to detect emergency warning sounds. For example, it identifies highly urgent sounds such as fire alarms and ambulance sirens and notifies the user. By analyzing the frequency spectrum of sounds and identifying specific patterns, the generative AI can quickly detect important warning sounds. Furthermore, the generative AI can monitor ambient sounds in real time and detect abnormal sounds. For example, if an unusual sound occurs, the generative AI analyzes the sound and notifies the user of the potential anomaly. In addition, the alert unit can select the notification method according to the user's situation. For example, it can select the method that the user can receive most effectively, such as visual, audio, or vibration notifications. This allows the alert unit to provide users with important information quickly and reliably, supporting their response in emergencies.

[0035] The transcription unit can transcribe conversations in real time and display them on a smartphone. For example, the transcription unit can use generative AI to convert audio to text and display it on a smartphone. The transcription unit can use generative AI and speech recognition technology to accurately transcribe the content of conversations and display it on a smartphone. The transcription unit can also use generative AI to analyze the characteristics of the audio, identify the speaker's voice, transcribe it, and display it on a smartphone. This prevents missed information by transcribing conversations in real time and displaying them on a smartphone. Some or all of the above-described processes in the transcription unit may be performed using generative AI or not. For example, the transcription unit can use generative AI to convert audio to text and display it on a smartphone.

[0036] The summarization section can summarize long conversations and display only the important points concisely. For example, the summarization section can use generative AI to summarize long conversations and display only the important points concisely. The summarization section can use generative AI to extract and summarize the main points of a conversation using natural language processing techniques. The summarization section can also use generative AI to understand the context of the conversation and prioritize the summarization of important information. This allows for efficient information acquisition by summarizing long conversations and displaying only the important points concisely. Some or all of the above processing in the summarization section may be performed using generative AI or not. For example, the summarization section can use generative AI to summarize long conversations and display only the important points concisely.

[0037] The emotion recognition unit can analyze the tone of voice and facial expressions of the person it is talking to and convey their emotions and intentions to the user. For example, the emotion recognition unit can use generative AI to analyze the tone of voice and facial expressions of the person it is talking to and convey their emotions and intentions to the user. The emotion recognition unit can use generative AI to analyze the pitch and volume of speech and detect changes in emotion. The emotion recognition unit can also use generative AI to analyze the movement of facial feature points and estimate emotions from facial expressions. This improves the quality of communication by analyzing the emotions and intentions of the person it is talking to and conveying them to the user. Some or all of the above processing in the emotion recognition unit may be performed using generative AI or not. For example, the emotion recognition unit can use generative AI to analyze the tone of voice and facial expressions of the person it is talking to and convey their emotions and intentions to the user.

[0038] The alert unit can monitor ambient sounds and quickly notify of important warning sounds and sirens. For example, the alert unit can use a generation AI to analyze ambient sounds and quickly notify of important warning sounds and sirens. The alert unit's generation AI can analyze the type and volume of sounds and detect emergency warning sounds. The alert unit's generation AI can also monitor ambient sounds in real time and detect abnormal sounds. This enables emergency response by monitoring ambient sounds and quickly notifying of important warning sounds and sirens. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can use a generation AI to analyze ambient sounds and quickly notify of important warning sounds and sirens.

[0039] The feedback unit can provide feedback based on the emotions analyzed by the emotion recognition unit. For example, the feedback unit can use generative AI to provide appropriate feedback to the user. The generative AI can provide encouraging messages and advice according to the user's emotional state. The feedback unit can also use the generative AI to provide real-time feedback and respond immediately to changes in the user's emotions. This allows the feedback unit to provide appropriate feedback to the user based on the emotions analyzed by the emotion recognition unit. Some or all of the above-described processes in the feedback unit may be performed using generative AI, or they may not. For example, the feedback unit can use generative AI to provide appropriate feedback to the user.

[0040] The transcription unit can automatically recognize and appropriately transcribe technical terms and slang according to the content of the conversation. For example, the transcription unit can use generative AI to analyze the content of the conversation, automatically recognize and transcribe technical terms and slang. The transcription unit can use generative AI to accurately recognize and transcribe medical terms in conversations in medical settings. The transcription unit can also use generative AI to appropriately recognize and transcribe slang and abbreviations in conversations among young people. The transcription unit can also use generative AI to recognize and transcribe industry-specific technical terms in business conversations. This allows for the provision of accurate information by appropriately transcribing technical terms and slang according to the content of the conversation. Some or all of the above-described processes in the transcription unit may be performed using generative AI, or they may not be performed using generative AI. For example, the transcription unit can use generative AI to analyze the content of the conversation, automatically recognize and transcribe technical terms and slang.

[0041] The transcription unit can automatically adjust its transcription speed according to the speed of the conversation. For example, the transcription unit can analyze the speed of the conversation using a generative AI and automatically adjust the transcription speed. If the generative AI is generating a fast conversation, the transcription unit can perform transcription at high speed to maintain real-time performance. If the generative AI is generating a slow conversation, the transcription unit can also perform transcription slowly to prioritize accuracy. The transcription unit can also adaptively adjust its speed when the speed of the conversation fluctuates and perform transcription. This allows for accurate transcription while maintaining real-time performance by adjusting the transcription speed according to the speed of the conversation. Some or all of the above processing in the transcription unit may be performed using a generative AI or not. For example, the transcription unit can analyze the speed of the conversation using a generative AI and automatically adjust the transcription speed.

[0042] The transcription unit can filter out background noise from conversations and transcribe only the important parts. For example, the transcription unit can use generative AI to analyze background noise from conversations and transcribe only the important parts. The transcription unit can use generative AI to filter out background noise in noisy environments and transcribe only the important parts. The transcription unit can also use generative AI to prioritize the voices of speakers during meetings. The transcription unit can also use generative AI to prioritize the content of conversations in noisy places such as cafes. By filtering out background noise from conversations, only the important parts can be accurately transcribed. Some or all of the above processing in the transcription unit may be performed using generative AI or not. For example, the transcription unit can use generative AI to analyze background noise from conversations and transcribe only the important parts.

[0043] The transcription unit can understand the context of a conversation and transcribe it by appropriately converting synonyms and related words. For example, the transcription unit can use generative AI to analyze the context of a conversation and transcribe it by appropriately converting synonyms and related words. The transcription unit can use generative AI to understand the context of a conversation and transcribe it by appropriately converting synonyms. The transcription unit can also use generative AI to appropriately convert related words in order to maintain the flow of the conversation. The transcription unit can also use generative AI to convert synonyms and related words according to the context in order to accurately convey the meaning of the conversation. This makes accurate transcription possible by understanding the context of the conversation and appropriately converting synonyms and related words. Some or all of the above processing in the transcription unit may be performed using generative AI or not. For example, the transcription unit can use generative AI to analyze the context of a conversation and transcribe it by appropriately converting synonyms and related words.

[0044] The summarization unit can determine the priority of summaries based on the importance of the conversation. For example, the summarization unit can use generative AI to analyze the importance of the conversation and determine the priority of summaries. The summarization unit can summarize by prioritizing important conversation content using generative AI. The summarization unit can also summarize by using generative AI to summarize based on keywords that frequently appear in the conversation. The summarization unit can also summarize by using generative AI to understand the flow of the conversation and prioritize important points. This allows for the priority provision of important information by determining the priority of summaries based on the importance of the conversation. Some or all of the above processing in the summarization unit may be performed using generative AI or not. For example, the summarization unit can use generative AI to analyze the importance of the conversation and determine the priority of summaries.

[0045] The summarization unit can apply different summarization algorithms depending on the category of the conversation. For example, the summarization unit can use generative AI to analyze the category of the conversation and apply a different summarization algorithm. The summarization unit can use generative AI to apply a business-specific summarization algorithm to business conversations. The summarization unit can also use generative AI to apply a daily-use summarization algorithm to daily-use conversations. The summarization unit can also use generative AI to apply a medical-use summarization algorithm to medical settings. By applying different summarization algorithms depending on the category of the conversation, a more appropriate summary is provided. Some or all of the above processing in the summarization unit may be performed using generative AI or not. For example, the summarization unit can use generative AI to analyze the category of the conversation and apply a different summarization algorithm.

[0046] The summarization unit can automatically adjust the length of the summary according to the length of the conversation. For example, the summarization unit can analyze the length of the conversation using a generative AI and automatically adjust the length of the summary. The summarization unit can use the generative AI to provide a concise summary that captures the main points in long conversations. The summarization unit can also use the generative AI to provide a detailed summary in short conversations. The summarization unit can also adaptively adjust the length of the summary when the length of the conversation fluctuates. This ensures that an appropriate summary is provided by adjusting the length of the summary according to the length of the conversation. Some or all of the above processing in the summarization unit may be performed using a generative AI or not. For example, the summarization unit can use a generative AI to analyze the length of the conversation and automatically adjust the length of the summary.

[0047] The summarization unit can adjust the order of summaries based on the relevance of the conversation. For example, the summarization unit can use generative AI to analyze the relevance of the conversation and adjust the order of the summaries. The summarization unit can have the generative AI prioritize important conversational content and display it at the beginning of the summary. The summarization unit can also have the generative AI understand the flow of the conversation and summarize highly relevant information in an orderly manner. The summarization unit can also have the generative AI adjust the order of summaries based on keywords that frequently appear in the conversation. This allows for the priority provision of important information by adjusting the order of summaries based on the relevance of the conversation. Some or all of the above processing in the summarization unit may be performed using generative AI or not. For example, the summarization unit can use generative AI to analyze the relevance of the conversation and adjust the order of the summaries.

[0048] The emotion recognition unit can analyze the tone of voice and facial expressions of the person it is talking to and track changes in emotion in real time. For example, the emotion recognition unit can use a generative AI to analyze the tone of voice and facial expressions of the person it is talking to and track changes in emotion in real time. The emotion recognition unit can also use a generative AI to analyze the facial expressions of the person it is talking to and track changes in emotion in real time. The emotion recognition unit can also use a generative AI to analyze the tone of voice and facial expressions of the person it is talking to and track changes in emotion in real time. This allows for more accurate emotion recognition by analyzing the tone of voice and facial expressions of the person it is talking to and tracking changes in emotion in real time. Some or all of the above processing in the emotion recognition unit may be performed using a generative AI or not. For example, the emotion recognition unit can use a generative AI to analyze the tone of voice and facial expressions of the person it is talking to and track changes in emotion in real time.

[0049] The emotion recognition unit can understand the context of a conversation and analyze the emotional intent more accurately. For example, the emotion recognition unit can use generative AI to analyze the context of a conversation and analyze the emotional intent more accurately. The emotion recognition unit can enable the generative AI to understand the context of a conversation and analyze the emotional intent accurately. The emotion recognition unit can also enable the generative AI to accurately analyze the emotional intent in order to maintain the flow of the conversation. The emotion recognition unit can also enable the generative AI to accurately convey the meaning of the conversation in order to analyze the emotional intent according to the context. This allows for more appropriate emotion recognition by understanding the context of the conversation and analyzing the emotional intent more accurately. Some or all of the above-described processes in the emotion recognition unit may be performed using generative AI or not. For example, the emotion recognition unit can use generative AI to analyze the context of a conversation and analyze the emotional intent more accurately.

[0050] The emotion recognition unit can improve the accuracy of emotion recognition by analyzing the gestures and body language of the conversation partner. For example, the emotion recognition unit can improve the accuracy of emotion recognition by using generative AI to analyze the gestures and body language of the conversation partner. The emotion recognition unit can also improve the accuracy of emotion recognition by having the generative AI analyze the gestures of the conversation partner. Furthermore, the emotion recognition unit can improve the accuracy of emotion recognition by having the generative AI analyze the body language of the conversation partner. The emotion recognition unit can also improve the accuracy of emotion recognition by having the generative AI analyze a combination of the gestures and body language of the conversation partner. This improves the accuracy of emotion recognition by analyzing the gestures and body language of the conversation partner. Some or all of the above-described processes in the emotion recognition unit may be performed using generative AI, or they may not. For example, the emotion recognition unit can improve the accuracy of emotion recognition by using generative AI to analyze the gestures and body language of the conversation partner.

[0051] The emotion recognition unit can improve the accuracy of emotion recognition by considering background sounds and ambient sounds in the conversation. For example, the emotion recognition unit can improve the accuracy of emotion recognition by using a generative AI to analyze background sounds and ambient sounds in the conversation. The emotion recognition unit can improve the accuracy of emotion recognition by having the generative AI filter out background sounds in the conversation. The emotion recognition unit can also improve the accuracy of emotion recognition by having the generative AI consider ambient sounds in the conversation. The emotion recognition unit can also improve the accuracy of emotion recognition by having the generative AI analyze a combination of background sounds and ambient sounds in the conversation. In this way, the accuracy of emotion recognition is improved by considering background sounds and ambient sounds in the conversation. Some or all of the above processing in the emotion recognition unit may be performed using a generative AI or not. For example, the emotion recognition unit can improve the accuracy of emotion recognition by using a generative AI to analyze background sounds and ambient sounds in the conversation.

[0052] The feedback unit can provide appropriate feedback in real time based on the content of the conversation. For example, the feedback unit can analyze the content of the conversation using generative AI and provide appropriate feedback in real time. The feedback unit can use generative AI to understand the flow of the conversation and provide feedback at the appropriate time. The feedback unit can also use generative AI to grasp the important points of the conversation and provide appropriate feedback in real time. This improves the quality of communication by providing appropriate feedback in real time based on the content of the conversation. Some or all of the above processing in the feedback unit may be performed using generative AI or not. For example, the feedback unit can use generative AI to analyze the content of the conversation and provide appropriate feedback in real time.

[0053] The feedback unit can understand the context of the conversation and optimize the timing of feedback. For example, the feedback unit can use generative AI to analyze the context of the conversation and optimize the timing of feedback. The feedback unit can use generative AI to understand the context of the conversation and provide the optimal timing for feedback. The feedback unit can also optimize the timing of feedback so that the generative AI maintains the flow of the conversation. The feedback unit can also optimize the timing of feedback according to the context so that the generative AI accurately conveys the meaning of the conversation. This allows for more appropriate feedback to be provided by understanding the context of the conversation and optimizing the timing of feedback. Some or all of the above-described processes in the feedback unit may be performed using generative AI or not. For example, the feedback unit can use generative AI to analyze the context of the conversation and optimize the timing of feedback.

[0054] The feedback unit can determine the priority of feedback based on the importance of the conversation. For example, the feedback unit can use generative AI to analyze the importance of the conversation and determine the priority of feedback. The feedback unit can provide feedback prioritizing important conversation content generated by the generative AI. The feedback unit can also provide feedback based on keywords that frequently appear in the conversation generated by the generative AI. The feedback unit can also use generative AI to understand the flow of the conversation and provide feedback prioritizing important points. This allows important information to be provided preferentially by determining the priority of feedback based on the importance of the conversation. Some or all of the above processing in the feedback unit may be performed using generative AI or not. For example, the feedback unit can use generative AI to analyze the importance of the conversation and determine the priority of feedback.

[0055] The feedback unit can apply different feedback algorithms depending on the category of the conversation. For example, the feedback unit can use generative AI to analyze the category of the conversation and apply a different feedback algorithm. The feedback unit can use generative AI to apply a business-oriented feedback algorithm in business conversations. The feedback unit can also use generative AI to apply a daily-use feedback algorithm in everyday conversations. The feedback unit can also use generative AI to apply a medical-use feedback algorithm in medical settings. By applying different feedback algorithms depending on the category of the conversation, more appropriate feedback is provided. Some or all of the above processing in the feedback unit may be performed using generative AI or not. For example, the feedback unit can use generative AI to analyze the category of the conversation and apply a different feedback algorithm.

[0056] The alert unit can analyze ambient sounds in real time and quickly detect important warning sounds. For example, the alert unit can use a generation AI to analyze ambient sounds in real time and quickly detect important warning sounds. The alert unit's generation AI can prioritize the detection of important warning sounds in noisy environments. The alert unit's generation AI can also detect subtle warning sounds in quiet environments. The alert unit's generation AI can analyze ambient sounds in real time and quickly detect important warning sounds. This enables emergency response by analyzing ambient sounds in real time and quickly detecting important warning sounds. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can use a generation AI to analyze ambient sounds in real time and quickly detect important warning sounds.

[0057] The alert unit can determine notification priorities based on the importance of the alerts. For example, the alert unit can use a generation AI to analyze the importance of the alerts and determine notification priorities. The alert unit can have the generation AI prioritize notifications for important warning sounds. The alert unit can also have the generation AI analyze ambient sounds and prioritize notifications for high-importance warning sounds. The alert unit can also have the generation AI determine notification priorities based on importance when multiple warning sounds occur simultaneously. This allows important information to be provided preferentially by determining notification priorities based on the importance of the alerts. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can use a generation AI to analyze the importance of the alerts and determine notification priorities.

[0058] The alert unit can filter ambient noise and notify only important sounds. For example, the alert unit can analyze ambient noise using a generation AI and notify only important sounds. In noisy environments, the alert unit can filter background noise using a generation AI and notify only important sounds. In quiet environments, the alert unit can also detect even subtle sounds and notify only important sounds. The alert unit can analyze ambient noise using a generation AI and notify only important sounds. This allows for accurate notification of only important sounds by filtering ambient noise. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can analyze ambient noise using a generation AI and notify only important sounds.

[0059] The alert unit can apply different notification methods depending on the alert category. For example, the alert unit can use a generation AI to analyze the alert category and apply different notification methods. The alert unit can use the generation AI to provide a fast and noticeable notification method for urgent alerts. The alert unit can also use the generation AI to provide a milder notification method for general alerts. The alert unit can also use the generation AI to provide an appropriate notification method for alerts of a specific category. This allows for more appropriate notifications by applying different notification methods depending on the alert category. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can use a generation AI to analyze the alert category and apply different notification methods.

[0060] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0061] The KikoYell system can also be equipped with an activity monitoring unit that monitors the user's activity status and provides feedback at the appropriate time. The activity monitoring unit can, for example, detect different activity statuses such as when the user is exercising or resting, and provide feedback accordingly. During exercise, it can notify the user of encouraging messages and exercise progress. During rest, it can also provide messages to encourage relaxation and suggest stretches. Furthermore, the activity monitoring unit can record the user's activity history and use it for long-term health management. This provides appropriate feedback according to the user's activity status, improving the quality of daily life.

[0062] The KikoYell system can also include a schedule management unit that manages the user's schedule and reminds them of important appointments. For example, the schedule management unit can integrate with the user's calendar app to provide reminders before important meetings and events. It can also notify users in advance of key points during meetings to ensure they don't miss important details. Furthermore, the schedule management unit can provide reminders to encourage breaks at appropriate times based on the user's schedule. This allows users to efficiently manage their schedules and avoid missing important appointments.

[0063] The Kiko-Yell system can also be equipped with a health monitoring unit that monitors the user's health status and issues alerts if abnormalities are detected. For example, the health monitoring unit can measure the user's heart rate and blood pressure in real time and quickly notify the user if abnormalities are detected. It can also provide messages encouraging relaxation if the user is experiencing stress. Furthermore, the health monitoring unit can record the user's health data over the long term and share it with medical institutions. This allows for constant monitoring of the user's health status and enables appropriate responses.

[0064] The KikoYell system can also include a learning support unit that monitors the user's learning progress and provides appropriate learning plans. For example, the learning support unit can provide challenging tasks when the user is focused and easy review questions when the user is tired. It can also recommend relevant learning resources when the user wants to acquire new knowledge. Furthermore, the learning support unit can improve learning motivation by recording the user's learning history and visualizing their progress. This ensures that appropriate learning plans are provided according to the user's learning situation, improving learning effectiveness.

[0065] The Kiko-Yell system can also be equipped with a sleep support unit that monitors the user's sleep state and provides an appropriate sleep environment. For example, the sleep support unit can analyze the user's sleep patterns and suggest the optimal sleep duration. It can also provide relaxing music and lighting to help the user relax and fall asleep. Furthermore, the sleep support unit can record the user's sleep data and create a long-term sleep improvement plan. This ensures that an appropriate sleep environment is provided according to the user's sleep state, improving sleep quality.

[0066] The Kiko-Yell system can also include a stress management unit that monitors the user's stress level and provides appropriate stress management methods. For example, the stress management unit analyzes the user's heart rate and breathing patterns to estimate their stress level. If the user is experiencing stress, it can suggest relaxing breathing techniques or meditation. Furthermore, the stress management unit can record the user's stress history and create a long-term stress management plan. This ensures that appropriate stress management methods are provided according to the user's stress level, thereby reducing stress.

[0067] The following briefly describes the processing flow for example form 1.

[0068] Step 1: The transcription unit transcribes the conversation in real time. For example, it uses generative AI to convert speech into text and speech recognition technology to accurately transcribe the content of the conversation. It can also identify and transcribe the voices of the speakers. Step 2: The summarization unit summarizes the conversation transcribed by the transcription unit. For example, it might use generative AI to summarize a long conversation and display only the important points concisely. It might use natural language processing technology to extract the main points of the conversation, understand the context, and prioritize and summarize the most important information. Step 3: The emotion recognition unit analyzes the emotions of the conversation partner based on the information summarized by the summarization unit. For example, it uses generative AI to analyze the tone of voice and facial expressions of the conversation partner and conveys their emotions and intentions to the user. It can analyze the pitch and volume of the voice to detect changes in emotion. It can also analyze the movement of facial feature points to estimate emotions from facial expressions. Step 4: The feedback unit provides feedback based on the emotions analyzed by the emotion recognition unit. For example, it uses generative AI to provide appropriate feedback to the user, offering encouraging messages and advice according to the user's emotional state. It can also provide feedback in real time and respond immediately to changes in the user's emotions. Step 5: The alert unit monitors ambient sounds and notifies of important warning sounds. For example, it uses a generation AI to analyze ambient sounds and quickly notifies of important warning sounds and sirens. It can analyze the type and volume of sounds to detect emergency warning sounds. It can also monitor ambient sounds in real time and detect abnormal sounds.

[0069] (Example of form 2) The Kiko-Yell System, according to an embodiment of the present invention, is a system that supports daily life while assisting hearing by utilizing hearing aids, smartphones, and generative AI. The Kiko-Yell System provides various functions to enable users to enjoy communication comfortably and with peace of mind. The Kiko-Yell System primarily solves communication problems for people with hearing difficulties (such as the elderly, the hearing impaired, and workers in noisy environments). For example, the Kiko-Yell System provides a function to analyze long conversations in everyday life and business settings in real time and summarize the main points concisely. This allows for efficient capture of important information that might otherwise be missed. Furthermore, the Kiko-Yell System improves the quality of communication by using emotion recognition to infer the emotions of the person being spoken to and providing feedback based on the analysis of voice and facial expressions. In addition, the Kiko-Yell System makes important sounds easier to hear even in noisy environments through advanced voice filtering using generative AI. It also provides an alert function that quickly detects abnormal sounds and warning sounds in emergencies and notifies the user. The Kiko-Yell System provides the following functions: 1. Conversation transcription function: Transcribes conversations in real time and displays them on the smartphone. 1. This allows users to review parts they missed or had difficulty hearing in text form. 2. Conversation summarization function: The generating AI summarizes long conversations, displaying only the important points concisely. This allows for efficient information acquisition. 3. Emotion recognition function: The generating AI analyzes the tone of voice and facial expressions of the person being spoken to, conveying their emotions and intentions to the user. 4. Emergency alert function: The generating AI monitors ambient noise and quickly notifies the user of important warning sounds and sirens. These functions enable the KikoYell system to achieve not only "hearing" but also "clearly understood" communication. For example, using the conversation transcription function, users can review parts they missed during a meeting on their smartphone. The emotion recognition function makes it easier to understand the emotions of the person being spoken to, enabling better communication. The emergency alert function ensures that important warning sounds are not missed even in noisy environments. In this way, the KikoYell system supports the user's hearing and makes daily life more comfortable.

[0070] The Kiko-Yell system according to this embodiment comprises a transcription unit, a summarization unit, an emotion recognition unit, a feedback unit, and an alert unit. The transcription unit transcribes conversations in real time. The transcription unit converts speech to text using, for example, a generative AI. The transcription unit can accurately transcribe the content of conversations using speech recognition technology with the generative AI. The transcription unit can also transcribe by analyzing the characteristics of speech and identifying the voices of the speakers using the generative AI. The summarization unit summarizes the conversations transcribed by the transcription unit. The summarization unit can summarize long conversations using, for example, a generative AI and concisely display only the important points. The summarization unit can extract and summarize the main points of a conversation using natural language processing technology with the generative AI. The summarization unit can also prioritize summarizing important information by understanding the context of the conversation with the generative AI. The emotion recognition unit analyzes the emotions of the conversation partners based on the information summarized by the summarization unit. The emotion recognition unit, for example, uses generative AI to analyze the tone of voice and facial expressions of the person it is talking to, and conveys emotions and intentions to the user. The emotion recognition unit can detect changes in emotion by having the generative AI analyze the pitch and volume of the voice. The emotion recognition unit can also estimate emotions from facial expressions by having the generative AI analyze the movement of facial feature points. The feedback unit provides feedback based on the emotions analyzed by the emotion recognition unit. The feedback unit, for example, uses generative AI to provide appropriate feedback to the user. The feedback unit can provide encouraging messages and advice according to the user's emotional state using the generative AI. The feedback unit can also provide feedback in real time using the generative AI and respond immediately to changes in the user's emotions. The alert unit monitors ambient sounds and notifies important warning sounds. The alert unit, for example, uses generative AI to analyze ambient sounds and quickly notifies important warning sounds and sirens. The alert unit can detect emergency warning sounds by having the generative AI analyze the type and volume of sounds. The alert unit can also use a generating AI to monitor ambient sounds in real time and detect abnormal sounds. As a result, the Kiko-Yell system according to this embodiment can transcribe conversations, summarize them, recognize emotions, provide feedback, and send alert notifications.

[0071] The transcription unit transcribes conversations in real time. For example, it uses generative AI to convert speech to text. Specifically, the generative AI utilizes advanced speech recognition technology to analyze the audio signal and accurately transcribe the spoken content. The generative AI analyzes the frequency spectrum of the speech to identify phonemes and words. Furthermore, the generative AI learns the characteristics of each speaker's voice, allowing it to identify and transcribe individual utterances even when multiple speakers are present. For example, in meetings or interviews, each speaker's utterance is recorded as a separate text, making later review and analysis easier. The generative AI also incorporates noise reduction technology, effectively removing background noise and other unwanted sounds to obtain clear audio data. This improves transcription accuracy and reduces the risk of misrecognition. Additionally, the generative AI considers context during the speech recognition process and automatically corrects homonyms and grammatical errors. This results in natural and easy-to-read generated text. Furthermore, the transcription unit provides an interface that displays the generated text in real time, allowing users to immediately review the content. This allows you to grasp important points without missing anything, even as the conversation is in progress.

[0072] The summarization unit summarizes conversations transcribed by the transcription unit. For example, the summarization unit uses generative AI to summarize long conversations, displaying only the most important points concisely. Specifically, the generative AI uses natural language processing technology to extract important information from text data and generate a summary. The generative AI understands the context of the conversation and identifies key keywords and phrases. For example, when summarizing meeting minutes, the generative AI prioritizes extracting agenda items, decisions, and action items to create a concise summary. The generative AI analyzes the text structure and omits redundant parts and repetitive expressions to improve the accuracy of the summary. Furthermore, the generative AI can adjust the level of detail in the summary according to the user's needs. For example, if a detailed summary is required, it generates a summary containing more information; if a concise summary is needed, it extracts only the most important points. In addition, the summarization unit provides a visual interface to display the generated summary, making it easy for users to understand. This allows for efficient comprehension of long conversations and large amounts of text data, and quick identification of important information.

[0073] The emotion recognition unit analyzes the emotions of the conversation partner based on the information summarized by the summarization unit. For example, the emotion recognition unit uses generative AI to analyze the tone of voice and facial expressions of the conversation partner and conveys their emotions and intentions to the user. Specifically, the generative AI analyzes features such as pitch, volume, and rhythm of the voice to detect changes in emotion. For example, a higher tone of voice may indicate excitement or joy, while a lower tone may indicate calmness or sadness. The generative AI learns these voice features and identifies patterns of emotion. The generative AI also analyzes the movement of facial feature points and estimates emotions from facial expressions. For example, it analyzes eyebrow movements and the degree to which the corners of the mouth are turned up to identify emotions such as joy, anger, and sadness. The generative AI integrates this information and can analyze the emotional state of the conversation partner in real time. Furthermore, the emotion recognition unit provides an interface that visually displays the analysis results, allowing the user to intuitively understand the emotions of the conversation partner. This enables the user to grasp the other person's emotions during a conversation and take appropriate action.

[0074] The feedback unit provides feedback based on emotions analyzed by the emotion recognition unit. For example, the feedback unit uses generative AI to provide appropriate feedback to the user. Specifically, the generative AI provides encouraging messages and advice according to the user's emotional state. For instance, if the user is stressed, the generative AI provides advice and words of encouragement to help them relax. If the user is feeling joy or excitement, the generative AI shares that emotion and provides positive feedback. The generative AI can monitor changes in the user's emotions in real time and respond immediately. Furthermore, the feedback unit collects user feedback and uses it as training data for the generative AI. This allows the generative AI to learn the user's preferences and tendencies, enabling it to provide more personalized feedback. The feedback unit combines visual and auditory feedback to communicate effectively with the user. This allows the user to receive appropriate support tailored to their emotional state.

[0075] The alert unit monitors ambient sounds and notifies users of important warning sounds. For example, it uses a generative AI to analyze ambient sounds and quickly notify users of important warning sounds and sirens. Specifically, the generative AI analyzes the type and volume of sounds to detect emergency warning sounds. For example, it identifies highly urgent sounds such as fire alarms and ambulance sirens and notifies the user. By analyzing the frequency spectrum of sounds and identifying specific patterns, the generative AI can quickly detect important warning sounds. Furthermore, the generative AI can monitor ambient sounds in real time and detect abnormal sounds. For example, if an unusual sound occurs, the generative AI analyzes the sound and notifies the user of the potential anomaly. In addition, the alert unit can select the notification method according to the user's situation. For example, it can select the method that the user can receive most effectively, such as visual, audio, or vibration notifications. This allows the alert unit to provide users with important information quickly and reliably, supporting their response in emergencies.

[0076] The transcription unit can transcribe conversations in real time and display them on a smartphone. For example, the transcription unit can use generative AI to convert audio to text and display it on a smartphone. The transcription unit can use generative AI and speech recognition technology to accurately transcribe the content of conversations and display it on a smartphone. The transcription unit can also use generative AI to analyze the characteristics of the audio, identify the speaker's voice, transcribe it, and display it on a smartphone. This prevents missed information by transcribing conversations in real time and displaying them on a smartphone. Some or all of the above-described processes in the transcription unit may be performed using generative AI or not. For example, the transcription unit can use generative AI to convert audio to text and display it on a smartphone.

[0077] The summarization section can summarize long conversations and display only the important points concisely. For example, the summarization section can use generative AI to summarize long conversations and display only the important points concisely. The summarization section can use generative AI to extract and summarize the main points of a conversation using natural language processing techniques. The summarization section can also use generative AI to understand the context of the conversation and prioritize the summarization of important information. This allows for efficient information acquisition by summarizing long conversations and displaying only the important points concisely. Some or all of the above processing in the summarization section may be performed using generative AI or not. For example, the summarization section can use generative AI to summarize long conversations and display only the important points concisely.

[0078] The emotion recognition unit can analyze the tone of voice and facial expressions of the person it is talking to and convey their emotions and intentions to the user. For example, the emotion recognition unit can use generative AI to analyze the tone of voice and facial expressions of the person it is talking to and convey their emotions and intentions to the user. The emotion recognition unit can use generative AI to analyze the pitch and volume of speech and detect changes in emotion. The emotion recognition unit can also use generative AI to analyze the movement of facial feature points and estimate emotions from facial expressions. This improves the quality of communication by analyzing the emotions and intentions of the person it is talking to and conveying them to the user. Some or all of the above processing in the emotion recognition unit may be performed using generative AI or not. For example, the emotion recognition unit can use generative AI to analyze the tone of voice and facial expressions of the person it is talking to and convey their emotions and intentions to the user.

[0079] The alert unit can monitor ambient sounds and quickly notify of important warning sounds and sirens. For example, the alert unit can use a generation AI to analyze ambient sounds and quickly notify of important warning sounds and sirens. The alert unit's generation AI can analyze the type and volume of sounds and detect emergency warning sounds. The alert unit's generation AI can also monitor ambient sounds in real time and detect abnormal sounds. This enables emergency response by monitoring ambient sounds and quickly notifying of important warning sounds and sirens. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can use a generation AI to analyze ambient sounds and quickly notify of important warning sounds and sirens.

[0080] The feedback unit can provide feedback based on the emotions analyzed by the emotion recognition unit. For example, the feedback unit can use generative AI to provide appropriate feedback to the user. The generative AI can provide encouraging messages and advice according to the user's emotional state. The feedback unit can also use the generative AI to provide real-time feedback and respond immediately to changes in the user's emotions. This allows the feedback unit to provide appropriate feedback to the user based on the emotions analyzed by the emotion recognition unit. Some or all of the above-described processes in the feedback unit may be performed using generative AI, or they may not. For example, the feedback unit can use generative AI to provide appropriate feedback to the user.

[0081] The transcription unit can estimate the user's emotions and adjust the accuracy of the transcription based on the estimated emotions. For example, the transcription unit can use a generative AI to estimate the user's emotions and adjust the accuracy of the transcription based on the estimated emotions. The transcription unit can analyze the user's emotional state using the generative AI and transcribe carefully if the user is tense to reduce misrecognition. The transcription unit can also transcribe at a natural speed to maintain the flow of conversation if the generative AI is relaxed. The transcription unit can also transcribe at high speed if the generative AI is in a hurry, prioritizing real-time performance. By adjusting the accuracy of the transcription based on the user's emotions, more accurate transcription becomes possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the transcription unit may be performed using or without the generative AI. For example, the transcription unit can use a generation AI to estimate the user's emotions and adjust the accuracy of the transcription based on the estimated emotions.

[0082] The transcription unit can automatically recognize and appropriately transcribe technical terms and slang according to the content of the conversation. For example, the transcription unit can use generative AI to analyze the content of the conversation, automatically recognize and transcribe technical terms and slang. The transcription unit can use generative AI to accurately recognize and transcribe medical terms in conversations in medical settings. The transcription unit can also use generative AI to appropriately recognize and transcribe slang and abbreviations in conversations among young people. The transcription unit can also use generative AI to recognize and transcribe industry-specific technical terms in business conversations. This allows for the provision of accurate information by appropriately transcribing technical terms and slang according to the content of the conversation. Some or all of the above-described processes in the transcription unit may be performed using generative AI, or they may not be performed using generative AI. For example, the transcription unit can use generative AI to analyze the content of the conversation, automatically recognize and transcribe technical terms and slang.

[0083] The transcription unit can automatically adjust its transcription speed according to the speed of the conversation. For example, the transcription unit can analyze the speed of the conversation using a generative AI and automatically adjust the transcription speed. If the generative AI is generating a fast conversation, the transcription unit can perform transcription at high speed to maintain real-time performance. If the generative AI is generating a slow conversation, the transcription unit can also perform transcription slowly to prioritize accuracy. The transcription unit can also adaptively adjust its speed when the speed of the conversation fluctuates and perform transcription. This allows for accurate transcription while maintaining real-time performance by adjusting the transcription speed according to the speed of the conversation. Some or all of the above processing in the transcription unit may be performed using a generative AI or not. For example, the transcription unit can analyze the speed of the conversation using a generative AI and automatically adjust the transcription speed.

[0084] The transcription unit can estimate the user's emotions and adjust the display method of the transcript based on the estimated emotions. For example, the transcription unit can use a generative AI to estimate the user's emotions and adjust the display method of the transcript based on the estimated emotions. The transcription unit can use the generative AI to analyze the user's emotional state and provide a simple and highly visible display method if the user is tense. The transcription unit can also provide a display method that includes detailed information if the user is relaxed. The transcription unit can also provide a display method that gets to the point if the user is in a hurry. By adjusting the display method of the transcript based on the user's emotions, a highly visible display becomes possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the transcription unit may be performed using a generative AI or not using a generative AI. For example, the transcription unit can use a generation AI to estimate the user's emotions and adjust the way the transcription is displayed based on the estimated emotions.

[0085] The transcription unit can filter out background noise from conversations and transcribe only the important parts. For example, the transcription unit can use generative AI to analyze background noise from conversations and transcribe only the important parts. The transcription unit can use generative AI to filter out background noise in noisy environments and transcribe only the important parts. The transcription unit can also use generative AI to prioritize the voices of speakers during meetings. The transcription unit can also use generative AI to prioritize the content of conversations in noisy places such as cafes. By filtering out background noise from conversations, only the important parts can be accurately transcribed. Some or all of the above processing in the transcription unit may be performed using generative AI or not. For example, the transcription unit can use generative AI to analyze background noise from conversations and transcribe only the important parts.

[0086] The transcription unit can understand the context of a conversation and transcribe it by appropriately converting synonyms and related words. For example, the transcription unit can use generative AI to analyze the context of a conversation and transcribe it by appropriately converting synonyms and related words. The transcription unit can use generative AI to understand the context of a conversation and transcribe it by appropriately converting synonyms. The transcription unit can also use generative AI to appropriately convert related words in order to maintain the flow of the conversation. The transcription unit can also use generative AI to convert synonyms and related words according to the context in order to accurately convey the meaning of the conversation. This makes accurate transcription possible by understanding the context of the conversation and appropriately converting synonyms and related words. Some or all of the above processing in the transcription unit may be performed using generative AI or not. For example, the transcription unit can use generative AI to analyze the context of a conversation and transcribe it by appropriately converting synonyms and related words.

[0087] The summarization unit can estimate the user's emotions and adjust the level of detail in the summary based on the estimated emotions. For example, the summarization unit can use a generative AI to estimate the user's emotions and adjust the level of detail in the summary based on the estimated emotions. The summarization unit can use a generative AI to analyze the user's emotional state and provide a detailed summary if the user is relaxed. The summarization unit can also provide a concise summary if the generative AI is in a hurry. The summarization unit can also provide a visually stimulating summary if the generative AI is excited. By adjusting the level of detail in the summary based on the user's emotions, a more appropriate summary is provided. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or a generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the summarization unit may be performed using a generative AI or not. For example, the summarization unit can use a generative AI to estimate the user's emotions and adjust the level of detail in the summary based on the estimated emotions.

[0088] The summarization unit can determine the priority of summaries based on the importance of the conversation. For example, the summarization unit can use generative AI to analyze the importance of the conversation and determine the priority of summaries. The summarization unit can summarize by prioritizing important conversation content using generative AI. The summarization unit can also summarize by using generative AI to summarize based on keywords that frequently appear in the conversation. The summarization unit can also summarize by using generative AI to understand the flow of the conversation and prioritize important points. This allows for the priority provision of important information by determining the priority of summaries based on the importance of the conversation. Some or all of the above processing in the summarization unit may be performed using generative AI or not. For example, the summarization unit can use generative AI to analyze the importance of the conversation and determine the priority of summaries.

[0089] The summarization unit can apply different summarization algorithms depending on the category of the conversation. For example, the summarization unit can use generative AI to analyze the category of the conversation and apply a different summarization algorithm. The summarization unit can use generative AI to apply a business-specific summarization algorithm to business conversations. The summarization unit can also use generative AI to apply a daily-use summarization algorithm to daily-use conversations. The summarization unit can also use generative AI to apply a medical-use summarization algorithm to medical settings. By applying different summarization algorithms depending on the category of the conversation, a more appropriate summary is provided. Some or all of the above processing in the summarization unit may be performed using generative AI or not. For example, the summarization unit can use generative AI to analyze the category of the conversation and apply a different summarization algorithm.

[0090] The summarization unit can estimate the user's emotions and adjust the way the summary is displayed based on the estimated emotions. For example, the summarization unit can use a generative AI to estimate the user's emotions and adjust the way the summary is displayed based on the estimated emotions. The summarization unit can use the generative AI to analyze the user's emotional state and provide a simple and highly visible display if the user is tense. The summarization unit can also provide a display that includes detailed information if the user is relaxed. The summarization unit can also provide a concise display if the user is in a hurry. By adjusting the way the summary is displayed based on the user's emotions, a highly visible display is possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to such examples. Some or all of the above processing in the summarization unit may be performed using a generative AI or not. For example, the summarization unit can use a generative AI to estimate the user's emotions and adjust the way the summary is displayed based on the estimated emotions.

[0091] The summarization unit can automatically adjust the length of the summary according to the length of the conversation. For example, the summarization unit can analyze the length of the conversation using a generative AI and automatically adjust the length of the summary. The summarization unit can use the generative AI to provide a concise summary that captures the main points in long conversations. The summarization unit can also use the generative AI to provide a detailed summary in short conversations. The summarization unit can also adaptively adjust the length of the summary when the length of the conversation fluctuates. This ensures that an appropriate summary is provided by adjusting the length of the summary according to the length of the conversation. Some or all of the above processing in the summarization unit may be performed using a generative AI or not. For example, the summarization unit can use a generative AI to analyze the length of the conversation and automatically adjust the length of the summary.

[0092] The summarization unit can adjust the order of summaries based on the relevance of the conversation. For example, the summarization unit can use generative AI to analyze the relevance of the conversation and adjust the order of the summaries. The summarization unit can have the generative AI prioritize important conversational content and display it at the beginning of the summary. The summarization unit can also have the generative AI understand the flow of the conversation and summarize highly relevant information in an orderly manner. The summarization unit can also have the generative AI adjust the order of summaries based on keywords that frequently appear in the conversation. This allows for the priority provision of important information by adjusting the order of summaries based on the relevance of the conversation. Some or all of the above processing in the summarization unit may be performed using generative AI or not. For example, the summarization unit can use generative AI to analyze the relevance of the conversation and adjust the order of the summaries.

[0093] The emotion recognition unit can estimate the user's emotions and adjust the accuracy of emotion recognition based on the estimated emotions. For example, the emotion recognition unit can estimate the user's emotions using a generative AI and adjust the accuracy of emotion recognition based on the estimated emotions. The emotion recognition unit can analyze the user's emotional state using the generative AI and perform emotion recognition carefully if the user is tense, thereby reducing misrecognition. If the generative AI is relaxed, the emotion recognition unit can perform emotion recognition at a natural speed, maintaining the flow of conversation. If the generative AI is in a hurry, the emotion recognition unit can perform emotion recognition at high speed, prioritizing real-time performance. By adjusting the accuracy of emotion recognition based on the user's emotions, more accurate emotion recognition becomes possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the emotion recognition unit may be performed using or without the generative AI. For example, the emotion recognition unit can use generative AI to estimate the user's emotions and adjust the accuracy of emotion recognition based on the estimated emotions.

[0094] The emotion recognition unit can analyze the tone of voice and facial expressions of the person it is talking to and track changes in emotion in real time. For example, the emotion recognition unit can use a generative AI to analyze the tone of voice and facial expressions of the person it is talking to and track changes in emotion in real time. The emotion recognition unit can also use a generative AI to analyze the facial expressions of the person it is talking to and track changes in emotion in real time. The emotion recognition unit can also use a generative AI to analyze the tone of voice and facial expressions of the person it is talking to and track changes in emotion in real time. This allows for more accurate emotion recognition by analyzing the tone of voice and facial expressions of the person it is talking to and tracking changes in emotion in real time. Some or all of the above processing in the emotion recognition unit may be performed using a generative AI or not. For example, the emotion recognition unit can use a generative AI to analyze the tone of voice and facial expressions of the person it is talking to and track changes in emotion in real time.

[0095] The emotion recognition unit can understand the context of a conversation and analyze the emotional intent more accurately. For example, the emotion recognition unit can use generative AI to analyze the context of a conversation and analyze the emotional intent more accurately. The emotion recognition unit can enable the generative AI to understand the context of a conversation and analyze the emotional intent accurately. The emotion recognition unit can also enable the generative AI to accurately analyze the emotional intent in order to maintain the flow of the conversation. The emotion recognition unit can also enable the generative AI to accurately convey the meaning of the conversation in order to analyze the emotional intent according to the context. This allows for more appropriate emotion recognition by understanding the context of the conversation and analyzing the emotional intent more accurately. Some or all of the above-described processes in the emotion recognition unit may be performed using generative AI or not. For example, the emotion recognition unit can use generative AI to analyze the context of a conversation and analyze the emotional intent more accurately.

[0096] The emotion recognition unit can estimate the user's emotions and adjust the method of displaying the emotion recognition results based on the estimated user emotions. For example, the emotion recognition unit can estimate the user's emotions using a generative AI and adjust the method of displaying the emotion recognition results based on the estimated emotions. The emotion recognition unit can have the generative AI analyze the user's emotional state and provide a simple and highly visible display method if the user is tense. The emotion recognition unit can also provide a display method that includes detailed information if the generative AI is relaxed. The emotion recognition unit can also provide a concise display method if the generative AI is in a hurry. By adjusting the method of displaying the emotion recognition results based on the user's emotions, a highly visible display becomes possible. Emotion estimation is achieved using an emotion estimation function, for example, with an emotion engine or a generative AI. The generative AI is, but is not limited to, a text generation AI (e.g., LLM) or a multimodal generation AI. Some or all of the above-described processing in the emotion recognition unit may be performed using a generative AI or not. For example, the emotion recognition unit can use generative AI to estimate the user's emotions and adjust the method of displaying the emotion recognition results based on the estimated emotions.

[0097] The emotion recognition unit can improve the accuracy of emotion recognition by analyzing the gestures and body language of the conversation partner. For example, the emotion recognition unit can improve the accuracy of emotion recognition by using generative AI to analyze the gestures and body language of the conversation partner. The emotion recognition unit can also improve the accuracy of emotion recognition by having the generative AI analyze the gestures of the conversation partner. Furthermore, the emotion recognition unit can improve the accuracy of emotion recognition by having the generative AI analyze the body language of the conversation partner. The emotion recognition unit can also improve the accuracy of emotion recognition by having the generative AI analyze a combination of the gestures and body language of the conversation partner. This improves the accuracy of emotion recognition by analyzing the gestures and body language of the conversation partner. Some or all of the above-described processes in the emotion recognition unit may be performed using generative AI, or they may not. For example, the emotion recognition unit can improve the accuracy of emotion recognition by using generative AI to analyze the gestures and body language of the conversation partner.

[0098] The emotion recognition unit can improve the accuracy of emotion recognition by considering background sounds and ambient sounds in the conversation. For example, the emotion recognition unit can improve the accuracy of emotion recognition by using a generative AI to analyze background sounds and ambient sounds in the conversation. The emotion recognition unit can improve the accuracy of emotion recognition by having the generative AI filter out background sounds in the conversation. The emotion recognition unit can also improve the accuracy of emotion recognition by having the generative AI consider ambient sounds in the conversation. The emotion recognition unit can also improve the accuracy of emotion recognition by having the generative AI analyze a combination of background sounds and ambient sounds in the conversation. In this way, the accuracy of emotion recognition is improved by considering background sounds and ambient sounds in the conversation. Some or all of the above processing in the emotion recognition unit may be performed using a generative AI or not. For example, the emotion recognition unit can improve the accuracy of emotion recognition by using a generative AI to analyze background sounds and ambient sounds in the conversation.

[0099] The feedback unit can estimate the user's emotions and adjust the content of the feedback based on the estimated emotions. For example, the feedback unit can use generative AI to estimate the user's emotions and adjust the content of the feedback based on the estimated emotions. The feedback unit can have the generative AI analyze the user's emotional state and provide relaxing feedback if the user is tense. The feedback unit can also provide detailed feedback if the generative AI is relaxed. The feedback unit can also provide quick and concise feedback if the generative AI is in a hurry. This allows for more appropriate feedback to be provided by adjusting the content of the feedback based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above-described processes in the feedback unit may be performed using or without the generative AI. For example, the feedback unit can use generative AI to estimate the user's emotions and adjust the content of the feedback based on the estimated emotions.

[0100] The feedback unit can provide appropriate feedback in real time based on the content of the conversation. For example, the feedback unit can analyze the content of the conversation using generative AI and provide appropriate feedback in real time. The feedback unit can use generative AI to understand the flow of the conversation and provide feedback at the appropriate time. The feedback unit can also use generative AI to grasp the important points of the conversation and provide appropriate feedback in real time. This improves the quality of communication by providing appropriate feedback in real time based on the content of the conversation. Some or all of the above processing in the feedback unit may be performed using generative AI or not. For example, the feedback unit can use generative AI to analyze the content of the conversation and provide appropriate feedback in real time.

[0101] The feedback unit can understand the context of the conversation and optimize the timing of feedback. For example, the feedback unit can use generative AI to analyze the context of the conversation and optimize the timing of feedback. The feedback unit can use generative AI to understand the context of the conversation and provide the optimal timing for feedback. The feedback unit can also optimize the timing of feedback so that the generative AI maintains the flow of the conversation. The feedback unit can also optimize the timing of feedback according to the context so that the generative AI accurately conveys the meaning of the conversation. This allows for more appropriate feedback to be provided by understanding the context of the conversation and optimizing the timing of feedback. Some or all of the above-described processes in the feedback unit may be performed using generative AI or not. For example, the feedback unit can use generative AI to analyze the context of the conversation and optimize the timing of feedback.

[0102] The feedback unit can estimate the user's emotions and adjust the way feedback is displayed based on the estimated emotions. For example, the feedback unit can use generative AI to estimate the user's emotions and adjust the way feedback is displayed based on the estimated emotions. The feedback unit can use the generative AI to analyze the user's emotional state and provide a simple, highly visible display if the user is tense. The feedback unit can also provide a display that includes detailed information if the user is relaxed. The feedback unit can also provide a concise display if the user is in a hurry. This allows for highly visible displays by adjusting the feedback display based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, such as an emotion engine or generative AI. The generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the feedback unit may be performed using or without the generative AI. For example, the feedback unit can use generative AI to estimate the user's emotions and adjust the way feedback is displayed based on the estimated emotions.

[0103] The feedback unit can determine the priority of feedback based on the importance of the conversation. For example, the feedback unit can use generative AI to analyze the importance of the conversation and determine the priority of feedback. The feedback unit can provide feedback prioritizing important conversation content generated by the generative AI. The feedback unit can also provide feedback based on keywords that frequently appear in the conversation generated by the generative AI. The feedback unit can also use generative AI to understand the flow of the conversation and provide feedback prioritizing important points. This allows important information to be provided preferentially by determining the priority of feedback based on the importance of the conversation. Some or all of the above processing in the feedback unit may be performed using generative AI or not. For example, the feedback unit can use generative AI to analyze the importance of the conversation and determine the priority of feedback.

[0104] The feedback unit can apply different feedback algorithms depending on the category of the conversation. For example, the feedback unit can use generative AI to analyze the category of the conversation and apply a different feedback algorithm. The feedback unit can use generative AI to apply a business-oriented feedback algorithm in business conversations. The feedback unit can also use generative AI to apply a daily-use feedback algorithm in everyday conversations. The feedback unit can also use generative AI to apply a medical-use feedback algorithm in medical settings. By applying different feedback algorithms depending on the category of the conversation, more appropriate feedback is provided. Some or all of the above processing in the feedback unit may be performed using generative AI or not. For example, the feedback unit can use generative AI to analyze the category of the conversation and apply a different feedback algorithm.

[0105] The alert unit can estimate the user's emotions and adjust the alert notification method based on the estimated emotions. For example, the alert unit can use generative AI to estimate the user's emotions and adjust the alert notification method based on the estimated emotions. The alert unit can use generative AI to analyze the user's emotional state and provide a gentle notification method if the user is tense. The alert unit can also provide a detailed notification method if the user is relaxed. The alert unit can also provide a quick and concise notification method if the user is in a hurry. This allows for more appropriate notifications by adjusting the alert notification method based on the user's emotions. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. Generative AI is, but is not limited to, text generation AI (e.g., LLM) or multimodal generation AI. Some or all of the above processing in the alert unit may be performed using generative AI or not. For example, the alert unit can use generative AI to estimate the user's emotions and adjust the alert notification method based on the estimated emotions.

[0106] The alert unit can analyze ambient sounds in real time and quickly detect important warning sounds. For example, the alert unit can use a generation AI to analyze ambient sounds in real time and quickly detect important warning sounds. The alert unit's generation AI can prioritize the detection of important warning sounds in noisy environments. The alert unit's generation AI can also detect subtle warning sounds in quiet environments. The alert unit's generation AI can analyze ambient sounds in real time and quickly detect important warning sounds. This enables emergency response by analyzing ambient sounds in real time and quickly detecting important warning sounds. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can use a generation AI to analyze ambient sounds in real time and quickly detect important warning sounds.

[0107] The alert unit can determine notification priorities based on the importance of the alerts. For example, the alert unit can use a generation AI to analyze the importance of the alerts and determine notification priorities. The alert unit can have the generation AI prioritize notifications for important warning sounds. The alert unit can also have the generation AI analyze ambient sounds and prioritize notifications for high-importance warning sounds. The alert unit can also have the generation AI determine notification priorities based on importance when multiple warning sounds occur simultaneously. This allows important information to be provided preferentially by determining notification priorities based on the importance of the alerts. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can use a generation AI to analyze the importance of the alerts and determine notification priorities.

[0108] The alert unit can estimate the user's emotions and adjust the way the alert is displayed based on the estimated emotions. For example, the alert unit can use a generative AI to estimate the user's emotions and adjust the way the alert is displayed based on the estimated emotions. The alert unit can use the generative AI to analyze the user's emotional state and provide a simple and highly visible display if the user is tense. The alert unit can also provide a display that includes detailed information if the user is relaxed. The alert unit can also provide a concise display if the user is in a hurry. By adjusting the way the alert is displayed based on the user's emotions, highly visible displays are possible. Emotion estimation is achieved using an emotion estimation function, for example, using an emotion engine or generative AI. The generative AI is a text generation AI (e.g., LLM) or a multimodal generation AI, but is not limited to these examples. Some or all of the above processing in the alert unit may be performed using a generative AI or not. For example, the alert unit can use a generative AI to estimate the user's emotions and adjust the way the alert is displayed based on the estimated emotions.

[0109] The alert unit can filter ambient noise and notify only important sounds. For example, the alert unit can analyze ambient noise using a generation AI and notify only important sounds. In noisy environments, the alert unit can filter background noise using a generation AI and notify only important sounds. In quiet environments, the alert unit can also detect even subtle sounds and notify only important sounds. The alert unit can analyze ambient noise using a generation AI and notify only important sounds. This allows for accurate notification of only important sounds by filtering ambient noise. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can analyze ambient noise using a generation AI and notify only important sounds.

[0110] The alert unit can apply different notification methods depending on the alert category. For example, the alert unit can use a generation AI to analyze the alert category and apply different notification methods. The alert unit can use the generation AI to provide a fast and noticeable notification method for urgent alerts. The alert unit can also use the generation AI to provide a milder notification method for general alerts. The alert unit can also use the generation AI to provide an appropriate notification method for alerts of a specific category. This allows for more appropriate notifications by applying different notification methods depending on the alert category. Some or all of the above processing in the alert unit may be performed using a generation AI or not. For example, the alert unit can use a generation AI to analyze the alert category and apply different notification methods.

[0111] The system according to the embodiment is not limited to the example described above, and various modifications are possible, for example, as follows.

[0112] The KikoYell system can also be equipped with an activity monitoring unit that monitors the user's activity status and provides feedback at the appropriate time. The activity monitoring unit can, for example, detect different activity statuses such as when the user is exercising or resting, and provide feedback accordingly. During exercise, it can notify the user of encouraging messages and exercise progress. During rest, it can also provide messages to encourage relaxation and suggest stretches. Furthermore, the activity monitoring unit can record the user's activity history and use it for long-term health management. This provides appropriate feedback according to the user's activity status, improving the quality of daily life.

[0113] The KikoYell system can also include a schedule management unit that manages the user's schedule and reminds them of important appointments. For example, the schedule management unit can integrate with the user's calendar app to provide reminders before important meetings and events. It can also notify users in advance of key points during meetings to ensure they don't miss important details. Furthermore, the schedule management unit can provide reminders to encourage breaks at appropriate times based on the user's schedule. This allows users to efficiently manage their schedules and avoid missing important appointments.

[0114] The Kiko-Yell system can also be equipped with a health monitoring unit that monitors the user's health status and issues alerts if abnormalities are detected. For example, the health monitoring unit can measure the user's heart rate and blood pressure in real time and quickly notify the user if abnormalities are detected. It can also provide messages encouraging relaxation if the user is experiencing stress. Furthermore, the health monitoring unit can record the user's health data over the long term and share it with medical institutions. This allows for constant monitoring of the user's health status and enables appropriate responses.

[0115] The KikoYell system can also include a music recommendation unit that estimates the user's emotions and selects music based on those emotions. For example, the music recommendation unit will recommend relaxing music if the user is relaxed, or uplifting music if the user wants to feel energized. If the user is stressed, it can also provide music that helps reduce stress. Furthermore, the music recommendation unit can create individually customized playlists based on the user's musical preferences and past playback history. This provides music that matches the user's emotions, improving the quality of their daily life.

[0116] The Kiko-Yell system can also be equipped with a lighting adjustment unit that estimates the user's emotions and adjusts the color and brightness of the lighting based on those emotions. For example, the lighting adjustment unit can provide warm, soft lighting when the user is relaxed and bright white lighting when the user wants to concentrate. If the user is feeling stressed, it can also adjust the lighting settings to help reduce stress. Furthermore, the lighting adjustment unit can automatically adjust the lighting according to the user's activity level and the time of day. This provides an optimal lighting environment that matches the user's emotions, improving the quality of daily life.

[0117] The KikoYell system can also be equipped with an exercise suggestion unit that estimates the user's emotions and proposes appropriate exercises based on those emotions. For example, if the user is relaxed, the exercise suggestion unit will suggest light stretching or yoga, and if the user wants to feel energized, it will suggest energetic exercises. If the user is feeling stressed, it can also suggest exercises that help reduce stress. Furthermore, the exercise suggestion unit can create individually customized exercise plans based on the user's exercise history and health status. This allows for the suggestion of appropriate exercises that match the user's emotions, improving health management.

[0118] The KikoYell system can also include a meal suggestion unit that estimates the user's emotions and proposes appropriate meals based on those emotions. For example, the meal suggestion unit might suggest a light meal if the user is relaxed, or an energetic meal if the user wants to feel more energetic. If the user is stressed, it can also suggest meals that help reduce stress. Furthermore, the meal suggestion unit can create individually customized meal plans based on the user's eating history and health condition. This allows for the suggestion of appropriate meals that match the user's emotions, improving health management.

[0119] The KikoYell system can also include a learning support unit that monitors the user's learning progress and provides appropriate learning plans. For example, the learning support unit can provide challenging tasks when the user is focused and easy review questions when the user is tired. It can also recommend relevant learning resources when the user wants to acquire new knowledge. Furthermore, the learning support unit can improve learning motivation by recording the user's learning history and visualizing their progress. This ensures that appropriate learning plans are provided according to the user's learning situation, improving learning effectiveness.

[0120] The Kiko-Yell system can also be equipped with a sleep support unit that monitors the user's sleep state and provides an appropriate sleep environment. For example, the sleep support unit can analyze the user's sleep patterns and suggest the optimal sleep duration. It can also provide relaxing music and lighting to help the user relax and fall asleep. Furthermore, the sleep support unit can record the user's sleep data and create a long-term sleep improvement plan. This ensures that an appropriate sleep environment is provided according to the user's sleep state, improving sleep quality.

[0121] The Kiko-Yell system can also include a stress management unit that monitors the user's stress level and provides appropriate stress management methods. For example, the stress management unit analyzes the user's heart rate and breathing patterns to estimate their stress level. If the user is experiencing stress, it can suggest relaxing breathing techniques or meditation. Furthermore, the stress management unit can record the user's stress history and create a long-term stress management plan. This ensures that appropriate stress management methods are provided according to the user's stress level, thereby reducing stress.

[0122] The following briefly describes the processing flow for example form 2.

[0123] Step 1: The transcription unit transcribes the conversation in real time. For example, it uses generative AI to convert speech into text and speech recognition technology to accurately transcribe the content of the conversation. It can also identify and transcribe the voices of the speakers. Step 2: The summarization unit summarizes the conversation transcribed by the transcription unit. For example, it might use generative AI to summarize a long conversation and display only the important points concisely. It might use natural language processing technology to extract the main points of the conversation, understand the context, and prioritize and summarize the most important information. Step 3: The emotion recognition unit analyzes the emotions of the conversation partner based on the information summarized by the summarization unit. For example, it uses generative AI to analyze the tone of voice and facial expressions of the conversation partner and conveys their emotions and intentions to the user. It can analyze the pitch and volume of the voice to detect changes in emotion. It can also analyze the movement of facial feature points to estimate emotions from facial expressions. Step 4: The feedback unit provides feedback based on the emotions analyzed by the emotion recognition unit. For example, it uses generative AI to provide appropriate feedback to the user, offering encouraging messages and advice according to the user's emotional state. It can also provide feedback in real time and respond immediately to changes in the user's emotions. Step 5: The alert unit monitors ambient sounds and notifies of important warning sounds. For example, it uses a generation AI to analyze ambient sounds and quickly notifies of important warning sounds and sirens. It can analyze the type and volume of sounds to detect emergency warning sounds. It can also monitor ambient sounds in real time and detect abnormal sounds.

[0124] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0125] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include text generation AI, image generation AI, and multimodal generation AI. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats from audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVMs), k-means clustering, convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each of the above parts is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example.Furthermore, processing performed by AI, including generative AI, may be replaced with rule-based processing, and rule-based processing may be replaced with processing performed by AI, including generative AI.

[0126] Furthermore, the processing performed by the data processing system 10 described above is carried out by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but it may also be carried out by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0127] Each of the multiple elements described above, including the transcription unit, summarization unit, emotion recognition unit, feedback unit, and alert unit, is implemented in at least one of the smart device 14 and the data processing unit 12. For example, the transcription unit is implemented by the processor 46 of the smart device 14 and converts speech to text using a generation AI. The summarization unit is implemented by the specific processing unit 290 of the data processing unit 12 and summarizes the conversation using a generation AI and natural language processing technology. The emotion recognition unit is implemented by the control unit 46A of the smart device 14 and estimates emotions by analyzing speech and facial expressions using a generation AI. The feedback unit is implemented by the specific processing unit 290 of the data processing unit 12 and provides appropriate feedback to the user using a generation AI. The alert unit is implemented by the processor 46 of the smart device 14 and notifies important warning sounds by analyzing ambient sounds using a generation AI. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0128] [Second Embodiment] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0129] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0130] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0131] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0132] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0133] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0134] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0135] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing by the processor 28. The storage 32 stores the specific processing program 56.

[0136] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0137] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0138] In the smart glasses 214, specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0139] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0140] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0141] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0142] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart glasses 214 or an external device, and the smart glasses 214 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0143] Each of the multiple elements described above, including the transcription unit, summarization unit, emotion recognition unit, feedback unit, and alert unit, is implemented in at least one of the smart glasses 214 and the data processing unit 12. For example, the transcription unit is implemented by the processor 46 of the smart glasses 214 and converts speech to text using a generation AI. The summarization unit is implemented by the identification processing unit 290 of the data processing unit 12 and summarizes the conversation using a generation AI and natural language processing technology. The emotion recognition unit is implemented by the control unit 46A of the smart glasses 214 and estimates emotions by analyzing speech and facial expressions using a generation AI. The feedback unit is implemented by the identification processing unit 290 of the data processing unit 12 and provides appropriate feedback to the user using a generation AI. The alert unit is implemented by the processor 46 of the smart glasses 214 and notifies important warning sounds by analyzing ambient sounds using a generation AI. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0144] [Third Embodiment] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0145] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0146] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0147] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0148] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0149] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0150] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0151] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0152] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0153] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0154] In the headset terminal 314, specific processing is performed by the processor 46. The storage 50 stores a specific program 60. The processor 46 reads the specific program 60 from the storage 50 and executes the read specific program 60 on the RAM 48. The specific processing is realized by the processor 46 acting as a control unit 46A according to the specific program 60 executed on the RAM 48. The headset terminal 314 also has a data generation model 58 and an emotion identification model 59, similar to the data generation model and emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.

[0155] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0156] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0157] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0158] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset terminal 314, but may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset terminal 314. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the headset terminal 314 or an external device, and the headset terminal 314 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0159] Each of the multiple elements described above, including the transcription unit, summarization unit, emotion recognition unit, feedback unit, and alert unit, is implemented in at least one of the headset terminal 314 and the data processing unit 12. For example, the transcription unit is implemented by the processor 46 of the headset terminal 314 and converts speech to text using a generation AI. The summarization unit is implemented by the specific processing unit 290 of the data processing unit 12 and uses a generation AI to summarize the conversation using natural language processing technology. The emotion recognition unit is implemented by the control unit 46A of the headset terminal 314 and uses a generation AI to estimate emotions by analyzing speech and facial expressions. The feedback unit is implemented by the specific processing unit 290 of the data processing unit 12 and uses a generation AI to provide appropriate feedback to the user. The alert unit is implemented by the processor 46 of the headset terminal 314 and uses a generation AI to analyze ambient sounds and notify important warning sounds. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0160] [Fourth Embodiment] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0161] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0162] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN and / or LAN.

[0163] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0164] The microphone 238 receives voice signals from the user and accepts instructions from the user. The microphone 238 captures the voice signals from the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0165] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS image sensor or CCD image sensor, which captures images of the area around the user (for example, an imaging range defined by a field of view equivalent to the field of vision of a typical healthy person).

[0166] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0167] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. The robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0168] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0169] The processor 28 reads a specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 acting as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0170] Storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate the user's emotions using the emotion identification model 59 and perform identification processing using the user's emotions. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotions, including but not limited to these examples. Furthermore, emotion estimation and prediction also include, for example, emotion analysis.

[0171] In robot 414, specific processing is performed by processor 46. A specific program 60 is stored in storage 50. Processor 46 reads the specific program 60 from storage 50 and executes it on RAM 48. The specific processing is achieved by processor 46 acting as a control unit 46A according to the specific program 60 executed on RAM 48. Robot 414 also has data generation model 58 and emotion identification model 59, similar to those of the robot, and can perform processing similar to that of the specific processing unit 290 using these models.

[0172] Furthermore, other devices besides the data processing device 12 may also have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 obtains processing results (such as prediction results) using the data generation model 58 by communicating with the server device that has the data generation model 58. Also, the data processing device 12 may be a server device or a terminal device owned by the user (for example, a mobile phone, robot, home appliance, etc.).

[0173] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0174] The data generation model 58 is a so-called generative AI. An example of a data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and inference data such as audio data representing speech, text data representing text, and image data representing images (e.g., still image data or video data). The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference result in one or more data formats such as audio data, text data, and image data. The data generation model 58 includes, for example, text generation AI, image generation AI, and multimodal generation AI. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specific processing unit 290 performs the specific processing described above using the data generation model 58. The data generation model 58 may be a fine-tuned model that outputs inference results from prompts that do not contain instructions, in which case the data generation model 58 can output inference results from prompts that do not contain instructions. In the data processing device 12, etc., there are multiple types of data generation models 58, and the data generation model 58 includes AI other than generative AI. AI other than generative AI includes, for example, linear regression, logistic regression, decision trees, random forests, support vector machines (SVM), k-means clustering, convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. Also, the AI ​​may be an AI agent. Furthermore, when the processing of each part described above is performed by the AI, the processing may be performed by the AI ​​in part or in whole, but is not limited to this example. Also, processing performed by an AI including a generative AI may be replaced by rule-based processing, and rule-based processing may be replaced by processing performed by an AI including a generative AI.

[0175] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is performed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but it may also be performed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. In addition, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the robot 414 or an external device, and the robot 414 acquires or collects information necessary for processing from the data processing device 12 or an external device.

[0176] Each of the multiple elements described above, including the transcription unit, summarization unit, emotion recognition unit, feedback unit, and alert unit, is implemented in at least one of the robot 414 and the data processing unit 12. For example, the transcription unit is implemented by the processor 46 of the robot 414 and converts speech to text using a generation AI. The summarization unit is implemented by the specific processing unit 290 of the data processing unit 12 and summarizes the conversation using a generation AI and natural language processing techniques. The emotion recognition unit is implemented by the control unit 46A of the robot 414 and estimates emotions by analyzing speech and facial expressions using a generation AI. The feedback unit is implemented by the specific processing unit 290 of the data processing unit 12 and provides appropriate feedback to the user using a generation AI. The alert unit is implemented by the processor 46 of the robot 414 and notifies important warning sounds by analyzing ambient sounds using a generation AI. The correspondence between each unit and the device or control unit is not limited to the examples described above and can be modified in various ways.

[0177] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0178] Figure 9 shows the emotion map 400, in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0179] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0180] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0181] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, and motorcycles, emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated based, for example, on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0182] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0183] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0184] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0185] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0186] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0187] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0188] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0189] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0190] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0191] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0192] Furthermore, although the above-described examples were divided into four embodiments, some or all of these embodiments may be combined. Also, the smart device 14, smart glasses 214, headset terminal 314, and robot 414 are just examples, and they may be combined, or other devices may be used. Also, although the above-described examples were divided into two embodiments, Embodiment 1 and Embodiment 2, these may be combined.

[0193] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and other things that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0194] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0195] (Note 1) The transcription team transcribes conversations in real time, A summarization unit that summarizes the conversation transcribed by the aforementioned transcription unit, An emotion recognition unit analyzes the emotions of the conversation partner based on the information summarized by the summarization unit, A feedback unit provides feedback based on the emotions analyzed by the emotion recognition unit, It includes an alert unit that monitors ambient sounds and notifies important warning sounds. A system characterized by the following features. (Note 2) The aforementioned transcription section is, The conversation is transcribed in real time and displayed on the smartphone. The system described in Appendix 1, characterized by the features described herein. (Note 3) The summary section above is, Summarize long conversations and display only the key points concisely. The system described in Appendix 1, characterized by the features described herein. (Note 4) The emotion recognition unit, It analyzes the tone of voice and facial expressions of the person you're talking to, and conveys their emotions and intentions to the user. The system described in Appendix 1, characterized by the features described herein. (Note 5) The alert unit is, It monitors ambient sounds and quickly notifies users of important warning sounds and sirens. The system described in Appendix 1, characterized by the features described herein. (Note 6) The aforementioned feedback unit is The emotion recognition unit provides feedback based on the emotions it analyzes. The system described in Appendix 1, characterized by the features described herein. (Note 7) The aforementioned transcription section is, It estimates the user's emotions and adjusts the accuracy of the transcription based on the estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 8) The aforementioned transcription section is, It automatically recognizes technical terms and slang based on the content of the conversation and transcribes them appropriately. The system described in Appendix 1, characterized by the features described herein. (Note 9) The aforementioned transcription section is, The transcription speed is automatically adjusted according to the speed of the conversation. The system described in Appendix 1, characterized by the features described herein. (Note 10) The aforementioned transcription section is, It estimates the user's emotions and adjusts how the transcript is displayed based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 11) The aforementioned transcription section is, Filter out background noise from conversations and transcribe only the important parts. The system described in Appendix 1, characterized by the features described herein. (Note 12) The aforementioned transcription section is, Understand the context of the conversation and transcribe it by appropriately converting synonyms and related words. The system described in Appendix 1, characterized by the features described herein. (Note 13) The summary section above is, It estimates the user's sentiment and adjusts the level of detail in the summary based on the estimated user sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 14) The summary section above is, Prioritize summaries based on the importance of the conversation. The system described in Appendix 1, characterized by the features described herein. (Note 15) The summary section above is, Apply different summarization algorithms depending on the category of the conversation. The system described in Appendix 1, characterized by the features described herein. (Note 16) The summary section above is, It estimates the user's sentiment and adjusts how the summary is displayed based on the estimated user sentiment. The system described in Appendix 1, characterized by the features described herein. (Note 17) The summary section above is, The length of the summary is automatically adjusted according to the length of the conversation. The system described in Appendix 1, characterized by the features described herein. (Note 18) The summary section above is, Adjust the order of the summaries based on the relevance of the conversation. The system described in Appendix 1, characterized by the features described herein. (Note 19) The emotion recognition unit, It estimates the user's emotions and adjusts the accuracy of emotion recognition based on the estimated user emotions. The system described in Appendix 1, characterized by the features described herein. (Note 20) The emotion recognition unit, It analyzes the tone of voice and facial expressions of the person you're talking to, and tracks changes in their emotions in real time. The system described in Appendix 1, characterized by the features described herein. (Note 21) The emotion recognition unit, Understanding the context of a conversation allows for a more accurate analysis of emotional intent. The system described in Appendix 1, characterized by the features described herein. (Note 22) The emotion recognition unit, Adjusting how we estimate user emotions and display emotion recognition results based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 23) The emotion recognition unit, Analyzing the gestures and body language of the person you're talking to improves the accuracy of emotion recognition. The system described in Appendix 1, characterized by the features described herein. (Note 24) The emotion recognition unit, Improving the accuracy of emotion recognition by taking into account background noise and ambient sounds in conversations. The system described in Appendix 1, characterized by the features described herein. (Note 25) The aforementioned feedback unit is It estimates the user's emotions and adjusts the content of the feedback based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 26) The aforementioned feedback unit is Based on the conversation, we provide appropriate feedback in real time. The system described in Appendix 1, characterized by the features described herein. (Note 27) The aforementioned feedback unit is Understand the context of the conversation and optimize the timing of feedback. The system described in Appendix 1, characterized by the features described herein. (Note 28) The aforementioned feedback unit is It estimates the user's emotions and adjusts how feedback is displayed based on those estimated emotions. The system described in Appendix 1, characterized by the features described herein. (Note 29) The aforementioned feedback unit is Prioritize feedback based on the importance of the conversation. The system described in Appendix 1, characterized by the features described herein. (Note 30) The aforementioned feedback unit is Apply different feedback algorithms depending on the category of the conversation. The system described in Appendix 1, characterized by the features described herein. (Note 31) The alert unit is, It estimates the user's emotions and adjusts how alerts are notified based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 32) The alert unit is, It analyzes ambient sounds in real time and quickly detects important warning sounds. The system described in Appendix 1, characterized by the features described herein. (Note 33) The alert unit is, Prioritize notifications based on the importance of the alert. The system described in Appendix 1, characterized by the features described herein. (Note 34) The alert unit is, It estimates the user's emotions and adjusts how alerts are displayed based on those emotions. The system described in Appendix 1, characterized by the features described herein. (Note 35) The alert unit is, Filters out ambient noise and notifies only of important sounds. The system described in Appendix 1, characterized by the features described herein. (Note 36) The alert unit is, Apply different notification methods depending on the alert category. The system described in Appendix 1, characterized by the features described herein. [Explanation of Symbols]

[0196] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots

Claims

1. The transcription team transcribes conversations in real time, A summarization unit that summarizes the conversation transcribed by the aforementioned transcription unit, An emotion recognition unit analyzes the emotions of the conversation partner based on the information summarized by the summarization unit, A feedback unit provides feedback based on the emotions analyzed by the emotion recognition unit, It includes an alert unit that monitors ambient sounds and notifies important warning sounds. A system characterized by the following features.

2. The aforementioned transcription section is, The conversation is transcribed in real time and displayed on the smartphone. The system according to feature 1.

3. The summary section above is, Summarize long conversations and display only the key points concisely. The system according to feature 1.

4. The emotion recognition unit, It analyzes the tone of voice and facial expressions of the person you're talking to, and conveys their emotions and intentions to the user. The system according to feature 1.

5. The alert unit is, It monitors ambient sounds and quickly notifies users of important warning sounds and sirens. The system according to feature 1.

6. The aforementioned feedback unit is The emotion recognition unit provides feedback based on the emotions it has analyzed. The system according to feature 1.

7. The aforementioned transcription section is, It estimates the user's emotions and adjusts the accuracy of the transcription based on the estimated emotions. The system according to feature 1.

8. The aforementioned transcription section is, It automatically recognizes technical terms and slang based on the content of the conversation and transcribes them appropriately. The system according to feature 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A