system

A system that converts and analyzes elderly speech to set personalized reminders, addressing memory issues and caregiving challenges, improves elderly independence and reduces caregiving burdens.

JP2026071637APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024181675
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

In an aging society, elderly individuals often forget important information, leading to difficulties in independent living and increased caregiving burdens, exacerbated by manpower shortages in the caregiving field.

Method used

A system that acquires elderly speech, converts it into text, extracts important information, and sets reminders using generative artificial intelligence, tailored to individual needs, while incorporating user feedback for improvement.

Benefits of technology

Enhances self-management abilities of the elderly, reduces caregiving burdens, and alleviates personnel shortages by providing flexible and appropriate support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071637000001_ABST
    Figure 2026071637000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of acquiring speech sounds from elderly people, A means of converting acquired audio into text, A means of extracting important information from the converted text, Means for recording and organizing the extracted information, A means of setting reminders based on recorded information, A means of notifying about set reminders, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In an aging society, the decline in the quality of life due to the elderly forgetting important information in daily life has become a problem. Also, the shortage of manpower in the caregiving field makes it difficult to provide sufficient support for the elderly to appropriately manage this information. As a result, it may become difficult for the elderly to live an independent life, and the burden on caregiving staff may also increase. Furthermore, with the increasing number of elderly people living alone, these problems are becoming increasingly serious.

Means for Solving the Problems

[0005] The system according to the present invention provides a series of means for acquiring the speech of elderly people, converting it into text, and further extracting, recording, and organizing important information. By using generative artificial intelligence, it has the ability to accurately extract important information and set it as a reminder for the user. Furthermore, by acquiring feedback from the user and reflecting that information in the system, it becomes possible to provide flexible and appropriate support tailored to each elderly person. This system can improve the self-management abilities of the elderly, reduce the burden on caregivers, and alleviate the shortage of personnel in the medical field.

[0006] The term "elderly" generally refers to people who have reached the age of 65 or older, and especially includes those who may require care or special support in their daily lives.

[0007] "Speech signals" refer to the acoustic signals generated when a person speaks.

[0008] "Means of acquisition" refers to a device or method for collecting and recording specific information, such as voice or data.

[0009] "Converting to text" refers to the process of converting audio information into a visually representable form of text.

[0010] "Important information" refers to matters, data, or events that require special attention in the user's life or health management.

[0011] "Means of extraction" refers to a method or process for selecting specific elements or patterns from a set of data.

[0012] "Recording and organization" refers to the process of storing information in a system and organizing it so that it can be efficiently searched and used later.

[0013] "Generative artificial intelligence" is a type of AI technology that has the ability to predict and generate functions and patterns based on training data.

[0014] A "reminder" refers to a message or alert that notifies a user at a specific time or under specific conditions, prompting them to take action.

[0015] "Means of notification" refers to mechanisms and methods for communicating important information and reminders to users.

[0016] "User feedback" refers to opinions, satisfaction levels, and information about usability provided by system users, and is collected to help improve the system. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be described.

[0020] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] To implement this invention, a system is constructed that uses terminals placed in the living environment of elderly people and a server for processing. The terminals are installed around the elderly person and are equipped with voice input devices such as microphones to acquire everyday conversations as voice data. As a result, each time the user speaks, the terminal acquires the voice and transmits it to the server in real time.

[0039] The server uses speech recognition technology to convert received audio data into text data. This text data is then analyzed by generative artificial intelligence, and important information is extracted. For example, if a user says, "I will take my medicine tomorrow morning," the region, time, and details of that action are extracted.

[0040] The extracted information is recorded in a database on the server and organized by category. Based on this organized information, the server sets reminders for elderly individuals. Reminders are notified via voice, screen display, or vibration. In particular, the timing and method of reminders are customized based on each user's lifestyle and past responses.

[0041] For example, if a user has a hospital appointment the next day, a reminder such as "You have a hospital appointment tomorrow" is set on the device. The server then determines the optimal time for the reminder based on the user's previously set behavioral patterns. The device then notifies the user a certain amount of time before the scheduled arrival time and provides route guidance if necessary.

[0042] This process also incorporates user feedback; for example, it can receive feedback on whether reminders were appropriate and incorporate that into the system as learning data. The goal is for the system to improve its accuracy and effectiveness over time, thereby enhancing the quality of life for the elderly.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The device continuously monitors the speech of elderly individuals and acquires audio data using noise cancellation technology. This audio data is temporarily stored within the device.

[0046] Step 2:

[0047] The terminal packages audio data at regular intervals and securely sends it to the server using an encryption protocol.

[0048] Step 3:

[0049] The server passes the received audio data to the speech recognition system, which then converts the acquired audio into text data. A language model is used in this process to ensure highly accurate text conversion.

[0050] Step 4:

[0051] The server uses generative artificial intelligence to analyze text data and extract the intent of the conversation and important information. This information is categorized into date, time, and details of the activity.

[0052] Step 5:

[0053] The server records the extracted important information in a database and organizes it by category. This allows for efficient access in subsequent processing.

[0054] Step 6:

[0055] The server generates reminders based on the organized information. The reminders use a scheduling algorithm to set the optimal notification time.

[0056] Step 7:

[0057] Once the reminder setup is complete, the server sends the reminder to the device. The device then keeps this information in standby mode.

[0058] Step 8:

[0059] When the user reaches the reminder time, the device notifies the user of the reminder using a pre-configured method (voice alert, screen display).

[0060] Step 9:

[0061] Users review the reminder content and take action as needed. If a user provides feedback about a reminder, the device collects that information.

[0062] Step 10:

[0063] The device sends the collected feedback to the server. The server analyzes this feedback and uses it as data to improve system performance.

[0064] (Example 1)

[0065] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0066] Improving the quality of life for the elderly requires them to properly manage their daily schedules and important matters and to act without forgetting. However, complex schedule management and declining memory are major challenges for the elderly. Therefore, there is a need for technology that allows for easy and effective schedule management and provides reminders tailored to individual needs.

[0067] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0068] In this invention, the server includes means for converting acquired speech into text, means for analyzing information from the converted text to extract important information, and means for setting notifications based on the stored information. This makes it possible to instantly analyze the speech content of elderly people, accurately grasp important matters, and provide reminders at the optimal time.

[0069] "Elderly people" refers to individuals who, as a result of aging, may require physical or cognitive assistance in their daily lives.

[0070] "Spoken words" refer to audio messages and words that a person utters orally.

[0071] "Device" refers to a machine or system designed and configured to perform a specific function.

[0072] "Text" refers to written information in a format that allows for reading and analysis of audio data, representing it as written text.

[0073] "Information analysis" refers to the process of collecting data, examining its content, and extracting necessary elements and meanings.

[0074] "Extracting important information" refers to the process of identifying and separating particularly noteworthy or highly relevant information from the obtained data.

[0075] "Storage" refers to the act of retaining or recording data or information so that it can be used later.

[0076] "Classification" refers to the process of organizing collected information based on specific criteria and dividing it into different categories or sections.

[0077] "Notification" refers to a message or alert used to communicate an event or information to others.

[0078] "Learning" refers to the process of adapting and improving knowledge and behavior based on new data and experiences.

[0079] "Optimization" refers to the act of improving a system or procedure in light of specific objectives or conditions to achieve the best possible results.

[0080] "Generative artificial intelligence" refers to artificial intelligence technology that has the ability to learn from large amounts of existing data and generate new text, images, and other data.

[0081] "Response" refers to a reaction or answer to an action or question.

[0082] To implement this invention, a system is constructed using a terminal equipped with a voice input device and a server for data processing. The terminal is installed in the living environment of the elderly person and has the function of acquiring spoken voice through a microphone. This terminal transmits the voice data to the server in real time. This allows the elderly person's speech to be recorded at the necessary time and provided to the server in a format that can be immediately analyzed.

[0083] The server utilizes speech recognition technology to convert audio data into text data. Specifically, it uses existing speech recognition software (for example, commonly used speech-to-text APIs). Then, it uses generative artificial intelligence technology to extract important information from the text. In this process, it identifies key elements within the data and clarifies the information that elderly people need.

[0084] For example, if a user says, "I'm going shopping tomorrow afternoon," the server automatically extracts the date, time, and details of the activity and organizes it as reminder information. The extracted information is then recorded in a database on the server, and notifications are set at the optimal time, taking into account the user's past behavior patterns. Notifications are conveyed to the elderly through their device via voice, screen display, or vibration.

[0085] Furthermore, users can provide feedback on whether the notifications were appropriate. This feedback is sent to the server as training data to continuously improve the accuracy and effectiveness of reminder settings. This system aims to provide optimal life support for the elderly by utilizing generative artificial intelligence.

[0086] As a concrete example, the prompt to the generating AI model could be "Set necessary reminders based on the user's utterances." This prompt allows the AI ​​model to accurately analyze the needs of the elderly and provide appropriate support.

[0087] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0088] Step 1:

[0089] The terminal uses a microphone to capture the user's spoken audio. The input is the user's conversational audio data. This audio is converted into clean audio data that can be processed in the next step by applying noise filtering technology. After that, the clean audio data is prepared to be sent to the server.

[0090] Step 2:

[0091] The terminal transmits the acquired clean audio data to the server in real time. The input is the audio data processed on the terminal side, and the output is the data sent to the server. This transmission uses a secure and efficient communication protocol and is designed to guarantee data integrity at all times.

[0092] Step 3:

[0093] The server converts the received audio data into text. The input is audio data sent from the terminal, and the output is text data. This conversion process uses high-precision speech recognition software and includes calculations to correct for the effects of intonation and background noise.

[0094] Step 4:

[0095] The server uses a generative AI model to analyze text data and extract important information. The input is the transformed text data, and the output is the extracted information. This process applies natural language processing algorithms to identify important elements related to date, time, and actions.

[0096] Step 5:

[0097] The server records the extracted information in a database, further categorizing and storing it. The input is the key information extracted in step 4, and the output is the data recorded as structured information. The information is efficiently organized based on specific parameters.

[0098] Step 6:

[0099] The server sets reminders for the user and determines the timing based on recorded information. Inputs are information from the database and the user's past behavior patterns, while output is reminder setting information. An AI model references past responses to determine the optimal notification method and timing.

[0100] Step 7:

[0101] The device receives reminder information from the server and notifies the user. The input is the reminder information set on the server side, and the output is the user receiving the notification. Notifications are made using audio output, screen display, or vibration function.

[0102] Step 8:

[0103] Users provide feedback on notifications and their timing. The input is the content of the notification received by the user, and the output is the user's feedback information. This information is sent from the device to the server and reflected in the next reminder setting.

[0104] Step 9:

[0105] The server analyzes user feedback and adjusts the generated AI model to improve the accuracy of reminders. The input is feedback data, and the output is the updated AI model and reminder setting algorithm. Over time, the system's adaptability improves, providing further assistance to the user.

[0106] (Application Example 1)

[0107] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0108] In today's information-saturated society, elderly people and customers struggle with everyday information processing and making appropriate product choices. They also need to manage important schedules in their daily lives and receive services efficiently. However, existing technologies have not adequately addressed the individual needs of users by providing information and suggestions tailored to their specific requirements.

[0109] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0110] In this invention, the server includes means for acquiring the speech of elderly people, means for converting the acquired speech into data, and means for extracting important information from the converted data. This makes it possible to provide action suggestions and product recommendations tailored to the individual needs of elderly people and customers based on their speech.

[0111] "Means for acquiring speech from elderly people" refers to devices that include microphones, sensor devices, etc., for collecting everyday conversations from elderly people.

[0112] "Means of converting acquired audio into data" refers to a process that uses speech recognition technology to convert audio data into text information or other data formats in real time.

[0113] "Means for extracting important information from transformed data" refers to a system that uses technologies such as generative artificial intelligence to analyze and identify user-related information from transformed data.

[0114] "Means for recording and organizing extracted information" refers to the process of storing the analyzed information in a database and organizing it based on categories and tags so that it can be easily used later.

[0115] "Means of suggesting actions based on recorded information" refers to a system that presents users with beneficial actions or options based on information stored in a database.

[0116] "Means of notifying the proposed action" refers to screen displays or audio message delivery functions that visually or audibly inform the user of the proposed content.

[0117] "Means for acquiring customer speech" refers to acoustic devices used to collect words spoken by customers within a store.

[0118] "Methods for analyzing consumer intent from acquired customer voice data" refers to a process that utilizes artificial intelligence technology to analyze customer speech data and evaluate purchasing intent and needs.

[0119] "A means of recommending products based on analyzed information" refers to an algorithm or system that proposes the most suitable products or services according to customer needs.

[0120] The system that realizes this application combines a voice input device, a cloud server, and generative artificial intelligence. The speech of elderly people and customers is acquired as digital audio through voice input devices installed in stores and living spaces. This audio data is transmitted via a network to a cloud server for processing.

[0121] The server converts audio data into text using speech recognition technologies such as Google® Speech-to-Text API. The converted text is then analyzed using OpenAI® generative AI models to extract important information. This analysis process allows for the detection of important schedules of elderly individuals and customers' purchasing intentions.

[0122] This extracted information is recorded and organized in the server's database, generating action suggestions and product recommendations based on user needs. These suggestions are presented to the user via smartphones or other devices as audio or visual notifications. For example, if a user says, "I want a new smartphone," the server will provide information on the latest smartphone models and related campaigns.

[0123] By utilizing generative AI models, it is possible to customize suggestions according to the user's individual behavior patterns and consumption trends. The server also receives feedback from users, learns from it as data, and improves the accuracy of its suggestions.

[0124] An example of a prompt message would be, "Based on what the user said, please suggest related products currently in stock at the store and their characteristics."

[0125] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0126] Step 1:

[0127] The device uses a voice input device to acquire speech spoken around the user. The input is an analog audio signal, including ambient noise, which is captured by the device's microphone. The output is digitized audio data.

[0128] Step 2:

[0129] The device acquires digital audio data and sends it to a cloud server. The input is digital audio data, which is transferred to the server via Wi-Fi or a wired network. The output is the audio data received by the server.

[0130] Step 3:

[0131] The server converts received audio data into text data using the Google Speech-to-Text API. The input is digitized audio data. The output is the corresponding text data, providing the content of the audio as written information.

[0132] Step 4:

[0133] The server inputs text data into an OpenAI generative AI model, which then extracts important information. This process identifies the user's needs and schedule. The input is text data, and the output is the data structure of the extracted information.

[0134] Step 5:

[0135] The server records and organizes the extracted information in a database. The input is structured data, stored in data storage. The output is a database containing organized information.

[0136] Step 6:

[0137] The server generates specific action suggestions and product recommendations for the user based on recorded information. The input is user information from the database, and the output is a list of generated action suggestions and products.

[0138] Step 7:

[0139] The server notifies the terminal of the proposed content and presents it to the user. The input is the generated proposed content, which is output as audio or screen display. The user receives the proposed content and decides on their next action.

[0140] Step 8:

[0141] Users provide feedback on the proposed ideas, which the server receives and incorporates into future proposal generation. The input is user feedback data, and the output is an updated generative model.

[0142] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0143] To implement this invention, the system is configured to include a terminal equipped with an emotion engine and a server for performing multiple processes in the living environment of an elderly person. The terminal includes a voice input device, a display, and a speaker, and acquires the elderly person's voice in real time, with the emotion engine recognizing emotions from that voice. The content of what the user says is collected by the terminal and transmitted to the server as voice data after emotion analysis.

[0144] The server receives this audio data and first converts it into text data using a speech recognition module. Simultaneously, the emotion recognition results are also processed. Generative artificial intelligence analyzes this text data to extract important information from the utterance and further evaluates the nuances of that information, taking into account the emotional state. For example, if a user says, "I don't want to go tomorrow, but I have to go to the doctor," the emotion engine recognizes this negative emotion, and the server determines that the user does not want to go to the doctor.

[0145] The extracted information and emotions are not only recorded in a database, but are also used to generate reminders tailored to the user's emotional state. For example, if a user is feeling negative about an important appointment, the content and notification method of the reminder are adjusted to alleviate the user's anxiety. The device then delivers the notification at the user's scheduled time, accompanied by a gentle voice message and additional information.

[0146] For example, if a user is feeling more depressed than usual and a reminder is set saying "It's time to take your medicine soon," the reminder will be updated with encouraging words such as "It's time for your important medicine, but if you have any concerns, please ask me later."

[0147] Furthermore, the system receives user feedback and learns which reminders were effective or which adjustments were preferred based on the user's emotions. This feedback information is stored on the server and reflected in future reminder generation, enabling more personalized support. By introducing an emotion engine, the system goes beyond simple time management to provide comprehensive support that incorporates consideration for the user's emotional well-being.

[0148] The following describes the processing flow.

[0149] Step 1:

[0150] The device acquires the elderly person's voice in real time, removes noise, and processes the audio data. At this stage, the emotion engine activates, analyzing the tone of voice and word choice to recognize emotions.

[0151] Step 2:

[0152] The device sends voice data along with the recognized emotion to the server. The data is encrypted for security reasons.

[0153] Step 3:

[0154] The server inputs the received audio data into a speech recognition module and converts it into text data. This conversion enables accurate understanding of the user's speech.

[0155] Step 4:

[0156] The server uses generative artificial intelligence to analyze text data and extract important information. For example, it can find information about when a user should take their medication.

[0157] Step 5:

[0158] The server references the emotion recognition results from the emotion engine to obtain the emotional context of the user's utterances. This information is used to create reminders.

[0159] Step 6:

[0160] The server generates user-optimized reminders based on important information and emotional data. For example, if a user is feeling stressed, it adds encouraging words to the reminder.

[0161] Step 7:

[0162] The generated reminder is sent to the device. The device notifies the user of the reminder at the specified time or under the specified circumstances.

[0163] Step 8:

[0164] The user receives the reminder and takes the necessary action. At this time, the user can input their reaction or thoughts on the reminder as feedback on their device.

[0165] Step 9:

[0166] The device sends feedback to the server, which then analyzes it. This feedback is used to improve the accuracy of future reminder generation and sentiment analysis.

[0167] (Example 2)

[0168] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0169] There is a need to effectively support the emotional and time management of elderly people in their living environments. Conventional systems have struggled to provide appropriate reminders that reflect their emotional state, and have not adequately supported the emotional well-being of users.

[0170] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0171] In this invention, the server includes means for performing emotion analysis, means for converting speech to text and processing it together with emotion information, and means for extracting and evaluating information using generative artificial intelligence. This makes it possible to generate and notify personalized reminders according to the user's emotional state.

[0172] "Elderly people" refers to a user group that has unique needs as a result of aging.

[0173] "Spoken voice" refers to the patterns of sounds that a user conveys orally, and communication is conducted using these sounds.

[0174] "Emotional analysis" refers to the process of identifying and interpreting emotions and sensibilities from audio or text.

[0175] "Text conversion" refers to the process of converting acquired audio data into text information.

[0176] "Important information" refers to meaningful data and content relevant to the user, and decisions are made based on this information.

[0177] "Generative artificial intelligence" refers to a computer program that has the ability to analyze given data and generate new information.

[0178] A "reminder" refers to a notification or warning set to attract the user's attention, and its purpose is schedule management.

[0179] "Feedback" refers to user reactions and opinions after using a product or service, and this feedback is used to improve the system.

[0180] To implement this invention, a system is used that consists of a terminal placed in the living environment of an elderly person and a server that performs information processing. The terminal is equipped with a voice input device, a display, and a speaker, and includes hardware for acquiring the elderly person's speech in real time. The terminal is also equipped with an emotion engine and has software implemented to recognize emotions from the user's voice.

[0181] The terminal sends the acquired audio data, along with the emotion recognition results, to the server. The server uses a speech recognition module to convert the received audio data into text data. Furthermore, a generative AI model analyzes this text data and emotion information to extract important information from the utterance. The generative AI model utilizes prompts to extract and process information according to the user's needs.

[0182] For example, if a user says, "I have plans to meet a friend tomorrow, but I don't feel like it," the emotion engine recognizes that the user is feeling down. Based on this information, the server generates and sends a reminder as needed, such as, "When you don't feel like it, don't force yourself; just relax." An example of this prompt message would be, "How should the next reminder be adjusted based on what the user said and their emotions?"

[0183] By providing feedback on the reminders users receive, the system can use this feedback to create more efficient reminders. This makes it possible to support the emotional well-being of older adults while improving their quality of life.

[0184] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0185] Step 1:

[0186] The user speaks into the device. The user's voice is input and captured by the device's voice input device. As a result, the user's speech is obtained as digital audio data.

[0187] Step 2:

[0188] The device analyzes the captured audio data using an emotion engine to recognize the user's emotions. The input audio data is analyzed, and the emotion engine outputs the user's emotional state. For example, it can recognize that the user is feeling down based on their tone of voice and word choice.

[0189] Step 3:

[0190] The device sends voice data and emotion recognition results to the server. In this process, the voice data and its emotion analysis results are sent to the server as input. The output is that the server waits for processing.

[0191] Step 4:

[0192] The server uses a speech recognition module to convert audio data into text data. It receives audio data as input, analyzes it with a speech recognition algorithm, and generates text data as output.

[0193] Step 5:

[0194] The generative AI model analyzes text data and sentiment information to extract important information. Text data and sentiment information are provided as input, and the generative AI analyzes the data based on the prompt sentence and outputs important information. For example, if a user says, "I'm dreading going to the hospital tomorrow," the generative AI extracts both the physical appointment and the emotional nuance.

[0195] Step 6:

[0196] The server generates optimized reminders based on the user's emotional state, using extracted information and emotional data. It receives important information and emotions as input, processes them using a reminder generation algorithm, and generates customized notification content as output.

[0197] Step 7:

[0198] The device notifies the user of the generated reminder at the specified time. It receives the generated reminder information as input and outputs an emotionally sensitive audio or visual notification through the speaker or display.

[0199] Step 8:

[0200] The user provides feedback based on the received reminder. This feedback is sent to the server via the terminal as input and is reflected in future prompts and generated content as output.

[0201] (Application Example 2)

[0202] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0203] Traditionally, real-time emotional recognition and the subsequent provision of optimal product recommendations to improve the customer purchasing experience in physical stores have been insufficient. This has made it difficult to provide individualized service tailored to each customer's emotional state, leading to decreased purchasing intent. Furthermore, it has been challenging to identify and respond quickly to situations where customers are in a hurry or have specific needs. Solving these challenges is crucial to improving customer satisfaction and increasing store sales.

[0204] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0205] In this invention, the server includes means for acquiring the spoken voice of a consumer, means for converting the acquired voice into text, and means for extracting important information and performing sentiment analysis using generative artificial intelligence. This makes it possible to adjust appropriate product suggestions and service provision based on the emotional state of the customer.

[0206] "Means for acquiring the spoken voice of consumers" refers to a device or sensor for collecting voice information emitted by consumers in real time.

[0207] "Means of converting acquired audio into text" refers to a speech recognition system that analyzes audio data and converts it into text data.

[0208] "Means for extracting important information" refers to algorithms that analyze important elements related to consumers' intentions and needs from collected and converted text data.

[0209] "Means for recording and organizing information" refers to a system that stores extracted information in a database and manages it systematically so that it can be easily searched later.

[0210] "Methods for setting notifications" refers to a system that schedules reminders and notifications to consumers at appropriate times based on recorded information.

[0211] "Means of providing notifications in a form adjusted according to emotional state" refers to a process that provides notifications with content tailored to the psychological state of the consumer, based on the results of emotional analysis.

[0212] "A means of analyzing consumers' emotional states using an emotion engine and utilizing the analysis results" refers to a process that uses a specific algorithm to determine consumers' emotional states in real time and uses the results to help respond to consumers.

[0213] "Means for extracting important information and performing sentiment analysis using generative artificial intelligence" refers to artificial intelligence technology that analyzes and evaluates consumers' emotions and important information through the analysis of text data.

[0214] "Means of incorporating feedback to improve notification content and methods" refers to the process of collecting responses from consumers and using them to adjust the content and improve the format of future notifications.

[0215] To implement this invention, a system is constructed to improve the shopping experience for customers. First, sensors are installed in the physical store to acquire the spoken voices of customers. These sensors acquire customer voice information in real time and transmit it to a server. The server converts the acquired voice into text using a speech recognition module and extracts important information and emotions using generative artificial intelligence.

[0216] Generative artificial intelligence, such as OpenAI's GPT model, analyzes the emotional state of customers based on their speech and records the results in a database. Furthermore, based on the analysis results, appropriate product suggestions and response methods tailored to the customer's emotions are presented to store staff in real time.

[0217] For example, if a customer says, "I want to finish quickly today," this statement is immediately understood by the system, and that information is transmitted to the staff, enabling them to provide the customer with prompt service.

[0218] An example of a prompt for the generative AI model is: "Generate ways to improve the purchase experience when the customer indicates 'I don't have time' in the emotion engine."

[0219] The emotion engine identifies emotions from customers' facial expressions and tone of voice, and reflects them in sales strategies and customer service attitudes. The system also continuously collects feedback from visitors, using it to improve notification content and further enhance customer satisfaction.

[0220] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0221] Step 1:

[0222] The terminal acquires the customer's spoken voice in real time via sensors. This voice data is then transmitted directly to the server. The input is a voice signal, and the output is the transmission of voice data to the server.

[0223] Step 2:

[0224] The server converts the received audio data into text data using a speech recognition module. In this process, the audio signal is analyzed and converted into character data. The input is the audio data sent to the server, and the output is the text representation of the audio.

[0225] Step 3:

[0226] The server uses generative AI to analyze text data for important information and emotional states. The generative AI model extracts information and evaluates emotions based on prompt text. The input is text data, and the output is the result of important information and emotion recognition.

[0227] Step 4:

[0228] The server records the analyzed key information and emotion recognition results in a database. This allows for the accumulation of analysis data, which can then be used for future analysis and improvement of responses. The input is the analysis results, and the output is the recording in the database.

[0229] Step 5:

[0230] The server sends real-time instructions to store staff to provide appropriate service to customers based on recorded emotional data. This process enables staff to take appropriate actions in the situation. The input is the analysis and recording results, and the output is specific action instructions for store staff.

[0231] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0232] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0233] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0234] [Second Embodiment]

[0235] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0236] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0237] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0238] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0239] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0240] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0241] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0242] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0243] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0244] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0245] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0246] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0247] To implement this invention, a system is constructed that uses terminals placed in the living environment of elderly people and a server for processing. The terminals are installed around the elderly person and are equipped with voice input devices such as microphones to acquire everyday conversations as voice data. As a result, each time the user speaks, the terminal acquires the voice and transmits it to the server in real time.

[0248] The server uses speech recognition technology to convert received audio data into text data. This text data is then analyzed by generative artificial intelligence, and important information is extracted. For example, if a user says, "I will take my medicine tomorrow morning," the region, time, and details of that action are extracted.

[0249] The extracted information is recorded in a database on the server and organized by category. Based on this organized information, the server sets reminders for elderly individuals. Reminders are notified via voice, screen display, or vibration. In particular, the timing and method of reminders are customized based on each user's lifestyle and past responses.

[0250] For example, if a user has a hospital appointment the next day, a reminder such as "You have a hospital appointment tomorrow" is set on the device. The server then determines the optimal time for the reminder based on the user's previously set behavioral patterns. The device then notifies the user a certain amount of time before the scheduled arrival time and provides route guidance if necessary.

[0251] This process also incorporates user feedback; for example, it can receive feedback on whether reminders were appropriate and incorporate that into the system as learning data. The goal is for the system to improve its accuracy and effectiveness over time, thereby enhancing the quality of life for the elderly.

[0252] The following describes the processing flow.

[0253] Step 1:

[0254] The device continuously monitors the speech of elderly individuals and acquires audio data using noise cancellation technology. This audio data is temporarily stored within the device.

[0255] Step 2:

[0256] The terminal packages audio data at regular intervals and securely sends it to the server using an encryption protocol.

[0257] Step 3:

[0258] The server passes the received audio data to the speech recognition system, which then converts the acquired audio into text data. A language model is used in this process to ensure highly accurate text conversion.

[0259] Step 4:

[0260] The server uses generative artificial intelligence to analyze text data and extract the intent of the conversation and important information. This information is categorized into date, time, and details of the activity.

[0261] Step 5:

[0262] The server records the extracted important information in a database and organizes it by category. This allows for efficient access in subsequent processing.

[0263] Step 6:

[0264] The server generates reminders based on the organized information. The reminders use a scheduling algorithm to set the optimal notification time.

[0265] Step 7:

[0266] Once the reminder setup is complete, the server sends the reminder to the device. The device then keeps this information in standby mode.

[0267] Step 8:

[0268] When the user reaches the reminder time, the device notifies the user of the reminder using a pre-configured method (voice alert, screen display).

[0269] Step 9:

[0270] Users review the reminder content and take action as needed. If a user provides feedback about a reminder, the device collects that information.

[0271] Step 10:

[0272] The device sends the collected feedback to the server. The server analyzes this feedback and uses it as data to improve system performance.

[0273] (Example 1)

[0274] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0275] Improving the quality of life for the elderly requires them to properly manage their daily schedules and important matters and to act without forgetting. However, complex schedule management and declining memory are major challenges for the elderly. Therefore, there is a need for technology that allows for easy and effective schedule management and provides reminders tailored to individual needs.

[0276] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0277] In this invention, the server includes means for converting acquired speech into text, means for analyzing information from the converted text to extract important information, and means for setting notifications based on the stored information. This makes it possible to instantly analyze the speech content of elderly people, accurately grasp important matters, and provide reminders at the optimal time.

[0278] "Elderly people" refers to individuals who, as a result of aging, may require physical or cognitive assistance in their daily lives.

[0279] "Spoken voice" refers to the voice messages or words uttered orally by a person.

[0280] "Device" refers to a machine or system designed and configured to perform a specific function.

[0281] "Text" refers to the representation of voice data as character information, and refers to text information in a format that enables reading and analysis.

[0282] "Analysis of information" refers to the process of collecting data, examining the content, and extracting necessary elements and meanings.

[0283] "Extraction of important information" refers to the operation of identifying and separating particularly notable information or highly relevant items from the obtained data.

[0284] "Storage" refers to the act of retaining or recording data and information so that they can be used later.

[0285] "Classification" refers to the operation of organizing the collected information based on specific criteria and dividing it into different categories or sections.

[0286] "Notification" refers to a message or alert for transmitting an event or information to others.

[0287] "Learning" refers to the process of adapting and improving knowledge and behavior based on new data and experiences.

[0288] "Optimization" refers to the act of improving a system or procedure in light of specific objectives and conditions to achieve the best results.

[0289] "Generative artificial intelligence" refers to an artificial intelligence technology that has the ability to learn a large amount of existing data and newly generate text, images, etc.

[0290] "Response" refers to a reaction or reply to an action or question.

[0291] To implement this invention, a system is constructed using a terminal equipped with a voice input device and a server for data processing. The terminal is installed in the living environment of the elderly person and has the function of acquiring spoken voice through a microphone. This terminal transmits the voice data to the server in real time. This allows the elderly person's speech to be recorded at the necessary time and provided to the server in a format that can be immediately analyzed.

[0292] The server utilizes speech recognition technology to convert audio data into text data. Specifically, it uses existing speech recognition software (for example, commonly used speech-to-text APIs). Then, it uses generative artificial intelligence technology to extract important information from the text. In this process, it identifies key elements within the data and clarifies the information that elderly people need.

[0293] For example, if a user says, "I'm going shopping tomorrow afternoon," the server automatically extracts the date, time, and details of the activity and organizes it as reminder information. The extracted information is then recorded in a database on the server, and notifications are set at the optimal time, taking into account the user's past behavior patterns. Notifications are conveyed to the elderly through their device via voice, screen display, or vibration.

[0294] Furthermore, users can provide feedback on whether the notifications were appropriate. This feedback is sent to the server as training data to continuously improve the accuracy and effectiveness of reminder settings. This system aims to provide optimal life support for the elderly by utilizing generative artificial intelligence.

[0295] As a concrete example, the prompt to the generating AI model could be "Set necessary reminders based on the user's utterances." This prompt allows the AI ​​model to accurately analyze the needs of the elderly and provide appropriate support.

[0296] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0297] Step 1:

[0298] The terminal uses a microphone to capture the user's spoken audio. The input is the user's conversational audio data. This audio is converted into clean audio data that can be processed in the next step by applying noise filtering technology. After that, the clean audio data is prepared to be sent to the server.

[0299] Step 2:

[0300] The terminal transmits the acquired clean audio data to the server in real time. The input is the audio data processed on the terminal side, and the output is the data sent to the server. This transmission uses a secure and efficient communication protocol and is designed to guarantee data integrity at all times.

[0301] Step 3:

[0302] The server converts the received audio data into text. The input is audio data sent from the terminal, and the output is text data. This conversion process uses high-precision speech recognition software and includes calculations to correct for the effects of intonation and background noise.

[0303] Step 4:

[0304] The server uses a generative AI model to analyze text data and extract important information. The input is the transformed text data, and the output is the extracted information. This process applies natural language processing algorithms to identify important elements related to date, time, and actions.

[0305] Step 5:

[0306] The server records the extracted information in a database, further categorizes it, and stores it. The input is the important information extracted in Step 4, and the output is the data recorded as structured information. Based on specific parameters, the information is efficiently organized.

[0307] Step 6:

[0308] Based on the recorded information, the server sets a reminder for the user and determines its timing. The input is the information from the database and the user's past behavior patterns, and the output is the reminder setting information. The AI model refers to past responses to determine the optimal notification method and time.

[0309] Step 7:

[0310] The terminal receives the reminder information from the server and notifies the user. The input is the reminder information set on the server side, and the output is the receipt of the notification by the user. The notification is made using voice output, screen display, or the vibration function.

[0311] Step 8:

[0312] The user provides feedback on the notification and its timing. The input is the content of the notification received by the user, and the output is the user's feedback information. This information is sent from the terminal to the server and reflected in the next reminder setting.

[0313] Step 9:

[0314] The server analyzes the feedback from the user, adjusts the generated AI model, and improves the accuracy of the reminder. The input is the feedback data, and the output is the updated AI model and the reminder setting algorithm. Over time, the adaptability of the system is improved, and further assistance to the user is provided.

[0315] (Application Example 1)

[0316] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0317] In today's information-saturated society, elderly people and customers struggle with everyday information processing and making appropriate product choices. They also need to manage important schedules in their daily lives and receive services efficiently. However, existing technologies have not adequately addressed the individual needs of users by providing information and suggestions tailored to their specific requirements.

[0318] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0319] In this invention, the server includes means for acquiring the speech of elderly people, means for converting the acquired speech into data, and means for extracting important information from the converted data. This makes it possible to provide action suggestions and product recommendations tailored to the individual needs of elderly people and customers based on their speech.

[0320] "Means for acquiring speech from elderly people" refers to devices that include microphones, sensor devices, etc., for collecting everyday conversations from elderly people.

[0321] "Means of converting acquired audio into data" refers to a process that uses speech recognition technology to convert audio data into text information or other data formats in real time.

[0322] "Means for extracting important information from transformed data" refers to a system that uses technologies such as generative artificial intelligence to analyze and identify user-related information from transformed data.

[0323] "Means for recording and organizing extracted information" refers to the process of storing the analyzed information in a database and organizing it based on categories and tags so that it can be easily used later.

[0324] "Means of suggesting actions based on recorded information" refers to a system that presents users with beneficial actions or options based on information stored in a database.

[0325] "Means of notifying the proposed action" refers to screen displays or audio message delivery functions that visually or audibly inform the user of the proposed content.

[0326] "Means for acquiring customer speech" refers to acoustic devices used to collect words spoken by customers within a store.

[0327] "Methods for analyzing consumer intent from acquired customer voice data" refers to a process that utilizes artificial intelligence technology to analyze customer speech data and evaluate purchasing intent and needs.

[0328] "A means of recommending products based on analyzed information" refers to an algorithm or system that proposes the most suitable products or services according to customer needs.

[0329] The system that realizes this application combines a voice input device, a cloud server, and generative artificial intelligence. The speech of elderly people and customers is acquired as digital audio through voice input devices installed in stores and living spaces. This audio data is transmitted via a network to a cloud server for processing.

[0330] The server converts audio data into text using speech recognition technologies such as the Google Speech-to-Text API. The converted text is then analyzed using OpenAI's generative AI model to extract important information. This analysis process allows for the detection of important schedules of elderly individuals and customers' purchasing intentions.

[0331] This extracted information is recorded and organized in the server's database, generating action suggestions and product recommendations based on user needs. These suggestions are presented to the user via smartphones or other devices as audio or visual notifications. For example, if a user says, "I want a new smartphone," the server will provide information on the latest smartphone models and related campaigns.

[0332] By utilizing generative AI models, it is possible to customize suggestions according to the user's individual behavior patterns and consumption trends. The server also receives feedback from users, learns from it as data, and improves the accuracy of its suggestions.

[0333] An example of a prompt message would be, "Based on what the user said, please suggest related products currently in stock at the store and their characteristics."

[0334] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0335] Step 1:

[0336] The device uses a voice input device to acquire speech spoken around the user. The input is an analog audio signal, including ambient noise, which is captured by the device's microphone. The output is digitized audio data.

[0337] Step 2:

[0338] The device acquires digital audio data and sends it to a cloud server. The input is digital audio data, which is transferred to the server via Wi-Fi or a wired network. The output is the audio data received by the server.

[0339] Step 3:

[0340] The server converts received audio data into text data using the Google Speech-to-Text API. The input is digitized audio data. The output is the corresponding text data, providing the content of the audio as written information.

[0341] Step 4:

[0342] The server inputs text data into an OpenAI generative AI model, which then extracts important information. This process identifies the user's needs and schedule. The input is text data, and the output is the data structure of the extracted information.

[0343] Step 5:

[0344] The server records and organizes the extracted information in a database. The input is structured data, stored in data storage. The output is a database containing organized information.

[0345] Step 6:

[0346] The server generates specific action suggestions and product recommendations for the user based on recorded information. The input is user information from the database, and the output is a list of generated action suggestions and products.

[0347] Step 7:

[0348] The server notifies the terminal of the proposed content and presents it to the user. The input is the generated proposed content, which is output as audio or screen display. The user receives the proposed content and decides on their next action.

[0349] Step 8:

[0350] Users provide feedback on the proposed ideas, which the server receives and incorporates into future proposal generation. The input is user feedback data, and the output is an updated generative model.

[0351] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0352] To implement this invention, the system is configured to include a terminal equipped with an emotion engine and a server for performing multiple processes in the living environment of an elderly person. The terminal includes a voice input device, a display, and a speaker, and acquires the elderly person's voice in real time, with the emotion engine recognizing emotions from that voice. The content of what the user says is collected by the terminal and transmitted to the server as voice data after emotion analysis.

[0353] The server receives this audio data and first converts it into text data using a speech recognition module. Simultaneously, the emotion recognition results are also processed. Generative artificial intelligence analyzes this text data to extract important information from the utterance and further evaluates the nuances of that information, taking into account the emotional state. For example, if a user says, "I don't want to go tomorrow, but I have to go to the doctor," the emotion engine recognizes this negative emotion, and the server determines that the user does not want to go to the doctor.

[0354] The extracted information and emotions are not only recorded in a database, but are also used to generate reminders tailored to the user's emotional state. For example, if a user is feeling negative about an important appointment, the content and notification method of the reminder are adjusted to alleviate the user's anxiety. The device then delivers the notification at the user's scheduled time, accompanied by a gentle voice message and additional information.

[0355] For example, if a user is feeling more depressed than usual and a reminder is set saying "It's time to take your medicine soon," the reminder will be updated with encouraging words such as "It's time for your important medicine, but if you have any concerns, please ask me later."

[0356] Furthermore, the system receives user feedback and learns which reminders were effective or which adjustments were preferred based on the user's emotions. This feedback information is stored on the server and reflected in future reminder generation, enabling more personalized support. By introducing an emotion engine, the system goes beyond simple time management to provide comprehensive support that incorporates consideration for the user's emotional well-being.

[0357] The following describes the processing flow.

[0358] Step 1:

[0359] The device acquires the elderly person's voice in real time, removes noise, and processes the audio data. At this stage, the emotion engine activates, analyzing the tone of voice and word choice to recognize emotions.

[0360] Step 2:

[0361] The device sends voice data along with the recognized emotion to the server. The data is encrypted for security reasons.

[0362] Step 3:

[0363] The server inputs the received audio data into a speech recognition module and converts it into text data. This conversion enables accurate understanding of the user's speech.

[0364] Step 4:

[0365] The server uses generative artificial intelligence to analyze text data and extract important information. For example, it can find information about when a user should take their medication.

[0366] Step 5:

[0367] The server references the emotion recognition results from the emotion engine to obtain the emotional context of the user's utterances. This information is used to create reminders.

[0368] Step 6:

[0369] The server generates user-optimized reminders based on important information and emotional data. For example, if a user is feeling stressed, it adds encouraging words to the reminder.

[0370] Step 7:

[0371] The generated reminder is sent to the device. The device notifies the user of the reminder at the specified time or under the specified circumstances.

[0372] Step 8:

[0373] The user receives the reminder and takes the necessary action. At this time, the user can input their reaction or thoughts on the reminder as feedback on their device.

[0374] Step 9:

[0375] The device sends feedback to the server, which then analyzes it. This feedback is used to improve the accuracy of future reminder generation and sentiment analysis.

[0376] (Example 2)

[0377] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0378] There is a need to effectively support the emotional and time management of elderly people in their living environments. Conventional systems have struggled to provide appropriate reminders that reflect their emotional state, and have not adequately supported the emotional well-being of users.

[0379] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0380] In this invention, the server includes means for performing emotion analysis, means for converting speech to text and processing it together with emotion information, and means for extracting and evaluating information using generative artificial intelligence. This makes it possible to generate and notify personalized reminders according to the user's emotional state.

[0381] "Elderly people" refers to a user group that has unique needs as a result of aging.

[0382] "Spoken voice" refers to the patterns of sounds that a user conveys orally, and communication is conducted using these sounds.

[0383] "Emotional analysis" refers to the process of identifying and interpreting emotions and sensibilities from audio or text.

[0384] "Text conversion" refers to the process of converting acquired audio data into text information.

[0385] "Important information" refers to meaningful data and content relevant to the user, and decisions are made based on this information.

[0386] "Generative artificial intelligence" refers to a computer program that has the ability to analyze given data and generate new information.

[0387] A "reminder" refers to a notification or warning set to attract the user's attention, and its purpose is schedule management.

[0388] "Feedback" refers to user reactions and opinions after using a product or service, and this feedback is used to improve the system.

[0389] To implement this invention, a system is used that consists of a terminal placed in the living environment of an elderly person and a server that performs information processing. The terminal is equipped with a voice input device, a display, and a speaker, and includes hardware for acquiring the elderly person's speech in real time. The terminal is also equipped with an emotion engine and has software implemented to recognize emotions from the user's voice.

[0390] The terminal sends the acquired audio data, along with the emotion recognition results, to the server. The server uses a speech recognition module to convert the received audio data into text data. Furthermore, a generative AI model analyzes this text data and emotion information to extract important information from the utterance. The generative AI model utilizes prompts to extract and process information according to the user's needs.

[0391] For example, if a user says, "I have plans to meet a friend tomorrow, but I don't feel like it," the emotion engine recognizes that the user is feeling down. Based on this information, the server generates and sends a reminder as needed, such as, "When you don't feel like it, don't force yourself; just relax." An example of this prompt message would be, "How should the next reminder be adjusted based on what the user said and their emotions?"

[0392] By providing feedback on the reminders users receive, the system can use this feedback to create more efficient reminders. This makes it possible to support the emotional well-being of older adults while improving their quality of life.

[0393] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0394] Step 1:

[0395] The user speaks into the device. The user's voice is input and captured by the device's voice input device. As a result, the user's speech is obtained as digital audio data.

[0396] Step 2:

[0397] The device analyzes the captured audio data using an emotion engine to recognize the user's emotions. The input audio data is analyzed, and the emotion engine outputs the user's emotional state. For example, it can recognize that the user is feeling down based on their tone of voice and word choice.

[0398] Step 3:

[0399] The device sends voice data and emotion recognition results to the server. In this process, the voice data and its emotion analysis results are sent to the server as input. The output is that the server waits for processing.

[0400] Step 4:

[0401] The server uses a speech recognition module to convert audio data into text data. It receives audio data as input, analyzes it with a speech recognition algorithm, and generates text data as output.

[0402] Step 5:

[0403] The generative AI model analyzes text data and sentiment information to extract important information. Text data and sentiment information are provided as input, and the generative AI analyzes the data based on the prompt sentence and outputs important information. For example, if a user says, "I'm dreading going to the hospital tomorrow," the generative AI extracts both the physical appointment and the emotional nuance.

[0404] Step 6:

[0405] The server generates optimized reminders based on the user's emotional state, using extracted information and emotional data. It receives important information and emotions as input, processes them using a reminder generation algorithm, and generates customized notification content as output.

[0406] Step 7:

[0407] The device notifies the user of the generated reminder at the specified time. It receives the generated reminder information as input and outputs an emotionally sensitive audio or visual notification through the speaker or display.

[0408] Step 8:

[0409] The user provides feedback based on the received reminder. This feedback is sent to the server via the terminal as input and is reflected in future prompts and generated content as output.

[0410] (Application Example 2)

[0411] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0412] Traditionally, real-time emotional recognition and the subsequent provision of optimal product recommendations to improve the customer purchasing experience in physical stores have been insufficient. This has made it difficult to provide individualized service tailored to each customer's emotional state, leading to decreased purchasing intent. Furthermore, it has been challenging to identify and respond quickly to situations where customers are in a hurry or have specific needs. Solving these challenges is crucial to improving customer satisfaction and increasing store sales.

[0413] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0414] In this invention, the server includes means for acquiring the spoken voice of a consumer, means for converting the acquired voice into text, and means for extracting important information and performing sentiment analysis using generative artificial intelligence. This makes it possible to adjust appropriate product suggestions and service provision based on the emotional state of the customer.

[0415] "Means for acquiring the spoken voice of consumers" refers to a device or sensor for collecting voice information emitted by consumers in real time.

[0416] "Means of converting acquired audio into text" refers to a speech recognition system that analyzes audio data and converts it into text data.

[0417] "Means for extracting important information" refers to algorithms that analyze important elements related to consumers' intentions and needs from collected and converted text data.

[0418] "Means for recording and organizing information" refers to a system that stores extracted information in a database and manages it systematically so that it can be easily searched later.

[0419] "Methods for setting notifications" refers to a system that schedules reminders and notifications to consumers at appropriate times based on recorded information.

[0420] "Means of providing notifications in a form adjusted according to emotional state" refers to a process that provides notifications with content tailored to the psychological state of the consumer, based on the results of emotional analysis.

[0421] "A means of analyzing consumers' emotional states using an emotion engine and utilizing the analysis results" refers to a process that uses a specific algorithm to determine consumers' emotional states in real time and uses the results to help respond to consumers.

[0422] "Means for extracting important information and performing sentiment analysis using generative artificial intelligence" refers to artificial intelligence technology that analyzes and evaluates consumers' emotions and important information through the analysis of text data.

[0423] "Means of incorporating feedback to improve notification content and methods" refers to the process of collecting responses from consumers and using them to adjust the content and improve the format of future notifications.

[0424] To implement this invention, a system is constructed to improve the shopping experience for customers. First, sensors are installed in the physical store to acquire the spoken voices of customers. These sensors acquire customer voice information in real time and transmit it to a server. The server converts the acquired voice into text using a speech recognition module and extracts important information and emotions using generative artificial intelligence.

[0425] Generative artificial intelligence, such as OpenAI's GPT model, analyzes the emotional state of customers based on their speech and records the results in a database. Furthermore, based on the analysis results, appropriate product suggestions and response methods tailored to the customer's emotions are presented to store staff in real time.

[0426] For example, if a customer says, "I want to finish quickly today," this statement is immediately understood by the system, and that information is transmitted to the staff, enabling them to provide the customer with prompt service.

[0427] An example of a prompt for the generative AI model is: "Generate ways to improve the purchase experience when the customer indicates 'I don't have time' in the emotion engine."

[0428] The emotion engine identifies emotions from customers' facial expressions and tone of voice, and reflects them in sales strategies and customer service attitudes. The system also continuously collects feedback from visitors, using it to improve notification content and further enhance customer satisfaction.

[0429] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0430] Step 1:

[0431] The terminal acquires the customer's spoken voice in real time via sensors. This voice data is then transmitted directly to the server. The input is a voice signal, and the output is the transmission of voice data to the server.

[0432] Step 2:

[0433] The server converts the received audio data into text data using a speech recognition module. In this process, the audio signal is analyzed and converted into character data. The input is the audio data sent to the server, and the output is the text representation of the audio.

[0434] Step 3:

[0435] The server uses generative AI to analyze text data for important information and emotional states. The generative AI model extracts information and evaluates emotions based on prompt text. The input is text data, and the output is the result of important information and emotion recognition.

[0436] Step 4:

[0437] The server records the analyzed key information and emotion recognition results in a database. This allows for the accumulation of analysis data, which can then be used for future analysis and improvement of responses. The input is the analysis results, and the output is the recording in the database.

[0438] Step 5:

[0439] The server sends real-time instructions to store staff to provide appropriate service to customers based on recorded emotional data. This process enables staff to take appropriate actions in the situation. The input is the analysis and recording results, and the output is specific action instructions for store staff.

[0440] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0441] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0442] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0443] [Third Embodiment]

[0444] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0445] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0446] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0447] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0448] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0449] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0450] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0451] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0452] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0453] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0454] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0455] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0456] To implement this invention, a system is constructed that uses terminals placed in the living environment of elderly people and a server for processing. The terminals are installed around the elderly person and are equipped with voice input devices such as microphones to acquire everyday conversations as voice data. As a result, each time the user speaks, the terminal acquires the voice and transmits it to the server in real time.

[0457] The server uses speech recognition technology to convert received audio data into text data. This text data is then analyzed by generative artificial intelligence, and important information is extracted. For example, if a user says, "I will take my medicine tomorrow morning," the region, time, and details of that action are extracted.

[0458] The extracted information is recorded in a database on the server and organized by category. Based on this organized information, the server sets reminders for elderly individuals. Reminders are notified via voice, screen display, or vibration. In particular, the timing and method of reminders are customized based on each user's lifestyle and past responses.

[0459] For example, if a user has a hospital appointment the next day, a reminder such as "You have a hospital appointment tomorrow" is set on the device. The server then determines the optimal time for the reminder based on the user's previously set behavioral patterns. The device then notifies the user a certain amount of time before the scheduled arrival time and provides route guidance if necessary.

[0460] This process also incorporates user feedback; for example, it can receive feedback on whether reminders were appropriate and incorporate that into the system as learning data. The goal is for the system to improve its accuracy and effectiveness over time, thereby enhancing the quality of life for the elderly.

[0461] The following describes the processing flow.

[0462] Step 1:

[0463] The device continuously monitors the speech of elderly individuals and acquires audio data using noise cancellation technology. This audio data is temporarily stored within the device.

[0464] Step 2:

[0465] The terminal packages audio data at regular intervals and securely sends it to the server using an encryption protocol.

[0466] Step 3:

[0467] The server passes the received audio data to the speech recognition system, which then converts the acquired audio into text data. A language model is used in this process to ensure highly accurate text conversion.

[0468] Step 4:

[0469] The server uses generative artificial intelligence to analyze text data and extract the intent of the conversation and important information. This information is categorized into date, time, and details of the activity.

[0470] Step 5:

[0471] The server records the extracted important information in a database and organizes it by category. This allows for efficient access in subsequent processing.

[0472] Step 6:

[0473] The server generates reminders based on the organized information. The reminders use a scheduling algorithm to set the optimal notification time.

[0474] Step 7:

[0475] Once the reminder setup is complete, the server sends the reminder to the device. The device then keeps this information in standby mode.

[0476] Step 8:

[0477] When the user reaches the reminder time, the device notifies the user of the reminder using a pre-configured method (voice alert, screen display).

[0478] Step 9:

[0479] Users review the reminder content and take action as needed. If a user provides feedback about a reminder, the device collects that information.

[0480] Step 10:

[0481] The device sends the collected feedback to the server. The server analyzes this feedback and uses it as data to improve system performance.

[0482] (Example 1)

[0483] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0484] Improving the quality of life for the elderly requires them to properly manage their daily schedules and important matters and to act without forgetting. However, complex schedule management and declining memory are major challenges for the elderly. Therefore, there is a need for technology that allows for easy and effective schedule management and provides reminders tailored to individual needs.

[0485] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0486] In this invention, the server includes means for converting acquired speech into text, means for analyzing information from the converted text to extract important information, and means for setting notifications based on the stored information. This makes it possible to instantly analyze the speech content of elderly people, accurately grasp important matters, and provide reminders at the optimal time.

[0487] "Elderly people" refers to individuals who, as a result of aging, may require physical or cognitive assistance in their daily lives.

[0488] "Spoken words" refer to audio messages and words that a person utters orally.

[0489] "Device" refers to a machine or system designed and configured to perform a specific function.

[0490] "Text" refers to written information in a format that allows for reading and analysis of audio data, representing it as written text.

[0491] "Information analysis" refers to the process of collecting data, examining its content, and extracting necessary elements and meanings.

[0492] "Extracting important information" refers to the process of identifying and separating particularly noteworthy or highly relevant information from the obtained data.

[0493] "Storage" refers to the act of retaining or recording data or information so that it can be used later.

[0494] "Classification" refers to the process of organizing collected information based on specific criteria and dividing it into different categories or sections.

[0495] "Notification" refers to a message or alert used to communicate an event or information to others.

[0496] "Learning" refers to the process of adapting and improving knowledge and behavior based on new data and experiences.

[0497] "Optimization" refers to the act of improving a system or procedure in light of specific objectives or conditions to achieve the best possible results.

[0498] "Generative artificial intelligence" refers to artificial intelligence technology that has the ability to learn from large amounts of existing data and generate new text, images, and other data.

[0499] "Response" refers to a reaction or answer to an action or question.

[0500] To implement this invention, a system is constructed using a terminal equipped with a voice input device and a server for data processing. The terminal is installed in the living environment of the elderly person and has the function of acquiring spoken voice through a microphone. This terminal transmits the voice data to the server in real time. This allows the elderly person's speech to be recorded at the necessary time and provided to the server in a format that can be immediately analyzed.

[0501] The server utilizes speech recognition technology to convert audio data into text data. Specifically, it uses existing speech recognition software (for example, commonly used speech-to-text APIs). Then, it uses generative artificial intelligence technology to extract important information from the text. In this process, it identifies key elements within the data and clarifies the information that elderly people need.

[0502] For example, if a user says, "I'm going shopping tomorrow afternoon," the server automatically extracts the date, time, and details of the activity and organizes it as reminder information. The extracted information is then recorded in a database on the server, and notifications are set at the optimal time, taking into account the user's past behavior patterns. Notifications are conveyed to the elderly through their device via voice, screen display, or vibration.

[0503] Furthermore, users can provide feedback on whether the notifications were appropriate. This feedback is sent to the server as training data to continuously improve the accuracy and effectiveness of reminder settings. This system aims to provide optimal life support for the elderly by utilizing generative artificial intelligence.

[0504] As a concrete example, the prompt to the generating AI model could be "Set necessary reminders based on the user's utterances." This prompt allows the AI ​​model to accurately analyze the needs of the elderly and provide appropriate support.

[0505] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0506] Step 1:

[0507] The terminal uses a microphone to capture the user's spoken audio. The input is the user's conversational audio data. This audio is converted into clean audio data that can be processed in the next step by applying noise filtering technology. After that, the clean audio data is prepared to be sent to the server.

[0508] Step 2:

[0509] The terminal transmits the acquired clean audio data to the server in real time. The input is the audio data processed on the terminal side, and the output is the data sent to the server. This transmission uses a secure and efficient communication protocol and is designed to guarantee data integrity at all times.

[0510] Step 3:

[0511] The server converts the received audio data into text. The input is audio data sent from the terminal, and the output is text data. This conversion process uses high-precision speech recognition software and includes calculations to correct for the effects of intonation and background noise.

[0512] Step 4:

[0513] The server uses a generative AI model to analyze text data and extract important information. The input is the transformed text data, and the output is the extracted information. This process applies natural language processing algorithms to identify important elements related to date, time, and actions.

[0514] Step 5:

[0515] The server records the extracted information in a database, further categorizing and storing it. The input is the key information extracted in step 4, and the output is the data recorded as structured information. The information is efficiently organized based on specific parameters.

[0516] Step 6:

[0517] The server sets reminders for the user and determines the timing based on recorded information. Inputs are information from the database and the user's past behavior patterns, while output is reminder setting information. An AI model references past responses to determine the optimal notification method and timing.

[0518] Step 7:

[0519] The device receives reminder information from the server and notifies the user. The input is the reminder information set on the server side, and the output is the user receiving the notification. Notifications are made using audio output, screen display, or vibration function.

[0520] Step 8:

[0521] Users provide feedback on notifications and their timing. The input is the content of the notification received by the user, and the output is the user's feedback information. This information is sent from the device to the server and reflected in the next reminder setting.

[0522] Step 9:

[0523] The server analyzes user feedback and adjusts the generated AI model to improve the accuracy of reminders. The input is feedback data, and the output is the updated AI model and reminder setting algorithm. Over time, the system's adaptability improves, providing further assistance to the user.

[0524] (Application Example 1)

[0525] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0526] In today's information-saturated society, elderly people and customers struggle with everyday information processing and making appropriate product choices. They also need to manage important schedules in their daily lives and receive services efficiently. However, existing technologies have not adequately addressed the individual needs of users by providing information and suggestions tailored to their specific requirements.

[0527] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0528] In this invention, the server includes means for acquiring the speech of elderly people, means for converting the acquired speech into data, and means for extracting important information from the converted data. This makes it possible to provide action suggestions and product recommendations tailored to the individual needs of elderly people and customers based on their speech.

[0529] "Means for acquiring speech from elderly people" refers to devices that include microphones, sensor devices, etc., for collecting everyday conversations from elderly people.

[0530] "Means of converting acquired audio into data" refers to a process that uses speech recognition technology to convert audio data into text information or other data formats in real time.

[0531] "Means for extracting important information from transformed data" refers to a system that uses technologies such as generative artificial intelligence to analyze and identify user-related information from transformed data.

[0532] "Means for recording and organizing extracted information" refers to the process of storing the analyzed information in a database and organizing it based on categories and tags so that it can be easily used later.

[0533] "Means of suggesting actions based on recorded information" refers to a system that presents users with beneficial actions or options based on information stored in a database.

[0534] "Means of notifying the proposed action" refers to screen displays or audio message delivery functions that visually or audibly inform the user of the proposed content.

[0535] "Means for acquiring customer speech" refers to acoustic devices used to collect words spoken by customers within a store.

[0536] "Methods for analyzing consumer intent from acquired customer voice data" refers to a process that utilizes artificial intelligence technology to analyze customer speech data and evaluate purchasing intent and needs.

[0537] "A means of recommending products based on analyzed information" refers to an algorithm or system that proposes the most suitable products or services according to customer needs.

[0538] The system that realizes this application combines a voice input device, a cloud server, and generative artificial intelligence. The speech of elderly people and customers is acquired as digital audio through voice input devices installed in stores and living spaces. This audio data is transmitted via a network to a cloud server for processing.

[0539] The server converts audio data into text using speech recognition technologies such as the Google Speech-to-Text API. The converted text is then analyzed using OpenAI's generative AI model to extract important information. This analysis process allows for the detection of important schedules of elderly individuals and customers' purchasing intentions.

[0540] This extracted information is recorded and organized in the server's database, generating action suggestions and product recommendations based on user needs. These suggestions are presented to the user via smartphones or other devices as audio or visual notifications. For example, if a user says, "I want a new smartphone," the server will provide information on the latest smartphone models and related campaigns.

[0541] By utilizing generative AI models, it is possible to customize suggestions according to the user's individual behavior patterns and consumption trends. The server also receives feedback from users, learns from it as data, and improves the accuracy of its suggestions.

[0542] An example of a prompt message would be, "Based on what the user said, please suggest related products currently in stock at the store and their characteristics."

[0543] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0544] Step 1:

[0545] The device uses a voice input device to acquire speech spoken around the user. The input is an analog audio signal, including ambient noise, which is captured by the device's microphone. The output is digitized audio data.

[0546] Step 2:

[0547] The device acquires digital audio data and sends it to a cloud server. The input is digital audio data, which is transferred to the server via Wi-Fi or a wired network. The output is the audio data received by the server.

[0548] Step 3:

[0549] The server converts received audio data into text data using the Google Speech-to-Text API. The input is digitized audio data. The output is the corresponding text data, providing the content of the audio as written information.

[0550] Step 4:

[0551] The server inputs text data into an OpenAI generative AI model, which then extracts important information. This process identifies the user's needs and schedule. The input is text data, and the output is the data structure of the extracted information.

[0552] Step 5:

[0553] The server records and organizes the extracted information in a database. The input is structured data, stored in data storage. The output is a database containing organized information.

[0554] Step 6:

[0555] The server generates specific action suggestions and product recommendations for the user based on recorded information. The input is user information from the database, and the output is a list of generated action suggestions and products.

[0556] Step 7:

[0557] The server notifies the terminal of the proposed content and presents it to the user. The input is the generated proposed content, which is output as audio or screen display. The user receives the proposed content and decides on their next action.

[0558] Step 8:

[0559] Users provide feedback on the proposed ideas, which the server receives and incorporates into future proposal generation. The input is user feedback data, and the output is an updated generative model.

[0560] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0561] To implement this invention, the system is configured to include a terminal equipped with an emotion engine and a server for performing multiple processes in the living environment of an elderly person. The terminal includes a voice input device, a display, and a speaker, and acquires the elderly person's voice in real time, with the emotion engine recognizing emotions from that voice. The content of what the user says is collected by the terminal and transmitted to the server as voice data after emotion analysis.

[0562] The server receives this audio data and first converts it into text data using a speech recognition module. Simultaneously, the emotion recognition results are also processed. Generative artificial intelligence analyzes this text data to extract important information from the utterance and further evaluates the nuances of that information, taking into account the emotional state. For example, if a user says, "I don't want to go tomorrow, but I have to go to the doctor," the emotion engine recognizes this negative emotion, and the server determines that the user does not want to go to the doctor.

[0563] The extracted information and emotions are not only recorded in a database, but are also used to generate reminders tailored to the user's emotional state. For example, if a user is feeling negative about an important appointment, the content and notification method of the reminder are adjusted to alleviate the user's anxiety. The device then delivers the notification at the user's scheduled time, accompanied by a gentle voice message and additional information.

[0564] For example, if a user is feeling more depressed than usual and a reminder is set saying "It's time to take your medicine soon," the reminder will be updated with encouraging words such as "It's time for your important medicine, but if you have any concerns, please ask me later."

[0565] Furthermore, the system receives user feedback and learns which reminders were effective or which adjustments were preferred based on the user's emotions. This feedback information is stored on the server and reflected in future reminder generation, enabling more personalized support. By introducing an emotion engine, the system goes beyond simple time management to provide comprehensive support that incorporates consideration for the user's emotional well-being.

[0566] The following describes the processing flow.

[0567] Step 1:

[0568] The device acquires the elderly person's voice in real time, removes noise, and processes the audio data. At this stage, the emotion engine activates, analyzing the tone of voice and word choice to recognize emotions.

[0569] Step 2:

[0570] The device sends voice data along with the recognized emotion to the server. The data is encrypted for security reasons.

[0571] Step 3:

[0572] The server inputs the received audio data into a speech recognition module and converts it into text data. This conversion enables accurate understanding of the user's speech.

[0573] Step 4:

[0574] The server uses generative artificial intelligence to analyze text data and extract important information. For example, it can find information about when a user should take their medication.

[0575] Step 5:

[0576] The server references the emotion recognition results from the emotion engine to obtain the emotional context of the user's utterances. This information is used to create reminders.

[0577] Step 6:

[0578] The server generates user-optimized reminders based on important information and emotional data. For example, if a user is feeling stressed, it adds encouraging words to the reminder.

[0579] Step 7:

[0580] The generated reminder is sent to the device. The device notifies the user of the reminder at the specified time or under the specified circumstances.

[0581] Step 8:

[0582] The user receives the reminder and takes the necessary action. At this time, the user can input their reaction or thoughts on the reminder as feedback on their device.

[0583] Step 9:

[0584] The device sends feedback to the server, which then analyzes it. This feedback is used to improve the accuracy of future reminder generation and sentiment analysis.

[0585] (Example 2)

[0586] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0587] There is a need to effectively support the emotional and time management of elderly people in their living environments. Conventional systems have struggled to provide appropriate reminders that reflect their emotional state, and have not adequately supported the emotional well-being of users.

[0588] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0589] In this invention, the server includes means for performing emotion analysis, means for converting speech to text and processing it together with emotion information, and means for extracting and evaluating information using generative artificial intelligence. This makes it possible to generate and notify personalized reminders according to the user's emotional state.

[0590] "Elderly people" refers to a user group that has unique needs as a result of aging.

[0591] "Spoken voice" refers to the patterns of sounds that a user conveys orally, and communication is conducted using these sounds.

[0592] "Emotional analysis" refers to the process of identifying and interpreting emotions and sensibilities from audio or text.

[0593] "Text conversion" refers to the process of converting acquired audio data into text information.

[0594] "Important information" refers to meaningful data and content relevant to the user, and decisions are made based on this information.

[0595] "Generative artificial intelligence" refers to a computer program that has the ability to analyze given data and generate new information.

[0596] A "reminder" refers to a notification or warning set to attract the user's attention, and its purpose is schedule management.

[0597] "Feedback" refers to user reactions and opinions after using a product or service, and this feedback is used to improve the system.

[0598] To implement this invention, a system is used that consists of a terminal placed in the living environment of an elderly person and a server that performs information processing. The terminal is equipped with a voice input device, a display, and a speaker, and includes hardware for acquiring the elderly person's speech in real time. The terminal is also equipped with an emotion engine and has software implemented to recognize emotions from the user's voice.

[0599] The terminal sends the acquired audio data, along with the emotion recognition results, to the server. The server uses a speech recognition module to convert the received audio data into text data. Furthermore, a generative AI model analyzes this text data and emotion information to extract important information from the utterance. The generative AI model utilizes prompts to extract and process information according to the user's needs.

[0600] For example, if a user says, "I have plans to meet a friend tomorrow, but I don't feel like it," the emotion engine recognizes that the user is feeling down. Based on this information, the server generates and sends a reminder as needed, such as, "When you don't feel like it, don't force yourself; just relax." An example of this prompt message would be, "How should the next reminder be adjusted based on what the user said and their emotions?"

[0601] By providing feedback on the reminders users receive, the system can use this feedback to create more efficient reminders. This makes it possible to support the emotional well-being of older adults while improving their quality of life.

[0602] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0603] Step 1:

[0604] The user speaks into the device. The user's voice is input and captured by the device's voice input device. As a result, the user's speech is obtained as digital audio data.

[0605] Step 2:

[0606] The device analyzes the captured audio data using an emotion engine to recognize the user's emotions. The input audio data is analyzed, and the emotion engine outputs the user's emotional state. For example, it can recognize that the user is feeling down based on their tone of voice and word choice.

[0607] Step 3:

[0608] The device sends voice data and emotion recognition results to the server. In this process, the voice data and its emotion analysis results are sent to the server as input. The output is that the server waits for processing.

[0609] Step 4:

[0610] The server uses a speech recognition module to convert audio data into text data. It receives audio data as input, analyzes it with a speech recognition algorithm, and generates text data as output.

[0611] Step 5:

[0612] The generative AI model analyzes text data and sentiment information to extract important information. Text data and sentiment information are provided as input, and the generative AI analyzes the data based on the prompt sentence and outputs important information. For example, if a user says, "I'm dreading going to the hospital tomorrow," the generative AI extracts both the physical appointment and the emotional nuance.

[0613] Step 6:

[0614] The server generates optimized reminders based on the user's emotional state, using extracted information and emotional data. It receives important information and emotions as input, processes them using a reminder generation algorithm, and generates customized notification content as output.

[0615] Step 7:

[0616] The device notifies the user of the generated reminder at the specified time. It receives the generated reminder information as input and outputs an emotionally sensitive audio or visual notification through the speaker or display.

[0617] Step 8:

[0618] The user provides feedback based on the received reminder. This feedback is sent to the server via the terminal as input and is reflected in future prompts and generated content as output.

[0619] (Application Example 2)

[0620] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0621] Traditionally, real-time emotional recognition and the subsequent provision of optimal product recommendations to improve the customer purchasing experience in physical stores have been insufficient. This has made it difficult to provide individualized service tailored to each customer's emotional state, leading to decreased purchasing intent. Furthermore, it has been challenging to identify and respond quickly to situations where customers are in a hurry or have specific needs. Solving these challenges is crucial to improving customer satisfaction and increasing store sales.

[0622] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0623] In this invention, the server includes means for acquiring the spoken voice of a consumer, means for converting the acquired voice into text, and means for extracting important information and performing sentiment analysis using generative artificial intelligence. This makes it possible to adjust appropriate product suggestions and service provision based on the emotional state of the customer.

[0624] "Means for acquiring the spoken voice of consumers" refers to a device or sensor for collecting voice information emitted by consumers in real time.

[0625] "Means of converting acquired audio into text" refers to a speech recognition system that analyzes audio data and converts it into text data.

[0626] "Means for extracting important information" refers to algorithms that analyze important elements related to consumers' intentions and needs from collected and converted text data.

[0627] "Means for recording and organizing information" refers to a system that stores extracted information in a database and manages it systematically so that it can be easily searched later.

[0628] "Methods for setting notifications" refers to a system that schedules reminders and notifications to consumers at appropriate times based on recorded information.

[0629] "Means of providing notifications in a form adjusted according to emotional state" refers to a process that provides notifications with content tailored to the psychological state of the consumer, based on the results of emotional analysis.

[0630] "A means of analyzing consumers' emotional states using an emotion engine and utilizing the analysis results" refers to a process that uses a specific algorithm to determine consumers' emotional states in real time and uses the results to help respond to consumers.

[0631] "Means for extracting important information and performing sentiment analysis using generative artificial intelligence" refers to artificial intelligence technology that analyzes and evaluates consumers' emotions and important information through the analysis of text data.

[0632] "Means of incorporating feedback to improve notification content and methods" refers to the process of collecting responses from consumers and using them to adjust the content and improve the format of future notifications.

[0633] To implement this invention, a system is constructed to improve the shopping experience for customers. First, sensors are installed in the physical store to acquire the spoken voices of customers. These sensors acquire customer voice information in real time and transmit it to a server. The server converts the acquired voice into text using a speech recognition module and extracts important information and emotions using generative artificial intelligence.

[0634] Generative artificial intelligence, such as OpenAI's GPT model, analyzes the emotional state of customers based on their speech and records the results in a database. Furthermore, based on the analysis results, appropriate product suggestions and response methods tailored to the customer's emotions are presented to store staff in real time.

[0635] For example, if a customer says, "I want to finish quickly today," this statement is immediately understood by the system, and that information is transmitted to the staff, enabling them to provide the customer with prompt service.

[0636] An example of a prompt for the generative AI model is: "Generate ways to improve the purchase experience when the customer indicates 'I don't have time' in the emotion engine."

[0637] The emotion engine identifies emotions from customers' facial expressions and tone of voice, and reflects them in sales strategies and customer service attitudes. The system also continuously collects feedback from visitors, using it to improve notification content and further enhance customer satisfaction.

[0638] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0639] Step 1:

[0640] The terminal acquires the customer's spoken voice in real time via sensors. This voice data is then transmitted directly to the server. The input is a voice signal, and the output is the transmission of voice data to the server.

[0641] Step 2:

[0642] The server converts the received audio data into text data using a speech recognition module. In this process, the audio signal is analyzed and converted into character data. The input is the audio data sent to the server, and the output is the text representation of the audio.

[0643] Step 3:

[0644] The server uses generative AI to analyze text data for important information and emotional states. The generative AI model extracts information and evaluates emotions based on prompt text. The input is text data, and the output is the result of important information and emotion recognition.

[0645] Step 4:

[0646] The server records the analyzed key information and emotion recognition results in a database. This allows for the accumulation of analysis data, which can then be used for future analysis and improvement of responses. The input is the analysis results, and the output is the recording in the database.

[0647] Step 5:

[0648] The server sends real-time instructions to store staff to provide appropriate service to customers based on recorded emotional data. This process enables staff to take appropriate actions in the situation. The input is the analysis and recording results, and the output is specific action instructions for store staff.

[0649] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0650] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0651] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0652] [Fourth Embodiment]

[0653] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0654] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0655] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0656] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0657] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0658] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0659] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0660] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0661] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0662] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0663] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0664] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0665] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0666] To implement this invention, a system is constructed that uses terminals placed in the living environment of elderly people and a server for processing. The terminals are installed around the elderly person and are equipped with voice input devices such as microphones to acquire everyday conversations as voice data. As a result, each time the user speaks, the terminal acquires the voice and transmits it to the server in real time.

[0667] The server uses speech recognition technology to convert received audio data into text data. This text data is then analyzed by generative artificial intelligence, and important information is extracted. For example, if a user says, "I will take my medicine tomorrow morning," the region, time, and details of that action are extracted.

[0668] The extracted information is recorded in a database on the server and organized by category. Based on this organized information, the server sets reminders for elderly individuals. Reminders are notified via voice, screen display, or vibration. In particular, the timing and method of reminders are customized based on each user's lifestyle and past responses.

[0669] For example, if a user has a hospital appointment the next day, a reminder such as "You have a hospital appointment tomorrow" is set on the device. The server then determines the optimal time for the reminder based on the user's previously set behavioral patterns. The device then notifies the user a certain amount of time before the scheduled arrival time and provides route guidance if necessary.

[0670] This process also incorporates user feedback; for example, it can receive feedback on whether reminders were appropriate and incorporate that into the system as learning data. The goal is for the system to improve its accuracy and effectiveness over time, thereby enhancing the quality of life for the elderly.

[0671] The following describes the processing flow.

[0672] Step 1:

[0673] The device continuously monitors the speech of elderly individuals and acquires audio data using noise cancellation technology. This audio data is temporarily stored within the device.

[0674] Step 2:

[0675] The terminal packages audio data at regular intervals and securely sends it to the server using an encryption protocol.

[0676] Step 3:

[0677] The server passes the received audio data to the speech recognition system, which then converts the acquired audio into text data. A language model is used in this process to ensure highly accurate text conversion.

[0678] Step 4:

[0679] The server uses generative artificial intelligence to analyze text data and extract the intent of the conversation and important information. This information is categorized into date, time, and details of the activity.

[0680] Step 5:

[0681] The server records the extracted important information in a database and organizes it by category. This allows for efficient access in subsequent processing.

[0682] Step 6:

[0683] The server generates reminders based on the organized information. The reminders use a scheduling algorithm to set the optimal notification time.

[0684] Step 7:

[0685] Once the reminder setup is complete, the server sends the reminder to the device. The device then keeps this information in standby mode.

[0686] Step 8:

[0687] When the user reaches the reminder time, the device notifies the user of the reminder using a pre-configured method (voice alert, screen display).

[0688] Step 9:

[0689] Users review the reminder content and take action as needed. If a user provides feedback about a reminder, the device collects that information.

[0690] Step 10:

[0691] The device sends the collected feedback to the server. The server analyzes this feedback and uses it as data to improve system performance.

[0692] (Example 1)

[0693] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0694] Improving the quality of life for the elderly requires them to properly manage their daily schedules and important matters and to act without forgetting. However, complex schedule management and declining memory are major challenges for the elderly. Therefore, there is a need for technology that allows for easy and effective schedule management and provides reminders tailored to individual needs.

[0695] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0696] In this invention, the server includes means for converting acquired speech into text, means for analyzing information from the converted text to extract important information, and means for setting notifications based on the stored information. This makes it possible to instantly analyze the speech content of elderly people, accurately grasp important matters, and provide reminders at the optimal time.

[0697] "Elderly people" refers to individuals who, as a result of aging, may require physical or cognitive assistance in their daily lives.

[0698] "Spoken words" refer to audio messages and words that a person utters orally.

[0699] "Device" refers to a machine or system designed and configured to perform a specific function.

[0700] "Text" refers to written information in a format that allows for reading and analysis of audio data, representing it as written text.

[0701] "Information analysis" refers to the process of collecting data, examining its content, and extracting necessary elements and meanings.

[0702] "Extracting important information" refers to the process of identifying and separating particularly noteworthy or highly relevant information from the obtained data.

[0703] "Storage" refers to the act of retaining or recording data or information so that it can be used later.

[0704] "Classification" refers to the process of organizing collected information based on specific criteria and dividing it into different categories or sections.

[0705] "Notification" refers to a message or alert used to communicate an event or information to others.

[0706] "Learning" refers to the process of adapting and improving knowledge and behavior based on new data and experiences.

[0707] "Optimization" refers to the act of improving a system or procedure in light of specific objectives or conditions to achieve the best possible results.

[0708] "Generative artificial intelligence" refers to artificial intelligence technology that has the ability to learn from large amounts of existing data and generate new text, images, and other data.

[0709] "Response" refers to a reaction or answer to an action or question.

[0710] To implement this invention, a system is constructed using a terminal equipped with a voice input device and a server for data processing. The terminal is installed in the living environment of the elderly person and has the function of acquiring spoken voice through a microphone. This terminal transmits the voice data to the server in real time. This allows the elderly person's speech to be recorded at the necessary time and provided to the server in a format that can be immediately analyzed.

[0711] The server utilizes speech recognition technology to convert audio data into text data. Specifically, it uses existing speech recognition software (for example, commonly used speech-to-text APIs). Then, it uses generative artificial intelligence technology to extract important information from the text. In this process, it identifies key elements within the data and clarifies the information that elderly people need.

[0712] For example, if a user says, "I'm going shopping tomorrow afternoon," the server automatically extracts the date, time, and details of the activity and organizes it as reminder information. The extracted information is then recorded in a database on the server, and notifications are set at the optimal time, taking into account the user's past behavior patterns. Notifications are conveyed to the elderly through their device via voice, screen display, or vibration.

[0713] Furthermore, users can provide feedback on whether the notifications were appropriate. This feedback is sent to the server as training data to continuously improve the accuracy and effectiveness of reminder settings. This system aims to provide optimal life support for the elderly by utilizing generative artificial intelligence.

[0714] As a concrete example, the prompt to the generating AI model could be "Set necessary reminders based on the user's utterances." This prompt allows the AI ​​model to accurately analyze the needs of the elderly and provide appropriate support.

[0715] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0716] Step 1:

[0717] The terminal uses a microphone to capture the user's spoken audio. The input is the user's conversational audio data. This audio is converted into clean audio data that can be processed in the next step by applying noise filtering technology. After that, the clean audio data is prepared to be sent to the server.

[0718] Step 2:

[0719] The terminal transmits the acquired clean audio data to the server in real time. The input is the audio data processed on the terminal side, and the output is the data sent to the server. This transmission uses a secure and efficient communication protocol and is designed to guarantee data integrity at all times.

[0720] Step 3:

[0721] The server converts the received audio data into text. The input is audio data sent from the terminal, and the output is text data. This conversion process uses high-precision speech recognition software and includes calculations to correct for the effects of intonation and background noise.

[0722] Step 4:

[0723] The server uses a generative AI model to analyze text data and extract important information. The input is the transformed text data, and the output is the extracted information. This process applies natural language processing algorithms to identify important elements related to date, time, and actions.

[0724] Step 5:

[0725] The server records the extracted information in a database, further categorizing and storing it. The input is the key information extracted in step 4, and the output is the data recorded as structured information. The information is efficiently organized based on specific parameters.

[0726] Step 6:

[0727] The server sets reminders for the user and determines the timing based on recorded information. Inputs are information from the database and the user's past behavior patterns, while output is reminder setting information. An AI model references past responses to determine the optimal notification method and timing.

[0728] Step 7:

[0729] The device receives reminder information from the server and notifies the user. The input is the reminder information set on the server side, and the output is the user receiving the notification. Notifications are made using audio output, screen display, or vibration function.

[0730] Step 8:

[0731] Users provide feedback on notifications and their timing. The input is the content of the notification received by the user, and the output is the user's feedback information. This information is sent from the device to the server and reflected in the next reminder setting.

[0732] Step 9:

[0733] The server analyzes user feedback and adjusts the generated AI model to improve the accuracy of reminders. The input is feedback data, and the output is the updated AI model and reminder setting algorithm. Over time, the system's adaptability improves, providing further assistance to the user.

[0734] (Application Example 1)

[0735] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0736] In today's information-saturated society, elderly people and customers struggle with everyday information processing and making appropriate product choices. They also need to manage important schedules in their daily lives and receive services efficiently. However, existing technologies have not adequately addressed the individual needs of users by providing information and suggestions tailored to their specific requirements.

[0737] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0738] In this invention, the server includes means for acquiring the speech of elderly people, means for converting the acquired speech into data, and means for extracting important information from the converted data. This makes it possible to provide action suggestions and product recommendations tailored to the individual needs of elderly people and customers based on their speech.

[0739] "Means for acquiring speech from elderly people" refers to devices that include microphones, sensor devices, etc., for collecting everyday conversations from elderly people.

[0740] "Means of converting acquired audio into data" refers to a process that uses speech recognition technology to convert audio data into text information or other data formats in real time.

[0741] "Means for extracting important information from transformed data" refers to a system that uses technologies such as generative artificial intelligence to analyze and identify user-related information from transformed data.

[0742] "Means for recording and organizing extracted information" refers to the process of storing the analyzed information in a database and organizing it based on categories and tags so that it can be easily used later.

[0743] "Means of suggesting actions based on recorded information" refers to a system that presents users with beneficial actions or options based on information stored in a database.

[0744] "Means of notifying the proposed action" refers to screen displays or audio message delivery functions that visually or audibly inform the user of the proposed content.

[0745] "Means for acquiring customer speech" refers to acoustic devices used to collect words spoken by customers within a store.

[0746] "Methods for analyzing consumer intent from acquired customer voice data" refers to a process that utilizes artificial intelligence technology to analyze customer speech data and evaluate purchasing intent and needs.

[0747] "A means of recommending products based on analyzed information" refers to an algorithm or system that proposes the most suitable products or services according to customer needs.

[0748] The system that realizes this application combines a voice input device, a cloud server, and generative artificial intelligence. The speech of elderly people and customers is acquired as digital audio through voice input devices installed in stores and living spaces. This audio data is transmitted via a network to a cloud server for processing.

[0749] The server converts audio data into text using speech recognition technologies such as the Google Speech-to-Text API. The converted text is then analyzed using OpenAI's generative AI model to extract important information. This analysis process allows for the detection of important schedules of elderly individuals and customers' purchasing intentions.

[0750] This extracted information is recorded and organized in the server's database, generating action suggestions and product recommendations based on user needs. These suggestions are presented to the user via smartphones or other devices as audio or visual notifications. For example, if a user says, "I want a new smartphone," the server will provide information on the latest smartphone models and related campaigns.

[0751] By utilizing generative AI models, it is possible to customize suggestions according to the user's individual behavior patterns and consumption trends. The server also receives feedback from users, learns from it as data, and improves the accuracy of its suggestions.

[0752] An example of a prompt message would be, "Based on what the user said, please suggest related products currently in stock at the store and their characteristics."

[0753] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0754] Step 1:

[0755] The device uses a voice input device to acquire speech spoken around the user. The input is an analog audio signal, including ambient noise, which is captured by the device's microphone. The output is digitized audio data.

[0756] Step 2:

[0757] The device acquires digital audio data and sends it to a cloud server. The input is digital audio data, which is transferred to the server via Wi-Fi or a wired network. The output is the audio data received by the server.

[0758] Step 3:

[0759] The server converts received audio data into text data using the Google Speech-to-Text API. The input is digitized audio data. The output is the corresponding text data, providing the content of the audio as written information.

[0760] Step 4:

[0761] The server inputs text data into an OpenAI generative AI model, which then extracts important information. This process identifies the user's needs and schedule. The input is text data, and the output is the data structure of the extracted information.

[0762] Step 5:

[0763] The server records and organizes the extracted information in a database. The input is structured data, stored in data storage. The output is a database containing organized information.

[0764] Step 6:

[0765] The server generates specific action suggestions and product recommendations for the user based on recorded information. The input is user information from the database, and the output is a list of generated action suggestions and products.

[0766] Step 7:

[0767] The server notifies the terminal of the proposed content and presents it to the user. The input is the generated proposed content, which is output as audio or screen display. The user receives the proposed content and decides on their next action.

[0768] Step 8:

[0769] Users provide feedback on the proposed ideas, which the server receives and incorporates into future proposal generation. The input is user feedback data, and the output is an updated generative model.

[0770] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0771] To implement this invention, the system is configured to include a terminal equipped with an emotion engine and a server for performing multiple processes in the living environment of an elderly person. The terminal includes a voice input device, a display, and a speaker, and acquires the elderly person's voice in real time, with the emotion engine recognizing emotions from that voice. The content of what the user says is collected by the terminal and transmitted to the server as voice data after emotion analysis.

[0772] The server receives this audio data and first converts it into text data using a speech recognition module. Simultaneously, the emotion recognition results are also processed. Generative artificial intelligence analyzes this text data to extract important information from the utterance and further evaluates the nuances of that information, taking into account the emotional state. For example, if a user says, "I don't want to go tomorrow, but I have to go to the doctor," the emotion engine recognizes this negative emotion, and the server determines that the user does not want to go to the doctor.

[0773] The extracted information and emotions are not only recorded in a database, but are also used to generate reminders tailored to the user's emotional state. For example, if a user is feeling negative about an important appointment, the content and notification method of the reminder are adjusted to alleviate the user's anxiety. The device then delivers the notification at the user's scheduled time, accompanied by a gentle voice message and additional information.

[0774] For example, if a user is feeling more depressed than usual and a reminder is set saying "It's time to take your medicine soon," the reminder will be updated with encouraging words such as "It's time for your important medicine, but if you have any concerns, please ask me later."

[0775] Furthermore, the system receives user feedback and learns which reminders were effective or which adjustments were preferred based on the user's emotions. This feedback information is stored on the server and reflected in future reminder generation, enabling more personalized support. By introducing an emotion engine, the system goes beyond simple time management to provide comprehensive support that incorporates consideration for the user's emotional well-being.

[0776] The following describes the processing flow.

[0777] Step 1:

[0778] The device acquires the elderly person's voice in real time, removes noise, and processes the audio data. At this stage, the emotion engine activates, analyzing the tone of voice and word choice to recognize emotions.

[0779] Step 2:

[0780] The device sends voice data along with the recognized emotion to the server. The data is encrypted for security reasons.

[0781] Step 3:

[0782] The server inputs the received audio data into a speech recognition module and converts it into text data. This conversion enables accurate understanding of the user's speech.

[0783] Step 4:

[0784] The server uses generative artificial intelligence to analyze text data and extract important information. For example, it can find information about when a user should take their medication.

[0785] Step 5:

[0786] The server references the emotion recognition results from the emotion engine to obtain the emotional context of the user's utterances. This information is used to create reminders.

[0787] Step 6:

[0788] The server generates user-optimized reminders based on important information and emotional data. For example, if a user is feeling stressed, it adds encouraging words to the reminder.

[0789] Step 7:

[0790] The generated reminder is sent to the device. The device notifies the user of the reminder at the specified time or under the specified circumstances.

[0791] Step 8:

[0792] The user receives the reminder and takes the necessary action. At this time, the user can input their reaction or thoughts on the reminder as feedback on their device.

[0793] Step 9:

[0794] The device sends feedback to the server, which then analyzes it. This feedback is used to improve the accuracy of future reminder generation and sentiment analysis.

[0795] (Example 2)

[0796] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0797] There is a need to effectively support the emotional and time management of elderly people in their living environments. Conventional systems have struggled to provide appropriate reminders that reflect their emotional state, and have not adequately supported the emotional well-being of users.

[0798] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0799] In this invention, the server includes means for performing emotion analysis, means for converting speech to text and processing it together with emotion information, and means for extracting and evaluating information using generative artificial intelligence. This makes it possible to generate and notify personalized reminders according to the user's emotional state.

[0800] "Elderly people" refers to a user group that has unique needs as a result of aging.

[0801] "Spoken voice" refers to the patterns of sounds that a user conveys orally, and communication is conducted using these sounds.

[0802] "Emotional analysis" refers to the process of identifying and interpreting emotions and sensibilities from audio or text.

[0803] "Text conversion" refers to the process of converting acquired audio data into text information.

[0804] "Important information" refers to meaningful data and content relevant to the user, and decisions are made based on this information.

[0805] "Generative artificial intelligence" refers to a computer program that has the ability to analyze given data and generate new information.

[0806] A "reminder" refers to a notification or warning set to attract the user's attention, and its purpose is schedule management.

[0807] "Feedback" refers to user reactions and opinions after using a product or service, and this feedback is used to improve the system.

[0808] To implement this invention, a system is used that consists of a terminal placed in the living environment of an elderly person and a server that performs information processing. The terminal is equipped with a voice input device, a display, and a speaker, and includes hardware for acquiring the elderly person's speech in real time. The terminal is also equipped with an emotion engine and has software implemented to recognize emotions from the user's voice.

[0809] The terminal sends the acquired audio data, along with the emotion recognition results, to the server. The server uses a speech recognition module to convert the received audio data into text data. Furthermore, a generative AI model analyzes this text data and emotion information to extract important information from the utterance. The generative AI model utilizes prompts to extract and process information according to the user's needs.

[0810] For example, if a user says, "I have plans to meet a friend tomorrow, but I don't feel like it," the emotion engine recognizes that the user is feeling down. Based on this information, the server generates and sends a reminder as needed, such as, "When you don't feel like it, don't force yourself; just relax." An example of this prompt message would be, "How should the next reminder be adjusted based on what the user said and their emotions?"

[0811] By providing feedback on the reminders users receive, the system can use this feedback to create more efficient reminders. This makes it possible to support the emotional well-being of older adults while improving their quality of life.

[0812] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0813] Step 1:

[0814] The user speaks into the device. The user's voice is input and captured by the device's voice input device. As a result, the user's speech is obtained as digital audio data.

[0815] Step 2:

[0816] The device analyzes the captured audio data using an emotion engine to recognize the user's emotions. The input audio data is analyzed, and the emotion engine outputs the user's emotional state. For example, it can recognize that the user is feeling down based on their tone of voice and word choice.

[0817] Step 3:

[0818] The device sends voice data and emotion recognition results to the server. In this process, the voice data and its emotion analysis results are sent to the server as input. The output is that the server waits for processing.

[0819] Step 4:

[0820] The server uses a speech recognition module to convert audio data into text data. It receives audio data as input, analyzes it with a speech recognition algorithm, and generates text data as output.

[0821] Step 5:

[0822] The generative AI model analyzes text data and sentiment information to extract important information. Text data and sentiment information are provided as input, and the generative AI analyzes the data based on the prompt sentence and outputs important information. For example, if a user says, "I'm dreading going to the hospital tomorrow," the generative AI extracts both the physical appointment and the emotional nuance.

[0823] Step 6:

[0824] The server generates optimized reminders based on the user's emotional state, using extracted information and emotional data. It receives important information and emotions as input, processes them using a reminder generation algorithm, and generates customized notification content as output.

[0825] Step 7:

[0826] The device notifies the user of the generated reminder at the specified time. It receives the generated reminder information as input and outputs an emotionally sensitive audio or visual notification through the speaker or display.

[0827] Step 8:

[0828] The user provides feedback based on the received reminder. This feedback is sent to the server via the terminal as input and is reflected in future prompts and generated content as output.

[0829] (Application Example 2)

[0830] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0831] Traditionally, real-time emotional recognition and the subsequent provision of optimal product recommendations to improve the customer purchasing experience in physical stores have been insufficient. This has made it difficult to provide individualized service tailored to each customer's emotional state, leading to decreased purchasing intent. Furthermore, it has been challenging to identify and respond quickly to situations where customers are in a hurry or have specific needs. Solving these challenges is crucial to improving customer satisfaction and increasing store sales.

[0832] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0833] In this invention, the server includes means for acquiring the spoken voice of a consumer, means for converting the acquired voice into text, and means for extracting important information and performing sentiment analysis using generative artificial intelligence. This makes it possible to adjust appropriate product suggestions and service provision based on the emotional state of the customer.

[0834] "Means for acquiring the spoken voice of consumers" refers to a device or sensor for collecting voice information emitted by consumers in real time.

[0835] "Means of converting acquired audio into text" refers to a speech recognition system that analyzes audio data and converts it into text data.

[0836] "Means for extracting important information" refers to algorithms that analyze important elements related to consumers' intentions and needs from collected and converted text data.

[0837] "Means for recording and organizing information" refers to a system that stores extracted information in a database and manages it systematically so that it can be easily searched later.

[0838] "Methods for setting notifications" refers to a system that schedules reminders and notifications to consumers at appropriate times based on recorded information.

[0839] "Means of providing notifications in a form adjusted according to emotional state" refers to a process that provides notifications with content tailored to the psychological state of the consumer, based on the results of emotional analysis.

[0840] "A means of analyzing consumers' emotional states using an emotion engine and utilizing the analysis results" refers to a process that uses a specific algorithm to determine consumers' emotional states in real time and uses the results to help respond to consumers.

[0841] "Means for extracting important information and performing sentiment analysis using generative artificial intelligence" refers to artificial intelligence technology that analyzes and evaluates consumers' emotions and important information through the analysis of text data.

[0842] "Means of incorporating feedback to improve notification content and methods" refers to the process of collecting responses from consumers and using them to adjust the content and improve the format of future notifications.

[0843] To implement this invention, a system is constructed to improve the shopping experience for customers. First, sensors are installed in the physical store to acquire the spoken voices of customers. These sensors acquire customer voice information in real time and transmit it to a server. The server converts the acquired voice into text using a speech recognition module and extracts important information and emotions using generative artificial intelligence.

[0844] Generative artificial intelligence, such as OpenAI's GPT model, analyzes the emotional state of customers based on their speech and records the results in a database. Furthermore, based on the analysis results, appropriate product suggestions and response methods tailored to the customer's emotions are presented to store staff in real time.

[0845] For example, if a customer says, "I want to finish quickly today," this statement is immediately understood by the system, and that information is transmitted to the staff, enabling them to provide the customer with prompt service.

[0846] An example of a prompt for the generative AI model is: "Generate ways to improve the purchase experience when the customer indicates 'I don't have time' in the emotion engine."

[0847] The emotion engine identifies emotions from customers' facial expressions and tone of voice, and reflects them in sales strategies and customer service attitudes. The system also continuously collects feedback from visitors, using it to improve notification content and further enhance customer satisfaction.

[0848] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0849] Step 1:

[0850] The terminal acquires the customer's spoken voice in real time via sensors. This voice data is then transmitted directly to the server. The input is a voice signal, and the output is the transmission of voice data to the server.

[0851] Step 2:

[0852] The server converts the received audio data into text data using a speech recognition module. In this process, the audio signal is analyzed and converted into character data. The input is the audio data sent to the server, and the output is the text representation of the audio.

[0853] Step 3:

[0854] The server uses generative AI to analyze text data for important information and emotional states. The generative AI model extracts information and evaluates emotions based on prompt text. The input is text data, and the output is the result of important information and emotion recognition.

[0855] Step 4:

[0856] The server records the analyzed key information and emotion recognition results in a database. This allows for the accumulation of analysis data, which can then be used for future analysis and improvement of responses. The input is the analysis results, and the output is the recording in the database.

[0857] Step 5:

[0858] The server sends real-time instructions to store staff to provide appropriate service to customers based on recorded emotional data. This process enables staff to take appropriate actions in the situation. The input is the analysis and recording results, and the output is specific action instructions for store staff.

[0859] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0860] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0861] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0862] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0863] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0864] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0865] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0866] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0867] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0868] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0869] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0870] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0871] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0872] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0873] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0874] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0875] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0876] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0877] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0878] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0879] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0880] The following is further disclosed regarding the embodiments described above.

[0881] (Claim 1)

[0882] A means of acquiring speech sounds from elderly people,

[0883] A means of converting acquired audio into text,

[0884] A means of extracting important information from the converted text,

[0885] Means for recording and organizing the extracted information,

[0886] A means of setting reminders based on recorded information,

[0887] A means of notifying about set reminders,

[0888] A system that includes this.

[0889] (Claim 2)

[0890] The system according to claim 1, comprising means for using generative artificial intelligence to extract important information.

[0891] (Claim 3)

[0892] The system according to claim 1, comprising means for obtaining user feedback and reflecting it in reminder settings.

[0893] "Example 1"

[0894] (Claim 1)

[0895] A device for acquiring speech from elderly people,

[0896] A device that converts acquired audio into text,

[0897] A device that analyzes information from converted text and extracts important information,

[0898] A device for storing and classifying the extracted information,

[0899] A device that sets up notifications based on stored information,

[0900] A device that outputs the configured notification,

[0901] A device that learns from the user's past behavior data to optimize notifications,

[0902] A system that includes this.

[0903] (Claim 2)

[0904] The system according to claim 1, comprising a device for extracting important information from text using generative artificial intelligence.

[0905] (Claim 3)

[0906] The system according to claim 1, further comprising a device for collecting user responses and reflecting them in notification settings.

[0907] "Application Example 1"

[0908] (Claim 1)

[0909] A means of acquiring speech sounds from elderly people,

[0910] A means of converting acquired audio into data,

[0911] A means of extracting important information from the converted data,

[0912] Means for recording and organizing the extracted information,

[0913] A means of proposing actions based on recorded information,

[0914] Means of notifying the proposed action,

[0915] A means of acquiring customer speech,

[0916] A method for analyzing consumer intent from acquired customer voice data,

[0917] A means of recommending products based on the analyzed information,

[0918] A system that includes this.

[0919] (Claim 2)

[0920] The system according to claim 1, comprising means for using generative artificial intelligence to extract important information.

[0921] (Claim 3)

[0922] The system according to claim 1, comprising means for obtaining user feedback and incorporating it into the proposed content.

[0923] "Example 2 of combining an emotion engine"

[0924] (Claim 1)

[0925] A method for acquiring speech voices of elderly people and performing emotion analysis,

[0926] A means for converting acquired audio into text and processing it together with the emotion recognition results,

[0927] A means of extracting important information from converted text and sentiment information,

[0928] A means of recording and organizing the extracted information and emotional states, and utilizing them for future analysis,

[0929] A means of generating and adjusting reminders based on recorded information and emotions,

[0930] A means of notifying users of generated reminders with a gentle voice,

[0931] A means of receiving user feedback and reflecting it in reminders,

[0932] A system that includes this.

[0933] (Claim 2)

[0934] The system according to claim 1, comprising means for using generative artificial intelligence to extract important information and evaluate emotions.

[0935] (Claim 3)

[0936] The system according to claim 1, comprising means for optimizing the content of a reminder while taking into account the user's emotional state.

[0937] "Application example 2 when combining with an emotional engine"

[0938] (Claim 1)

[0939] Means for acquiring the speech of consumers,

[0940] A means of converting acquired audio into text,

[0941] A means of extracting important information from the converted text,

[0942] Means for recording and organizing the extracted information,

[0943] A means of setting up notifications based on recorded information,

[0944] A means of notifying users of pre-configured notifications in a form adjusted according to their emotional state,

[0945] A means of analyzing consumers' emotional states using an emotion engine and utilizing the analysis results,

[0946] A system that includes this.

[0947] (Claim 2)

[0948] The system according to claim 1, comprising means for extracting important information and performing sentiment analysis using generative artificial intelligence.

[0949] (Claim 3)

[0950] The system according to claim 1, comprising means for obtaining feedback from consumers and incorporating that feedback to improve the content and method of notifications. [Explanation of Symbols]

[0951] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of acquiring speech sounds from elderly people, A means of converting acquired audio into text, A means of extracting important information from the converted text, Means for recording and organizing the extracted information, A means of setting reminders based on recorded information, A means of notifying about set reminders, A system that includes this.

2. The system according to claim 1, comprising means for using generative artificial intelligence to extract important information.

3. The system according to claim 1, comprising means for obtaining user feedback and reflecting it in reminder settings.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A