system

A voice-based system addressing loneliness and health management challenges for the elderly through integrated speech recognition, natural language processing, and nutritional support enhances daily life engagement and mental well-being.

JP2026074870APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Elderly individuals face challenges with loneliness, health management, and difficulties in maintaining balanced nutrition and daily activities due to thinning family connections, leading to issues in physical and mental health.

Method used

A system integrating speech recognition, natural language processing, conversation generation, speech synthesis, nutritional evaluation, and schedule management to provide voice-based support for daily life tasks, including nutritional advice and reminders.

Benefits of technology

The system alleviates feelings of loneliness and improves health management by engaging in natural dialogue, providing nutritional guidance, and managing daily schedules through voice interaction, enhancing the quality of life for elderly users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074870000001_ABST
    Figure 2026074870000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A speech recognition means that receives voice input and converts the voice into text data, A natural language processing means that analyzes the aforementioned text data and understands the user's intent, A conversation generation means that generates a response based on the understood user intent, A speech synthesis means that converts the generated response into audio data and outputs it to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, many elderly people are suffering from a sense of isolation and difficulties in health management. Especially in the current situation where the connection with family is thinning, daily conversations and life management are lacking, and it is becoming difficult to maintain physical and mental health. Also, the management of a diet with a balanced nutrition and daily actions such as taking medicine have become major issues.

Means for Solving the Problems

[0005] To solve the above problems, the present invention provides a system equipped with speech recognition means, natural language processing means, conversation generation means, and speech synthesis means. This system aims to alleviate feelings of loneliness by engaging in natural dialogue with the user, understanding the user's intentions, and providing responses. Furthermore, it utilizes a nutrition database reference means and an advice generation means to perform nutritional evaluation of meals and support a healthy diet. In addition, it is equipped with a schedule management means and a voice notification means, enabling the user to smoothly manage their daily life through a reminder function.

[0006] "Speech recognition means" refers to a technology that analyzes input speech and converts it into corresponding text data.

[0007] "Natural language processing means" refers to technologies that interpret user intent from text data and convert it into a format that a computer can understand.

[0008] A "conversation generation method" is a technology that generates appropriate responses based on intentions interpreted through natural language processing.

[0009] "Speech synthesis means" refers to a technology that converts generated text responses into speech data and outputs it to the user as speech.

[0010] A "nutrition database referencing method" is a technology that obtains nutritional information based on the contents of a meal and references a database for evaluating nutritional balance.

[0011] "Advice generation means" refers to technology that generates nutritional advice for users based on nutritional assessment results.

[0012] A "schedule management method" is a technology for managing reminders and appointments and executing them at a predetermined time.

[0013] "Voice notification means" refers to a technology that notifies users of reminders by voice based on a schedule. [Brief explanation of the drawing]

[0014] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is a multi-functional system that supports the daily lives of the elderly by utilizing voice input and output. Specifically, it integrates voice recognition, natural language processing, conversation generation, speech synthesis, nutritional management, and reminder functions. The aim of this system is to alleviate feelings of loneliness and health management problems faced by the elderly and to provide them with a comfortable life.

[0036] When a user speaks to this system, the terminal sends voice data to the server. The server analyzes the received voice data using speech recognition and converts it into text data. The converted text data is then interpreted by natural language processing to understand the user's intent. Based on that intent, a conversation generation system creates an appropriate response, and a speech synthesis system converts it back into voice data and returns it to the user.

[0037] For example, if a user asks their device, "What should I have for breakfast today?", the server analyzes this and makes a specific suggestion, such as, "For breakfast, I recommend yogurt and fruit, which are nutritious and easy to eat." This response is conveyed to the user via voice.

[0038] Furthermore, when a user reports the food they have eaten, the server uses a nutrition database reference to perform a nutritional assessment based on the meal content. The results are then analyzed by an advice generation tool to provide the user with recommended nutritional supplements and health advice. Specifically, if a user inputs "I ate a rice ball today," the server will advise, "You may be lacking protein today. Try incorporating tofu or chicken into your lunch."

[0039] Furthermore, to support users' daily activities, a reminder function is provided using scheduling tools. For example, if a user has set a reminder to take their medication at 8 a.m. every day, the server will send a notification to the device based on the setting, and the device will remind the user with a voice message saying, "It's time to take your medication."

[0040] In this way, the present invention is a system that provides support to the elderly through voice interaction, thereby reducing feelings of loneliness and improving health management.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The user speaks questions or requests into the device. This audio is captured in real time by the device's microphone.

[0044] Step 2:

[0045] The terminal converts the captured audio data into a digital format and sends it to the server. During this process, the audio data is properly encoded and securely transferred to the server.

[0046] Step 3:

[0047] The server converts the received audio data into text data using speech recognition technology. The speech recognition engine analyzes the audio waveform and generates the corresponding text.

[0048] Step 4:

[0049] The server analyzes the converted text data using natural language processing techniques and interprets the data to understand the user's intent and the content of their question.

[0050] Step 5:

[0051] Using conversation generation tools, natural-sounding response text is created based on the user's intent. A pre-trained AI model is utilized in this process.

[0052] Step 6:

[0053] The generated response text is converted into speech data using a speech synthesis system. The synthesized speech is then adjusted to be output in a natural and easy-to-listen-to form for the user.

[0054] Step 7:

[0055] The server sends audio data to the terminal, and the terminal plays this audio data to deliver a response to the user.

[0056] Step 8:

[0057] When a user reports their meal, they provide the meal information to the server via voice. The server analyzes the meal content and evaluates the nutritional balance by referring to a nutrition database.

[0058] Step 9:

[0059] Based on the nutritional assessment, the server generates nutritional advice to support the user's health through an advice generation mechanism and communicates it to the user using the same procedure as the speech synthesis described above.

[0060] Step 10:

[0061] When using the reminder function, the scheduling system prepares a notification based on the time specified by the user. When the specified time arrives, the server sends audio data to the device, and the device announces the user's reminder aloud.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] Loneliness and health management issues are serious problems in the lives of the elderly, and support to alleviate these is needed. However, existing technologies lack an integrated system to support all aspects of the elderly's lives. Therefore, the challenge is to provide a system that integrates the conversation, nutritional management, and reminder functions that the elderly need on a daily basis.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language processing means for analyzing the text data and understanding the user's intent, and conversation generation means for generating a response based on the understood user's intent. This enables a multi-functional integrated system that supports the daily lives of elderly people through voice interaction.

[0067] "Speech recognition means" refers to a technology or device that analyzes speech input and converts it into corresponding text data.

[0068] "Natural language processing means" refers to technologies or devices that analyze text data and interpret the meaning and emotions intended by humans.

[0069] "Conversation generation means" refers to a technology or device for generating an appropriate response based on the analyzed user intent.

[0070] "Speech synthesis means" refers to a technology or device that converts a generated text-based response into speech data and outputs it to the user.

[0071] "Nutritional data reference means" refers to technology or devices that access databases or information sources used to evaluate the content of meals consumed by a user and to analyze nutritional balance.

[0072] "Advice generation means" refers to a technology or device that provides appropriate nutritional advice to the user based on the results of a nutritional assessment.

[0073] "Management means" refers to technology or devices that manage and provide notifications and reminders based on the user's schedule and settings.

[0074] A "voice generation device" is a technology or device that outputs managed schedules and reminders to the user as voice messages.

[0075] This invention is an integrated voice support system that provides multifaceted support for the lives of the elderly, and includes speech recognition, natural language processing, conversation generation, speech synthesis, nutritional management, and schedule management. It is composed of three main components: a server, a terminal, and a user, and a series of processes are carried out using specific hardware and software.

[0076] The server uses cloud-based speech recognition software to convert user voice input into text data. This utilizes common speech recognition technologies such as Google® Cloud Speech-to-Text API. The converted text data is then analyzed by a generative AI model (e.g., a model capable of broad natural language processing) to interpret the user's intent. OpenAI®'s natural language processing technology is one example of such a model.

[0077] Let's say a user asks, "What should I have for breakfast today?" The server then converts the voice data into text and uses a generative AI model to generate an appropriate response. A suggestion like, "A balanced breakfast of fruit and yogurt is recommended," might be generated. This response is then spoken using speech synthesis technology such as Amazon Polly and delivered to the user via their device.

[0078] The terminal plays the role of sending user input to the server and communicating responses from the server to the user. The terminal is equipped with a microphone and speaker, enabling audio input and output.

[0079] Nutritional management for users is done by reporting their meals to the server. For example, in response to a report such as "I ate a rice ball today," the server consults a database and generates advice that takes into account the necessary nutrients. Advice such as "You may be lacking protein. Try incorporating tofu or chicken into your lunch" might be given.

[0080] Furthermore, the server manages user-set reminders using a scheduling system and notifies the user via their device as a voice message. For example, based on a medication reminder set for 8:00 AM, a voice message saying "It's time for your medication" will be sent.

[0081] An example of a prompt using a generative AI model is: "In a health management support system for the elderly, generate a sample conversation that suggests appropriate meals when a user asks about today's meals." This allows the system to provide support tailored to the user's daily life.

[0082] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0083] Step 1:

[0084] The user speaks to the system through a voice input device. The voice input is received as an analog audio signal via a microphone and then converted into digital data. The input can range from user questions to instructions.

[0085] Step 2:

[0086] The terminal transmits audio data received from the user to a server in the cloud via the internet. The data is typically transferred using a specific compression format (e.g., a codec). The input is digital audio data, and the output is a confirmation that the transfer to the server is complete.

[0087] Step 3:

[0088] The server uses speech recognition software to convert digital speech data into text data. For example, it detects pitch and phonemes through speech waveform analysis and identifies words. The input is the received speech data, and the output is the analyzed text data.

[0089] Step 4:

[0090] The server uses a generative AI model to perform natural language processing on text data and analyze the user's intent. It performs grammatical and semantic analysis to understand what the user is asking for. The input is text data, and the output is user intent information.

[0091] Step 5:

[0092] The server uses a conversation generation engine to generate appropriate responses based on the user's intent. It creates natural-sounding response sentences using template-based or AI-driven conversation models. The input is user intent information, and the output is the response text.

[0093] Step 6:

[0094] The server uses speech synthesis technology to convert the response text into audio data. The speech synthesis algorithm reproduces it as natural-sounding speech and saves it in an audio file format. The input is the response text, and the output is the generated audio data.

[0095] Step 7:

[0096] The terminal transmits audio data sent from the server to the user via a playback device. The audio is output through a speaker, which the user can then hear. The input is audio data, and the output is auditory feedback of the sound.

[0097] Step 8:

[0098] When a user reports their meal, the server analyzes the nutritional information by referring to a nutrition database. It calculates the nutritional value based on the reported ingredients and evaluates the balance. The input is the meal information from the user, and the output is the nutritional evaluation result.

[0099] Step 9:

[0100] The server generates nutritional advice for the user based on the nutritional assessment results. It suggests foods and meal plans to supplement any deficient nutrients as needed. The input is the nutritional assessment results, and the output is the advice provided.

[0101] Step 10:

[0102] The server manages reminders according to the user's schedule settings and sends notifications at the designated time. It uses a timer function to trigger set alarms. The input is the user's schedule information, and the output is notification data.

[0103] Step 11:

[0104] The device delivers a voice reminder to the user based on notification data from the server. It utilizes voice notification functionality to prompt the user for the next action. The input is notification data, and the output is a voice reminder message.

[0105] (Application Example 1)

[0106] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0107] In facilities for the elderly, users often experience difficulty finding products or obtaining appropriate information during shopping and product searches. Furthermore, a lack of assistance with selecting nutritionally balanced products and managing shopping lists makes efficient and healthy shopping challenging. Difficulty managing schedules within a set timeframe is another significant issue.

[0108] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0109] In this invention, the server includes: speech recognition means for receiving voice input and converting the voice into text data; natural language processing means for analyzing the text data and understanding the user's intent; conversation generation means for generating a response based on the understood user's intent; speech synthesis means for converting the generated response into voice data and outputting it to the user; information retrieval means for searching for information requested by the user within the facility and returning that information; and location guidance generation means for guiding users to the location of products within the facility and generating related information. This enables elderly people and facility users to efficiently search for products, obtain information that takes nutritional balance into consideration, and have a comfortable shopping experience. Furthermore, it enables users to effectively manage their schedules, which is expected to improve their overall quality of life.

[0110] A "speech recognition means" is a device that receives speech input from a user and converts that speech into text data.

[0111] "Natural language processing means" refers to technologies that analyze text data obtained by speech recognition means to understand the user's intentions and requests.

[0112] A "conversation generation means" is a technology that has the function of generating an appropriate response based on the user's intent understood by a natural language processing means.

[0113] "Speech synthesis means" refers to a function that converts responses generated by conversation generation means into speech data and outputs it to the user.

[0114] "Information retrieval means" refers to technology that searches for information requested by users within a facility and provides the results.

[0115] A "location guidance generation means" is a technology that has the function of guiding users to the location of products within a facility and generating and providing related information to the user.

[0116] "Nutritional information database reference means" refers to a function that allows users to refer to data on the food they have consumed and evaluate its nutritional balance.

[0117] "Guidance generation means" refers to technology for generating nutritional guidance for users based on the evaluation results of a nutritional information database reference means.

[0118] A "time management system" is a technology that has the function of sending notifications to users according to a predetermined time.

[0119] "Voice notification means" refers to a function that outputs notifications set by the time management means in voice format.

[0120] The system designed to realize this application will provide support to users in efficiently searching for products within a facility and obtaining information that takes nutritional balance into consideration. Furthermore, by appropriately communicating necessary information and assistance via voice, it will enable a comfortable shopping experience.

[0121] The server uses speech recognition means to receive voice input and convert that voice into text data. Specifically, voice data entered via a smartphone or dedicated terminal is sent to a cloud server and converted into text data using the Google Speech-to-Text API. The converted text data is analyzed for the user's intent using natural language processing means (e.g., NLTK or SpaCy). Based on the analyzed intent, an appropriate response is generated using conversation generation means (e.g., GPT-3®). The generated response is converted into voice data by speech synthesis means (e.g., Amazon Polly) and output to the user.

[0122] Furthermore, the information retrieval means searches for the information requested by the user within the store and refers to a database to return that information. In addition, the location guidance generation means uses a database that manages product location information within the facility to guide the user to where that product is located.

[0123] For example, if a user asks the terminal, "Where is the milk?", the server analyzes the voice, retrieves information from the product database such as "Milk is in the dairy section," and can then provide voice guidance. In this way, it becomes possible to provide information tailored to the user's needs.

[0124] An example of a prompt message to the generative AI model would be: "The user wants to search for a product. Please return section information for the product specified by the user. The product name is 'milk'."

[0125] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0126] Step 1:

[0127] Voice input begins when the user speaks into the device. The device acquires the user's voice using its microphone and sends it to the server as digital audio data. The input is analog audio, and the output is digital audio data.

[0128] Step 2:

[0129] The server passes the received digital audio data to the Google Speech-to-Text API, which converts the audio into text data. During this process, the audio data is converted into a string based on a specific algorithm. The input is digital audio data, and the output is text data.

[0130] Step 3:

[0131] The server analyzes text data using natural language processing tools (such as NLTK and SpaCy) to extract the user's intent. This is done through grammatical analysis and word semantic analysis. The input is text data, and the output is analyzed data that includes the user's intent.

[0132] Step 4:

[0133] Based on the analyzed user intent, the server generates appropriate responses as prompts using a conversation generation tool (such as GPT-3). At this stage, prompts are constructed and passed to the generation AI model. The input is the analyzed data, and the output is the prompts.

[0134] Step 5:

[0135] The generated prompt sentence is analyzed by a generative AI model to obtain an appropriate response. This response is the final conversational text. The input is the prompt sentence, and the output is the response text.

[0136] Step 6:

[0137] The server converts the response text into speech using a speech synthesis tool (such as Amazon Polly). Here, text data is transformed into natural-sounding speech by a synthesis algorithm. The input is the response text, and the output is speech data.

[0138] Step 7:

[0139] The device receives audio data and outputs it to the user via its speaker. This allows the user to receive answers via voice. The input is audio data, and the output is the presentation of information via voice.

[0140] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0141] This invention is a multi-functional system that combines speech recognition technology and emotion recognition technology to support the daily lives of the elderly. This system analyzes the user's voice input in real time and provides optimal responses and support based on the content and emotional state of the input. In addition to its main voice input / output functions, the system is equipped with an emotion engine, enabling it to address the user's psychological needs.

[0142] First, when a user speaks into the main device, the audio data is collected by the device's microphone. The audio data is sent to a server, where it is converted into text using speech recognition technology. Next, this text information is analyzed by natural language processing technology to understand the user's intent. Then, an emotion engine extracts emotions from the user's voice and compares this with data stored in an emotion database.

[0143] Based on the information obtained in this way, the conversation generation means generates an appropriate response, which is then converted back into natural-sounding speech by the speech synthesis means and output to the user through the terminal. This entire process allows for not only simple verbal exchange but also responses that are tailored to the user's mood.

[0144] For example, if a user says, "I'm feeling a bit down today," the server converts the speech to text and extracts emotional components such as "sadness" through an emotion engine. Based on these results, the system generates an empathetic response for the user, such as, "That's tough. Do you want to talk about it?" and delivers it to the user via voice.

[0145] Furthermore, by monitoring the user's emotional state over the long term, specific patterns and fluctuations can be recognized, and if a state of emotional outburst persists, support such as encouraging consultation with a specialist can be proposed. Thus, the present invention aims to help the mental stability of the elderly through emotional recognition.

[0146] The following describes the processing flow.

[0147] Step 1:

[0148] The user speaks into the device, making questions and statements via voice. The device acquires this voice as digital audio data.

[0149] Step 2:

[0150] The terminal sends the acquired audio data to the server. The server receives the audio data and converts it into text data using speech recognition technology. Noise is removed during this process to ensure accurate text conversion.

[0151] Step 3:

[0152] The server analyzes the converted text data using natural language processing techniques to extract the user's intent. During this process, it understands the intent of the question by interpreting the text content within its context.

[0153] Step 4:

[0154] Text data is passed to an emotion engine, which determines the user's emotions based on speech intonation and text expression. Emotional components are extracted and compared with an emotion database.

[0155] Step 5:

[0156] Based on the analyzed intent and emotion data, the server uses a conversation generation mechanism to create an appropriate response to the user. This response is generated in a way that resonates with the user's emotions.

[0157] Step 6:

[0158] The generated response text is converted into speech data by a speech synthesis system. Here, the tone of the voice is adjusted to produce a natural and gentle voice.

[0159] Step 7:

[0160] The server converts the audio data and sends it to the terminal. The terminal plays a voice response back to the user and observes whether the user's emotions remain stable after the conversation ends.

[0161] Step 8:

[0162] User interaction history and emotional data are anonymized and stored in a database for long-term monitoring and analysis of the user's mental state. This data is used for regular feedback and support services as needed.

[0163] (Example 2)

[0164] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0165] In the daily lives of the elderly and those requiring assistance, there is a need to provide interactive life support, including psychological support. However, simple voice recognition and digital information provision cannot adequately improve the quality of communication, and providing support that takes emotions into consideration is particularly difficult. Furthermore, there are challenges in providing support tailored to individual circumstances, such as managing diet and schedules.

[0166] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0167] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into digital text, language analysis means for analyzing the digital text and understanding the communicator's intent, and emotion analysis means for extracting emotional information from the digital text and recognizing the emotional state. This makes it possible to generate responses that take the user's emotions into consideration, thereby improving the psychological stability and quality of life of elderly people.

[0168] "Voice input" is a means for users to communicate information and instructions to a system through sound or words.

[0169] "Digital text" refers to data that is generated by analyzing voice input and representing it as textual information.

[0170] "Speech recognition means" refers to technology for receiving speech input and converting it into digital text.

[0171] "Linguistic analysis tools" are technologies that analyze digital text to understand the user's intent and the content of their questions.

[0172] "Emotional analysis methods" are technologies for extracting and recognizing a user's emotional state from digital text or audio.

[0173] A "response generation means" is a technology for creating an appropriate response based on the user's intentions and emotional state.

[0174] "Speech conversion means" refers to a technology that converts text into speech in order to output the generated response as sound.

[0175] A "nutritional information reference method" is a technology for evaluating the nutritional composition of food based on data from foods consumed by the user.

[0176] A "guidance generation method" is a technology that provides appropriate nutritional guidance to users based on the evaluation results of nutritional information.

[0177] A "schedule management system" is a technology that notifies users of pre-set reminders and appointments at the appropriate time.

[0178] "Voice notification means" refers to a technology that communicates schedules and reminders to users through voice output.

[0179] This invention is a system that provides interactive support in daily life for the elderly and individuals who require assistance. This system combines speech recognition technology and emotion recognition technology, aiming to improve the user's mental well-being and quality of life.

[0180] The user provides voice input through the main device's microphone. This voice input is converted into digital data on the device and sent to the server. The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the voice data into digital text. This digital text is then analyzed by a language analysis engine (e.g., Google Cloud Natural Language API) to understand the user's intent.

[0181] Next, emotion recognition is the process of extracting emotional information from received audio and text data. This is done using an emotion analysis engine (e.g., IBM Watson® Tone Analyzer). This process identifies the user's emotional state and compares it against an emotion database.

[0182] The server uses a generative AI model (e.g., GPT-3) based on the analysis results to generate an appropriate response that matches the user's intent and emotions. This generated response is then converted into audio format using a speech synthesis engine (e.g., Amazon Polly). Finally, this audio data is sent to the device and delivered to the user through the speaker.

[0183] For example, if a user says, "I'm feeling a bit down today," the server converts this speech into text and recognizes an emotion such as "sadness." Based on this, it generates an empathetic response such as, "That's tough. Do you want to talk about it?" and conveys it to the user verbally.

[0184] An example of a prompt might be, "How would you respond if the user said, 'I'm feeling a bit down today'?" This prompt allows the AI ​​to provide a response that takes the user's emotions into consideration.

[0185] This allows users to receive personalized support tailored to their individual emotions, while also contributing to the psychological well-being and problem-solving of elderly individuals.

[0186] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0187] Step 1:

[0188] The user provides voice input through the main terminal's microphone. The terminal converts the voice into a digital signal and sends this signal to the server. The input is the user's spoken voice, and the output is digitized voice data. By collecting the voice signal, preparations are made for the next process.

[0189] Step 2:

[0190] The server uses a speech recognition engine to convert received digital audio data into text format. In this process, it identifies the content of the speech from the audio signal and represents it in text. The input is digital audio data, and the output is text data. Specifically, it analyzes phonemes and converts them into strings.

[0191] Step 3:

[0192] The server performs natural language processing to analyze text data. The language analysis engine performs syntactic and semantic analysis to extract the communicator's intent from the text data and understand the user's needs. The input is text data, and the output is data corresponding to an understanding of the user's intent. This enables action selection based on the content of the dialogue.

[0193] Step 4:

[0194] The server uses an emotion analysis engine to extract emotional information from text and audio data. By analyzing the type and intensity of emotions, it recognizes the user's emotional state. Input is text or audio data, and output is data indicating the characteristics of the emotion. This evaluation is based on matching against an emotion database.

[0195] Step 5:

[0196] The server generates responses using a generative AI model based on the user's intent and recognized sentiment information. In this process, it selects appropriate responses from input data and constructs personalized messages. The input is the user's intent and sentiment data, and the output is the generated response text. The generative AI combines the information to create natural conversation.

[0197] Step 6:

[0198] The server converts the generated response text into speech data using a speech synthesis engine. The speech synthesis process generates natural and understandable speech from the text data and pronounces it. The input is the response text, and the output is synthesized speech data. Computer synthesis enables playback in a natural voice.

[0199] Step 7:

[0200] The server sends synthesized speech data to the terminal. The terminal completes the interaction by delivering this speech data to the user through its speaker. The input is synthesized speech data, and the output is the sound that reaches the user's ears. As a result, the user can hear the system's response and continue the interaction.

[0201] (Application Example 2)

[0202] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0203] In the lives of the elderly, there is a need for support systems that alleviate daily anxieties and feelings of loneliness, and enable them to live safely and securely. However, existing technologies make it difficult to provide real-time support that is tailored to the user's emotional state. To solve this problem, a system is needed that integrates speech recognition and emotion recognition to provide appropriate security support.

[0204] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0205] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language processing means for analyzing the text data and understanding the user's intent, and emotion recognition means for extracting emotion information from the user's voice data. This makes it possible to provide appropriate responses and security support in real time according to the user's emotional state.

[0206] "Voice input" refers to voice data from the user, which is information that the system receives and processes.

[0207] "Text data" refers to the character information of voice input converted by a speech recognition system.

[0208] "Speech recognition means" refers to a mechanism or device for converting speech input into text data.

[0209] "Natural language processing methods" refer to technologies and processes for analyzing text data and understanding the user's intent and meaning.

[0210] "Conversation generation means" refers to technologies and functions for generating appropriate responses based on the user's intent.

[0211] "Speech synthesis means" refers to a device or technology that converts responses generated by a conversation generation means into speech data and outputs it to the user.

[0212] "Emotion recognition means" refers to technologies and devices for extracting emotional information from a user's voice data.

[0213] "Response generation means" refers to technologies and mechanisms for providing appropriate security support to users based on emotional information.

[0214] "Security support" refers to assistance and services provided to enhance safety and peace of mind, tailored to the user's emotional state.

[0215] The system for realizing this invention comprises speech recognition means, natural language processing means, emotion recognition means, conversation generation means, speech synthesis means, and response generation means.

[0216] The terminal receives voice input from the user and sends it to the server as digital data. This voice data is converted into text data on the server by speech recognition. Next, this text data is analyzed by natural language processing to determine the user's intent.

[0217] Subsequently, the emotion recognition system extracts emotion information from the user's voice, which is then compared with an emotion database. Using prompt sentences, a generative AI model generates an appropriate response, and the response to the user is constructed through the conversation generation system. The generated response is converted into natural-sounding speech by the speech synthesis system and output to the user through the terminal.

[0218] This system utilizes Python's speech_recognition library and Hugging Face's transformers to perform speech-to-text conversion and emotion recognition. It also uses Google's gTTS (Google Text-to-Speech) for speech synthesis. This enables the provision of real-time security support tailored to the user's emotional state.

[0219] For example, if a user says, "I've been worried a lot lately," the system will analyze that emotion and respond in a calm voice, "Why don't you try taking a break?" Examples of prompt messages include the following:

[0220] Input: I've been feeling stressed lately.

[0221] Expected output: I recommend taking a short rest.

[0222] In this way, the system enables interactive conversations that reassure users while being sensitive to their emotions. As a security support system, it can help reduce anxieties in daily life for the elderly, enabling them to live with greater peace of mind.

[0223] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0224] Step 1:

[0225] The user provides voice input to the terminal. The terminal records the voice as digital data and sends it to the server. The input is raw voice data, and the output is the digital format of that voice data. The terminal captures the voice using a microphone and sends it to the server using its data communication function.

[0226] Step 2:

[0227] The server converts received audio data into text data using speech recognition technology. The input is audio data, and the output is the converted text data. The Python `speech_recognition` library is used for this process. This allows the audio data to be analyzed as linguistic information.

[0228] Step 3:

[0229] The server analyzes text data using natural language processing techniques to understand the user's intent. The input is text data, and the output is the analysis result representing the user's intent. This process utilizes the Hugging Face transformers library, with a generative AI model assisting the analysis.

[0230] Step 4:

[0231] Emotional information is extracted from the user's voice using emotion recognition technology. The input is voice data, and the output is emotional information. For emotion recognition, the transformers library is also used to analyze the type and intensity of the emotion.

[0232] Step 5:

[0233] The server generates conversations based on the user's intent and sentiment information. The input is intent and sentiment information, and the output is the text of the generated conversation. In this process, a generative AI model uses prompt sentences to determine the content of the response.

[0234] Step 6:

[0235] The generated conversation text is converted into speech data using a speech synthesis method and sent to the terminal. The input is conversation text, and the output is speech data. Google's gTTS is used for this process to convert it into natural-sounding speech.

[0236] Step 7:

[0237] The terminal plays audio data received from the server for the user. The input is audio data, and the output is audio playback through the terminal. The terminal physically plays the audio using a speaker and responds to the user.

[0238] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0239] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0240] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0241] [Second Embodiment]

[0242] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0243] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0244] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0245] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0246] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0247] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0248] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0249] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0250] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0251] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0252] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0253] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0254] This invention is a multi-functional system that supports the daily lives of the elderly by utilizing voice input and output. Specifically, it integrates voice recognition, natural language processing, conversation generation, speech synthesis, nutritional management, and reminder functions. The aim of this system is to alleviate feelings of loneliness and health management problems faced by the elderly and to provide them with a comfortable life.

[0255] When a user speaks to this system, the terminal sends voice data to the server. The server analyzes the received voice data using speech recognition and converts it into text data. The converted text data is then interpreted by natural language processing to understand the user's intent. Based on that intent, a conversation generation system creates an appropriate response, and a speech synthesis system converts it back into voice data and returns it to the user.

[0256] For example, if a user asks their device, "What should I have for breakfast today?", the server analyzes this and makes a specific suggestion, such as, "For breakfast, I recommend yogurt and fruit, which are nutritious and easy to eat." This response is conveyed to the user via voice.

[0257] Furthermore, when a user reports the food they have eaten, the server uses a nutrition database reference to perform a nutritional assessment based on the meal content. The results are then analyzed by an advice generation tool to provide the user with recommended nutritional supplements and health advice. Specifically, if a user inputs "I ate a rice ball today," the server will advise, "You may be lacking protein today. Try incorporating tofu or chicken into your lunch."

[0258] Furthermore, to support users' daily activities, a reminder function is provided using scheduling tools. For example, if a user has set a reminder to take their medication at 8 a.m. every day, the server will send a notification to the device based on the setting, and the device will remind the user with a voice message saying, "It's time to take your medication."

[0259] In this way, the present invention is a system that provides support to the elderly through voice interaction, thereby reducing feelings of loneliness and improving health management.

[0260] The following describes the processing flow.

[0261] Step 1:

[0262] The user speaks questions or requests into the device. This audio is captured in real time by the device's microphone.

[0263] Step 2:

[0264] The terminal converts the captured audio data into a digital format and sends it to the server. During this process, the audio data is properly encoded and securely transferred to the server.

[0265] Step 3:

[0266] The server converts the received audio data into text data using speech recognition technology. The speech recognition engine analyzes the audio waveform and generates the corresponding text.

[0267] Step 4:

[0268] The server analyzes the converted text data using natural language processing techniques and interprets the data to understand the user's intent and the content of their question.

[0269] Step 5:

[0270] Using conversation generation tools, natural-sounding response text is created based on the user's intent. A pre-trained AI model is utilized in this process.

[0271] Step 6:

[0272] The generated response text is converted into speech data using a speech synthesis system. The synthesized speech is then adjusted to be output in a natural and easy-to-listen-to form for the user.

[0273] Step 7:

[0274] The server sends audio data to the terminal, and the terminal plays this audio data to deliver a response to the user.

[0275] Step 8:

[0276] When the user reports the content of a meal, the information about the meal is provided to the server through voice as well. The server analyzes the meal content and evaluates the nutritional balance by referring to the nutritional database.

[0277] Step 9:

[0278] Based on the nutritional evaluation, the server generates nutritional advice for supporting the user's health through the advice generation means and conveys it to the user in the same procedure as the above-mentioned voice synthesis.

[0279] Step 10:

[0280] When using the reminder function, the schedule management means prepares a notification based on the time specified by the user. When the specified time arrives, voice data is transmitted from the server to the terminal, and the terminal conveys the user's reminder by voice.

[0281] (Example 1)

[0282] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0283] In the life of the elderly, the problems of loneliness and health management are serious, and support for reducing these is required. However, in existing technologies, there is a lack of an integrated system for supporting the entire life of the elderly. Therefore, it is an issue to provide a system that integrates the conversation, nutritional management, and reminder functions that the elderly need daily.

[0284] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 1 is realized by the following respective means.

[0285] In this invention, the server includes voice recognition means for receiving voice input and converting the voice into text data, natural language processing means for analyzing the text data to understand the user's intention, and conversation generation means for generating a response based on the understood user's intention. As a result, a multifunctional integrated system that supports daily life through voice interaction with the elderly becomes possible.

[0286] The "voice recognition means" is a technology or device that analyzes voice input and converts it into corresponding text data.

[0287] The "natural language processing means" is a technology or device that analyzes text data and interprets the meaning and emotion intended by humans.

[0288] The "conversation generation means" is a technology or device for generating an appropriate response based on the analyzed user's intention.

[0289] The "voice synthesis means" is a technology or device that converts the generated text-based response into voice data and outputs it to the user.

[0290] The "nutritional data reference means" is a technology or device that accesses a database or information source used to evaluate the content of the meals consumed by the user and analyze the nutritional balance.

[0291] The "advice creation means" is a technology or device that provides appropriate nutritional advice to the user based on the nutritional evaluation results.

[0292] The "management means" is a technology or device that manages and provides notifications and reminders based on the user's schedule and settings.

[0293] The "voice generation device" is a technology or device that outputs the managed schedule and reminders to the user as voice messages.

[0294] This invention is an integrated voice support system that provides multifaceted support for the lives of the elderly, and includes speech recognition, natural language processing, conversation generation, speech synthesis, nutritional management, and schedule management. It is composed of three main components: a server, a terminal, and a user, and a series of processes are carried out using specific hardware and software.

[0295] The server uses cloud-based speech recognition software to convert user voice input into text data. This utilizes common speech recognition technologies such as the Google Cloud Speech-to-Text API. The converted text data is then analyzed by a generative AI model (e.g., a model capable of broad natural language processing) to interpret the user's intent. OpenAI's natural language processing technology is one example of this.

[0296] Let's say a user asks, "What should I have for breakfast today?" The server then converts the voice data into text and uses a generative AI model to generate an appropriate response. A suggestion like, "A balanced breakfast of fruit and yogurt is recommended," might be generated. This response is then spoken using speech synthesis technology such as Amazon Polly and delivered to the user via their device.

[0297] The terminal plays the role of sending user input to the server and communicating responses from the server to the user. The terminal is equipped with a microphone and speaker, enabling audio input and output.

[0298] Nutritional management for users is done by reporting their meals to the server. For example, in response to a report such as "I ate a rice ball today," the server consults a database and generates advice that takes into account the necessary nutrients. Advice such as "You may be lacking protein. Try incorporating tofu or chicken into your lunch" might be given.

[0299] In addition, the server manages the reminders set by the user using the schedule management means and notifies the user through the terminal as a voice message. For example, based on the reminder setting for taking medicine at 8:00 am, the user is notified by voice that "It's time to take your medicine."

[0300] As an example of a prompt sentence using the generation AI model, there is a form such as "In a health management support system for the elderly, when the user asks a question about today's meal, please generate a conversation example that provides appropriate meal suggestions." This enables the system to provide support tailored to the user's daily life.

[0301] The flow of the specific process in Example 1 will be described using FIG. 11.

[0302] Step 1:

[0303] The user addresses the system through the voice input device. The voice input is received as an analog voice signal through the microphone and then converted into digital data. The input covers a wide range, such as the user's questions and instructions.

[0304] Step 2:

[0305] The terminal transmits the voice data received from the user to the server on the cloud via the Internet. The data is usually transferred using a certain compression format (e.g., Codec). The input is digital voice data, and the output is the confirmation of the transfer completion to the server.

[0306] Step 3:

[0307] The server uses voice recognition software to convert the digital voice data into text data. For example, the pitch and phonemes are detected by analyzing the waveform of the voice to identify words. The input is the received voice data, and the output is the analyzed text data.

[0308] Step 4:

[0309] The server uses a generative AI model to perform natural language processing on text data and analyze the user's intent. It performs grammatical and semantic analysis to understand what the user is asking for. The input is text data, and the output is user intent information.

[0310] Step 5:

[0311] The server uses a conversation generation engine to generate appropriate responses based on the user's intent. It creates natural-sounding response sentences using template-based or AI-driven conversation models. The input is user intent information, and the output is the response text.

[0312] Step 6:

[0313] The server uses speech synthesis technology to convert the response text into audio data. The speech synthesis algorithm reproduces it as natural-sounding speech and saves it in an audio file format. The input is the response text, and the output is the generated audio data.

[0314] Step 7:

[0315] The terminal transmits audio data sent from the server to the user via a playback device. The audio is output through a speaker, which the user can then hear. The input is audio data, and the output is auditory feedback of the sound.

[0316] Step 8:

[0317] When a user reports their meal, the server analyzes the nutritional information by referring to a nutrition database. It calculates the nutritional value based on the reported ingredients and evaluates the balance. The input is the meal information from the user, and the output is the nutritional evaluation result.

[0318] Step 9:

[0319] The server generates nutritional advice for the user based on the nutritional assessment results. It suggests foods and meal plans to supplement any deficient nutrients as needed. The input is the nutritional assessment results, and the output is the advice provided.

[0320] Step 10:

[0321] The server manages reminders according to the user's schedule settings and sends notifications at the designated time. It uses a timer function to trigger set alarms. The input is the user's schedule information, and the output is notification data.

[0322] Step 11:

[0323] The device delivers a voice reminder to the user based on notification data from the server. It utilizes voice notification functionality to prompt the user for the next action. The input is notification data, and the output is a voice reminder message.

[0324] (Application Example 1)

[0325] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0326] In facilities for the elderly, users often experience difficulty finding products or obtaining appropriate information during shopping and product searches. Furthermore, a lack of assistance with selecting nutritionally balanced products and managing shopping lists makes efficient and healthy shopping challenging. Difficulty managing schedules within a set timeframe is another significant issue.

[0327] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0328] In this invention, the server includes: speech recognition means for receiving voice input and converting the voice into text data; natural language processing means for analyzing the text data and understanding the user's intent; conversation generation means for generating a response based on the understood user's intent; speech synthesis means for converting the generated response into voice data and outputting it to the user; information retrieval means for searching for information requested by the user within the facility and returning that information; and location guidance generation means for guiding users to the location of products within the facility and generating related information. This enables elderly people and facility users to efficiently search for products, obtain information that takes nutritional balance into consideration, and have a comfortable shopping experience. Furthermore, it enables users to effectively manage their schedules, which is expected to improve their overall quality of life.

[0329] A "speech recognition means" is a device that receives speech input from a user and converts that speech into text data.

[0330] "Natural language processing means" refers to technologies that analyze text data obtained by speech recognition means to understand the user's intentions and requests.

[0331] A "conversation generation means" is a technology that has the function of generating an appropriate response based on the user's intent understood by a natural language processing means.

[0332] "Speech synthesis means" refers to a function that converts responses generated by conversation generation means into speech data and outputs it to the user.

[0333] "Information retrieval means" refers to technology that searches for information requested by users within a facility and provides the results.

[0334] A "location guidance generation means" is a technology that has the function of guiding users to the location of products within a facility and generating and providing related information to the user.

[0335] "Nutritional information database reference means" refers to a function that allows users to refer to data on the food they have consumed and evaluate its nutritional balance.

[0336] "Guidance generation means" refers to technology for generating nutritional guidance for users based on the evaluation results of a nutritional information database reference means.

[0337] A "time management system" is a technology that has the function of sending notifications to users according to a predetermined time.

[0338] "Voice notification means" refers to a function that outputs notifications set by the time management means in voice format.

[0339] The system designed to realize this application will provide support to users in efficiently searching for products within a facility and obtaining information that takes nutritional balance into consideration. Furthermore, by appropriately communicating necessary information and assistance via voice, it will enable a comfortable shopping experience.

[0340] The server uses speech recognition technology to receive voice input and convert it into text data. Specifically, voice data entered via a smartphone or dedicated terminal is sent to a cloud server and converted into text data using the Google Speech-to-Text API. The converted text data is then analyzed for the user's intent using natural language processing technology (e.g., NLTK or SpaCy). Based on the analyzed intent, an appropriate response is generated using conversation generation technology (e.g., GPT-3). The generated response is then converted into voice data using speech synthesis technology (e.g., Amazon Polly) and output to the user.

[0341] Furthermore, the information retrieval means searches for the information requested by the user within the store and refers to a database to return that information. In addition, the location guidance generation means uses a database that manages product location information within the facility to guide the user to where that product is located.

[0342] For example, if a user asks the terminal, "Where is the milk?", the server analyzes the voice, retrieves information from the product database such as "Milk is in the dairy section," and can then provide voice guidance. In this way, it becomes possible to provide information tailored to the user's needs.

[0343] An example of a prompt message to the generative AI model would be: "The user wants to search for a product. Please return section information for the product specified by the user. The product name is 'milk'."

[0344] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0345] Step 1:

[0346] Voice input begins when the user speaks into the device. The device acquires the user's voice using its microphone and sends it to the server as digital audio data. The input is analog audio, and the output is digital audio data.

[0347] Step 2:

[0348] The server passes the received digital audio data to the Google Speech-to-Text API, which converts the audio into text data. During this process, the audio data is converted into a string based on a specific algorithm. The input is digital audio data, and the output is text data.

[0349] Step 3:

[0350] The server analyzes text data using natural language processing tools (such as NLTK and SpaCy) to extract the user's intent. This is done through grammatical analysis and word semantic analysis. The input is text data, and the output is analyzed data that includes the user's intent.

[0351] Step 4:

[0352] Based on the analyzed user intent, the server generates appropriate responses as prompts using a conversation generation tool (such as GPT-3). At this stage, prompts are constructed and passed to the generation AI model. The input is the analyzed data, and the output is the prompts.

[0353] Step 5:

[0354] The generated prompt sentence is analyzed by a generative AI model to obtain an appropriate response. This response is the final conversational text. The input is the prompt sentence, and the output is the response text.

[0355] Step 6:

[0356] The server converts the response text into speech using a speech synthesis tool (such as Amazon Polly). Here, text data is transformed into natural-sounding speech by a synthesis algorithm. The input is the response text, and the output is speech data.

[0357] Step 7:

[0358] The device receives audio data and outputs it to the user via its speaker. This allows the user to receive answers via voice. The input is audio data, and the output is the presentation of information via voice.

[0359] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0360] This invention is a multi-functional system that combines speech recognition technology and emotion recognition technology to support the daily lives of the elderly. This system analyzes the user's voice input in real time and provides optimal responses and support based on the content and emotional state of the input. In addition to its main voice input / output functions, the system is equipped with an emotion engine, enabling it to address the user's psychological needs.

[0361] First, when a user speaks into the main device, the audio data is collected by the device's microphone. The audio data is sent to a server, where it is converted into text using speech recognition technology. Next, this text information is analyzed by natural language processing technology to understand the user's intent. Then, an emotion engine extracts emotions from the user's voice and compares this with data stored in an emotion database.

[0362] Based on the information obtained in this way, the conversation generation means generates an appropriate response, which is then converted back into natural-sounding speech by the speech synthesis means and output to the user through the terminal. This entire process allows for not only simple verbal exchange but also responses that are tailored to the user's mood.

[0363] For example, if a user says, "I'm feeling a bit down today," the server converts the speech to text and extracts emotional components such as "sadness" through an emotion engine. Based on these results, the system generates an empathetic response for the user, such as, "That's tough. Do you want to talk about it?" and delivers it to the user via voice.

[0364] Furthermore, by monitoring the user's emotional state over the long term, specific patterns and fluctuations can be recognized, and if a state of emotional outburst persists, support such as encouraging consultation with a specialist can be proposed. Thus, the present invention aims to help the mental stability of the elderly through emotional recognition.

[0365] The following describes the processing flow.

[0366] Step 1:

[0367] The user speaks into the device, making questions and statements via voice. The device acquires this voice as digital audio data.

[0368] Step 2:

[0369] The terminal sends the acquired audio data to the server. The server receives the audio data and converts it into text data using speech recognition technology. Noise is removed during this process to ensure accurate text conversion.

[0370] Step 3:

[0371] The server analyzes the converted text data using natural language processing techniques to extract the user's intent. During this process, it understands the intent of the question by interpreting the text content within its context.

[0372] Step 4:

[0373] Text data is passed to an emotion engine, which determines the user's emotions based on speech intonation and text expression. Emotional components are extracted and compared with an emotion database.

[0374] Step 5:

[0375] Based on the analyzed intent and emotion data, the server uses a conversation generation mechanism to create an appropriate response to the user. This response is generated in a way that resonates with the user's emotions.

[0376] Step 6:

[0377] The generated response text is converted into speech data by a speech synthesis system. Here, the tone of the voice is adjusted to produce a natural and gentle voice.

[0378] Step 7:

[0379] The server converts the audio data and sends it to the terminal. The terminal plays a voice response back to the user and observes whether the user's emotions remain stable after the conversation ends.

[0380] Step 8:

[0381] User interaction history and emotional data are anonymized and stored in a database for long-term monitoring and analysis of the user's mental state. This data is used for regular feedback and support services as needed.

[0382] (Example 2)

[0383] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0384] In the daily lives of the elderly and those requiring assistance, there is a need to provide interactive life support, including psychological support. However, simple voice recognition and digital information provision cannot adequately improve the quality of communication, and providing support that takes emotions into consideration is particularly difficult. Furthermore, there are challenges in providing support tailored to individual circumstances, such as managing diet and schedules.

[0385] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0386] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into digital text, language analysis means for analyzing the digital text and understanding the communicator's intent, and emotion analysis means for extracting emotional information from the digital text and recognizing the emotional state. This makes it possible to generate responses that take the user's emotions into consideration, thereby improving the psychological stability and quality of life of elderly people.

[0387] "Voice input" is a means for users to communicate information and instructions to a system through sound or words.

[0388] "Digital text" refers to data that is generated by analyzing voice input and representing it as textual information.

[0389] "Speech recognition means" refers to technology for receiving speech input and converting it into digital text.

[0390] "Linguistic analysis tools" are technologies that analyze digital text to understand the user's intent and the content of their questions.

[0391] "Emotional analysis methods" are technologies for extracting and recognizing a user's emotional state from digital text or audio.

[0392] A "response generation means" is a technology for creating an appropriate response based on the user's intentions and emotional state.

[0393] "Speech conversion means" refers to a technology that converts text into speech in order to output the generated response as sound.

[0394] A "nutritional information reference method" is a technology for evaluating the nutritional composition of food based on data from foods consumed by the user.

[0395] A "guidance generation method" is a technology that provides appropriate nutritional guidance to users based on the evaluation results of nutritional information.

[0396] A "schedule management system" is a technology that notifies users of pre-set reminders and appointments at the appropriate time.

[0397] "Voice notification means" refers to a technology that communicates schedules and reminders to users through voice output.

[0398] This invention is a system that provides interactive support in daily life for the elderly and individuals who require assistance. This system combines speech recognition technology and emotion recognition technology, aiming to improve the user's mental well-being and quality of life.

[0399] The user provides voice input through the main device's microphone. This voice input is converted into digital data on the device and sent to the server. The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the voice data into digital text. This digital text is then analyzed by a language analysis engine (e.g., Google Cloud Natural Language API) to understand the user's intent.

[0400] Next, emotion recognition is the process of extracting emotional information from received audio and text data. This is done using an emotion analysis engine (e.g., IBM Watson Tone Analyzer). This process identifies the user's emotional state and compares it against an emotion database.

[0401] The server uses a generative AI model (e.g., GPT-3) based on the analysis results to generate an appropriate response that matches the user's intent and emotions. This generated response is then converted into audio format using a speech synthesis engine (e.g., Amazon Polly). Finally, this audio data is sent to the device and delivered to the user through the speaker.

[0402] For example, if a user says, "I'm feeling a bit down today," the server converts this speech into text and recognizes an emotion such as "sadness." Based on this, it generates an empathetic response such as, "That's tough. Do you want to talk about it?" and conveys it to the user verbally.

[0403] An example of a prompt might be, "How would you respond if the user said, 'I'm feeling a bit down today'?" This prompt allows the AI ​​to provide a response that takes the user's emotions into consideration.

[0404] This allows users to receive personalized support tailored to their individual emotions, while also contributing to the psychological well-being and problem-solving of elderly individuals.

[0405] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0406] Step 1:

[0407] The user provides voice input through the main terminal's microphone. The terminal converts the voice into a digital signal and sends this signal to the server. The input is the user's spoken voice, and the output is digitized voice data. By collecting the voice signal, preparations are made for the next process.

[0408] Step 2:

[0409] The server uses a speech recognition engine to convert received digital audio data into text format. In this process, it identifies the content of the speech from the audio signal and represents it in text. The input is digital audio data, and the output is text data. Specifically, it analyzes phonemes and converts them into strings.

[0410] Step 3:

[0411] The server performs natural language processing to analyze text data. The language analysis engine performs syntactic and semantic analysis to extract the communicator's intent from the text data and understand the user's needs. The input is text data, and the output is data corresponding to an understanding of the user's intent. This enables action selection based on the content of the dialogue.

[0412] Step 4:

[0413] The server uses an emotion analysis engine to extract emotional information from text and audio data. By analyzing the type and intensity of emotions, it recognizes the user's emotional state. Input is text or audio data, and output is data indicating the characteristics of the emotion. This evaluation is based on matching against an emotion database.

[0414] Step 5:

[0415] The server generates responses using a generative AI model based on the user's intent and recognized sentiment information. In this process, it selects appropriate responses from input data and constructs personalized messages. The input is the user's intent and sentiment data, and the output is the generated response text. The generative AI combines the information to create natural conversation.

[0416] Step 6:

[0417] The server converts the generated response text into speech data using a speech synthesis engine. The speech synthesis process generates natural and understandable speech from the text data and pronounces it. The input is the response text, and the output is synthesized speech data. Computer synthesis enables playback in a natural voice.

[0418] Step 7:

[0419] The server sends synthesized speech data to the terminal. The terminal completes the interaction by delivering this speech data to the user through its speaker. The input is synthesized speech data, and the output is the sound that reaches the user's ears. As a result, the user can hear the system's response and continue the interaction.

[0420] (Application Example 2)

[0421] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0422] In the lives of the elderly, there is a need for support systems that alleviate daily anxieties and feelings of loneliness, and enable them to live safely and securely. However, existing technologies make it difficult to provide real-time support that is tailored to the user's emotional state. To solve this problem, a system is needed that integrates speech recognition and emotion recognition to provide appropriate security support.

[0423] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0424] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language processing means for analyzing the text data and understanding the user's intent, and emotion recognition means for extracting emotion information from the user's voice data. This makes it possible to provide appropriate responses and security support in real time according to the user's emotional state.

[0425] "Voice input" refers to voice data from the user, which is information that the system receives and processes.

[0426] "Text data" refers to the character information of voice input converted by a speech recognition system.

[0427] "Speech recognition means" refers to a mechanism or device for converting speech input into text data.

[0428] "Natural language processing methods" refer to technologies and processes for analyzing text data and understanding the user's intent and meaning.

[0429] "Conversation generation means" refers to technologies and functions for generating appropriate responses based on the user's intent.

[0430] "Speech synthesis means" refers to a device or technology that converts responses generated by a conversation generation means into speech data and outputs it to the user.

[0431] "Emotion recognition means" refers to technologies and devices for extracting emotional information from a user's voice data.

[0432] "Response generation means" refers to technologies and mechanisms for providing appropriate security support to users based on emotional information.

[0433] "Security support" refers to assistance and services provided to enhance safety and peace of mind, tailored to the user's emotional state.

[0434] The system for realizing this invention comprises speech recognition means, natural language processing means, emotion recognition means, conversation generation means, speech synthesis means, and response generation means.

[0435] The terminal receives voice input from the user and sends it to the server as digital data. This voice data is converted into text data on the server by speech recognition. Next, this text data is analyzed by natural language processing to determine the user's intent.

[0436] Subsequently, the emotion recognition system extracts emotion information from the user's voice, which is then compared with an emotion database. Using prompt sentences, a generative AI model generates an appropriate response, and the response to the user is constructed through the conversation generation system. The generated response is converted into natural-sounding speech by the speech synthesis system and output to the user through the terminal.

[0437] This system utilizes Python's speech_recognition library and Hugging Face's transformers to perform speech-to-text conversion and emotion recognition. It also uses Google's gTTS (Google Text-to-Speech) for speech synthesis. This enables the provision of real-time security support tailored to the user's emotional state.

[0438] For example, if a user says, "I've been worried a lot lately," the system will analyze that emotion and respond in a calm voice, "Why don't you try taking a break?" Examples of prompt messages include the following:

[0439] Input: I've been feeling stressed lately.

[0440] Expected output: I recommend taking a short rest.

[0441] In this way, the system enables interactive conversations that reassure users while being sensitive to their emotions. As a security support system, it can help reduce anxieties in daily life for the elderly, enabling them to live with greater peace of mind.

[0442] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0443] Step 1:

[0444] The user provides voice input to the terminal. The terminal records the voice as digital data and sends it to the server. The input is raw voice data, and the output is the digital format of that voice data. The terminal captures the voice using a microphone and sends it to the server using its data communication function.

[0445] Step 2:

[0446] The server converts received audio data into text data using speech recognition technology. The input is audio data, and the output is the converted text data. The Python `speech_recognition` library is used for this process. This allows the audio data to be analyzed as linguistic information.

[0447] Step 3:

[0448] The server analyzes text data using natural language processing techniques to understand the user's intent. The input is text data, and the output is the analysis result representing the user's intent. This process utilizes the Hugging Face transformers library, with a generative AI model assisting the analysis.

[0449] Step 4:

[0450] Emotional information is extracted from the user's voice using emotion recognition technology. The input is voice data, and the output is emotional information. For emotion recognition, the transformers library is also used to analyze the type and intensity of the emotion.

[0451] Step 5:

[0452] The server generates conversations based on the user's intent and sentiment information. The input is intent and sentiment information, and the output is the text of the generated conversation. In this process, a generative AI model uses prompt sentences to determine the content of the response.

[0453] Step 6:

[0454] The generated conversation text is converted into speech data using a speech synthesis method and sent to the terminal. The input is conversation text, and the output is speech data. Google's gTTS is used for this process to convert it into natural-sounding speech.

[0455] Step 7:

[0456] The terminal plays audio data received from the server for the user. The input is audio data, and the output is audio playback through the terminal. The terminal physically plays the audio using a speaker and responds to the user.

[0457] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0458] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0459] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0460] [Third Embodiment]

[0461] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0462] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0463] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0464] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0465] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0466] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0467] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0468] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0469] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0470] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0471] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0472] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0473] This invention is a multi-functional system that supports the daily lives of the elderly by utilizing voice input and output. Specifically, it integrates voice recognition, natural language processing, conversation generation, speech synthesis, nutritional management, and reminder functions. The aim of this system is to alleviate feelings of loneliness and health management problems faced by the elderly and to provide them with a comfortable life.

[0474] When a user speaks to this system, the terminal sends voice data to the server. The server analyzes the received voice data using speech recognition and converts it into text data. The converted text data is then interpreted by natural language processing to understand the user's intent. Based on that intent, a conversation generation system creates an appropriate response, and a speech synthesis system converts it back into voice data and returns it to the user.

[0475] For example, if a user asks their device, "What should I have for breakfast today?", the server analyzes this and makes a specific suggestion, such as, "For breakfast, I recommend yogurt and fruit, which are nutritious and easy to eat." This response is conveyed to the user via voice.

[0476] Furthermore, when a user reports the food they have eaten, the server uses a nutrition database reference to perform a nutritional assessment based on the meal content. The results are then analyzed by an advice generation tool to provide the user with recommended nutritional supplements and health advice. Specifically, if a user inputs "I ate a rice ball today," the server will advise, "You may be lacking protein today. Try incorporating tofu or chicken into your lunch."

[0477] Furthermore, to support users' daily activities, a reminder function is provided using scheduling tools. For example, if a user has set a reminder to take their medication at 8 a.m. every day, the server will send a notification to the device based on the setting, and the device will remind the user with a voice message saying, "It's time to take your medication."

[0478] In this way, the present invention is a system that provides support to the elderly through voice interaction, thereby reducing feelings of loneliness and improving health management.

[0479] The following describes the processing flow.

[0480] Step 1:

[0481] The user speaks questions or requests into the device. This audio is captured in real time by the device's microphone.

[0482] Step 2:

[0483] The terminal converts the captured audio data into a digital format and sends it to the server. During this process, the audio data is properly encoded and securely transferred to the server.

[0484] Step 3:

[0485] The server converts the received audio data into text data using speech recognition technology. The speech recognition engine analyzes the audio waveform and generates the corresponding text.

[0486] Step 4:

[0487] The server analyzes the converted text data using natural language processing techniques and interprets the data to understand the user's intent and the content of their question.

[0488] Step 5:

[0489] Using conversation generation tools, natural-sounding response text is created based on the user's intent. A pre-trained AI model is utilized in this process.

[0490] Step 6:

[0491] The generated response text is converted into speech data using a speech synthesis system. The synthesized speech is then adjusted to be output in a natural and easy-to-listen-to form for the user.

[0492] Step 7:

[0493] The server sends audio data to the terminal, and the terminal plays this audio data to deliver a response to the user.

[0494] Step 8:

[0495] When a user reports their meal, they provide the meal information to the server via voice. The server analyzes the meal content and evaluates the nutritional balance by referring to a nutrition database.

[0496] Step 9:

[0497] Based on the nutritional assessment, the server generates nutritional advice to support the user's health through an advice generation mechanism and communicates it to the user using the same procedure as the speech synthesis described above.

[0498] Step 10:

[0499] When using the reminder function, the scheduling system prepares a notification based on the time specified by the user. When the specified time arrives, the server sends audio data to the device, and the device announces the user's reminder aloud.

[0500] (Example 1)

[0501] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0502] Loneliness and health management issues are serious problems in the lives of the elderly, and support to alleviate these is needed. However, existing technologies lack an integrated system to support all aspects of the elderly's lives. Therefore, the challenge is to provide a system that integrates the conversation, nutritional management, and reminder functions that the elderly need on a daily basis.

[0503] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0504] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language processing means for analyzing the text data and understanding the user's intent, and conversation generation means for generating a response based on the understood user's intent. This enables a multi-functional integrated system that supports the daily lives of elderly people through voice interaction.

[0505] "Speech recognition means" refers to a technology or device that analyzes speech input and converts it into corresponding text data.

[0506] "Natural language processing means" refers to technologies or devices that analyze text data and interpret the meaning and emotions intended by humans.

[0507] "Conversation generation means" refers to a technology or device for generating an appropriate response based on the analyzed user intent.

[0508] "Speech synthesis means" refers to a technology or device that converts a generated text-based response into speech data and outputs it to the user.

[0509] "Nutritional data reference means" refers to technology or devices that access databases or information sources used to evaluate the content of meals consumed by a user and to analyze nutritional balance.

[0510] "Advice generation means" refers to a technology or device that provides appropriate nutritional advice to the user based on the results of a nutritional assessment.

[0511] "Management means" refers to technology or devices that manage and provide notifications and reminders based on the user's schedule and settings.

[0512] A "voice generation device" is a technology or device that outputs managed schedules and reminders to the user as voice messages.

[0513] This invention is an integrated voice support system that provides multifaceted support for the lives of the elderly, and includes speech recognition, natural language processing, conversation generation, speech synthesis, nutritional management, and schedule management. It is composed of three main components: a server, a terminal, and a user, and a series of processes are carried out using specific hardware and software.

[0514] The server uses cloud-based speech recognition software to convert user voice input into text data. This utilizes common speech recognition technologies such as the Google Cloud Speech-to-Text API. The converted text data is then analyzed by a generative AI model (e.g., a model capable of broad natural language processing) to interpret the user's intent. OpenAI's natural language processing technology is one example of this.

[0515] Let's say a user asks, "What should I have for breakfast today?" The server then converts the voice data into text and uses a generative AI model to generate an appropriate response. A suggestion like, "A balanced breakfast of fruit and yogurt is recommended," might be generated. This response is then spoken using speech synthesis technology such as Amazon Polly and delivered to the user via their device.

[0516] The terminal plays the role of sending user input to the server and communicating responses from the server to the user. The terminal is equipped with a microphone and speaker, enabling audio input and output.

[0517] Nutritional management for users is done by reporting their meals to the server. For example, in response to a report such as "I ate a rice ball today," the server consults a database and generates advice that takes into account the necessary nutrients. Advice such as "You may be lacking protein. Try incorporating tofu or chicken into your lunch" might be given.

[0518] Furthermore, the server manages user-set reminders using a scheduling system and notifies the user via their device as a voice message. For example, based on a medication reminder set for 8:00 AM, a voice message saying "It's time for your medication" will be sent.

[0519] An example of a prompt using a generative AI model is: "In a health management support system for the elderly, generate a sample conversation that suggests appropriate meals when a user asks about today's meals." This allows the system to provide support tailored to the user's daily life.

[0520] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0521] Step 1:

[0522] The user speaks to the system through a voice input device. The voice input is received as an analog audio signal via a microphone and then converted into digital data. The input can range from user questions to instructions.

[0523] Step 2:

[0524] The terminal transmits audio data received from the user to a server in the cloud via the internet. The data is typically transferred using a specific compression format (e.g., a codec). The input is digital audio data, and the output is a confirmation that the transfer to the server is complete.

[0525] Step 3:

[0526] The server uses speech recognition software to convert digital speech data into text data. For example, it detects pitch and phonemes through speech waveform analysis and identifies words. The input is the received speech data, and the output is the analyzed text data.

[0527] Step 4:

[0528] The server uses a generative AI model to perform natural language processing on text data and analyze the user's intent. It performs grammatical and semantic analysis to understand what the user is asking for. The input is text data, and the output is user intent information.

[0529] Step 5:

[0530] The server uses a conversation generation engine to generate appropriate responses based on the user's intent. It creates natural-sounding response sentences using template-based or AI-driven conversation models. The input is user intent information, and the output is the response text.

[0531] Step 6:

[0532] The server uses speech synthesis technology to convert the response text into audio data. The speech synthesis algorithm reproduces it as natural-sounding speech and saves it in an audio file format. The input is the response text, and the output is the generated audio data.

[0533] Step 7:

[0534] The terminal transmits audio data sent from the server to the user via a playback device. The audio is output through a speaker, which the user can then hear. The input is audio data, and the output is auditory feedback of the sound.

[0535] Step 8:

[0536] When a user reports their meal, the server analyzes the nutritional information by referring to a nutrition database. It calculates the nutritional value based on the reported ingredients and evaluates the balance. The input is the meal information from the user, and the output is the nutritional evaluation result.

[0537] Step 9:

[0538] The server generates nutritional advice for the user based on the nutritional assessment results. It suggests foods and meal plans to supplement any deficient nutrients as needed. The input is the nutritional assessment results, and the output is the advice provided.

[0539] Step 10:

[0540] The server manages reminders according to the user's schedule settings and sends notifications at the designated time. It uses a timer function to trigger set alarms. The input is the user's schedule information, and the output is notification data.

[0541] Step 11:

[0542] The device delivers a voice reminder to the user based on notification data from the server. It utilizes voice notification functionality to prompt the user for the next action. The input is notification data, and the output is a voice reminder message.

[0543] (Application Example 1)

[0544] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0545] In facilities for the elderly, users often experience difficulty finding products or obtaining appropriate information during shopping and product searches. Furthermore, a lack of assistance with selecting nutritionally balanced products and managing shopping lists makes efficient and healthy shopping challenging. Difficulty managing schedules within a set timeframe is another significant issue.

[0546] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0547] In this invention, the server includes: speech recognition means for receiving voice input and converting the voice into text data; natural language processing means for analyzing the text data and understanding the user's intent; conversation generation means for generating a response based on the understood user's intent; speech synthesis means for converting the generated response into voice data and outputting it to the user; information retrieval means for searching for information requested by the user within the facility and returning that information; and location guidance generation means for guiding users to the location of products within the facility and generating related information. This enables elderly people and facility users to efficiently search for products, obtain information that takes nutritional balance into consideration, and have a comfortable shopping experience. Furthermore, it enables users to effectively manage their schedules, which is expected to improve their overall quality of life.

[0548] A "speech recognition means" is a device that receives speech input from a user and converts that speech into text data.

[0549] "Natural language processing means" refers to technologies that analyze text data obtained by speech recognition means to understand the user's intentions and requests.

[0550] A "conversation generation means" is a technology that has the function of generating an appropriate response based on the user's intent understood by a natural language processing means.

[0551] "Speech synthesis means" refers to a function that converts responses generated by conversation generation means into speech data and outputs it to the user.

[0552] "Information retrieval means" refers to technology that searches for information requested by users within a facility and provides the results.

[0553] A "location guidance generation means" is a technology that has the function of guiding users to the location of products within a facility and generating and providing related information to the user.

[0554] "Nutritional information database reference means" refers to a function that allows users to refer to data on the food they have consumed and evaluate its nutritional balance.

[0555] "Guidance generation means" refers to technology for generating nutritional guidance for users based on the evaluation results of a nutritional information database reference means.

[0556] A "time management system" is a technology that has the function of sending notifications to users according to a predetermined time.

[0557] "Voice notification means" refers to a function that outputs notifications set by the time management means in voice format.

[0558] The system designed to realize this application will provide support to users in efficiently searching for products within a facility and obtaining information that takes nutritional balance into consideration. Furthermore, by appropriately communicating necessary information and assistance via voice, it will enable a comfortable shopping experience.

[0559] The server uses speech recognition technology to receive voice input and convert it into text data. Specifically, voice data entered via a smartphone or dedicated terminal is sent to a cloud server and converted into text data using the Google Speech-to-Text API. The converted text data is then analyzed for the user's intent using natural language processing technology (e.g., NLTK or SpaCy). Based on the analyzed intent, an appropriate response is generated using conversation generation technology (e.g., GPT-3). The generated response is then converted into voice data using speech synthesis technology (e.g., Amazon Polly) and output to the user.

[0560] Furthermore, the information retrieval means searches for the information requested by the user within the store and refers to a database to return that information. In addition, the location guidance generation means uses a database that manages product location information within the facility to guide the user to where that product is located.

[0561] For example, if a user asks the terminal, "Where is the milk?", the server analyzes the voice, retrieves information from the product database such as "Milk is in the dairy section," and can then provide voice guidance. In this way, it becomes possible to provide information tailored to the user's needs.

[0562] An example of a prompt message to the generative AI model would be: "The user wants to search for a product. Please return section information for the product specified by the user. The product name is 'milk'."

[0563] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0564] Step 1:

[0565] Voice input begins when the user speaks into the device. The device acquires the user's voice using its microphone and sends it to the server as digital audio data. The input is analog audio, and the output is digital audio data.

[0566] Step 2:

[0567] The server passes the received digital audio data to the Google Speech-to-Text API, which converts the audio into text data. During this process, the audio data is converted into a string based on a specific algorithm. The input is digital audio data, and the output is text data.

[0568] Step 3:

[0569] The server analyzes text data using natural language processing tools (such as NLTK and SpaCy) to extract the user's intent. This is done through grammatical analysis and word semantic analysis. The input is text data, and the output is analyzed data that includes the user's intent.

[0570] Step 4:

[0571] Based on the analyzed user intent, the server generates appropriate responses as prompts using a conversation generation tool (such as GPT-3). At this stage, prompts are constructed and passed to the generation AI model. The input is the analyzed data, and the output is the prompts.

[0572] Step 5:

[0573] The generated prompt sentence is analyzed by a generative AI model to obtain an appropriate response. This response is the final conversational text. The input is the prompt sentence, and the output is the response text.

[0574] Step 6:

[0575] The server converts the response text into speech using a speech synthesis tool (such as Amazon Polly). Here, text data is transformed into natural-sounding speech by a synthesis algorithm. The input is the response text, and the output is speech data.

[0576] Step 7:

[0577] The device receives audio data and outputs it to the user via its speaker. This allows the user to receive answers via voice. The input is audio data, and the output is the presentation of information via voice.

[0578] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0579] This invention is a multi-functional system that combines speech recognition technology and emotion recognition technology to support the daily lives of the elderly. This system analyzes the user's voice input in real time and provides optimal responses and support based on the content and emotional state of the input. In addition to its main voice input / output functions, the system is equipped with an emotion engine, enabling it to address the user's psychological needs.

[0580] First, when a user speaks into the main device, the audio data is collected by the device's microphone. The audio data is sent to a server, where it is converted into text using speech recognition technology. Next, this text information is analyzed by natural language processing technology to understand the user's intent. Then, an emotion engine extracts emotions from the user's voice and compares this with data stored in an emotion database.

[0581] Based on the information obtained in this way, the conversation generation means generates an appropriate response, which is then converted back into natural-sounding speech by the speech synthesis means and output to the user through the terminal. This entire process allows for not only simple verbal exchange but also responses that are tailored to the user's mood.

[0582] For example, if a user says, "I'm feeling a bit down today," the server converts the speech to text and extracts emotional components such as "sadness" through an emotion engine. Based on these results, the system generates an empathetic response for the user, such as, "That's tough. Do you want to talk about it?" and delivers it to the user via voice.

[0583] Furthermore, by monitoring the user's emotional state over the long term, specific patterns and fluctuations can be recognized, and if a state of emotional outburst persists, support such as encouraging consultation with a specialist can be proposed. Thus, the present invention aims to help the mental stability of the elderly through emotional recognition.

[0584] The following describes the processing flow.

[0585] Step 1:

[0586] The user speaks into the device, making questions and statements via voice. The device acquires this voice as digital audio data.

[0587] Step 2:

[0588] The terminal sends the acquired audio data to the server. The server receives the audio data and converts it into text data using speech recognition technology. Noise is removed during this process to ensure accurate text conversion.

[0589] Step 3:

[0590] The server analyzes the converted text data using natural language processing techniques to extract the user's intent. During this process, it understands the intent of the question by interpreting the text content within its context.

[0591] Step 4:

[0592] Text data is passed to an emotion engine, which determines the user's emotions based on speech intonation and text expression. Emotional components are extracted and compared with an emotion database.

[0593] Step 5:

[0594] Based on the analyzed intent and emotion data, the server uses a conversation generation mechanism to create an appropriate response to the user. This response is generated in a way that resonates with the user's emotions.

[0595] Step 6:

[0596] The generated response text is converted into speech data by a speech synthesis system. Here, the tone of the voice is adjusted to produce a natural and gentle voice.

[0597] Step 7:

[0598] The server converts the audio data and sends it to the terminal. The terminal plays a voice response back to the user and observes whether the user's emotions remain stable after the conversation ends.

[0599] Step 8:

[0600] User interaction history and emotional data are anonymized and stored in a database for long-term monitoring and analysis of the user's mental state. This data is used for regular feedback and support services as needed.

[0601] (Example 2)

[0602] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0603] In the daily lives of the elderly and those requiring assistance, there is a need to provide interactive life support, including psychological support. However, simple voice recognition and digital information provision cannot adequately improve the quality of communication, and providing support that takes emotions into consideration is particularly difficult. Furthermore, there are challenges in providing support tailored to individual circumstances, such as managing diet and schedules.

[0604] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0605] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into digital text, language analysis means for analyzing the digital text and understanding the communicator's intent, and emotion analysis means for extracting emotional information from the digital text and recognizing the emotional state. This makes it possible to generate responses that take the user's emotions into consideration, thereby improving the psychological stability and quality of life of elderly people.

[0606] "Voice input" is a means for users to communicate information and instructions to a system through sound or words.

[0607] "Digital text" refers to data that is generated by analyzing voice input and representing it as textual information.

[0608] "Speech recognition means" refers to technology for receiving speech input and converting it into digital text.

[0609] "Linguistic analysis tools" are technologies that analyze digital text to understand the user's intent and the content of their questions.

[0610] "Emotional analysis methods" are technologies for extracting and recognizing a user's emotional state from digital text or audio.

[0611] A "response generation means" is a technology for creating an appropriate response based on the user's intentions and emotional state.

[0612] "Speech conversion means" refers to a technology that converts text into speech in order to output the generated response as sound.

[0613] A "nutritional information reference method" is a technology for evaluating the nutritional composition of food based on data from foods consumed by the user.

[0614] A "guidance generation method" is a technology that provides appropriate nutritional guidance to users based on the evaluation results of nutritional information.

[0615] A "schedule management system" is a technology that notifies users of pre-set reminders and appointments at the appropriate time.

[0616] "Voice notification means" refers to a technology that communicates schedules and reminders to users through voice output.

[0617] This invention is a system that provides interactive support in daily life for the elderly and individuals who require assistance. This system combines speech recognition technology and emotion recognition technology, aiming to improve the user's mental well-being and quality of life.

[0618] The user provides voice input through the main device's microphone. This voice input is converted into digital data on the device and sent to the server. The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the voice data into digital text. This digital text is then analyzed by a language analysis engine (e.g., Google Cloud Natural Language API) to understand the user's intent.

[0619] Next, emotion recognition is the process of extracting emotional information from received audio and text data. This is done using an emotion analysis engine (e.g., IBM Watson Tone Analyzer). This process identifies the user's emotional state and compares it against an emotion database.

[0620] The server uses a generative AI model (e.g., GPT-3) based on the analysis results to generate an appropriate response that matches the user's intent and emotions. This generated response is then converted into audio format using a speech synthesis engine (e.g., Amazon Polly). Finally, this audio data is sent to the device and delivered to the user through the speaker.

[0621] For example, if a user says, "I'm feeling a bit down today," the server converts this speech into text and recognizes an emotion such as "sadness." Based on this, it generates an empathetic response such as, "That's tough. Do you want to talk about it?" and conveys it to the user verbally.

[0622] An example of a prompt might be, "How would you respond if the user said, 'I'm feeling a bit down today'?" This prompt allows the AI ​​to provide a response that takes the user's emotions into consideration.

[0623] This allows users to receive personalized support tailored to their individual emotions, while also contributing to the psychological well-being and problem-solving of elderly individuals.

[0624] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0625] Step 1:

[0626] The user provides voice input through the main terminal's microphone. The terminal converts the voice into a digital signal and sends this signal to the server. The input is the user's spoken voice, and the output is digitized voice data. By collecting the voice signal, preparations are made for the next process.

[0627] Step 2:

[0628] The server uses a speech recognition engine to convert received digital audio data into text format. In this process, it identifies the content of the speech from the audio signal and represents it in text. The input is digital audio data, and the output is text data. Specifically, it analyzes phonemes and converts them into strings.

[0629] Step 3:

[0630] The server performs natural language processing to analyze text data. The language analysis engine performs syntactic and semantic analysis to extract the communicator's intent from the text data and understand the user's needs. The input is text data, and the output is data corresponding to an understanding of the user's intent. This enables action selection based on the content of the dialogue.

[0631] Step 4:

[0632] The server uses an emotion analysis engine to extract emotional information from text and audio data. By analyzing the type and intensity of emotions, it recognizes the user's emotional state. Input is text or audio data, and output is data indicating the characteristics of the emotion. This evaluation is based on matching against an emotion database.

[0633] Step 5:

[0634] The server generates responses using a generative AI model based on the user's intent and recognized sentiment information. In this process, it selects appropriate responses from input data and constructs personalized messages. The input is the user's intent and sentiment data, and the output is the generated response text. The generative AI combines the information to create natural conversation.

[0635] Step 6:

[0636] The server converts the generated response text into speech data using a speech synthesis engine. The speech synthesis process generates natural and understandable speech from the text data and pronounces it. The input is the response text, and the output is synthesized speech data. Computer synthesis enables playback in a natural voice.

[0637] Step 7:

[0638] The server sends synthesized speech data to the terminal. The terminal completes the interaction by delivering this speech data to the user through its speaker. The input is synthesized speech data, and the output is the sound that reaches the user's ears. As a result, the user can hear the system's response and continue the interaction.

[0639] (Application Example 2)

[0640] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0641] In the lives of the elderly, there is a need for support systems that alleviate daily anxieties and feelings of loneliness, and enable them to live safely and securely. However, existing technologies make it difficult to provide real-time support that is tailored to the user's emotional state. To solve this problem, a system is needed that integrates speech recognition and emotion recognition to provide appropriate security support.

[0642] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0643] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language processing means for analyzing the text data and understanding the user's intent, and emotion recognition means for extracting emotion information from the user's voice data. This makes it possible to provide appropriate responses and security support in real time according to the user's emotional state.

[0644] "Voice input" refers to voice data from the user, which is information that the system receives and processes.

[0645] "Text data" refers to the character information of voice input converted by a speech recognition system.

[0646] "Speech recognition means" refers to a mechanism or device for converting speech input into text data.

[0647] "Natural language processing methods" refer to technologies and processes for analyzing text data and understanding the user's intent and meaning.

[0648] "Conversation generation means" refers to technologies and functions for generating appropriate responses based on the user's intent.

[0649] "Speech synthesis means" refers to a device or technology that converts responses generated by a conversation generation means into speech data and outputs it to the user.

[0650] "Emotion recognition means" refers to technologies and devices for extracting emotional information from a user's voice data.

[0651] "Response generation means" refers to technologies and mechanisms for providing appropriate security support to users based on emotional information.

[0652] "Security support" refers to assistance and services provided to enhance safety and peace of mind, tailored to the user's emotional state.

[0653] The system for realizing this invention comprises speech recognition means, natural language processing means, emotion recognition means, conversation generation means, speech synthesis means, and response generation means.

[0654] The terminal receives voice input from the user and sends it to the server as digital data. This voice data is converted into text data on the server by speech recognition. Next, this text data is analyzed by natural language processing to determine the user's intent.

[0655] Subsequently, the emotion recognition system extracts emotion information from the user's voice, which is then compared with an emotion database. Using prompt sentences, a generative AI model generates an appropriate response, and the response to the user is constructed through the conversation generation system. The generated response is converted into natural-sounding speech by the speech synthesis system and output to the user through the terminal.

[0656] This system utilizes Python's speech_recognition library and Hugging Face's transformers to perform speech-to-text conversion and emotion recognition. It also uses Google's gTTS (Google Text-to-Speech) for speech synthesis. This enables the provision of real-time security support tailored to the user's emotional state.

[0657] For example, if a user says, "I've been worried a lot lately," the system will analyze that emotion and respond in a calm voice, "Why don't you try taking a break?" Examples of prompt messages include the following:

[0658] Input: I've been feeling stressed lately.

[0659] Expected output: I recommend taking a short rest.

[0660] In this way, the system enables interactive conversations that reassure users while being sensitive to their emotions. As a security support system, it can help reduce anxieties in daily life for the elderly, enabling them to live with greater peace of mind.

[0661] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0662] Step 1:

[0663] The user provides voice input to the terminal. The terminal records the voice as digital data and sends it to the server. The input is raw voice data, and the output is the digital format of that voice data. The terminal captures the voice using a microphone and sends it to the server using its data communication function.

[0664] Step 2:

[0665] The server converts received audio data into text data using speech recognition technology. The input is audio data, and the output is the converted text data. The Python `speech_recognition` library is used for this process. This allows the audio data to be analyzed as linguistic information.

[0666] Step 3:

[0667] The server analyzes text data using natural language processing techniques to understand the user's intent. The input is text data, and the output is the analysis result representing the user's intent. This process utilizes the Hugging Face transformers library, with a generative AI model assisting the analysis.

[0668] Step 4:

[0669] Emotional information is extracted from the user's voice using emotion recognition technology. The input is voice data, and the output is emotional information. For emotion recognition, the transformers library is also used to analyze the type and intensity of the emotion.

[0670] Step 5:

[0671] The server generates conversations based on the user's intent and sentiment information. The input is intent and sentiment information, and the output is the text of the generated conversation. In this process, a generative AI model uses prompt sentences to determine the content of the response.

[0672] Step 6:

[0673] The generated conversation text is converted into speech data using a speech synthesis method and sent to the terminal. The input is conversation text, and the output is speech data. Google's gTTS is used for this process to convert it into natural-sounding speech.

[0674] Step 7:

[0675] The terminal plays audio data received from the server for the user. The input is audio data, and the output is audio playback through the terminal. The terminal physically plays the audio using a speaker and responds to the user.

[0676] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0677] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0678] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0679] [Fourth Embodiment]

[0680] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0681] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0682] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0683] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0684] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0685] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0686] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0687] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0688] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0689] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0690] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0691] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0692] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0693] This invention is a multi-functional system that supports the daily lives of the elderly by utilizing voice input and output. Specifically, it integrates voice recognition, natural language processing, conversation generation, speech synthesis, nutritional management, and reminder functions. The aim of this system is to alleviate feelings of loneliness and health management problems faced by the elderly and to provide them with a comfortable life.

[0694] When a user speaks to this system, the terminal sends voice data to the server. The server analyzes the received voice data using speech recognition and converts it into text data. The converted text data is then interpreted by natural language processing to understand the user's intent. Based on that intent, a conversation generation system creates an appropriate response, and a speech synthesis system converts it back into voice data and returns it to the user.

[0695] For example, if a user asks their device, "What should I have for breakfast today?", the server analyzes this and makes a specific suggestion, such as, "For breakfast, I recommend yogurt and fruit, which are nutritious and easy to eat." This response is conveyed to the user via voice.

[0696] Furthermore, when a user reports the food they have eaten, the server uses a nutrition database reference to perform a nutritional assessment based on the meal content. The results are then analyzed by an advice generation tool to provide the user with recommended nutritional supplements and health advice. Specifically, if a user inputs "I ate a rice ball today," the server will advise, "You may be lacking protein today. Try incorporating tofu or chicken into your lunch."

[0697] Furthermore, to support users' daily activities, a reminder function is provided using scheduling tools. For example, if a user has set a reminder to take their medication at 8 a.m. every day, the server will send a notification to the device based on the setting, and the device will remind the user with a voice message saying, "It's time to take your medication."

[0698] In this way, the present invention is a system that provides support to the elderly through voice interaction, thereby reducing feelings of loneliness and improving health management.

[0699] The following describes the processing flow.

[0700] Step 1:

[0701] The user speaks questions or requests into the device. This audio is captured in real time by the device's microphone.

[0702] Step 2:

[0703] The terminal converts the captured audio data into a digital format and sends it to the server. During this process, the audio data is properly encoded and securely transferred to the server.

[0704] Step 3:

[0705] The server converts the received audio data into text data using speech recognition technology. The speech recognition engine analyzes the audio waveform and generates the corresponding text.

[0706] Step 4:

[0707] The server analyzes the converted text data using natural language processing techniques and interprets the data to understand the user's intent and the content of their question.

[0708] Step 5:

[0709] Using conversation generation tools, natural-sounding response text is created based on the user's intent. A pre-trained AI model is utilized in this process.

[0710] Step 6:

[0711] The generated response text is converted into speech data using a speech synthesis system. The synthesized speech is then adjusted to be output in a natural and easy-to-listen-to form for the user.

[0712] Step 7:

[0713] The server sends audio data to the terminal, and the terminal plays this audio data to deliver a response to the user.

[0714] Step 8:

[0715] When a user reports their meal, they provide the meal information to the server via voice. The server analyzes the meal content and evaluates the nutritional balance by referring to a nutrition database.

[0716] Step 9:

[0717] Based on the nutritional assessment, the server generates nutritional advice to support the user's health through an advice generation mechanism and communicates it to the user using the same procedure as the speech synthesis described above.

[0718] Step 10:

[0719] When using the reminder function, the scheduling system prepares a notification based on the time specified by the user. When the specified time arrives, the server sends audio data to the device, and the device announces the user's reminder aloud.

[0720] (Example 1)

[0721] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0722] Loneliness and health management issues are serious problems in the lives of the elderly, and support to alleviate these is needed. However, existing technologies lack an integrated system to support all aspects of the elderly's lives. Therefore, the challenge is to provide a system that integrates the conversation, nutritional management, and reminder functions that the elderly need on a daily basis.

[0723] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0724] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language processing means for analyzing the text data and understanding the user's intent, and conversation generation means for generating a response based on the understood user's intent. This enables a multi-functional integrated system that supports the daily lives of elderly people through voice interaction.

[0725] "Speech recognition means" refers to a technology or device that analyzes speech input and converts it into corresponding text data.

[0726] "Natural language processing means" refers to technologies or devices that analyze text data and interpret the meaning and emotions intended by humans.

[0727] "Conversation generation means" refers to a technology or device for generating an appropriate response based on the analyzed user intent.

[0728] "Speech synthesis means" refers to a technology or device that converts a generated text-based response into speech data and outputs it to the user.

[0729] "Nutritional data reference means" refers to technology or devices that access databases or information sources used to evaluate the content of meals consumed by a user and to analyze nutritional balance.

[0730] "Advice generation means" refers to a technology or device that provides appropriate nutritional advice to the user based on the results of a nutritional assessment.

[0731] "Management means" refers to technology or devices that manage and provide notifications and reminders based on the user's schedule and settings.

[0732] A "voice generation device" is a technology or device that outputs managed schedules and reminders to the user as voice messages.

[0733] This invention is an integrated voice support system that provides multifaceted support for the lives of the elderly, and includes speech recognition, natural language processing, conversation generation, speech synthesis, nutritional management, and schedule management. It is composed of three main components: a server, a terminal, and a user, and a series of processes are carried out using specific hardware and software.

[0734] The server uses cloud-based speech recognition software to convert user voice input into text data. This utilizes common speech recognition technologies such as the Google Cloud Speech-to-Text API. The converted text data is then analyzed by a generative AI model (e.g., a model capable of broad natural language processing) to interpret the user's intent. OpenAI's natural language processing technology is one example of this.

[0735] Let's say a user asks, "What should I have for breakfast today?" The server then converts the voice data into text and uses a generative AI model to generate an appropriate response. A suggestion like, "A balanced breakfast of fruit and yogurt is recommended," might be generated. This response is then spoken using speech synthesis technology such as Amazon Polly and delivered to the user via their device.

[0736] The terminal plays the role of sending user input to the server and communicating responses from the server to the user. The terminal is equipped with a microphone and speaker, enabling audio input and output.

[0737] Nutritional management for users is done by reporting their meals to the server. For example, in response to a report such as "I ate a rice ball today," the server consults a database and generates advice that takes into account the necessary nutrients. Advice such as "You may be lacking protein. Try incorporating tofu or chicken into your lunch" might be given.

[0738] Furthermore, the server manages user-set reminders using a scheduling system and notifies the user via their device as a voice message. For example, based on a medication reminder set for 8:00 AM, a voice message saying "It's time for your medication" will be sent.

[0739] An example of a prompt using a generative AI model is: "In a health management support system for the elderly, generate a sample conversation that suggests appropriate meals when a user asks about today's meals." This allows the system to provide support tailored to the user's daily life.

[0740] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0741] Step 1:

[0742] The user speaks to the system through a voice input device. The voice input is received as an analog audio signal via a microphone and then converted into digital data. The input can range from user questions to instructions.

[0743] Step 2:

[0744] The terminal transmits audio data received from the user to a server in the cloud via the internet. The data is typically transferred using a specific compression format (e.g., a codec). The input is digital audio data, and the output is a confirmation that the transfer to the server is complete.

[0745] Step 3:

[0746] The server uses speech recognition software to convert digital speech data into text data. For example, it detects pitch and phonemes through speech waveform analysis and identifies words. The input is the received speech data, and the output is the analyzed text data.

[0747] Step 4:

[0748] The server uses a generative AI model to perform natural language processing on text data and analyze the user's intent. It performs grammatical and semantic analysis to understand what the user is asking for. The input is text data, and the output is user intent information.

[0749] Step 5:

[0750] The server uses a conversation generation engine to generate appropriate responses based on the user's intent. It creates natural-sounding response sentences using template-based or AI-driven conversation models. The input is user intent information, and the output is the response text.

[0751] Step 6:

[0752] The server uses speech synthesis technology to convert the response text into audio data. The speech synthesis algorithm reproduces it as natural-sounding speech and saves it in an audio file format. The input is the response text, and the output is the generated audio data.

[0753] Step 7:

[0754] The terminal transmits audio data sent from the server to the user via a playback device. The audio is output through a speaker, which the user can then hear. The input is audio data, and the output is auditory feedback of the sound.

[0755] Step 8:

[0756] When a user reports their meal, the server analyzes the nutritional information by referring to a nutrition database. It calculates the nutritional value based on the reported ingredients and evaluates the balance. The input is the meal information from the user, and the output is the nutritional evaluation result.

[0757] Step 9:

[0758] The server generates nutritional advice for the user based on the nutritional assessment results. It suggests foods and meal plans to supplement any deficient nutrients as needed. The input is the nutritional assessment results, and the output is the advice provided.

[0759] Step 10:

[0760] The server manages reminders according to the user's schedule settings and sends notifications at the designated time. It uses a timer function to trigger set alarms. The input is the user's schedule information, and the output is notification data.

[0761] Step 11:

[0762] The device delivers a voice reminder to the user based on notification data from the server. It utilizes voice notification functionality to prompt the user for the next action. The input is notification data, and the output is a voice reminder message.

[0763] (Application Example 1)

[0764] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0765] In facilities for the elderly, users often experience difficulty finding products or obtaining appropriate information during shopping and product searches. Furthermore, a lack of assistance with selecting nutritionally balanced products and managing shopping lists makes efficient and healthy shopping challenging. Difficulty managing schedules within a set timeframe is another significant issue.

[0766] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0767] In this invention, the server includes: speech recognition means for receiving voice input and converting the voice into text data; natural language processing means for analyzing the text data and understanding the user's intent; conversation generation means for generating a response based on the understood user's intent; speech synthesis means for converting the generated response into voice data and outputting it to the user; information retrieval means for searching for information requested by the user within the facility and returning that information; and location guidance generation means for guiding users to the location of products within the facility and generating related information. This enables elderly people and facility users to efficiently search for products, obtain information that takes nutritional balance into consideration, and have a comfortable shopping experience. Furthermore, it enables users to effectively manage their schedules, which is expected to improve their overall quality of life.

[0768] A "speech recognition means" is a device that receives speech input from a user and converts that speech into text data.

[0769] "Natural language processing means" refers to technologies that analyze text data obtained by speech recognition means to understand the user's intentions and requests.

[0770] A "conversation generation means" is a technology that has the function of generating an appropriate response based on the user's intent understood by a natural language processing means.

[0771] "Speech synthesis means" refers to a function that converts responses generated by conversation generation means into speech data and outputs it to the user.

[0772] "Information retrieval means" refers to technology that searches for information requested by users within a facility and provides the results.

[0773] A "location guidance generation means" is a technology that has the function of guiding users to the location of products within a facility and generating and providing related information to the user.

[0774] "Nutritional information database reference means" refers to a function that allows users to refer to data on the food they have consumed and evaluate its nutritional balance.

[0775] "Guidance generation means" refers to technology for generating nutritional guidance for users based on the evaluation results of a nutritional information database reference means.

[0776] A "time management system" is a technology that has the function of sending notifications to users according to a predetermined time.

[0777] "Voice notification means" refers to a function that outputs notifications set by the time management means in voice format.

[0778] The system designed to realize this application will provide support to users in efficiently searching for products within a facility and obtaining information that takes nutritional balance into consideration. Furthermore, by appropriately communicating necessary information and assistance via voice, it will enable a comfortable shopping experience.

[0779] The server uses speech recognition technology to receive voice input and convert it into text data. Specifically, voice data entered via a smartphone or dedicated terminal is sent to a cloud server and converted into text data using the Google Speech-to-Text API. The converted text data is then analyzed for the user's intent using natural language processing technology (e.g., NLTK or SpaCy). Based on the analyzed intent, an appropriate response is generated using conversation generation technology (e.g., GPT-3). The generated response is then converted into voice data using speech synthesis technology (e.g., Amazon Polly) and output to the user.

[0780] Furthermore, the information retrieval means searches for the information requested by the user within the store and refers to a database to return that information. In addition, the location guidance generation means uses a database that manages product location information within the facility to guide the user to where that product is located.

[0781] For example, if a user asks the terminal, "Where is the milk?", the server analyzes the voice, retrieves information from the product database such as "Milk is in the dairy section," and can then provide voice guidance. In this way, it becomes possible to provide information tailored to the user's needs.

[0782] An example of a prompt message to the generative AI model would be: "The user wants to search for a product. Please return section information for the product specified by the user. The product name is 'milk'."

[0783] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0784] Step 1:

[0785] Voice input begins when the user speaks into the device. The device acquires the user's voice using its microphone and sends it to the server as digital audio data. The input is analog audio, and the output is digital audio data.

[0786] Step 2:

[0787] The server passes the received digital audio data to the Google Speech-to-Text API, which converts the audio into text data. During this process, the audio data is converted into a string based on a specific algorithm. The input is digital audio data, and the output is text data.

[0788] Step 3:

[0789] The server analyzes text data using natural language processing tools (such as NLTK and SpaCy) to extract the user's intent. This is done through grammatical analysis and word semantic analysis. The input is text data, and the output is analyzed data that includes the user's intent.

[0790] Step 4:

[0791] Based on the analyzed user intent, the server generates appropriate responses as prompts using a conversation generation tool (such as GPT-3). At this stage, prompts are constructed and passed to the generation AI model. The input is the analyzed data, and the output is the prompts.

[0792] Step 5:

[0793] The generated prompt sentence is analyzed by a generative AI model to obtain an appropriate response. This response is the final conversational text. The input is the prompt sentence, and the output is the response text.

[0794] Step 6:

[0795] The server converts the response text into speech using a speech synthesis tool (such as Amazon Polly). Here, text data is transformed into natural-sounding speech by a synthesis algorithm. The input is the response text, and the output is speech data.

[0796] Step 7:

[0797] The device receives audio data and outputs it to the user via its speaker. This allows the user to receive answers via voice. The input is audio data, and the output is the presentation of information via voice.

[0798] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0799] This invention is a multi-functional system that combines speech recognition technology and emotion recognition technology to support the daily lives of the elderly. This system analyzes the user's voice input in real time and provides optimal responses and support based on the content and emotional state of the input. In addition to its main voice input / output functions, the system is equipped with an emotion engine, enabling it to address the user's psychological needs.

[0800] First, when a user speaks into the main device, the audio data is collected by the device's microphone. The audio data is sent to a server, where it is converted into text using speech recognition technology. Next, this text information is analyzed by natural language processing technology to understand the user's intent. Then, an emotion engine extracts emotions from the user's voice and compares this with data stored in an emotion database.

[0801] Based on the information obtained in this way, the conversation generation means generates an appropriate response, which is then converted back into natural-sounding speech by the speech synthesis means and output to the user through the terminal. This entire process allows for not only simple verbal exchange but also responses that are tailored to the user's mood.

[0802] For example, if a user says, "I'm feeling a bit down today," the server converts the speech to text and extracts emotional components such as "sadness" through an emotion engine. Based on these results, the system generates an empathetic response for the user, such as, "That's tough. Do you want to talk about it?" and delivers it to the user via voice.

[0803] Furthermore, by monitoring the user's emotional state over the long term, specific patterns and fluctuations can be recognized, and if a state of emotional outburst persists, support such as encouraging consultation with a specialist can be proposed. Thus, the present invention aims to help the mental stability of the elderly through emotional recognition.

[0804] The following describes the processing flow.

[0805] Step 1:

[0806] The user speaks into the device, making questions and statements via voice. The device acquires this voice as digital audio data.

[0807] Step 2:

[0808] The terminal sends the acquired audio data to the server. The server receives the audio data and converts it into text data using speech recognition technology. Noise is removed during this process to ensure accurate text conversion.

[0809] Step 3:

[0810] The server analyzes the converted text data using natural language processing techniques to extract the user's intent. During this process, it understands the intent of the question by interpreting the text content within its context.

[0811] Step 4:

[0812] Text data is passed to an emotion engine, which determines the user's emotions based on speech intonation and text expression. Emotional components are extracted and compared with an emotion database.

[0813] Step 5:

[0814] Based on the analyzed intent and emotion data, the server uses a conversation generation mechanism to create an appropriate response to the user. This response is generated in a way that resonates with the user's emotions.

[0815] Step 6:

[0816] The generated response text is converted into speech data by a speech synthesis system. Here, the tone of the voice is adjusted to produce a natural and gentle voice.

[0817] Step 7:

[0818] The server converts the audio data and sends it to the terminal. The terminal plays a voice response back to the user and observes whether the user's emotions remain stable after the conversation ends.

[0819] Step 8:

[0820] User interaction history and emotional data are anonymized and stored in a database for long-term monitoring and analysis of the user's mental state. This data is used for regular feedback and support services as needed.

[0821] (Example 2)

[0822] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0823] In the daily lives of the elderly and those requiring assistance, there is a need to provide interactive life support, including psychological support. However, simple voice recognition and digital information provision cannot adequately improve the quality of communication, and providing support that takes emotions into consideration is particularly difficult. Furthermore, there are challenges in providing support tailored to individual circumstances, such as managing diet and schedules.

[0824] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0825] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into digital text, language analysis means for analyzing the digital text and understanding the communicator's intent, and emotion analysis means for extracting emotional information from the digital text and recognizing the emotional state. This makes it possible to generate responses that take the user's emotions into consideration, thereby improving the psychological stability and quality of life of elderly people.

[0826] "Voice input" is a means for users to communicate information and instructions to a system through sound or words.

[0827] "Digital text" refers to data that is generated by analyzing voice input and representing it as textual information.

[0828] "Speech recognition means" refers to technology for receiving speech input and converting it into digital text.

[0829] "Linguistic analysis tools" are technologies that analyze digital text to understand the user's intent and the content of their questions.

[0830] "Emotional analysis methods" are technologies for extracting and recognizing a user's emotional state from digital text or audio.

[0831] A "response generation means" is a technology for creating an appropriate response based on the user's intentions and emotional state.

[0832] "Speech conversion means" refers to a technology that converts text into speech in order to output the generated response as sound.

[0833] A "nutritional information reference method" is a technology for evaluating the nutritional composition of food based on data from foods consumed by the user.

[0834] A "guidance generation method" is a technology that provides appropriate nutritional guidance to users based on the evaluation results of nutritional information.

[0835] A "schedule management system" is a technology that notifies users of pre-set reminders and appointments at the appropriate time.

[0836] "Voice notification means" refers to a technology that communicates schedules and reminders to users through voice output.

[0837] This invention is a system that provides interactive support in daily life for the elderly and individuals who require assistance. This system combines speech recognition technology and emotion recognition technology, aiming to improve the user's mental well-being and quality of life.

[0838] The user provides voice input through the main device's microphone. This voice input is converted into digital data on the device and sent to the server. The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the voice data into digital text. This digital text is then analyzed by a language analysis engine (e.g., Google Cloud Natural Language API) to understand the user's intent.

[0839] Next, emotion recognition is the process of extracting emotional information from received audio and text data. This is done using an emotion analysis engine (e.g., IBM Watson Tone Analyzer). This process identifies the user's emotional state and compares it against an emotion database.

[0840] The server uses a generative AI model (e.g., GPT-3) based on the analysis results to generate an appropriate response that matches the user's intent and emotions. This generated response is then converted into audio format using a speech synthesis engine (e.g., Amazon Polly). Finally, this audio data is sent to the device and delivered to the user through the speaker.

[0841] For example, if a user says, "I'm feeling a bit down today," the server converts this speech into text and recognizes an emotion such as "sadness." Based on this, it generates an empathetic response such as, "That's tough. Do you want to talk about it?" and conveys it to the user verbally.

[0842] An example of a prompt might be, "How would you respond if the user said, 'I'm feeling a bit down today'?" This prompt allows the AI ​​to provide a response that takes the user's emotions into consideration.

[0843] This allows users to receive personalized support tailored to their individual emotions, while also contributing to the psychological well-being and problem-solving of elderly individuals.

[0844] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0845] Step 1:

[0846] The user provides voice input through the main terminal's microphone. The terminal converts the voice into a digital signal and sends this signal to the server. The input is the user's spoken voice, and the output is digitized voice data. By collecting the voice signal, preparations are made for the next process.

[0847] Step 2:

[0848] The server uses a speech recognition engine to convert received digital audio data into text format. In this process, it identifies the content of the speech from the audio signal and represents it in text. The input is digital audio data, and the output is text data. Specifically, it analyzes phonemes and converts them into strings.

[0849] Step 3:

[0850] The server performs natural language processing to analyze text data. The language analysis engine performs syntactic and semantic analysis to extract the communicator's intent from the text data and understand the user's needs. The input is text data, and the output is data corresponding to an understanding of the user's intent. This enables action selection based on the content of the dialogue.

[0851] Step 4:

[0852] The server uses an emotion analysis engine to extract emotional information from text and audio data. By analyzing the type and intensity of emotions, it recognizes the user's emotional state. Input is text or audio data, and output is data indicating the characteristics of the emotion. This evaluation is based on matching against an emotion database.

[0853] Step 5:

[0854] The server generates responses using a generative AI model based on the user's intent and recognized sentiment information. In this process, it selects appropriate responses from input data and constructs personalized messages. The input is the user's intent and sentiment data, and the output is the generated response text. The generative AI combines the information to create natural conversation.

[0855] Step 6:

[0856] The server converts the generated response text into speech data using a speech synthesis engine. The speech synthesis process generates natural and understandable speech from the text data and pronounces it. The input is the response text, and the output is synthesized speech data. Computer synthesis enables playback in a natural voice.

[0857] Step 7:

[0858] The server sends synthesized speech data to the terminal. The terminal completes the interaction by delivering this speech data to the user through its speaker. The input is synthesized speech data, and the output is the sound that reaches the user's ears. As a result, the user can hear the system's response and continue the interaction.

[0859] (Application Example 2)

[0860] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0861] In the lives of the elderly, there is a need for support systems that alleviate daily anxieties and feelings of loneliness, and enable them to live safely and securely. However, existing technologies make it difficult to provide real-time support that is tailored to the user's emotional state. To solve this problem, a system is needed that integrates speech recognition and emotion recognition to provide appropriate security support.

[0862] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0863] In this invention, the server includes speech recognition means for receiving voice input and converting the voice into text data, natural language processing means for analyzing the text data and understanding the user's intent, and emotion recognition means for extracting emotion information from the user's voice data. This makes it possible to provide appropriate responses and security support in real time according to the user's emotional state.

[0864] "Voice input" refers to voice data from the user, which is information that the system receives and processes.

[0865] "Text data" refers to the character information of voice input converted by a speech recognition system.

[0866] "Speech recognition means" refers to a mechanism or device for converting speech input into text data.

[0867] "Natural language processing methods" refer to technologies and processes for analyzing text data and understanding the user's intent and meaning.

[0868] "Conversation generation means" refers to technologies and functions for generating appropriate responses based on the user's intent.

[0869] "Speech synthesis means" refers to a device or technology that converts responses generated by a conversation generation means into speech data and outputs it to the user.

[0870] "Emotion recognition means" refers to technologies and devices for extracting emotional information from a user's voice data.

[0871] "Response generation means" refers to technologies and mechanisms for providing appropriate security support to users based on emotional information.

[0872] "Security support" refers to assistance and services provided to enhance safety and peace of mind, tailored to the user's emotional state.

[0873] The system for realizing this invention comprises speech recognition means, natural language processing means, emotion recognition means, conversation generation means, speech synthesis means, and response generation means.

[0874] The terminal receives voice input from the user and sends it to the server as digital data. This voice data is converted into text data on the server by speech recognition. Next, this text data is analyzed by natural language processing to determine the user's intent.

[0875] Subsequently, the emotion recognition system extracts emotion information from the user's voice, which is then compared with an emotion database. Using prompt sentences, a generative AI model generates an appropriate response, and the response to the user is constructed through the conversation generation system. The generated response is converted into natural-sounding speech by the speech synthesis system and output to the user through the terminal.

[0876] This system utilizes Python's speech_recognition library and Hugging Face's transformers to perform speech-to-text conversion and emotion recognition. It also uses Google's gTTS (Google Text-to-Speech) for speech synthesis. This enables the provision of real-time security support tailored to the user's emotional state.

[0877] For example, if a user says, "I've been worried a lot lately," the system will analyze that emotion and respond in a calm voice, "Why don't you try taking a break?" Examples of prompt messages include the following:

[0878] Input: I've been feeling stressed lately.

[0879] Expected output: I recommend taking a short rest.

[0880] In this way, the system enables interactive conversations that reassure users while being sensitive to their emotions. As a security support system, it can help reduce anxieties in daily life for the elderly, enabling them to live with greater peace of mind.

[0881] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0882] Step 1:

[0883] The user provides voice input to the terminal. The terminal records the voice as digital data and sends it to the server. The input is raw voice data, and the output is the digital format of that voice data. The terminal captures the voice using a microphone and sends it to the server using its data communication function.

[0884] Step 2:

[0885] The server converts received audio data into text data using speech recognition technology. The input is audio data, and the output is the converted text data. The Python `speech_recognition` library is used for this process. This allows the audio data to be analyzed as linguistic information.

[0886] Step 3:

[0887] The server analyzes text data using natural language processing techniques to understand the user's intent. The input is text data, and the output is the analysis result representing the user's intent. This process utilizes the Hugging Face transformers library, with a generative AI model assisting the analysis.

[0888] Step 4:

[0889] Emotional information is extracted from the user's voice using emotion recognition technology. The input is voice data, and the output is emotional information. For emotion recognition, the transformers library is also used to analyze the type and intensity of the emotion.

[0890] Step 5:

[0891] The server generates conversations based on the user's intent and sentiment information. The input is intent and sentiment information, and the output is the text of the generated conversation. In this process, a generative AI model uses prompt sentences to determine the content of the response.

[0892] Step 6:

[0893] The generated conversation text is converted into speech data using a speech synthesis method and sent to the terminal. The input is conversation text, and the output is speech data. Google's gTTS is used for this process to convert it into natural-sounding speech.

[0894] Step 7:

[0895] The terminal plays audio data received from the server for the user. The input is audio data, and the output is audio playback through the terminal. The terminal physically plays the audio using a speaker and responds to the user.

[0896] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0897] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0898] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0899] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0900] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0901] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0902] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0903] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0904] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0905] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0906] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0907] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0908] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0909] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0910] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0911] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0912] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0913] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0914] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0915] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0916] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0917] The following is further disclosed regarding the embodiments described above.

[0918] (Claim 1)

[0919] A speech recognition means that receives voice input and converts the voice into text data,

[0920] A natural language processing means that analyzes the aforementioned text data and understands the user's intent,

[0921] A conversation generation means that generates a response based on the understood user intent,

[0922] A speech synthesis means that converts the generated response into audio data and outputs it to the user,

[0923] A system that includes this.

[0924] (Claim 2)

[0925] A means for referencing a nutritional database to input the contents of meals consumed by the user and evaluate the nutritional balance,

[0926] An advice generation means that generates nutritional advice based on the evaluation results,

[0927] The system according to claim 1, including the following:

[0928] (Claim 3)

[0929] A scheduling management system that notifies users of pre-set reminders at a predetermined time,

[0930] A voice notification means that outputs the aforementioned reminder notification as an audio signal,

[0931] The system according to claim 1, including the following:

[0932] "Example 1"

[0933] (Claim 1)

[0934] A speech recognition means that receives voice input and converts the voice into text data,

[0935] A natural language processing means that analyzes the aforementioned text data and understands the user's intent,

[0936] A conversation generation means that generates a response based on the understood user intent,

[0937] A speech synthesis means that converts the generated response into audio data and outputs it to the user,

[0938] A database for storing user information and managing diet and health,

[0939] A question-answering system that provides dynamic suggestions and reminders based on user voice input,

[0940] A management system that manages schedules according to user settings and generates notifications at predetermined times,

[0941] A voice generation means for outputting the above notification as audio,

[0942] A system that includes this.

[0943] (Claim 2)

[0944] A means for users to input the contents of the food and drinks they have consumed and to refer to nutritional data to evaluate the nutritional balance,

[0945] An advice generation means for generating nutritional advice based on the aforementioned evaluation results,

[0946] A means of providing customized feedback according to the user's health status,

[0947] The system according to claim 1, including the following:

[0948] (Claim 3)

[0949] A management means that notifies users of a pre-set schedule at a predetermined time,

[0950] A voice generation device for outputting the aforementioned schedule notification as voice,

[0951] A means of providing additional alerts based on anticipated user behavior and needs,

[0952] The system according to claim 1, including the following:

[0953] "Application Example 1"

[0954] (Claim 1)

[0955] A speech recognition means that receives voice input and converts the voice into text data,

[0956] A natural language processing means that analyzes the aforementioned text data and understands the user's intent,

[0957] A conversation generation means that generates a response based on the understood user's intent,

[0958] A speech synthesis means that converts the generated response into audio data and outputs it to the user,

[0959] An information retrieval method that searches for information requested by users within the facility and returns that information,

[0960] A location guidance generation means that provides information on the location of products within a facility and generates related information,

[0961] A system that includes this.

[0962] (Claim 2)

[0963] A means for referencing a nutritional information database to input the contents of the food consumed by the user and evaluate the nutritional balance,

[0964] A guidance generation means for generating nutritional guidance based on the evaluation results,

[0965] The system according to claim 1, including the following:

[0966] (Claim 3)

[0967] A time management means that transmits pre-set notifications to users at a predetermined time,

[0968] A voice notification means that outputs the aforementioned notification as an audio,

[0969] The system according to claim 1, including the following:

[0970] "Example 2 of combining an emotion engine"

[0971] (Claim 1)

[0972] A speech recognition means that receives voice input and converts the voice into digital text,

[0973] A language analysis means for analyzing the aforementioned digital text and understanding the intent of the communicator,

[0974] An emotion analysis means for extracting emotional information from the aforementioned digital text and recognizing the emotional state,

[0975] A dialogue generation means that generates a response based on the understood intention and emotional state,

[0976] A voice conversion means that converts the generated response into voice data and outputs it to the communicator,

[0977] A system that includes this.

[0978] (Claim 2)

[0979] A means for referencing nutritional information to input the contents of food consumed by the communicator and evaluate its nutritional composition,

[0980] A guidance generation means for generating nutritional guidance based on the evaluation results,

[0981] The system according to claim 1, including the following:

[0982] (Claim 3)

[0983] A scheduled management means that notifies the communicator of pre-set memorized information at a predetermined time,

[0984] A notification voice means for outputting the aforementioned notification of stored information in voice,

[0985] The system according to claim 1, including the following:

[0986] "Application example 2 of combining emotional engines"

[0987] (Claim 1)

[0988] A speech recognition means that receives voice input and converts the voice into text data,

[0989] A natural language processing means that analyzes the aforementioned text data and understands the user's intent,

[0990] A conversation generation means that generates a response based on the understood user's intentions and emotional state,

[0991] A speech synthesis means that converts the generated response into audio data and outputs it to the user,

[0992] An emotion recognition method that extracts emotional information from the user's voice data,

[0993] A response generation method that provides security support based on user sentiment information,

[0994] A system that includes this.

[0995] (Claim 2)

[0996] A means of referencing nutritional information to input the contents of meals consumed by the user and to evaluate nutritional balance,

[0997] An advice generation means for generating nutritional advice based on the evaluation results,

[0998] The system according to claim 1, including the following:

[0999] (Claim 3)

[1000] A scheduling management system that notifies users of pre-set reminders at a specified time,

[1001] A voice notification means that outputs the aforementioned reminder notification as an audio signal,

[1002] The system according to claim 1, including the following: [Explanation of symbols]

[1003] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A speech recognition means that receives voice input and converts the voice into text data, A natural language processing means that analyzes the aforementioned text data and understands the user's intent, A conversation generation means that generates a response based on the understood user intent, A speech synthesis means that converts the generated response into audio data and outputs it to the user, A system that includes this.

2. A means for referencing a nutritional database to input the contents of meals consumed by the user and evaluate the nutritional balance, An advice generation means that generates nutritional advice based on the evaluation results, The system according to claim 1, including the following:

3. A scheduling management system that notifies users of pre-set reminders at a predetermined time, A voice notification means that outputs the aforementioned reminder notification as an audio signal, The system according to claim 1, including the following:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A