system
A system processes text, voice, and photo data using generative AI to summarize and convert it into specified formats, addressing communication challenges for elderly and dementia patients.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
Elderly individuals and those with advanced dementia face challenges in effectively conveying their intentions and thoughts to family members due to communication difficulties, necessitating a technology that can record and transmit this information efficiently.
A system that receives text, voice, and photo data, processes it using a generative AI to summarize and convert it into a specified format, and provides the output in a viewable, downloadable, or shareable form.
The system accurately records and communicates the wishes of elderly and dementia patients by summarizing and converting their inputs into appropriate formats, ensuring their intentions are conveyed effectively.
Smart Images

Figure 2026064776000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Due to aging or the progression of dementia, there is a problem that it is difficult for elderly people to appropriately convey their own intentions and thoughts to family members and relatives. To address this issue, there are also time constraints, and there is a need for a technology that can appropriately record the information and thoughts of that person and make it easily available and transmissible later.
Means for Solving the Problems
[0005] This invention provides a system that includes means for receiving text, voice, and photo data entered by a user, means for passing the received data to a generating AI for summarization and conversion into a specified format, and means for providing the generated output to the user. With this system, users can input any data they think of at any time, and that data is summarized and output in an appropriate format, effectively solving the problem of elderly people or those with advanced dementia who have difficulty communicating their wishes. For example, voice data can be converted into a text-based will, or a summarized message can be generated along with a photograph, making the recorded information easily usable and transmitted.
[0006] A "user" refers to a person who uses a system to input data and requests its output.
[0007] "Text data" refers to information in text format that is entered by the user.
[0008] "Audio data" refers to recorded audio information entered by the user.
[0009] "Photo data" refers to images and photo format information uploaded by users.
[0010] "Means of receiving data" refers to the mechanism for importing text, audio, and photo data entered by the user into the system.
[0011] "Generative AI" refers to artificial intelligence that analyzes received data and converts it into a specified format.
[0012] "Summarization" refers to the process of shortening and extracting the main points of the input data, and presenting it in a concise form.
[0013] "Specified format" refers to the format of the output generated by the system, and includes wills, last wills, funeral speeches, etc.
[0014] "The means for conversion" refers to a mechanism that executes a process of summarizing the received data by a generative AI and converting it into a specified format.
[0015] "Output" refers to output information such as text, images, videos, etc. generated as the result of summarization and conversion.
[0016] "The means for providing" refers to a mechanism that provides the generated output in a form that can be browsed, downloaded, or shared by the user.
Brief Description of the Drawings
[0017] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and specified format. The specific operation method and program processing of this system will be described below.
[0039] 1. Data Input
[0040] Terminal-side processing
[0041] The user accesses a provided input form using their device. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user might enter the contents of a will in text format, record a voice message about memories, and upload related photos.
[0042] After a user enters or uploads data, the device temporarily stores this data and sends it to the server. This makes the entered data available within the system.
[0043] 2. Summarizing and Transformation
[0044] Server-side processing
[0045] The server receives data sent from the terminal and stores it in storage. The stored data is then prepared to be passed on to the generating AI later. Specifically, this involves preprocessing of text data, transcribing audio data into text, and extracting metadata from photo data.
[0046] The server passes the stored data to the generating AI, instructing it to summarize the data and convert it into a specified format. The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video message for a funeral eulogy.
[0047] 3. Output generation and provision
[0048] Server-side processing
[0049] The server receives the summary and output in the specified format created by the generation AI and saves it for each user. Next, it converts this output into a format that the user can view, download, or share. For example, a will video might be saved as an MP4 file.
[0050] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[0051] Terminal-side processing
[0052] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly communicated.
[0053] To illustrate with a concrete example, consider a scenario where a user uploads a memorable photo and records an audio recording of the memory associated with that photo. The device sends the audio data and photo data to a server, which then passes them to a generative AI. The generative AI extracts metadata from the photo, transcribes the audio data into text, and summarizes its content. Finally, an output combining the summarized text and the photo is generated and provided to the user.
[0054] In this way, this system accurately records the user's thoughts and provides them in the appropriate format when needed, thereby solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[0055] The following describes the processing flow.
[0056] Step 1:
[0057] The device displays an input form to the user. The form includes a text box, a voice recording button, and a photo upload button.
[0058] Step 2:
[0059] Users input text data, record voice messages, and upload photos.
[0060] Step 3:
[0061] The device temporarily stores data entered by the user. This temporary storage uses the device's memory.
[0062] Step 4:
[0063] The device sends the temporarily stored data to the server. Once the transmission is complete, the device displays a confirmation message.
[0064] Step 5:
[0065] The server receives data sent from the terminal and saves it to storage. The storage location is managed individually for each user.
[0066] Step 6:
[0067] The server preprocesses the stored data. Specifically, this includes verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[0068] Step 7:
[0069] The server passes the pre-processed data to the generating AI, instructing it to summarize and convert it into a specified format.
[0070] Step 8:
[0071] The generation AI analyzes the received data and generates summaries and outputs in a specified format. For example, it can convert audio data into a text-based will or create a message based on photographs.
[0072] Step 9:
[0073] The server receives the generated output and saves it to storage. The saved output is managed on a per-user basis.
[0074] Step 10:
[0075] The server converts the saved output into a format that users can view, download, or share.
[0076] Step 11:
[0077] The server notifies the user when the output is ready. This notification is sent via email or an in-app notification.
[0078] Step 12:
[0079] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output.
[0080] Step 13:
[0081] Users can view the output through their device and download or share it as needed. For example, they can play the generated will video and share it with their family.
[0082] Through these steps, the system enables users to record memories and important messages in the appropriate format and provide them at the appropriate time.
[0083] (Example 1)
[0084] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0085] For elderly users or those with advanced dementia who have difficulty communicating their wishes effectively, efficiently recording data such as text, audio, and photographs, summarizing it, and converting it to a specified format for delivery presents a significant challenge. Solving this problem requires a system that handles everything from data input to output delivery, but current technology lacks a suitable solution.
[0086] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0087] In this invention, the server includes means for temporarily storing and transmitting user-inputted text, voice, and photo data to the server; means for pre-processing the data stored on the server; means for passing the pre-processed data to a generation AI model for summarization and conversion into a specified format; means for storing the generated output for each user; and means for notifying the user of the stored output and providing it in a format that can be viewed, downloaded, or shared. This makes it possible for elderly users or users with advanced dementia who have difficulty communicating their wishes to accurately and concisely record their wishes and provide them in an appropriate format when needed.
[0088] "Text data" refers to the text information that a user enters into an input form.
[0089] "Audio data" refers to voice messages recorded by the user using the voice recording button.
[0090] "Photo data" refers to image files that users upload using the photo upload button.
[0091] "Temporary storage" refers to the process of temporarily storing data before it is sent from the device to the server.
[0092] A "server" is a central system that receives, stores, and processes data sent from terminals.
[0093] "Preprocessing" refers to the process by which a server formats received data for analysis and transformation. Specifically, this includes text cleaning, audio-to-text conversion, and image metadata extraction.
[0094] A "generative AI model" is an artificial intelligence system that summarizes received data and converts it into a specified format.
[0095] "Output" refers to the output data created by the generative AI model, which is summarized and formatted in a specified format.
[0096] "Storage" refers to the process where the server identifies and stores the processed output for each user.
[0097] "Notification" refers to the server informing the user that the output is ready.
[0098] "Viewing" means making the generated output available for the user to review.
[0099] "Downloading" refers to the act of a user saving the generated output to their own device.
[0100] "Sharing" refers to the act of a user sharing the output they have generated with others.
[0101] "Metadata" refers to additional information included in photo data, such as the date and time of shooting, location, and camera settings.
[0102] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and converted format. This system receives user input data, processes it on a server, and provides output that is summarized and converted by a generating AI model.
[0103] Specifically, this is achieved using terminals, servers, and generative AI models as follows:
[0104] Hardware and software to be used
[0105] terminal
[0106] Hardware: Smartphone, tablet, or PC
[0107] Software: Dedicated application or form on a web browser
[0108] server
[0109] Hardware: High-performance cloud servers
[0110] Software: Databases, storage systems, analysis libraries (e.g., OpenCV, NLP libraries)
[0111] Generative AI Models
[0112] Software: Natural language processing algorithms, speech recognition engines (e.g., Google® Speech-to-Text API)
[0113] Specific operating methods for the system
[0114] 1. Data Input
[0115] Terminal-side processing
[0116] The user launches a dedicated application on their device and accesses an input form displayed on the main screen. This form includes a text box, a voice recording button, and a photo upload button. The user enters the contents of their will in text format, records a voice message about their memories, and uploads related photos. After entering, recording, and uploading the data, the device temporarily stores the data and sends it to the server.
[0117] 2. Summarizing and Transformation
[0118] Server-side processing
[0119] The server receives data sent from the terminal and stores it in storage. After storage, the server preprocesses the data. Specifically, it cleans text data, transcribes audio data, and extracts metadata from photo data. Then, it passes the preprocessed data to a generative AI model and instructs it to summarize and transform it into a specified format. The generative AI model uses natural language processing algorithms to summarize and transform the data.
[0120] 3. Output generation and provision
[0121] Server-side processing
[0122] The system receives the summary and output in the specified format generated by the generative AI model and saves it for each user. Next, it converts this output into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file. The user is then notified when this information is ready. Notification methods include email and in-app notifications.
[0123] Terminal-side processing
[0124] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user reviews the generated output and downloads or shares it as needed.
[0125] Specific example
[0126] Consider a scenario where a user uploads a memorable photo and records an audio recording of the associated memories. In this case, the device sends the audio data and photo data to a server. The server passes this data to a generative AI model, which extracts the metadata from the photo and transcribes the audio data into text to summarize its content. Finally, an output combining the summarized text and the photo is generated and provided to the user.
[0127] Example of a prompt
[0128] Text Summary: "Please summarize this text: 'A message to my family: Thank you for everything. I look forward to your continued support.'"
[0129] Speech-to-text: "Please generate text from this audio. The audio file says, 'This is an important message to my family. This photo is from a trip we took together.'"
[0130] In this way, the system helps users effectively communicate their intentions by accurately understanding them and converting them into an appropriate format.
[0131] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0132] Step 1:
[0133] Displaying the input form
[0134] The user launches the application on their device and accesses the input form displayed on the main screen. The input form includes a text box, a voice recording button, and a photo upload button. Specifically, the device renders the input form and prompts the user for input. Input can be in the form of text, audio, or an image file.
[0135] Step 2:
[0136] Data entry, recording, and uploading.
[0137] The user enters text into a form, taps the voice recording button to record a voice message, and uses the photo upload button to select and upload a photo. Specifically, the device temporarily saves the entered text, the recorded voice data, and the uploaded photo data. As output, the device retains the saved data.
[0138] Step 3:
[0139] Sending data
[0140] After the user completes input, recording, and uploading, they tap the send button. The device then sends this data to the server. Specifically, the device sends the temporarily stored data to the server via an HTTP request. The input consists of temporarily stored text, audio, and image data, while the output is the transmission of data to the server.
[0141] Step 4:
[0142] Receiving and storing data
[0143] The server receives data sent from the terminal and stores it in storage. Specifically, the server stores the received data in a database and saves audio and images to file storage. The input is the data sent from the terminal, and the output is the data stored in storage.
[0144] Step 5:
[0145] Data preprocessing
[0146] The server performs preprocessing on the stored data. Specifically, the server uses an NLP (Natural Language Processing) library to clean text, a speech recognition engine (e.g., Google Speech-to-Text API) to convert speech to text, and an image analysis tool (e.g., OpenCV) to extract metadata from photos. The input consists of stored text, audio, and image data, and the output is preprocessed data.
[0147] Step 6:
[0148] Sending data to the generative AI model
[0149] The server passes pre-processed data to the generative AI model and instructs it to summarize and transform it into a specified format. Specifically, the server sends API requests to the generative AI model, inputting prompts for summarization and transformation. The input is pre-processed data, and the output is the output generated by the generative AI model.
[0150] Step 7:
[0151] Receiving and saving the generated output
[0152] The server receives the output from the generative AI model and saves it for each user. Specifically, the server saves the generated output, such as text and video, to a database. The input is the output from the generative AI model, and the output is the saved output.
[0153] Step 8:
[0154] Output format conversion
[0155] The server converts the generated output into a format that users can view, download, or share. Specifically, the server uses a file conversion tool to generate PDF or MP4 files. The input is the saved output, and the output is a file in a format usable by the user.
[0156] Step 9:
[0157] Notification to the user
[0158] The server notifies the user when the output is ready. Specifically, the server calls a notification API and sends an email or push notification to the user. The input is the output after format conversion, and the output is the notification to the user.
[0159] Step 10:
[0160] Receiving and displaying notifications
[0161] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output. Specifically, the terminal checks the notification and renders a dedicated confirmation screen. The input is the notification from the server, and the output is the confirmation screen displayed to the user.
[0162] Step 11:
[0163] Check the output and take action
[0164] The user reviews the generated output and downloads or shares it as needed. Specifically, the device displays an output viewing screen and provides download links and share buttons. The input is the output displayed on the confirmation screen, and the output is the download or sharing action performed by the user.
[0165] (Application Example 1)
[0166] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0167] Currently, there is no system that allows elderly users or those with advanced dementia, who may have difficulty communicating their wishes effectively, to easily and accurately order food delivery. In particular, there is a need for systems that understand the user's ordering intent using voice and photo data and provide appropriate suggestions based on that understanding. However, conventional systems struggle to meet these requirements. This creates a problem where elderly users and those with dementia find it difficult to easily use food delivery services.
[0168] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0169] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI to summarize and convert it into a specified format; means for providing the generated output to the user; means for converting voice data into text and extracting metadata from image data to generate food suggestions; and means for presenting the generated food suggestions to the user. This makes it possible for elderly users or users with advanced dementia who have difficulty communicating their wishes to use food delivery services easily and accurately.
[0170] "Users" refer to elderly people or individuals with dementia who have difficulty communicating their wishes appropriately and who use this system.
[0171] "Text data" refers to information in text format that is entered into the device used by the user.
[0172] "Audio data" refers to information in audio format provided by the user through voice input.
[0173] "Photo data" refers to image data uploaded by users via their devices.
[0174] "Means of receiving" refers to devices or programs that have the function of receiving text, audio, and photographic data provided by the user.
[0175] "Generative AI" refers to artificial intelligence models used to summarize received data and convert it into a specified format.
[0176] "Means for summarizing and converting to a specified format" refers to a processing mechanism that passes the received data to a generating AI, summarizes the data concisely, and converts it to the required format.
[0177] "Means of providing output" refers to devices or programs that have the function of presenting output data generated by the generation AI to the user in an appropriate format.
[0178] "Converting audio data to text" refers to the process of analyzing audio data and converting it into corresponding text data.
[0179] "Metadata" refers to information associated with photographic data, such as the date and time of shooting, location, and the type of object depicted.
[0180] "Food suggestions" refers to dishes and foods recommended by the generating AI based on the user's input data.
[0181] "Means for generating suggestions" refers to a technical system for creating suggestions tailored to the user based on the received data.
[0182] "Means of presenting proposals" refers to devices or programs that have the function of visually or audibly displaying the generated proposals to the user.
[0183] The system for realizing this invention is designed to allow users to easily utilize food delivery even if they are elderly or have difficulty communicating their wishes due to the progression of dementia. The system consists of several main components, including a terminal, a server, and a generative AI model.
[0184] Overview of program processing
[0185] 1. Processing on the terminal side
[0186] The terminal receives text, voice, and photo data entered by the user. Through an application installed on the terminal, the user can voice-input descriptions of desired dishes or favorite dishes, and upload photos of previously ordered dishes. The input data is temporarily stored on the terminal and then sent to the server.
[0187] 2. Server-side processing
[0188] The server receives data sent from the terminal and performs data preprocessing. Specifically, it converts audio data into text format (speech recognition) and extracts metadata from photo data (image analysis). After this, the data is passed to a generative AI model to generate a data summary and food suggestions. The generative AI model uses natural language processing algorithms and image recognition algorithms.
[0189] 3. Output generation and provision
[0190] The system receives suggestions generated by a generative AI model and generates output tailored to each user's requirements. A list of suggested foods is presented to the user through the application. The user can then select from this list and complete their order.
[0191] Specific example
[0192] If a user voice-inputs "I want to eat spaghetti today" and uploads a photo of spaghetti they previously ordered, the system will process it as follows:
[0193] 1. The device records the user's voice saying "I want to eat spaghetti today," retrieves a photo of spaghetti ordered previously, and sends it to the server.
[0194] 2. The server converts the audio data into text format and extracts the metadata from the photos.
[0195] 3. The generative AI model analyzes text and photo data to suggest food suitable for the user.
[0196] 4. The user selects spaghetti from the suggested food list and completes the order.
[0197] Hardware and software to be used
[0198] 1. Microphone: Used to input the user's voice.
[0199] 2. Camera: Used to take photos that users upload.
[0200] 3. Terminal (smartphone / tablet): A device used by the user to input data and communicate with the server.
[0201] 4. Server: Responsible for storing data and running the generated AI model.
[0202] 5. Generative AI Models: Algorithms that combine speech recognition, natural language processing, and image recognition.
[0203] Example of a prompt
[0204] In the user's own voice: "What I want to eat right now is, for example, curry rice or ramen, or any of my favorite dishes."
[0205] Example photo: "Please upload a photo of the curry rice you ordered previously."
[0206] This allows the system to provide an easy-to-use food delivery ordering system for elderly and dementia users, making access to meals easier.
[0207] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0208] Step 1:
[0209] Users use their devices to voice-input the characteristics of the dishes they want to eat or their favorite dishes, and upload photos of dishes they have ordered in the past.
[0210] Input: Audio data, photo data
[0211] Operation: The user inputs audio using the application's voice recording function and uploads photos using the camera function.
[0212] Output: The input audio and photo data are temporarily stored on the device.
[0213] Step 2:
[0214] The device sends the voice and photo data entered by the user to the server.
[0215] Input: Temporarily stored audio data, photo data
[0216] Operation: The terminal sends data to the server via the network.
[0217] Output: Audio and photo data transferred to the server
[0218] Step 3:
[0219] The server converts the audio data to text and extracts metadata from the photo data.
[0220] Input: Audio data, photo data
[0221] Data processing: Convert audio data into text data using speech recognition technology. Extract metadata from photo data using image analysis technology.
[0222] Operation: The server uses a specified speech recognition algorithm (e.g., Google Speech Recognition). Image recognition APIs (e.g., AWS® Rekognition) are used to extract metadata from photos.
[0223] Output: Text data, photo metadata
[0224] Step 4:
[0225] The server passes text data and photo metadata to a generating AI model to generate food suggestions.
[0226] Input: Text data, photo metadata
[0227] Data processing: A generative AI model analyzes text data and photo metadata to generate food suggestions.
[0228] Operation: The server uses a generative AI model (e.g., GPT-3® or BERT) to generate optimal food suggestions for the user.
[0229] Output: List of suggested foods generated
[0230] Step 5:
[0231] The server sends the generated food suggestions to the user's terminal.
[0232] Input: List of suggested foods generated
[0233] Operation: The server sends the suggestion list to the user's terminal via the network.
[0234] Output: List of food suggestions sent to the user's terminal
[0235] Step 6:
[0236] The terminal displays a list of suggested foods to the user. The user selects from the list and completes the order.
[0237] Input: Food suggestion list
[0238] Operation: The terminal displays a list of suggestions through the application's user interface, and the user selects and places an order.
[0239] Output: Order information for food selected by the user.
[0240] Step 7:
[0241] The server sends the user's order information to the food delivery service and confirms the order.
[0242] Input: Order information for food selected by the user
[0243] Operation: The server sends order information to the food delivery service's API.
[0244] Output: Confirmed order information for food delivery service
[0245] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0246] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and specified format. By combining this system with an emotion engine, it can recognize the user's emotions and adjust the content and tone of the generated output. The specific operation method and program processing of this system are described below.
[0247] 1. Data Input
[0248] Terminal-side processing
[0249] The user accesses a provided input form using their device. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user might enter the contents of a will in text format, record a voice message about memories, and upload related photos.
[0250] After a user enters or uploads data, the device temporarily stores this data and sends it to the server. This makes the entered data available within the system.
[0251] 2. Summarizing and Transformation
[0252] Server-side processing
[0253] The server receives data sent from the terminal and stores it in storage. The stored data is then prepared to be passed on to the generating AI later. Specifically, this involves verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[0254] The server passes the stored data to the generating AI, instructing it to summarize the data and convert it into a specified format. The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video message for a funeral eulogy.
[0255] 3. Emotion Recognition and Regulation
[0256] Emotional Engine Processing
[0257] Before the data is passed to the generative AI, the server passes the audio and photo data to the emotion engine for emotion recognition. For example, the emotion engine analyzes the user's emotional state from the audio data and provides the results to the generative AI. Similarly, for photo data, the emotion engine recognizes the user's emotions when analyzing the metadata.
[0258] The emotional information recognized by the emotion engine is used as a guideline for the generating AI to adjust the output content and tone. For example, if the content of a will is very moving, the tone of the text and the atmosphere of the video will be adjusted to emphasize that.
[0259] 4. Output generation and delivery
[0260] Server-side processing
[0261] The server receives the summary and output in the specified format created by the generation AI and saves it to storage. The saved output is managed on a per-user basis. Next, this output is converted into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file.
[0262] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[0263] Terminal-side processing
[0264] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly communicated.
[0265] As a concrete example, consider a scenario where a user uploads a memorable photo and records an emotionally charged audio recording of the memories associated with that photo. The device sends the audio and photo data to a server, which then passes this data to an emotion engine and a generative AI. The generative AI adjusts the output based on the photo's metadata and emotional information, ultimately generating an emotionally resonant output that combines summarized text and the photo, which is then delivered to the user.
[0266] In this way, this system accurately records the user's thoughts and provides them at the appropriate time in a form that reflects their emotions, effectively solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[0267] The following describes the processing flow.
[0268] Step 1:
[0269] The device displays an input form to the user. The form includes a text box, a voice recording button, and a photo upload button.
[0270] Step 2:
[0271] Users input text data, record voice messages, and upload photos. For example, they might enter the contents of their will in text format, record an emotionally charged voice message about a memory, and upload related photos.
[0272] Step 3:
[0273] The device temporarily stores data entered by the user. This temporary storage uses the device's memory.
[0274] Step 4:
[0275] The device sends the temporarily stored data to the server. If the transmission is successful, the device displays a confirmation message.
[0276] Step 5:
[0277] The server receives data sent from the terminal and saves it to storage. The storage location is managed individually for each user.
[0278] Step 6:
[0279] The server preprocesses the stored data. Specifically, this includes verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[0280] Step 7:
[0281] The server passes the pre-processed data to the emotion engine and instructs it to analyze the user's emotions.
[0282] Step 8:
[0283] The emotion engine recognizes the user's emotions from voice data and photo data and provides the results to the generative AI. For example, the emotion engine analyzes the user's emotional state from the voice data and identifies emotions such as "compassion", "sadness", "joy", etc.
[0284] Step 9:
[0285] The server passes the emotion information obtained from the emotion engine and the preprocessed data to the generative AI and instructs it to summarize and convert it into the specified format.
[0286] Step 10:
[0287] The generative AI analyzes the data and generates a summary and an output in the specified format that reflects the emotion information. For example, if the content of the will is very touching, the tone of the text and the atmosphere of the video are adjusted to emphasize it.
[0288] Step 11:
[0289] The server receives the generated output and saves it in storage. The saved output is managed for each user.
[0290] Step 12:
[0291] The server converts the output into a format that allows the user to view, download, or share it. For example, a will video that reflects emotions is saved as an MP4 file.
[0292] Step 13:
[0293] The server notifies the user that the output is ready. As notification means, emails or in-app notifications are used.
[0294] Step 14:
[0295] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output.
[0296] Step 15:
[0297] Users can view the output through their device and download or share it as needed. For example, they can play the generated will video and share it with their family.
[0298] Through the steps described above, this system can analyze user-input data using an emotion engine and generative AI, record it in a way that reflects emotions, and provide it at the appropriate time.
[0299] (Example 2)
[0300] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0301] For elderly users or those with advanced dementia who have difficulty communicating their wishes, there is a need to efficiently record data such as memories and wills in a way that reflects their emotions, and provide it as appropriate output. However, conventional systems cannot recognize and reflect the user's emotions, and the content and tone of the generated output may not match the user's wishes. To solve this problem, a system is needed that can recognize the user's emotions and adjust the content and tone of the generated output.
[0302] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0303] In this invention, the server includes means for receiving character, voice, and photo data input by the user, means for passing the received data to a generative AI for summarization and conversion into a specified format, means for providing the generated output to the user, means for recognizing emotions from voice data and photo data by an emotion engine, means for adjusting the content and tone of the output based on the recognized emotion information, and means for storing the final output in storage and notifying the user. Thereby, an output reflecting the user's emotions is generated, enabling the appropriate transmission of the user's intentions and memories.
[0304] "Character data" refers to text information input by the user through an input form or the like.
[0305] "Voice data" refers to voice messages input by the user using the recording function.
[0306] "Photo data" refers to still image files uploaded by the user.
[0307] "Means for receiving" is a function for acquiring character data, voice data, and photo data input by the user from the terminal to the server.
[0308] "Generative AI" refers to an artificial intelligence model used to summarize and convert the data received from the user into a specified format.
[0309] "Means for summarizing" is a function for summarizing the received data and performing processing to concisely organize the information.
[0310] "Means for converting into a specified format" is a function for converting the summarized data into a format (such as text, video, etc.) that is easy for the user to view.
[0311] "Means for providing" is a function for notifying the user of the output created by the generative AI and enabling viewing, downloading, and sharing.
[0312] The "emotion engine" is a processing system that analyzes a user's emotions from audio and photo data and provides the results to the generating AI.
[0313] "Means of recognition" refers to a function that uses an emotion engine to identify the user's emotions from audio and photo data.
[0314] "Means of adjustment" refers to a function that allows the generating AI to change the content and tone of its output based on emotional information.
[0315] "Storage" refers to a location where generated output is saved and made accessible to users as needed.
[0316] "Notification means" refers to a notification function that informs the user when the server has finished preparing the generated output.
[0317] Modes for carrying out the invention
[0318] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides them in a summarized and specified format. By combining this system with an emotion engine, it can recognize the user's emotions and adjust the content and tone of the generated output.
[0319] 1. Data Input
[0320] Terminal-side processing
[0321] The user accesses an input form provided by the system using their device. This input form includes a text box, a voice recording button, and a photo upload button. Through this form, the user can, for example, enter the contents of a will in text format, record a voice message about memories, and upload related photos. This allows the user to express their wishes in a variety of ways.
[0322] After data entry, the terminal temporarily stores this data and sends it to the server. This transmission process makes the entered data available within the system.
[0323] 2. Summarizing and Transformation
[0324] Server-side processing
[0325] The server receives data sent from the terminal and stores it in storage. The received data is processed as follows:
[0326] The text data undergoes grammatical and consistency checks.
[0327] The audio data is converted to text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[0328] Metadata (such as the date and time of shooting, and location) is extracted from the photo data.
[0329] The data, after the above processing is complete, is passed to a generating AI (e.g., a GPT model) for summarization and conversion to a specified format. The generating AI uses natural language processing algorithms to perform tasks such as summarizing the contents of a will into a concise text format or generating a video message for a funeral eulogy.
[0330] 3. Emotion Recognition and Regulation
[0331] Emotional Engine Processing
[0332] Before the data is passed to the generative AI, the server passes the audio and photo data to an emotion engine (e.g., Microsoft® Azure® Cognitive Services) to recognize emotions. For example, the emotion engine analyzes the user's emotional state (joy, sadness, anger, etc.) from the audio data and provides the result to the generative AI. Similarly, for photo data, the emotion engine recognizes the user's emotions by analyzing the metadata.
[0333] The emotional information recognized by the emotion engine is used as a guideline for the generating AI to adjust the output content and tone. For example, if the content of a will is emotional, the tone of the text and the atmosphere of the video will be adjusted to emphasize that emotion.
[0334] 4. Output generation and provision
[0335] Server-side processing
[0336] The server receives the summary and output in the specified format created by the generation AI and saves it to storage. The saved output is managed on a per-user basis. Next, the server converts the output into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file. The server then notifies the user when the output is ready. Notification methods include email notifications and in-app notifications.
[0337] Terminal-side processing
[0338] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly conveyed.
[0339] Specific example
[0340] As a concrete example, consider a scenario where a user uploads photos of travel memories and records an emotionally moving story related to those photos in audio. The device sends the photos and audio data to a server, which then passes them to an emotion engine and a generative AI. The emotion engine extracts emotional information from the audio, and the generative AI uses this to create an emotionally moving output. Finally, the user is provided with an output that combines summarized text and photos.
[0341] Example of a prompt
[0342] An example of a prompt to input into a generative AI model is, "Based on the audio message and photos provided by the user about past memories, please generate a summary text that evokes emotion." Using such prompts, the generative AI can produce output that accurately reflects the user's intentions.
[0343] As described above, this system accurately records the user's thoughts and provides output that reflects their emotions, thereby effectively solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[0344] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0345] Step 1:
[0346] User enters data
[0347] The user uses a device to access the system's input form. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user enters the contents of a will in text format, records an audio message about memories, and uploads related photos (input: text, audio, and photo data). The device temporarily stores this data (output: data with temporarily stored content).
[0348] Step 2:
[0349] The device temporarily stores and transmits data.
[0350] After the user completes data entry, the device sends the temporarily stored data to the server (Input: Temporarily stored text, audio, and photo data). This process ensures that the data is sent to the server securely and efficiently (Output: Data sent to the server).
[0351] Step 3:
[0352] The server saves the data to the receiving storage.
[0353] The server receives data sent from the terminal and stores it in storage (Input: Data sent from the terminal). The stored data is then ready for processing and transformation (Output: Data stored in storage).
[0354] Step 4:
[0355] The server performs proofreading and preprocessing.
[0356] The server classifies the stored data and performs the following actions on each:
[0357] Text data validation (checking grammar and consistency)
[0358] Converting audio data to text (using speech recognition software to convert speech to text)
[0359] Metadata extraction from photo data (extracting information such as date, time, and location of shooting)
[0360] (Input: Saved text, audio, and photo data). The data is then ready to proceed to the next processing step (Output: Preprocessed data).
[0361] Step 5:
[0362] The server passes the data to the generating AI and instructs it on summarization and transformation.
[0363] The server passes the pre-processed data to the generating AI, requesting it to summarize and convert it to a specified format (Input: Pre-processed data). The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video for a funeral eulogy (Output: Summarized and converted output).
[0364] Step 6:
[0365] The server passes the data to the emotion engine to perform emotion recognition.
[0366] Before the data is passed to the generating AI, the server passes the audio and photo data to the emotion engine to recognize emotions (input: audio data, photo data). The emotion engine analyzes the user's emotional state from the audio data and also analyzes the metadata of the photo data to recognize emotions (output: recognized emotion information).
[0367] Step 7:
[0368] The AI generates the output based on emotional information, adjusting the content and tone.
[0369] The server provides the recognized emotional information to the generating AI, which then adjusts the content and tone of the output based on that information (input: recognized emotional information). For example, it sets the tone for the text or video of an emotional will (output: output reflecting the emotion).
[0370] Step 8:
[0371] The server saves the final output and notifies the user.
[0372] The server receives the final output created by the generation AI and saves it to storage (Input: Final Output). Then, it notifies the user that the output is ready (Output: Saved Output, Notification to User).
[0373] Step 9:
[0374] The device prompts for output confirmation and supports sharing.
[0375] The device receives a notification from the server and displays a screen prompting the user to review the output (Input: Notification). The user reviews the generated output through this screen and downloads or shares it as needed (Output: Reviewed output, downloaded or shared).
[0376] In this way, the system can effectively manage user data and provide output that reflects emotions.
[0377] (Application Example 2)
[0378] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0379] With an increasing number of elderly users and those with advanced dementia who have difficulty communicating their wishes effectively, there is a problem in that these users have difficulty communicating their needs regarding the products and services they are looking for in physical stores. This also hinders smooth communication with store staff, resulting in a decline in the quality of the user experience. To solve this, a system is needed that can accurately convey the user's wishes in a way that reflects their emotions.
[0380] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0381] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI for summarization and conversion into a specified format; means for recognizing the user's emotions from the received data; means for adjusting the content and tone of the output based on the emotion recognition results; and means for providing the generated output to the user. This enables users to communicate their intentions accurately and emotionally in a physical store setting.
[0382] A "user" is a person who uses the system to input data, and in particular, in physical stores, this includes elderly people who have difficulty communicating their intentions and users whose dementia is progressing.
[0383] "Text data" refers to information in text format entered by users, including requests and questions about products and services.
[0384] "Audio data" refers to audio information recorded by users, used to convey opinions and requests regarding products and services.
[0385] "Photo data" refers to image information uploaded by users to the system, including visual data related to products and services.
[0386] "Generative AI" refers to artificial intelligence (AI) technology used to summarize and convert received data into a specified format, employing natural language processing and data generation models.
[0387] An "emotion engine" refers to a technology that recognizes a user's emotions from audio and photo data, analyzing emotional information and adjusting the output accordingly.
[0388] "Output" refers to summaries or results in a specified format created by the generating AI based on data entered by the user, and is provided in formats such as text, video, and audio.
[0389] "Emotion recognition" is the process of using an emotion engine to analyze a user's emotional state from audio or photos and providing that information to a generating AI.
[0390] "Tone" refers to the emotion and atmosphere in the content and format of the output, and is adjusted according to the user's emotional state.
[0391] System Configuration
[0392] This system consists of a terminal that receives text data, voice data, and photo data entered by the user, a server that processes and analyzes the data, and a means of providing the generated output. The main components are as follows:
[0393] 1. Device: Users input text data, audio data, and photo data through a device such as a smartphone or tablet. This device is connected to the internet and has network communication capabilities to send data to the server.
[0394] 2. Server: The server stores and analyzes the received data. Using a cloud server as a hosting service is recommended here. Programming languages such as Python and related libraries are used for data analysis.
[0395] 3. Generative AI: Uses natural language processing and data generation models (e.g., Hugging Face's transformers library) to summarize user input and convert it to a specified format.
[0396] 4. Emotion Engine: Uses technologies (e.g., DeepFace library, Google Speech Recognition API) to recognize the user's emotional state from photo and audio data.
[0397] Processing flow
[0398] The server primarily processes data in the following steps:
[0399] 1. Receiving and storing data:
[0400] This function temporarily stores text data, audio data, and photo data entered by the device.
[0401] The data will be stored in cloud storage with security in mind.
[0402] 2. Speech recognition:
[0403] The server uses the Google Speech Recognition API to convert the audio data into text data.
[0404] This text data will be used for subsequent processing in the generative AI and emotion engine.
[0405] 3. Emotion recognition:
[0406] Using the DeepFace library, we recognize the user's emotions from photographic data and extract emotional information.
[0407] Emotional information is passed to the generating AI and used to adjust the tone of the output.
[0408] 4. Text summarization and conversion:
[0409] Use models from the transformers library to perform summarization and conversion to a specified format.
[0410] It incorporates emotional information into tone adjustments to generate appropriate text and video messages.
[0411] 5. Output generation and delivery:
[0412] The generated output is provided to the user. Specifically, notifications are sent via in-app notifications or email.
[0413] Users can view, download, or share the output generated through their device.
[0414] Specific example
[0415] Users record a voice message on their device to describe the details of the products they want in a physical store and upload photos of the store. The voice data is converted to text using the Google Speech Recognition API, and then the photo data undergoes emotion recognition using the DeepFace library. Based on this information, a generative AI model summarizes and provides detailed product information in a friendly tone.
[0416] Example of a prompt:
[0417] Convert the audio file "customer_request.wav" recorded by the customer in the store into text and summarize the text. Also, recognize the customer's emotions from the photo "customer_photo.jpg" and generate a response in a tone appropriate to those emotions.
[0418] This will enable elderly users and those with dementia to communicate their wishes accurately and emotionally to store staff.
[0419] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0420] Step 1:
[0421] Data entry and reception
[0422] The device receives text data entered by the user, audio data recorded, and photo data uploaded by the user. Users input this data using a smartphone or tablet via a dedicated application. The device temporarily stores the received data in local storage and sends it to a cloud server via the internet. Input is performed by the user using input forms, recording buttons, and photo upload functions, and the output in these cases is the data sent to the cloud server.
[0423] Step 2:
[0424] Data storage and preprocessing
[0425] The server receives text, audio, and photo data sent from the terminal and saves it to cloud storage. Simultaneously with saving, it begins data preprocessing. Specifically, audio data is converted to text format using the Google Speech Recognition API. This process converts the audio data into text data, which is then passed on to the next processing step. For photo data, identifiable metadata is extracted.
[0426] Step 3:
[0427] emotion recognition
[0428] The server performs emotion recognition using pre-processed data. It uses the DeepFace library to recognize the user's emotions from photographic data. Given photographic data as input, the emotion recognition engine extracts the user's emotional state (e.g., joy, sadness, surprise) as depicted in the photograph. Emotional information is obtained as output. This prepares the user's emotional information for delivery to the generating AI model.
[0429] Step 4:
[0430] Text summarization and tone adjustment
[0431] The server uses a generative AI model based on emotional information to summarize and tone-adjust text data. In particular, it uses the Hugging Face transformers library to summarize text. The input is pre-processed text and emotional information, and the output text is generated with the summary result and tone adjusted according to the emotion. Here, the tone is adjusted to be friendly when the emotion is joyful, and gentle when the emotion is sad.
[0432] Step 5:
[0433] Output generation and delivery
[0434] The generated text or other specified output formats (e.g., video messages) are delivered to the user from the server. The user is notified when the output is ready via in-app notifications or email. The user can then view, download, or share the generated output through their device. For example, the AI generates detailed information about the product the user wants to purchase, summarized in a friendly tone, making it easy for the user to communicate their intentions.
[0435] Through all these processing steps, the system enables elderly users and those with advanced dementia who have difficulty communicating to smoothly convey their intentions in physical stores.
[0436] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0437] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0438] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0439] [Second Embodiment]
[0440] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0441] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0442] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0443] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0444] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0445] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0446] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0447] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0448] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0449] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0450] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0451] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0452] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and specified format. The specific operation method and program processing of this system will be described below.
[0453] 1. Data Input
[0454] Terminal-side processing
[0455] The user accesses a provided input form using their device. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user might enter the contents of a will in text format, record a voice message about memories, and upload related photos.
[0456] After a user enters or uploads data, the device temporarily stores this data and sends it to the server. This makes the entered data available within the system.
[0457] 2. Summarizing and Transformation
[0458] Server-side processing
[0459] The server receives data sent from the terminal and stores it in storage. The stored data is then prepared to be passed on to the generating AI later. Specifically, this involves preprocessing of text data, transcribing audio data into text, and extracting metadata from photo data.
[0460] The server passes the stored data to the generating AI, instructing it to summarize the data and convert it into a specified format. The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video message for a funeral eulogy.
[0461] 3. Output generation and provision
[0462] Server-side processing
[0463] The server receives the summary and output in the specified format created by the generation AI and saves it for each user. Next, it converts this output into a format that the user can view, download, or share. For example, a will video might be saved as an MP4 file.
[0464] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[0465] Terminal-side processing
[0466] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly communicated.
[0467] To illustrate with a concrete example, consider a scenario where a user uploads a memorable photo and records an audio recording of the memory associated with that photo. The device sends the audio data and photo data to a server, which then passes them to a generative AI. The generative AI extracts metadata from the photo, transcribes the audio data into text, and summarizes its content. Finally, an output combining the summarized text and the photo is generated and provided to the user.
[0468] In this way, this system accurately records the user's thoughts and provides them in the appropriate format when needed, thereby solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[0469] The following describes the processing flow.
[0470] Step 1:
[0471] The device displays an input form to the user. The form includes a text box, a voice recording button, and a photo upload button.
[0472] Step 2:
[0473] Users input text data, record voice messages, and upload photos.
[0474] Step 3:
[0475] The device temporarily stores data entered by the user. This temporary storage uses the device's memory.
[0476] Step 4:
[0477] The device sends the temporarily stored data to the server. Once the transmission is complete, the device displays a confirmation message.
[0478] Step 5:
[0479] The server receives data sent from the terminal and saves it to storage. The storage location is managed individually for each user.
[0480] Step 6:
[0481] The server preprocesses the stored data. Specifically, this includes verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[0482] Step 7:
[0483] The server passes the pre-processed data to the generating AI, instructing it to summarize and convert it into a specified format.
[0484] Step 8:
[0485] The generation AI analyzes the received data and generates summaries and outputs in a specified format. For example, it can convert audio data into a text-based will or create a message based on photographs.
[0486] Step 9:
[0487] The server receives the generated output and saves it to storage. The saved output is managed on a per-user basis.
[0488] Step 10:
[0489] The server converts the saved output into a format that users can view, download, or share.
[0490] Step 11:
[0491] The server notifies the user when the output is ready. This notification is sent via email or an in-app notification.
[0492] Step 12:
[0493] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output.
[0494] Step 13:
[0495] Users can view the output through their device and download or share it as needed. For example, they can play the generated will video and share it with their family.
[0496] Through these steps, the system enables users to record memories and important messages in the appropriate format and provide them at the appropriate time.
[0497] (Example 1)
[0498] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0499] For elderly users or those with advanced dementia who have difficulty communicating their wishes effectively, efficiently recording data such as text, audio, and photographs, summarizing it, and converting it to a specified format for delivery presents a significant challenge. Solving this problem requires a system that handles everything from data input to output delivery, but current technology lacks a suitable solution.
[0500] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0501] In this invention, the server includes means for temporarily storing and transmitting user-inputted text, voice, and photo data to the server; means for pre-processing the data stored on the server; means for passing the pre-processed data to a generation AI model for summarization and conversion into a specified format; means for storing the generated output for each user; and means for notifying the user of the stored output and providing it in a format that can be viewed, downloaded, or shared. This makes it possible for elderly users or users with advanced dementia who have difficulty communicating their wishes to accurately and concisely record their wishes and provide them in an appropriate format when needed.
[0502] "Text data" refers to the text information that a user enters into an input form.
[0503] "Audio data" refers to voice messages recorded by the user using the voice recording button.
[0504] "Photo data" refers to image files that users upload using the photo upload button.
[0505] "Temporary storage" refers to the process of temporarily storing data before it is sent from the device to the server.
[0506] A "server" is a central system that receives, stores, and processes data sent from terminals.
[0507] "Preprocessing" refers to the process by which a server formats received data for analysis and transformation. Specifically, this includes text cleaning, audio-to-text conversion, and image metadata extraction.
[0508] A "generative AI model" is an artificial intelligence system that summarizes received data and converts it into a specified format.
[0509] "Output" refers to the output data created by the generative AI model, which is summarized and formatted in a specified format.
[0510] "Storage" refers to the process where the server identifies and stores the processed output for each user.
[0511] "Notification" refers to the server informing the user that the output is ready.
[0512] "Viewing" means making the generated output available for the user to review.
[0513] "Downloading" refers to the act of a user saving the generated output to their own device.
[0514] "Sharing" refers to the act of a user sharing the output they have generated with others.
[0515] "Metadata" refers to additional information included in photo data, such as the date and time of shooting, location, and camera settings.
[0516] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and converted format. This system receives user input data, processes it on a server, and provides output that is summarized and converted by a generating AI model.
[0517] Specifically, this is achieved using terminals, servers, and generative AI models as follows:
[0518] Hardware and software to be used
[0519] terminal
[0520] Hardware: Smartphone, tablet, or PC
[0521] Software: Dedicated application or form on a web browser
[0522] server
[0523] Hardware: High-performance cloud servers
[0524] Software: Databases, storage systems, analysis libraries (e.g., OpenCV, NLP libraries)
[0525] Generative AI Models
[0526] Software: Natural language processing algorithms, speech recognition engines (e.g., Google Speech-to-Text API)
[0527] Specific operating methods for the system
[0528] 1. Data Input
[0529] Terminal-side processing
[0530] The user launches a dedicated application on their device and accesses an input form displayed on the main screen. This form includes a text box, a voice recording button, and a photo upload button. The user enters the contents of their will in text format, records a voice message about their memories, and uploads related photos. After entering, recording, and uploading the data, the device temporarily stores the data and sends it to the server.
[0531] 2. Summarizing and Transformation
[0532] Server-side processing
[0533] The server receives data sent from the terminal and stores it in storage. After storage, the server preprocesses the data. Specifically, it cleans text data, transcribes audio data, and extracts metadata from photo data. Then, it passes the preprocessed data to a generative AI model and instructs it to summarize and transform it into a specified format. The generative AI model uses natural language processing algorithms to summarize and transform the data.
[0534] 3. Output generation and provision
[0535] Server-side processing
[0536] The system receives the summary and output in the specified format generated by the generative AI model and saves it for each user. Next, it converts this output into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file. The user is then notified when this information is ready. Notification methods include email and in-app notifications.
[0537] Terminal-side processing
[0538] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user reviews the generated output and downloads or shares it as needed.
[0539] Specific example
[0540] Consider a scenario where a user uploads a memorable photo and records an audio recording of the associated memories. In this case, the device sends the audio data and photo data to a server. The server passes this data to a generative AI model, which extracts the metadata from the photo and transcribes the audio data into text to summarize its content. Finally, an output combining the summarized text and the photo is generated and provided to the user.
[0541] Example of a prompt
[0542] Text Summary: "Please summarize this text: 'A message to my family: Thank you for everything. I look forward to your continued support.'"
[0543] Speech-to-text: "Please generate text from this audio. The audio file says, 'This is an important message to my family. This photo is from a trip we took together.'"
[0544] In this way, the system helps users effectively communicate their intentions by accurately understanding them and converting them into an appropriate format.
[0545] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0546] Step 1:
[0547] Displaying the input form
[0548] The user launches the application on their device and accesses the input form displayed on the main screen. The input form includes a text box, a voice recording button, and a photo upload button. Specifically, the device renders the input form and prompts the user for input. Input can be in the form of text, audio, or an image file.
[0549] Step 2:
[0550] Data entry, recording, and uploading.
[0551] The user enters text into a form, taps the voice recording button to record a voice message, and uses the photo upload button to select and upload a photo. Specifically, the device temporarily saves the entered text, the recorded voice data, and the uploaded photo data. As output, the device retains the saved data.
[0552] Step 3:
[0553] Sending data
[0554] After the user completes input, recording, and uploading, they tap the send button. The device then sends this data to the server. Specifically, the device sends the temporarily stored data to the server via an HTTP request. The input consists of temporarily stored text, audio, and image data, while the output is the transmission of data to the server.
[0555] Step 4:
[0556] Receiving and storing data
[0557] The server receives data sent from the terminal and stores it in storage. Specifically, the server stores the received data in a database and saves audio and images to file storage. The input is the data sent from the terminal, and the output is the data stored in storage.
[0558] Step 5:
[0559] Data preprocessing
[0560] The server performs preprocessing on the stored data. Specifically, the server uses an NLP (Natural Language Processing) library to clean text, a speech recognition engine (e.g., Google Speech-to-Text API) to convert speech to text, and an image analysis tool (e.g., OpenCV) to extract metadata from photos. The input consists of stored text, audio, and image data, and the output is preprocessed data.
[0561] Step 6:
[0562] Sending data to the generative AI model
[0563] The server passes pre-processed data to the generative AI model and instructs it to summarize and transform it into a specified format. Specifically, the server sends API requests to the generative AI model, inputting prompts for summarization and transformation. The input is pre-processed data, and the output is the output generated by the generative AI model.
[0564] Step 7:
[0565] Receiving and saving the generated output
[0566] The server receives the output from the generative AI model and saves it for each user. Specifically, the server saves the generated output, such as text and video, to a database. The input is the output from the generative AI model, and the output is the saved output.
[0567] Step 8:
[0568] Output format conversion
[0569] The server converts the generated output into a format that users can view, download, or share. Specifically, the server uses a file conversion tool to generate PDF or MP4 files. The input is the saved output, and the output is a file in a format usable by the user.
[0570] Step 9:
[0571] Notification to the user
[0572] The server notifies the user when the output is ready. Specifically, the server calls a notification API and sends an email or push notification to the user. The input is the output after format conversion, and the output is the notification to the user.
[0573] Step 10:
[0574] Receiving and displaying notifications
[0575] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output. Specifically, the terminal checks the notification and renders a dedicated confirmation screen. The input is the notification from the server, and the output is the confirmation screen displayed to the user.
[0576] Step 11:
[0577] Check the output and take action
[0578] The user reviews the generated output and downloads or shares it as needed. Specifically, the device displays an output viewing screen and provides download links and share buttons. The input is the output displayed on the confirmation screen, and the output is the download or sharing action performed by the user.
[0579] (Application Example 1)
[0580] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0581] Currently, there is no system that allows elderly users or those with advanced dementia, who may have difficulty communicating their wishes effectively, to easily and accurately order food delivery. In particular, there is a need for systems that understand the user's ordering intent using voice and photo data and provide appropriate suggestions based on that understanding. However, conventional systems struggle to meet these requirements. This creates a problem where elderly users and those with dementia find it difficult to easily use food delivery services.
[0582] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0583] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI to summarize and convert it into a specified format; means for providing the generated output to the user; means for converting voice data into text and extracting metadata from image data to generate food suggestions; and means for presenting the generated food suggestions to the user. This makes it possible for elderly users or users with advanced dementia who have difficulty communicating their wishes to use food delivery services easily and accurately.
[0584] "Users" refer to elderly people or individuals with dementia who have difficulty communicating their wishes appropriately and who use this system.
[0585] "Text data" refers to information in text format that is entered into the device used by the user.
[0586] "Audio data" refers to information in audio format provided by the user through voice input.
[0587] "Photo data" refers to image data uploaded by users via their devices.
[0588] "Means of receiving" refers to devices or programs that have the function of receiving text, audio, and photographic data provided by the user.
[0589] "Generative AI" refers to artificial intelligence models used to summarize received data and convert it into a specified format.
[0590] "Means for summarizing and converting to a specified format" refers to a processing mechanism that passes the received data to a generating AI, summarizes the data concisely, and converts it to the required format.
[0591] "Means of providing output" refers to devices or programs that have the function of presenting output data generated by the generation AI to the user in an appropriate format.
[0592] "Converting audio data to text" refers to the process of analyzing audio data and converting it into corresponding text data.
[0593] "Metadata" refers to information associated with photographic data, such as the date and time of shooting, location, and the type of object depicted.
[0594] "Food suggestions" refers to dishes and foods recommended by the generating AI based on the user's input data.
[0595] "Means for generating suggestions" refers to a technical system for creating suggestions tailored to the user based on the received data.
[0596] "Means of presenting proposals" refers to devices or programs that have the function of visually or audibly displaying the generated proposals to the user.
[0597] The system for realizing this invention is designed to allow users to easily utilize food delivery even if they are elderly or have difficulty communicating their wishes due to the progression of dementia. The system consists of several main components, including a terminal, a server, and a generative AI model.
[0598] Overview of program processing
[0599] 1. Processing on the terminal side
[0600] The terminal receives text, voice, and photo data entered by the user. Through an application installed on the terminal, the user can voice-input descriptions of desired dishes or favorite dishes, and upload photos of previously ordered dishes. The input data is temporarily stored on the terminal and then sent to the server.
[0601] 2. Server-side processing
[0602] The server receives data sent from the terminal and performs data preprocessing. Specifically, it converts audio data into text format (speech recognition) and extracts metadata from photo data (image analysis). After this, the data is passed to a generative AI model to generate a data summary and food suggestions. The generative AI model uses natural language processing algorithms and image recognition algorithms.
[0603] 3. Output generation and provision
[0604] The system receives suggestions generated by a generative AI model and generates output tailored to each user's requirements. A list of suggested foods is presented to the user through the application. The user can then select from this list and complete their order.
[0605] Specific example
[0606] If a user voice-inputs "I want to eat spaghetti today" and uploads a photo of spaghetti they previously ordered, the system will process it as follows:
[0607] 1. The device records the user's voice saying "I want to eat spaghetti today," retrieves a photo of spaghetti ordered previously, and sends it to the server.
[0608] 2. The server converts the audio data into text format and extracts the metadata from the photos.
[0609] 3. The generative AI model analyzes text and photo data to suggest food suitable for the user.
[0610] 4. The user selects spaghetti from the suggested food list and completes the order.
[0611] Hardware and software to be used
[0612] 1. Microphone: Used to input the user's voice.
[0613] 2. Camera: Used to take photos that users upload.
[0614] 3. Terminal (smartphone / tablet): A device used by the user to input data and communicate with the server.
[0615] 4. Server: Responsible for storing data and running the generated AI model.
[0616] 5. Generative AI Models: Algorithms that combine speech recognition, natural language processing, and image recognition.
[0617] Example of a prompt
[0618] In the user's own voice: "What I want to eat right now is, for example, curry rice or ramen, or any of my favorite dishes."
[0619] Example photo: "Please upload a photo of the curry rice you ordered previously."
[0620] This allows the system to provide an easy-to-use food delivery ordering system for elderly and dementia users, making access to meals easier.
[0621] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0622] Step 1:
[0623] Users use their devices to voice-input the characteristics of the dishes they want to eat or their favorite dishes, and upload photos of dishes they have ordered in the past.
[0624] Input: Audio data, photo data
[0625] Operation: The user inputs audio using the application's voice recording function and uploads photos using the camera function.
[0626] Output: The input audio and photo data are temporarily stored on the device.
[0627] Step 2:
[0628] The device sends the voice and photo data entered by the user to the server.
[0629] Input: Temporarily stored audio data, photo data
[0630] Operation: The terminal sends data to the server via the network.
[0631] Output: Audio and photo data transferred to the server
[0632] Step 3:
[0633] The server converts the audio data to text and extracts metadata from the photo data.
[0634] Input: Audio data, photo data
[0635] Data processing: Convert audio data into text data using speech recognition technology. Extract metadata from photo data using image analysis technology.
[0636] Operation: The server uses a specified speech recognition algorithm (e.g., Google Speech Recognition). Image recognition APIs (e.g., AWS Rekognition) are used to extract metadata from photos.
[0637] Output: Text data, photo metadata
[0638] Step 4:
[0639] The server passes text data and photo metadata to a generating AI model to generate food suggestions.
[0640] Input: Text data, photo metadata
[0641] Data processing: A generative AI model analyzes text data and photo metadata to generate food suggestions.
[0642] Operation: The server uses a generative AI model (e.g., GPT-3 or BERT) to generate optimal food suggestions for the user.
[0643] Output: List of suggested foods generated
[0644] Step 5:
[0645] The server sends the generated food suggestions to the user's terminal.
[0646] Input: List of suggested foods generated
[0647] Operation: The server sends the suggestion list to the user's terminal via the network.
[0648] Output: List of food suggestions sent to the user's terminal
[0649] Step 6:
[0650] The terminal displays a list of suggested foods to the user. The user selects from the list and completes the order.
[0651] Input: Food suggestion list
[0652] Operation: The terminal displays a list of suggestions through the application's user interface, and the user selects and places an order.
[0653] Output: Order information for food selected by the user.
[0654] Step 7:
[0655] The server sends the user's order information to the food delivery service and confirms the order.
[0656] Input: Order information for food selected by the user
[0657] Operation: The server sends order information to the food delivery service's API.
[0658] Output: Confirmed order information for food delivery service
[0659] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0660] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and specified format. By combining this system with an emotion engine, it can recognize the user's emotions and adjust the content and tone of the generated output. The specific operation method and program processing of this system are described below.
[0661] 1. Data Input
[0662] Terminal-side processing
[0663] The user accesses a provided input form using their device. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user might enter the contents of a will in text format, record a voice message about memories, and upload related photos.
[0664] After a user enters or uploads data, the device temporarily stores this data and sends it to the server. This makes the entered data available within the system.
[0665] 2. Summarizing and Transformation
[0666] Server-side processing
[0667] The server receives data sent from the terminal and stores it in storage. The stored data is then prepared to be passed on to the generating AI later. Specifically, this involves verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[0668] The server passes the stored data to the generating AI, instructing it to summarize the data and convert it into a specified format. The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video message for a funeral eulogy.
[0669] 3. Emotion Recognition and Regulation
[0670] Emotional Engine Processing
[0671] Before the data is passed to the generative AI, the server passes the audio and photo data to the emotion engine for emotion recognition. For example, the emotion engine analyzes the user's emotional state from the audio data and provides the results to the generative AI. Similarly, for photo data, the emotion engine recognizes the user's emotions when analyzing the metadata.
[0672] The emotional information recognized by the emotion engine is used as a guideline for the generating AI to adjust the output content and tone. For example, if the content of a will is very moving, the tone of the text and the atmosphere of the video will be adjusted to emphasize that.
[0673] 4. Output generation and delivery
[0674] Server-side processing
[0675] The server receives the summary and output in the specified format created by the generation AI and saves it to storage. The saved output is managed on a per-user basis. Next, this output is converted into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file.
[0676] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[0677] Terminal-side processing
[0678] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly communicated.
[0679] As a concrete example, consider a scenario where a user uploads a memorable photo and records an emotionally charged audio recording of the memories associated with that photo. The device sends the audio and photo data to a server, which then passes this data to an emotion engine and a generative AI. The generative AI adjusts the output based on the photo's metadata and emotional information, ultimately generating an emotionally resonant output that combines summarized text and the photo, which is then delivered to the user.
[0680] In this way, this system accurately records the user's thoughts and provides them at the appropriate time in a form that reflects their emotions, effectively solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[0681] The following describes the processing flow.
[0682] Step 1:
[0683] The device displays an input form to the user. The form includes a text box, a voice recording button, and a photo upload button.
[0684] Step 2:
[0685] Users input text data, record voice messages, and upload photos. For example, they might enter the contents of their will in text format, record an emotionally charged voice message about a memory, and upload related photos.
[0686] Step 3:
[0687] The device temporarily stores data entered by the user. This temporary storage uses the device's memory.
[0688] Step 4:
[0689] The device sends the temporarily stored data to the server. If the transmission is successful, the device displays a confirmation message.
[0690] Step 5:
[0691] The server receives data sent from the terminal and saves it to storage. The storage location is managed individually for each user.
[0692] Step 6:
[0693] The server preprocesses the stored data. Specifically, this includes verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[0694] Step 7:
[0695] The server passes the pre-processed data to the emotion engine and instructs it to analyze the user's emotions.
[0696] Step 8:
[0697] The emotion engine recognizes the user's emotions from voice and photo data and provides the results to the generating AI. For example, the emotion engine analyzes the user's emotional state from voice data and identifies emotions such as "compassion," "sadness," and "joy."
[0698] Step 9:
[0699] The server passes the emotional information obtained from the emotion engine and pre-processed data to the generating AI, instructing it to summarize and convert it into a specified format.
[0700] Step 10:
[0701] The generative AI analyzes data and generates summaries and output in a specified format that reflect emotional information. For example, if the content of a will is very moving, it will adjust the tone of the text and the atmosphere of the video to emphasize that.
[0702] Step 11:
[0703] The server receives the generated output and saves it to storage. The saved output is managed on a per-user basis.
[0704] Step 12:
[0705] The server converts the output into a format that users can view, download, or share. For example, a will video reflecting emotions might be saved as an MP4 file.
[0706] Step 13:
[0707] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[0708] Step 14:
[0709] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output.
[0710] Step 15:
[0711] Users can view the output through their device and download or share it as needed. For example, they can play the generated will video and share it with their family.
[0712] Through the steps described above, this system can analyze user-input data using an emotion engine and generative AI, record it in a way that reflects emotions, and provide it at the appropriate time.
[0713] (Example 2)
[0714] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0715] For elderly users or those with advanced dementia who have difficulty communicating their wishes, there is a need to efficiently record data such as memories and wills in a way that reflects their emotions, and provide it as appropriate output. However, conventional systems cannot recognize and reflect the user's emotions, and the content and tone of the generated output may not match the user's wishes. To solve this problem, a system is needed that can recognize the user's emotions and adjust the content and tone of the generated output.
[0716] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0717] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI to summarize and convert it into a specified format; means for providing the generated output to the user; means for recognizing emotions from the voice and photo data using an emotion engine; means for adjusting the content and tone of the output based on the recognized emotion information; and means for saving the final output to storage and notifying the user. This enables the generation of output that reflects the user's emotions, allowing them to appropriately convey their intentions and memories.
[0718] "Text data" refers to text information entered by users through input forms or similar means.
[0719] "Audio data" refers to voice messages entered by the user using the recording function.
[0720] "Photo data" refers to still image files uploaded by users.
[0721] "Means of receiving" refers to the function of acquiring text data, audio data, and photo data entered by the user from the terminal to the server.
[0722] "Generative AI" refers to artificial intelligence models used to summarize and convert data received from users into a specified format.
[0723] A "summarizing method" is a function that processes received data to summarize it and concisely organize the information.
[0724] "Means of converting to a specified format" refers to a function that converts summarized data into a format (text, video, etc.) that is easy for users to view.
[0725] "Means of provision" refers to functions that notify users of the output created by the generative AI, and enable them to view, download, and share it.
[0726] The "emotion engine" is a processing system that analyzes a user's emotions from audio and photo data and provides the results to the generating AI.
[0727] "Means of recognition" refers to a function that uses an emotion engine to identify the user's emotions from audio and photo data.
[0728] "Means of adjustment" refers to a function that allows the generating AI to change the content and tone of its output based on emotional information.
[0729] "Storage" refers to a location where generated output is saved and made accessible to users as needed.
[0730] "Notification means" refers to a notification function that informs the user when the server has finished preparing the generated output.
[0731] Modes for carrying out the invention
[0732] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides them in a summarized and specified format. By combining this system with an emotion engine, it can recognize the user's emotions and adjust the content and tone of the generated output.
[0733] 1. Data Input
[0734] Terminal-side processing
[0735] The user accesses an input form provided by the system using their device. This input form includes a text box, a voice recording button, and a photo upload button. Through this form, the user can, for example, enter the contents of a will in text format, record a voice message about memories, and upload related photos. This allows the user to express their wishes in a variety of ways.
[0736] After data entry, the terminal temporarily stores this data and sends it to the server. This transmission process makes the entered data available within the system.
[0737] 2. Summarizing and Transformation
[0738] Server-side processing
[0739] The server receives data sent from the terminal and stores it in storage. The received data is processed as follows:
[0740] The text data undergoes grammatical and consistency checks.
[0741] The audio data is converted to text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[0742] Metadata (such as the date and time of shooting, and location) is extracted from the photo data.
[0743] The data, after the above processing is complete, is passed to a generating AI (e.g., a GPT model) for summarization and conversion to a specified format. The generating AI uses natural language processing algorithms to perform tasks such as summarizing the contents of a will into a concise text format or generating a video message for a funeral eulogy.
[0744] 3. Emotion Recognition and Regulation
[0745] Emotional Engine Processing
[0746] Before the data is passed to the generative AI, the server passes the audio and photo data to an emotion engine (e.g., Microsoft Azure Cognitive Services) to recognize emotions. For example, the emotion engine analyzes the user's emotional state (joy, sadness, anger, etc.) from the audio data and provides the result to the generative AI. Similarly, for photo data, the emotion engine recognizes the user's emotions by analyzing the metadata.
[0747] The emotional information recognized by the emotion engine is used as a guideline for the generating AI to adjust the output content and tone. For example, if the content of a will is emotional, the tone of the text and the atmosphere of the video will be adjusted to emphasize that emotion.
[0748] 4. Output generation and provision
[0749] Server-side processing
[0750] The server receives the summary and output in the specified format created by the generation AI and saves it to storage. The saved output is managed on a per-user basis. Next, the server converts the output into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file. The server then notifies the user when the output is ready. Notification methods include email notifications and in-app notifications.
[0751] Terminal-side processing
[0752] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly conveyed.
[0753] Specific example
[0754] As a concrete example, consider a scenario where a user uploads photos of travel memories and records an emotionally moving story related to those photos in audio. The device sends the photos and audio data to a server, which then passes them to an emotion engine and a generative AI. The emotion engine extracts emotional information from the audio, and the generative AI uses this to create an emotionally moving output. Finally, the user is provided with an output that combines summarized text and photos.
[0755] Example of a prompt
[0756] An example of a prompt to input into a generative AI model is, "Based on the audio message and photos provided by the user about past memories, please generate a summary text that evokes emotion." Using such prompts, the generative AI can produce output that accurately reflects the user's intentions.
[0757] As described above, this system accurately records the user's thoughts and provides output that reflects their emotions, thereby effectively solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[0758] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0759] Step 1:
[0760] User enters data
[0761] The user uses a device to access the system's input form. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user enters the contents of a will in text format, records an audio message about memories, and uploads related photos (input: text, audio, and photo data). The device temporarily stores this data (output: data with temporarily stored content).
[0762] Step 2:
[0763] The device temporarily stores and transmits data.
[0764] After the user completes data entry, the device sends the temporarily stored data to the server (Input: Temporarily stored text, audio, and photo data). This process ensures that the data is sent to the server securely and efficiently (Output: Data sent to the server).
[0765] Step 3:
[0766] The server saves the data to the receiving storage.
[0767] The server receives data sent from the terminal and stores it in storage (Input: Data sent from the terminal). The stored data is then ready for processing and transformation (Output: Data stored in storage).
[0768] Step 4:
[0769] The server performs proofreading and preprocessing.
[0770] The server classifies the stored data and performs the following actions on each:
[0771] Text data validation (checking grammar and consistency)
[0772] Converting audio data to text (using speech recognition software to convert speech to text)
[0773] Metadata extraction from photo data (extracting information such as date, time, and location of shooting)
[0774] (Input: Saved text, audio, and photo data). The data is then ready to proceed to the next processing step (Output: Preprocessed data).
[0775] Step 5:
[0776] The server passes the data to the generating AI and instructs it on summarization and transformation.
[0777] The server passes the pre-processed data to the generating AI, requesting it to summarize and convert it to a specified format (Input: Pre-processed data). The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video for a funeral eulogy (Output: Summarized and converted output).
[0778] Step 6:
[0779] The server passes the data to the emotion engine to perform emotion recognition.
[0780] Before the data is passed to the generating AI, the server passes the audio and photo data to the emotion engine to recognize emotions (input: audio data, photo data). The emotion engine analyzes the user's emotional state from the audio data and also analyzes the metadata of the photo data to recognize emotions (output: recognized emotion information).
[0781] Step 7:
[0782] The AI generates the output based on emotional information, adjusting the content and tone.
[0783] The server provides the recognized emotional information to the generating AI, which then adjusts the content and tone of the output based on that information (input: recognized emotional information). For example, it sets the tone for the text or video of an emotional will (output: output reflecting the emotion).
[0784] Step 8:
[0785] The server saves the final output and notifies the user.
[0786] The server receives the final output created by the generation AI and saves it to storage (Input: Final Output). Then, it notifies the user that the output is ready (Output: Saved Output, Notification to User).
[0787] Step 9:
[0788] The device prompts for output confirmation and supports sharing.
[0789] The device receives a notification from the server and displays a screen prompting the user to review the output (Input: Notification). The user reviews the generated output through this screen and downloads or shares it as needed (Output: Reviewed output, downloaded or shared).
[0790] In this way, the system can effectively manage user data and provide output that reflects emotions.
[0791] (Application Example 2)
[0792] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0793] With an increasing number of elderly users and those with advanced dementia who have difficulty communicating their wishes effectively, there is a problem in that these users have difficulty communicating their needs regarding the products and services they are looking for in physical stores. This also hinders smooth communication with store staff, resulting in a decline in the quality of the user experience. To solve this, a system is needed that can accurately convey the user's wishes in a way that reflects their emotions.
[0794] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0795] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI for summarization and conversion into a specified format; means for recognizing the user's emotions from the received data; means for adjusting the content and tone of the output based on the emotion recognition results; and means for providing the generated output to the user. This enables users to communicate their intentions accurately and emotionally in a physical store setting.
[0796] A "user" is a person who uses the system to input data, and in particular, in physical stores, this includes elderly people who have difficulty communicating their intentions and users whose dementia is progressing.
[0797] "Text data" refers to information in text format entered by users, including requests and questions about products and services.
[0798] "Audio data" refers to audio information recorded by users, used to convey opinions and requests regarding products and services.
[0799] "Photo data" refers to image information uploaded by users to the system, including visual data related to products and services.
[0800] "Generative AI" refers to artificial intelligence (AI) technology used to summarize and convert received data into a specified format, employing natural language processing and data generation models.
[0801] An "emotion engine" refers to a technology that recognizes a user's emotions from audio and photo data, analyzing emotional information and adjusting the output accordingly.
[0802] "Output" refers to summaries or results in a specified format created by the generating AI based on data entered by the user, and is provided in formats such as text, video, and audio.
[0803] "Emotion recognition" is the process of using an emotion engine to analyze a user's emotional state from audio or photos and providing that information to a generating AI.
[0804] "Tone" refers to the emotion and atmosphere in the content and format of the output, and is adjusted according to the user's emotional state.
[0805] System Configuration
[0806] This system consists of a terminal that receives text data, voice data, and photo data entered by the user, a server that processes and analyzes the data, and a means of providing the generated output. The main components are as follows:
[0807] 1. Device: Users input text data, audio data, and photo data through a device such as a smartphone or tablet. This device is connected to the internet and has network communication capabilities to send data to the server.
[0808] 2. Server: The server stores and analyzes the received data. Using a cloud server as a hosting service is recommended here. Programming languages such as Python and related libraries are used for data analysis.
[0809] 3. Generative AI: Uses natural language processing and data generation models (e.g., Hugging Face's transformers library) to summarize user input and convert it to a specified format.
[0810] 4. Emotion Engine: Uses technologies (e.g., DeepFace library, Google Speech Recognition API) to recognize the user's emotional state from photo and audio data.
[0811] Processing flow
[0812] The server primarily processes data in the following steps:
[0813] 1. Receiving and storing data:
[0814] This function temporarily stores text data, audio data, and photo data entered by the device.
[0815] The data will be stored in cloud storage with security in mind.
[0816] 2. Speech recognition:
[0817] The server uses the Google Speech Recognition API to convert the audio data into text data.
[0818] This text data will be used for subsequent processing in the generative AI and emotion engine.
[0819] 3. Emotion recognition:
[0820] Using the DeepFace library, we recognize the user's emotions from photographic data and extract emotional information.
[0821] Emotional information is passed to the generating AI and used to adjust the tone of the output.
[0822] 4. Text summarization and conversion:
[0823] Use models from the transformers library to perform summarization and conversion to a specified format.
[0824] It incorporates emotional information into tone adjustments to generate appropriate text and video messages.
[0825] 5. Output generation and delivery:
[0826] The generated output is provided to the user. Specifically, notifications are sent via in-app notifications or email.
[0827] Users can view, download, or share the output generated through their device.
[0828] Specific example
[0829] Users record a voice message on their device to describe the details of the products they want in a physical store and upload photos of the store. The voice data is converted to text using the Google Speech Recognition API, and then the photo data undergoes emotion recognition using the DeepFace library. Based on this information, a generative AI model summarizes and provides detailed product information in a friendly tone.
[0830] Example of a prompt:
[0831] Convert the audio file "customer_request.wav" recorded by the customer in the store into text and summarize the text. Also, recognize the customer's emotions from the photo "customer_photo.jpg" and generate a response in a tone appropriate to those emotions.
[0832] This will enable elderly users and those with dementia to communicate their wishes accurately and emotionally to store staff.
[0833] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0834] Step 1:
[0835] Data entry and reception
[0836] The device receives text data entered by the user, audio data recorded, and photo data uploaded by the user. Users input this data using a smartphone or tablet via a dedicated application. The device temporarily stores the received data in local storage and sends it to a cloud server via the internet. Input is performed by the user using input forms, recording buttons, and photo upload functions, and the output in these cases is the data sent to the cloud server.
[0837] Step 2:
[0838] Data storage and preprocessing
[0839] The server receives text, audio, and photo data sent from the terminal and saves it to cloud storage. Simultaneously with saving, it begins data preprocessing. Specifically, audio data is converted to text format using the Google Speech Recognition API. This process converts the audio data into text data, which is then passed on to the next processing step. For photo data, identifiable metadata is extracted.
[0840] Step 3:
[0841] emotion recognition
[0842] The server performs emotion recognition using pre-processed data. It uses the DeepFace library to recognize the user's emotions from photographic data. Given photographic data as input, the emotion recognition engine extracts the user's emotional state (e.g., joy, sadness, surprise) as depicted in the photograph. Emotional information is obtained as output. This prepares the user's emotional information for delivery to the generating AI model.
[0843] Step 4:
[0844] Text summarization and tone adjustment
[0845] The server uses a generative AI model based on emotional information to summarize and tone-adjust text data. In particular, it uses the Hugging Face transformers library to summarize text. The input is pre-processed text and emotional information, and the output text is generated with the summary result and tone adjusted according to the emotion. Here, the tone is adjusted to be friendly when the emotion is joyful, and gentle when the emotion is sad.
[0846] Step 5:
[0847] Output generation and delivery
[0848] The generated text or other specified output formats (e.g., video messages) are delivered to the user from the server. The user is notified when the output is ready via in-app notifications or email. The user can then view, download, or share the generated output through their device. For example, the AI generates detailed information about the product the user wants to purchase, summarized in a friendly tone, making it easy for the user to communicate their intentions.
[0849] Through all these processing steps, the system enables elderly users and those with advanced dementia who have difficulty communicating to smoothly convey their intentions in physical stores.
[0850] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0851] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0852] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0853] [Third Embodiment]
[0854] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0855] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0856] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0857] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0858] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0859] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0860] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0861] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0862] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0863] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0864] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0865] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0866] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and specified format. The specific operation method and program processing of this system will be described below.
[0867] 1. Data Input
[0868] Terminal-side processing
[0869] The user accesses a provided input form using their device. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user might enter the contents of a will in text format, record a voice message about memories, and upload related photos.
[0870] After a user enters or uploads data, the device temporarily stores this data and sends it to the server. This makes the entered data available within the system.
[0871] 2. Summarizing and Transformation
[0872] Server-side processing
[0873] The server receives data sent from the terminal and stores it in storage. The stored data is then prepared to be passed on to the generating AI later. Specifically, this involves preprocessing of text data, transcribing audio data into text, and extracting metadata from photo data.
[0874] The server passes the stored data to the generating AI, instructing it to summarize the data and convert it into a specified format. The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video message for a funeral eulogy.
[0875] 3. Output generation and provision
[0876] Server-side processing
[0877] The server receives the summary and output in the specified format created by the generation AI and saves it for each user. Next, it converts this output into a format that the user can view, download, or share. For example, a will video might be saved as an MP4 file.
[0878] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[0879] Terminal-side processing
[0880] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly communicated.
[0881] To illustrate with a concrete example, consider a scenario where a user uploads a memorable photo and records an audio recording of the memory associated with that photo. The device sends the audio data and photo data to a server, which then passes them to a generative AI. The generative AI extracts metadata from the photo, transcribes the audio data into text, and summarizes its content. Finally, an output combining the summarized text and the photo is generated and provided to the user.
[0882] In this way, this system accurately records the user's thoughts and provides them in the appropriate format when needed, thereby solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[0883] The following describes the processing flow.
[0884] Step 1:
[0885] The device displays an input form to the user. The form includes a text box, a voice recording button, and a photo upload button.
[0886] Step 2:
[0887] Users input text data, record voice messages, and upload photos.
[0888] Step 3:
[0889] The device temporarily stores data entered by the user. This temporary storage uses the device's memory.
[0890] Step 4:
[0891] The device sends the temporarily stored data to the server. Once the transmission is complete, the device displays a confirmation message.
[0892] Step 5:
[0893] The server receives data sent from the terminal and saves it to storage. The storage location is managed individually for each user.
[0894] Step 6:
[0895] The server preprocesses the stored data. Specifically, this includes verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[0896] Step 7:
[0897] The server passes the pre-processed data to the generating AI, instructing it to summarize and convert it into a specified format.
[0898] Step 8:
[0899] The generation AI analyzes the received data and generates summaries and outputs in a specified format. For example, it can convert audio data into a text-based will or create a message based on photographs.
[0900] Step 9:
[0901] The server receives the generated output and saves it to storage. The saved output is managed on a per-user basis.
[0902] Step 10:
[0903] The server converts the saved output into a format that users can view, download, or share.
[0904] Step 11:
[0905] The server notifies the user when the output is ready. This notification is sent via email or an in-app notification.
[0906] Step 12:
[0907] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output.
[0908] Step 13:
[0909] Users can view the output through their device and download or share it as needed. For example, they can play the generated will video and share it with their family.
[0910] Through these steps, the system enables users to record memories and important messages in the appropriate format and provide them at the appropriate time.
[0911] (Example 1)
[0912] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0913] For elderly users or those with advanced dementia who have difficulty communicating their wishes effectively, efficiently recording data such as text, audio, and photographs, summarizing it, and converting it to a specified format for delivery presents a significant challenge. Solving this problem requires a system that handles everything from data input to output delivery, but current technology lacks a suitable solution.
[0914] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0915] In this invention, the server includes means for temporarily storing and transmitting user-inputted text, voice, and photo data to the server; means for pre-processing the data stored on the server; means for passing the pre-processed data to a generation AI model for summarization and conversion into a specified format; means for storing the generated output for each user; and means for notifying the user of the stored output and providing it in a format that can be viewed, downloaded, or shared. This makes it possible for elderly users or users with advanced dementia who have difficulty communicating their wishes to accurately and concisely record their wishes and provide them in an appropriate format when needed.
[0916] "Text data" refers to the text information that a user enters into an input form.
[0917] "Audio data" refers to voice messages recorded by the user using the voice recording button.
[0918] "Photo data" refers to image files that users upload using the photo upload button.
[0919] "Temporary storage" refers to the process of temporarily storing data before it is sent from the device to the server.
[0920] A "server" is a central system that receives, stores, and processes data sent from terminals.
[0921] "Preprocessing" refers to the process by which a server formats received data for analysis and transformation. Specifically, this includes text cleaning, audio-to-text conversion, and image metadata extraction.
[0922] A "generative AI model" is an artificial intelligence system that summarizes received data and converts it into a specified format.
[0923] "Output" refers to the output data created by the generative AI model, which is summarized and formatted in a specified format.
[0924] "Storage" refers to the process where the server identifies and stores the processed output for each user.
[0925] "Notification" refers to the server informing the user that the output is ready.
[0926] "Viewing" means making the generated output available for the user to review.
[0927] "Downloading" refers to the act of a user saving the generated output to their own device.
[0928] "Sharing" refers to the act of a user sharing the output they have generated with others.
[0929] "Metadata" refers to additional information included in photo data, such as the date and time of shooting, location, and camera settings.
[0930] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and converted format. This system receives user input data, processes it on a server, and provides output that is summarized and converted by a generating AI model.
[0931] Specifically, this is achieved using terminals, servers, and generative AI models as follows:
[0932] Hardware and software to be used
[0933] terminal
[0934] Hardware: Smartphone, tablet, or PC
[0935] Software: Dedicated application or form on a web browser
[0936] server
[0937] Hardware: High-performance cloud servers
[0938] Software: Databases, storage systems, analysis libraries (e.g., OpenCV, NLP libraries)
[0939] Generative AI Models
[0940] Software: Natural language processing algorithms, speech recognition engines (e.g., Google Speech-to-Text API)
[0941] Specific operating methods for the system
[0942] 1. Data Input
[0943] Terminal-side processing
[0944] The user launches a dedicated application on their device and accesses an input form displayed on the main screen. This form includes a text box, a voice recording button, and a photo upload button. The user enters the contents of their will in text format, records a voice message about their memories, and uploads related photos. After entering, recording, and uploading the data, the device temporarily stores the data and sends it to the server.
[0945] 2. Summarizing and Transformation
[0946] Server-side processing
[0947] The server receives data sent from the terminal and stores it in storage. After storage, the server preprocesses the data. Specifically, it cleans text data, transcribes audio data, and extracts metadata from photo data. Then, it passes the preprocessed data to a generative AI model and instructs it to summarize and transform it into a specified format. The generative AI model uses natural language processing algorithms to summarize and transform the data.
[0948] 3. Output generation and provision
[0949] Server-side processing
[0950] The system receives the summary and output in the specified format generated by the generative AI model and saves it for each user. Next, it converts this output into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file. The user is then notified when this information is ready. Notification methods include email and in-app notifications.
[0951] Terminal-side processing
[0952] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user reviews the generated output and downloads or shares it as needed.
[0953] Specific example
[0954] Consider a scenario where a user uploads a memorable photo and records an audio recording of the associated memories. In this case, the device sends the audio data and photo data to a server. The server passes this data to a generative AI model, which extracts the metadata from the photo and transcribes the audio data into text to summarize its content. Finally, an output combining the summarized text and the photo is generated and provided to the user.
[0955] Example of a prompt
[0956] Text Summary: "Please summarize this text: 'A message to my family: Thank you for everything. I look forward to your continued support.'"
[0957] Speech-to-text: "Please generate text from this audio. The audio file says, 'This is an important message to my family. This photo is from a trip we took together.'"
[0958] In this way, the system helps users effectively communicate their intentions by accurately understanding them and converting them into an appropriate format.
[0959] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0960] Step 1:
[0961] Displaying the input form
[0962] The user launches the application on their device and accesses the input form displayed on the main screen. The input form includes a text box, a voice recording button, and a photo upload button. Specifically, the device renders the input form and prompts the user for input. Input can be in the form of text, audio, or an image file.
[0963] Step 2:
[0964] Data entry, recording, and uploading.
[0965] The user enters text into a form, taps the voice recording button to record a voice message, and uses the photo upload button to select and upload a photo. Specifically, the device temporarily saves the entered text, the recorded voice data, and the uploaded photo data. As output, the device retains the saved data.
[0966] Step 3:
[0967] Sending data
[0968] After the user completes input, recording, and uploading, they tap the send button. The device then sends this data to the server. Specifically, the device sends the temporarily stored data to the server via an HTTP request. The input consists of temporarily stored text, audio, and image data, while the output is the transmission of data to the server.
[0969] Step 4:
[0970] Receiving and storing data
[0971] The server receives data sent from the terminal and stores it in storage. Specifically, the server stores the received data in a database and saves audio and images to file storage. The input is the data sent from the terminal, and the output is the data stored in storage.
[0972] Step 5:
[0973] Data preprocessing
[0974] The server performs preprocessing on the stored data. Specifically, the server uses an NLP (Natural Language Processing) library to clean text, a speech recognition engine (e.g., Google Speech-to-Text API) to convert speech to text, and an image analysis tool (e.g., OpenCV) to extract metadata from photos. The input consists of stored text, audio, and image data, and the output is preprocessed data.
[0975] Step 6:
[0976] Sending data to the generative AI model
[0977] The server passes pre-processed data to the generative AI model and instructs it to summarize and transform it into a specified format. Specifically, the server sends API requests to the generative AI model, inputting prompts for summarization and transformation. The input is pre-processed data, and the output is the output generated by the generative AI model.
[0978] Step 7:
[0979] Receiving and saving the generated output
[0980] The server receives the output from the generative AI model and saves it for each user. Specifically, the server saves the generated output, such as text and video, to a database. The input is the output from the generative AI model, and the output is the saved output.
[0981] Step 8:
[0982] Output format conversion
[0983] The server converts the generated output into a format that users can view, download, or share. Specifically, the server uses a file conversion tool to generate PDF or MP4 files. The input is the saved output, and the output is a file in a format usable by the user.
[0984] Step 9:
[0985] Notification to the user
[0986] The server notifies the user when the output is ready. Specifically, the server calls a notification API and sends an email or push notification to the user. The input is the output after format conversion, and the output is the notification to the user.
[0987] Step 10:
[0988] Receiving and displaying notifications
[0989] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output. Specifically, the terminal checks the notification and renders a dedicated confirmation screen. The input is the notification from the server, and the output is the confirmation screen displayed to the user.
[0990] Step 11:
[0991] Check the output and take action
[0992] The user reviews the generated output and downloads or shares it as needed. Specifically, the device displays an output viewing screen and provides download links and share buttons. The input is the output displayed on the confirmation screen, and the output is the download or sharing action performed by the user.
[0993] (Application Example 1)
[0994] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0995] Currently, there is no system that allows elderly users or those with advanced dementia, who may have difficulty communicating their wishes effectively, to easily and accurately order food delivery. In particular, there is a need for systems that understand the user's ordering intent using voice and photo data and provide appropriate suggestions based on that understanding. However, conventional systems struggle to meet these requirements. This creates a problem where elderly users and those with dementia find it difficult to easily use food delivery services.
[0996] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0997] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI to summarize and convert it into a specified format; means for providing the generated output to the user; means for converting voice data into text and extracting metadata from image data to generate food suggestions; and means for presenting the generated food suggestions to the user. This makes it possible for elderly users or users with advanced dementia who have difficulty communicating their wishes to use food delivery services easily and accurately.
[0998] "Users" refer to elderly people or individuals with dementia who have difficulty communicating their wishes appropriately and who use this system.
[0999] "Text data" refers to information in text format that is entered into the device used by the user.
[1000] "Audio data" refers to information in audio format provided by the user through voice input.
[1001] "Photo data" refers to image data uploaded by users via their devices.
[1002] "Means of receiving" refers to devices or programs that have the function of receiving text, audio, and photographic data provided by the user.
[1003] "Generative AI" refers to artificial intelligence models used to summarize received data and convert it into a specified format.
[1004] "Means for summarizing and converting to a specified format" refers to a processing mechanism that passes the received data to a generating AI, summarizes the data concisely, and converts it to the required format.
[1005] "Means of providing output" refers to devices or programs that have the function of presenting output data generated by the generation AI to the user in an appropriate format.
[1006] "Converting audio data to text" refers to the process of analyzing audio data and converting it into corresponding text data.
[1007] "Metadata" refers to information associated with photographic data, such as the date and time of shooting, location, and the type of object depicted.
[1008] "Food suggestions" refers to dishes and foods recommended by the generating AI based on the user's input data.
[1009] "Means for generating suggestions" refers to a technical system for creating suggestions tailored to the user based on the received data.
[1010] "Means of presenting proposals" refers to devices or programs that have the function of visually or audibly displaying the generated proposals to the user.
[1011] The system for realizing this invention is designed to allow users to easily utilize food delivery even if they are elderly or have difficulty communicating their wishes due to the progression of dementia. The system consists of several main components, including a terminal, a server, and a generative AI model.
[1012] Overview of program processing
[1013] 1. Processing on the terminal side
[1014] The terminal receives text, voice, and photo data entered by the user. Through an application installed on the terminal, the user can voice-input descriptions of desired dishes or favorite dishes, and upload photos of previously ordered dishes. The input data is temporarily stored on the terminal and then sent to the server.
[1015] 2. Server-side processing
[1016] The server receives data sent from the terminal and performs data preprocessing. Specifically, it converts audio data into text format (speech recognition) and extracts metadata from photo data (image analysis). After this, the data is passed to a generative AI model to generate a data summary and food suggestions. The generative AI model uses natural language processing algorithms and image recognition algorithms.
[1017] 3. Output generation and provision
[1018] The system receives suggestions generated by a generative AI model and generates output tailored to each user's requirements. A list of suggested foods is presented to the user through the application. The user can then select from this list and complete their order.
[1019] Specific example
[1020] If a user voice-inputs "I want to eat spaghetti today" and uploads a photo of spaghetti they previously ordered, the system will process it as follows:
[1021] 1. The device records the user's voice saying "I want to eat spaghetti today," retrieves a photo of spaghetti ordered previously, and sends it to the server.
[1022] 2. The server converts the audio data into text format and extracts the metadata from the photos.
[1023] 3. The generative AI model analyzes text and photo data to suggest food suitable for the user.
[1024] 4. The user selects spaghetti from the suggested food list and completes the order.
[1025] Hardware and software to be used
[1026] 1. Microphone: Used to input the user's voice.
[1027] 2. Camera: Used to take photos that users upload.
[1028] 3. Terminal (smartphone / tablet): A device used by the user to input data and communicate with the server.
[1029] 4. Server: Responsible for storing data and running the generated AI model.
[1030] 5. Generative AI Models: Algorithms that combine speech recognition, natural language processing, and image recognition.
[1031] Example of a prompt
[1032] In the user's own voice: "What I want to eat right now is, for example, curry rice or ramen, or any of my favorite dishes."
[1033] Example photo: "Please upload a photo of the curry rice you ordered previously."
[1034] This allows the system to provide an easy-to-use food delivery ordering system for elderly and dementia users, making access to meals easier.
[1035] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1036] Step 1:
[1037] Users use their devices to voice-input the characteristics of the dishes they want to eat or their favorite dishes, and upload photos of dishes they have ordered in the past.
[1038] Input: Audio data, photo data
[1039] Operation: The user inputs audio using the application's voice recording function and uploads photos using the camera function.
[1040] Output: The input audio and photo data are temporarily stored on the device.
[1041] Step 2:
[1042] The device sends the voice and photo data entered by the user to the server.
[1043] Input: Temporarily stored audio data, photo data
[1044] Operation: The terminal sends data to the server via the network.
[1045] Output: Audio and photo data transferred to the server
[1046] Step 3:
[1047] The server converts the audio data to text and extracts metadata from the photo data.
[1048] Input: Audio data, photo data
[1049] Data processing: Convert audio data into text data using speech recognition technology. Extract metadata from photo data using image analysis technology.
[1050] Operation: The server uses a specified speech recognition algorithm (e.g., Google Speech Recognition). Image recognition APIs (e.g., AWS Rekognition) are used to extract metadata from photos.
[1051] Output: Text data, photo metadata
[1052] Step 4:
[1053] The server passes text data and photo metadata to a generating AI model to generate food suggestions.
[1054] Input: Text data, photo metadata
[1055] Data processing: A generative AI model analyzes text data and photo metadata to generate food suggestions.
[1056] Operation: The server uses a generative AI model (e.g., GPT-3 or BERT) to generate optimal food suggestions for the user.
[1057] Output: List of suggested foods generated
[1058] Step 5:
[1059] The server sends the generated food suggestions to the user's terminal.
[1060] Input: List of suggested foods generated
[1061] Operation: The server sends the suggestion list to the user's terminal via the network.
[1062] Output: List of food suggestions sent to the user's terminal
[1063] Step 6:
[1064] The terminal displays a list of suggested foods to the user. The user selects from the list and completes the order.
[1065] Input: Food suggestion list
[1066] Operation: The terminal displays a list of suggestions through the application's user interface, and the user selects and places an order.
[1067] Output: Order information for food selected by the user.
[1068] Step 7:
[1069] The server sends the user's order information to the food delivery service and confirms the order.
[1070] Input: Order information for food selected by the user
[1071] Operation: The server sends order information to the food delivery service's API.
[1072] Output: Confirmed order information for food delivery service
[1073] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1074] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and specified format. By combining this system with an emotion engine, it can recognize the user's emotions and adjust the content and tone of the generated output. The specific operation method and program processing of this system are described below.
[1075] 1. Data Input
[1076] Terminal-side processing
[1077] The user accesses a provided input form using their device. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user might enter the contents of a will in text format, record a voice message about memories, and upload related photos.
[1078] After a user enters or uploads data, the device temporarily stores this data and sends it to the server. This makes the entered data available within the system.
[1079] 2. Summarizing and Transformation
[1080] Server-side processing
[1081] The server receives data sent from the terminal and stores it in storage. The stored data is then prepared to be passed on to the generating AI later. Specifically, this involves verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[1082] The server passes the stored data to the generating AI, instructing it to summarize the data and convert it into a specified format. The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video message for a funeral eulogy.
[1083] 3. Emotion Recognition and Regulation
[1084] Emotional Engine Processing
[1085] Before the data is passed to the generative AI, the server passes the audio and photo data to the emotion engine for emotion recognition. For example, the emotion engine analyzes the user's emotional state from the audio data and provides the results to the generative AI. Similarly, for photo data, the emotion engine recognizes the user's emotions when analyzing the metadata.
[1086] The emotional information recognized by the emotion engine is used as a guideline for the generating AI to adjust the output content and tone. For example, if the content of a will is very moving, the tone of the text and the atmosphere of the video will be adjusted to emphasize that.
[1087] 4. Output generation and delivery
[1088] Server-side processing
[1089] The server receives the summary and output in the specified format created by the generation AI and saves it to storage. The saved output is managed on a per-user basis. Next, this output is converted into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file.
[1090] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[1091] Terminal-side processing
[1092] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly communicated.
[1093] As a concrete example, consider a scenario where a user uploads a memorable photo and records an emotionally charged audio recording of the memories associated with that photo. The device sends the audio and photo data to a server, which then passes this data to an emotion engine and a generative AI. The generative AI adjusts the output based on the photo's metadata and emotional information, ultimately generating an emotionally resonant output that combines summarized text and the photo, which is then delivered to the user.
[1094] In this way, this system accurately records the user's thoughts and provides them at the appropriate time in a form that reflects their emotions, effectively solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[1095] The following describes the processing flow.
[1096] Step 1:
[1097] The device displays an input form to the user. The form includes a text box, a voice recording button, and a photo upload button.
[1098] Step 2:
[1099] Users input text data, record voice messages, and upload photos. For example, they might enter the contents of their will in text format, record an emotionally charged voice message about a memory, and upload related photos.
[1100] Step 3:
[1101] The device temporarily stores data entered by the user. This temporary storage uses the device's memory.
[1102] Step 4:
[1103] The device sends the temporarily stored data to the server. If the transmission is successful, the device displays a confirmation message.
[1104] Step 5:
[1105] The server receives data sent from the terminal and saves it to storage. The storage location is managed individually for each user.
[1106] Step 6:
[1107] The server preprocesses the stored data. Specifically, this includes verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[1108] Step 7:
[1109] The server passes the pre-processed data to the emotion engine and instructs it to analyze the user's emotions.
[1110] Step 8:
[1111] The emotion engine recognizes the user's emotions from voice and photo data and provides the results to the generating AI. For example, the emotion engine analyzes the user's emotional state from voice data and identifies emotions such as "compassion," "sadness," and "joy."
[1112] Step 9:
[1113] The server passes the emotional information obtained from the emotion engine and pre-processed data to the generating AI, instructing it to summarize and convert it into a specified format.
[1114] Step 10:
[1115] The generative AI analyzes data and generates summaries and output in a specified format that reflect emotional information. For example, if the content of a will is very moving, it will adjust the tone of the text and the atmosphere of the video to emphasize that.
[1116] Step 11:
[1117] The server receives the generated output and saves it to storage. The saved output is managed on a per-user basis.
[1118] Step 12:
[1119] The server converts the output into a format that users can view, download, or share. For example, a will video reflecting emotions might be saved as an MP4 file.
[1120] Step 13:
[1121] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[1122] Step 14:
[1123] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output.
[1124] Step 15:
[1125] Users can view the output through their device and download or share it as needed. For example, they can play the generated will video and share it with their family.
[1126] Through the steps described above, this system can analyze user-input data using an emotion engine and generative AI, record it in a way that reflects emotions, and provide it at the appropriate time.
[1127] (Example 2)
[1128] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1129] For elderly users or those with advanced dementia who have difficulty communicating their wishes, there is a need to efficiently record data such as memories and wills in a way that reflects their emotions, and provide it as appropriate output. However, conventional systems cannot recognize and reflect the user's emotions, and the content and tone of the generated output may not match the user's wishes. To solve this problem, a system is needed that can recognize the user's emotions and adjust the content and tone of the generated output.
[1130] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1131] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI to summarize and convert it into a specified format; means for providing the generated output to the user; means for recognizing emotions from the voice and photo data using an emotion engine; means for adjusting the content and tone of the output based on the recognized emotion information; and means for saving the final output to storage and notifying the user. This enables the generation of output that reflects the user's emotions, allowing them to appropriately convey their intentions and memories.
[1132] "Text data" refers to text information entered by users through input forms or similar means.
[1133] "Audio data" refers to voice messages entered by the user using the recording function.
[1134] "Photo data" refers to still image files uploaded by users.
[1135] "Means of receiving" refers to the function of acquiring text data, audio data, and photo data entered by the user from the terminal to the server.
[1136] "Generative AI" refers to artificial intelligence models used to summarize and convert data received from users into a specified format.
[1137] A "summarizing method" is a function that processes received data to summarize it and concisely organize the information.
[1138] "Means of converting to a specified format" refers to a function that converts summarized data into a format (text, video, etc.) that is easy for users to view.
[1139] "Means of provision" refers to functions that notify users of the output created by the generative AI, and enable them to view, download, and share it.
[1140] The "emotion engine" is a processing system that analyzes a user's emotions from audio and photo data and provides the results to the generating AI.
[1141] "Means of recognition" refers to a function that uses an emotion engine to identify the user's emotions from audio and photo data.
[1142] "Means of adjustment" refers to a function that allows the generating AI to change the content and tone of its output based on emotional information.
[1143] "Storage" refers to a location where generated output is saved and made accessible to users as needed.
[1144] "Notification means" refers to a notification function that informs the user when the server has finished preparing the generated output.
[1145] Modes for carrying out the invention
[1146] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides them in a summarized and specified format. By combining this system with an emotion engine, it can recognize the user's emotions and adjust the content and tone of the generated output.
[1147] 1. Data Input
[1148] Terminal-side processing
[1149] The user accesses an input form provided by the system using their device. This input form includes a text box, a voice recording button, and a photo upload button. Through this form, the user can, for example, enter the contents of a will in text format, record a voice message about memories, and upload related photos. This allows the user to express their wishes in a variety of ways.
[1150] After data entry, the terminal temporarily stores this data and sends it to the server. This transmission process makes the entered data available within the system.
[1151] 2. Summarizing and Transformation
[1152] Server-side processing
[1153] The server receives data sent from the terminal and stores it in storage. The received data is processed as follows:
[1154] The text data undergoes grammatical and consistency checks.
[1155] The audio data is converted to text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[1156] Metadata (such as the date and time of shooting, and location) is extracted from the photo data.
[1157] The data, after the above processing is complete, is passed to a generating AI (e.g., a GPT model) for summarization and conversion to a specified format. The generating AI uses natural language processing algorithms to perform tasks such as summarizing the contents of a will into a concise text format or generating a video message for a funeral eulogy.
[1158] 3. Emotion Recognition and Regulation
[1159] Emotional Engine Processing
[1160] Before the data is passed to the generative AI, the server passes the audio and photo data to an emotion engine (e.g., Microsoft Azure Cognitive Services) to recognize emotions. For example, the emotion engine analyzes the user's emotional state (joy, sadness, anger, etc.) from the audio data and provides the result to the generative AI. Similarly, for photo data, the emotion engine recognizes the user's emotions by analyzing the metadata.
[1161] The emotional information recognized by the emotion engine is used as a guideline for the generating AI to adjust the output content and tone. For example, if the content of a will is emotional, the tone of the text and the atmosphere of the video will be adjusted to emphasize that emotion.
[1162] 4. Output generation and provision
[1163] Server-side processing
[1164] The server receives the summary and output in the specified format created by the generation AI and saves it to storage. The saved output is managed on a per-user basis. Next, the server converts the output into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file. The server then notifies the user when the output is ready. Notification methods include email notifications and in-app notifications.
[1165] Terminal-side processing
[1166] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly conveyed.
[1167] Specific example
[1168] As a concrete example, consider a scenario where a user uploads photos of travel memories and records an emotionally moving story related to those photos in audio. The device sends the photos and audio data to a server, which then passes them to an emotion engine and a generative AI. The emotion engine extracts emotional information from the audio, and the generative AI uses this to create an emotionally moving output. Finally, the user is provided with an output that combines summarized text and photos.
[1169] Example of a prompt
[1170] An example of a prompt to input into a generative AI model is, "Based on the audio message and photos provided by the user about past memories, please generate a summary text that evokes emotion." Using such prompts, the generative AI can produce output that accurately reflects the user's intentions.
[1171] As described above, this system accurately records the user's thoughts and provides output that reflects their emotions, thereby effectively solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[1172] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1173] Step 1:
[1174] User enters data
[1175] The user uses a device to access the system's input form. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user enters the contents of a will in text format, records an audio message about memories, and uploads related photos (input: text, audio, and photo data). The device temporarily stores this data (output: data with temporarily stored content).
[1176] Step 2:
[1177] The device temporarily stores and transmits data.
[1178] After the user completes data entry, the device sends the temporarily stored data to the server (Input: Temporarily stored text, audio, and photo data). This process ensures that the data is sent to the server securely and efficiently (Output: Data sent to the server).
[1179] Step 3:
[1180] The server saves the data to the receiving storage.
[1181] The server receives data sent from the terminal and stores it in storage (Input: Data sent from the terminal). The stored data is then ready for processing and transformation (Output: Data stored in storage).
[1182] Step 4:
[1183] The server performs proofreading and preprocessing.
[1184] The server classifies the stored data and performs the following actions on each:
[1185] Text data validation (checking grammar and consistency)
[1186] Converting audio data to text (using speech recognition software to convert speech to text)
[1187] Metadata extraction from photo data (extracting information such as date, time, and location of shooting)
[1188] (Input: Saved text, audio, and photo data). The data is then ready to proceed to the next processing step (Output: Preprocessed data).
[1189] Step 5:
[1190] The server passes the data to the generating AI and instructs it on summarization and transformation.
[1191] The server passes the pre-processed data to the generating AI, requesting it to summarize and convert it to a specified format (Input: Pre-processed data). The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video for a funeral eulogy (Output: Summarized and converted output).
[1192] Step 6:
[1193] The server passes the data to the emotion engine to perform emotion recognition.
[1194] Before the data is passed to the generating AI, the server passes the audio and photo data to the emotion engine to recognize emotions (input: audio data, photo data). The emotion engine analyzes the user's emotional state from the audio data and also analyzes the metadata of the photo data to recognize emotions (output: recognized emotion information).
[1195] Step 7:
[1196] The AI generates the output based on emotional information, adjusting the content and tone.
[1197] The server provides the recognized emotional information to the generating AI, which then adjusts the content and tone of the output based on that information (input: recognized emotional information). For example, it sets the tone for the text or video of an emotional will (output: output reflecting the emotion).
[1198] Step 8:
[1199] The server saves the final output and notifies the user.
[1200] The server receives the final output created by the generation AI and saves it to storage (Input: Final Output). Then, it notifies the user that the output is ready (Output: Saved Output, Notification to User).
[1201] Step 9:
[1202] The device prompts for output confirmation and supports sharing.
[1203] The device receives a notification from the server and displays a screen prompting the user to review the output (Input: Notification). The user reviews the generated output through this screen and downloads or shares it as needed (Output: Reviewed output, downloaded or shared).
[1204] In this way, the system can effectively manage user data and provide output that reflects emotions.
[1205] (Application Example 2)
[1206] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1207] With an increasing number of elderly users and those with advanced dementia who have difficulty communicating their wishes effectively, there is a problem in that these users have difficulty communicating their needs regarding the products and services they are looking for in physical stores. This also hinders smooth communication with store staff, resulting in a decline in the quality of the user experience. To solve this, a system is needed that can accurately convey the user's wishes in a way that reflects their emotions.
[1208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1209] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI for summarization and conversion into a specified format; means for recognizing the user's emotions from the received data; means for adjusting the content and tone of the output based on the emotion recognition results; and means for providing the generated output to the user. This enables users to communicate their intentions accurately and emotionally in a physical store setting.
[1210] A "user" is a person who uses the system to input data, and in particular, in physical stores, this includes elderly people who have difficulty communicating their intentions and users whose dementia is progressing.
[1211] "Text data" refers to information in text format entered by users, including requests and questions about products and services.
[1212] "Audio data" refers to audio information recorded by users, used to convey opinions and requests regarding products and services.
[1213] "Photo data" refers to image information uploaded by users to the system, including visual data related to products and services.
[1214] "Generative AI" refers to artificial intelligence (AI) technology used to summarize and convert received data into a specified format, employing natural language processing and data generation models.
[1215] An "emotion engine" refers to a technology that recognizes a user's emotions from audio and photo data, analyzing emotional information and adjusting the output accordingly.
[1216] "Output" refers to summaries or results in a specified format created by the generating AI based on data entered by the user, and is provided in formats such as text, video, and audio.
[1217] "Emotion recognition" is the process of using an emotion engine to analyze a user's emotional state from audio or photos and providing that information to a generating AI.
[1218] "Tone" refers to the emotion and atmosphere in the content and format of the output, and is adjusted according to the user's emotional state.
[1219] System Configuration
[1220] This system consists of a terminal that receives text data, voice data, and photo data entered by the user, a server that processes and analyzes the data, and a means of providing the generated output. The main components are as follows:
[1221] 1. Device: Users input text data, audio data, and photo data through a device such as a smartphone or tablet. This device is connected to the internet and has network communication capabilities to send data to the server.
[1222] 2. Server: The server stores and analyzes the received data. Using a cloud server as a hosting service is recommended here. Programming languages such as Python and related libraries are used for data analysis.
[1223] 3. Generative AI: Uses natural language processing and data generation models (e.g., Hugging Face's transformers library) to summarize user input and convert it to a specified format.
[1224] 4. Emotion Engine: Uses technologies (e.g., DeepFace library, Google Speech Recognition API) to recognize the user's emotional state from photo and audio data.
[1225] Processing flow
[1226] The server primarily processes data in the following steps:
[1227] 1. Receiving and storing data:
[1228] This function temporarily stores text data, audio data, and photo data entered by the device.
[1229] The data will be stored in cloud storage with security in mind.
[1230] 2. Speech recognition:
[1231] The server uses the Google Speech Recognition API to convert the audio data into text data.
[1232] This text data will be used for subsequent processing in the generative AI and emotion engine.
[1233] 3. Emotion recognition:
[1234] Using the DeepFace library, we recognize the user's emotions from photographic data and extract emotional information.
[1235] Emotional information is passed to the generating AI and used to adjust the tone of the output.
[1236] 4. Text summarization and conversion:
[1237] Use models from the transformers library to perform summarization and conversion to a specified format.
[1238] It incorporates emotional information into tone adjustments to generate appropriate text and video messages.
[1239] 5. Output generation and delivery:
[1240] The generated output is provided to the user. Specifically, notifications are sent via in-app notifications or email.
[1241] Users can view, download, or share the output generated through their device.
[1242] Specific example
[1243] Users record a voice message on their device to describe the details of the products they want in a physical store and upload photos of the store. The voice data is converted to text using the Google Speech Recognition API, and then the photo data undergoes emotion recognition using the DeepFace library. Based on this information, a generative AI model summarizes and provides detailed product information in a friendly tone.
[1244] Example of a prompt:
[1245] Convert the audio file "customer_request.wav" recorded by the customer in the store into text and summarize the text. Also, recognize the customer's emotions from the photo "customer_photo.jpg" and generate a response in a tone appropriate to those emotions.
[1246] This will enable elderly users and those with dementia to communicate their wishes accurately and emotionally to store staff.
[1247] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1248] Step 1:
[1249] Data entry and reception
[1250] The device receives text data entered by the user, audio data recorded, and photo data uploaded by the user. Users input this data using a smartphone or tablet via a dedicated application. The device temporarily stores the received data in local storage and sends it to a cloud server via the internet. Input is performed by the user using input forms, recording buttons, and photo upload functions, and the output in these cases is the data sent to the cloud server.
[1251] Step 2:
[1252] Data storage and preprocessing
[1253] The server receives text, audio, and photo data sent from the terminal and saves it to cloud storage. Simultaneously with saving, it begins data preprocessing. Specifically, audio data is converted to text format using the Google Speech Recognition API. This process converts the audio data into text data, which is then passed on to the next processing step. For photo data, identifiable metadata is extracted.
[1254] Step 3:
[1255] emotion recognition
[1256] The server performs emotion recognition using pre-processed data. It uses the DeepFace library to recognize the user's emotions from photographic data. Given photographic data as input, the emotion recognition engine extracts the user's emotional state (e.g., joy, sadness, surprise) as depicted in the photograph. Emotional information is obtained as output. This prepares the user's emotional information for delivery to the generating AI model.
[1257] Step 4:
[1258] Text summarization and tone adjustment
[1259] The server uses a generative AI model based on emotional information to summarize and tone-adjust text data. In particular, it uses the Hugging Face transformers library to summarize text. The input is pre-processed text and emotional information, and the output text is generated with the summary result and tone adjusted according to the emotion. Here, the tone is adjusted to be friendly when the emotion is joyful, and gentle when the emotion is sad.
[1260] Step 5:
[1261] Output generation and delivery
[1262] The generated text or other specified output formats (e.g., video messages) are delivered to the user from the server. The user is notified when the output is ready via in-app notifications or email. The user can then view, download, or share the generated output through their device. For example, the AI generates detailed information about the product the user wants to purchase, summarized in a friendly tone, making it easy for the user to communicate their intentions.
[1263] Through all these processing steps, the system enables elderly users and those with advanced dementia who have difficulty communicating to smoothly convey their intentions in physical stores.
[1264] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1265] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1266] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1267] [Fourth Embodiment]
[1268] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1269] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1270] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1271] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1272] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1273] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1274] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1275] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1276] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1277] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1278] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1279] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1280] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1281] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and specified format. The specific operation method and program processing of this system will be described below.
[1282] 1. Data Input
[1283] Terminal-side processing
[1284] The user accesses a provided input form using their device. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user might enter the contents of a will in text format, record a voice message about memories, and upload related photos.
[1285] After a user enters or uploads data, the device temporarily stores this data and sends it to the server. This makes the entered data available within the system.
[1286] 2. Summarizing and Transformation
[1287] Server-side processing
[1288] The server receives data sent from the terminal and stores it in storage. The stored data is then prepared to be passed on to the generating AI later. Specifically, this involves preprocessing of text data, transcribing audio data into text, and extracting metadata from photo data.
[1289] The server passes the stored data to the generating AI, instructing it to summarize the data and convert it into a specified format. The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video message for a funeral eulogy.
[1290] 3. Output generation and provision
[1291] Server-side processing
[1292] The server receives the summary and output in the specified format created by the generation AI and saves it for each user. Next, it converts this output into a format that the user can view, download, or share. For example, a will video might be saved as an MP4 file.
[1293] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[1294] Terminal-side processing
[1295] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly communicated.
[1296] To illustrate with a concrete example, consider a scenario where a user uploads a memorable photo and records an audio recording of the memory associated with that photo. The device sends the audio data and photo data to a server, which then passes them to a generative AI. The generative AI extracts metadata from the photo, transcribes the audio data into text, and summarizes its content. Finally, an output combining the summarized text and the photo is generated and provided to the user.
[1297] In this way, this system accurately records the user's thoughts and provides them in the appropriate format when needed, thereby solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[1298] The following describes the processing flow.
[1299] Step 1:
[1300] The device displays an input form to the user. The form includes a text box, a voice recording button, and a photo upload button.
[1301] Step 2:
[1302] Users input text data, record voice messages, and upload photos.
[1303] Step 3:
[1304] The device temporarily stores data entered by the user. This temporary storage uses the device's memory.
[1305] Step 4:
[1306] The device sends the temporarily stored data to the server. Once the transmission is complete, the device displays a confirmation message.
[1307] Step 5:
[1308] The server receives data sent from the terminal and saves it to storage. The storage location is managed individually for each user.
[1309] Step 6:
[1310] The server preprocesses the stored data. Specifically, this includes verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[1311] Step 7:
[1312] The server passes the pre-processed data to the generating AI, instructing it to summarize and convert it into a specified format.
[1313] Step 8:
[1314] The generation AI analyzes the received data and generates summaries and outputs in a specified format. For example, it can convert audio data into a text-based will or create a message based on photographs.
[1315] Step 9:
[1316] The server receives the generated output and saves it to storage. The saved output is managed on a per-user basis.
[1317] Step 10:
[1318] The server converts the saved output into a format that users can view, download, or share.
[1319] Step 11:
[1320] The server notifies the user when the output is ready. This notification is sent via email or an in-app notification.
[1321] Step 12:
[1322] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output.
[1323] Step 13:
[1324] Users can view the output through their device and download or share it as needed. For example, they can play the generated will video and share it with their family.
[1325] Through these steps, the system enables users to record memories and important messages in the appropriate format and provide them at the appropriate time.
[1326] (Example 1)
[1327] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1328] For elderly users or those with advanced dementia who have difficulty communicating their wishes effectively, efficiently recording data such as text, audio, and photographs, summarizing it, and converting it to a specified format for delivery presents a significant challenge. Solving this problem requires a system that handles everything from data input to output delivery, but current technology lacks a suitable solution.
[1329] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1330] In this invention, the server includes means for temporarily storing and transmitting user-inputted text, voice, and photo data to the server; means for pre-processing the data stored on the server; means for passing the pre-processed data to a generation AI model for summarization and conversion into a specified format; means for storing the generated output for each user; and means for notifying the user of the stored output and providing it in a format that can be viewed, downloaded, or shared. This makes it possible for elderly users or users with advanced dementia who have difficulty communicating their wishes to accurately and concisely record their wishes and provide them in an appropriate format when needed.
[1331] "Text data" refers to the text information that a user enters into an input form.
[1332] "Audio data" refers to voice messages recorded by the user using the voice recording button.
[1333] "Photo data" refers to image files that users upload using the photo upload button.
[1334] "Temporary storage" refers to the process of temporarily storing data before it is sent from the device to the server.
[1335] A "server" is a central system that receives, stores, and processes data sent from terminals.
[1336] "Preprocessing" refers to the process by which a server formats received data for analysis and transformation. Specifically, this includes text cleaning, audio-to-text conversion, and image metadata extraction.
[1337] A "generative AI model" is an artificial intelligence system that summarizes received data and converts it into a specified format.
[1338] "Output" refers to the output data created by the generative AI model, which is summarized and formatted in a specified format.
[1339] "Storage" refers to the process where the server identifies and stores the processed output for each user.
[1340] "Notification" refers to the server informing the user that the output is ready.
[1341] "Viewing" means making the generated output available for the user to review.
[1342] "Downloading" refers to the act of a user saving the generated output to their own device.
[1343] "Sharing" refers to the act of a user sharing the output they have generated with others.
[1344] "Metadata" refers to additional information included in photo data, such as the date and time of shooting, location, and camera settings.
[1345] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and converted format. This system receives user input data, processes it on a server, and provides output that is summarized and converted by a generating AI model.
[1346] Specifically, this is achieved using terminals, servers, and generative AI models as follows:
[1347] Hardware and software to be used
[1348] terminal
[1349] Hardware: Smartphone, tablet, or PC
[1350] Software: Dedicated application or form on a web browser
[1351] server
[1352] Hardware: High-performance cloud servers
[1353] Software: Databases, storage systems, analysis libraries (e.g., OpenCV, NLP libraries)
[1354] Generative AI Models
[1355] Software: Natural language processing algorithms, speech recognition engines (e.g., Google Speech-to-Text API)
[1356] Specific operating methods for the system
[1357] 1. Data Input
[1358] Terminal-side processing
[1359] The user launches a dedicated application on their device and accesses an input form displayed on the main screen. This form includes a text box, a voice recording button, and a photo upload button. The user enters the contents of their will in text format, records a voice message about their memories, and uploads related photos. After entering, recording, and uploading the data, the device temporarily stores the data and sends it to the server.
[1360] 2. Summarizing and Transformation
[1361] Server-side processing
[1362] The server receives data sent from the terminal and stores it in storage. After storage, the server preprocesses the data. Specifically, it cleans text data, transcribes audio data, and extracts metadata from photo data. Then, it passes the preprocessed data to a generative AI model and instructs it to summarize and transform it into a specified format. The generative AI model uses natural language processing algorithms to summarize and transform the data.
[1363] 3. Output generation and provision
[1364] Server-side processing
[1365] The system receives the summary and output in the specified format generated by the generative AI model and saves it for each user. Next, it converts this output into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file. The user is then notified when this information is ready. Notification methods include email and in-app notifications.
[1366] Terminal-side processing
[1367] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user reviews the generated output and downloads or shares it as needed.
[1368] Specific example
[1369] Consider a scenario where a user uploads a memorable photo and records an audio recording of the associated memories. In this case, the device sends the audio data and photo data to a server. The server passes this data to a generative AI model, which extracts the metadata from the photo and transcribes the audio data into text to summarize its content. Finally, an output combining the summarized text and the photo is generated and provided to the user.
[1370] Example of a prompt
[1371] Text Summary: "Please summarize this text: 'A message to my family: Thank you for everything. I look forward to your continued support.'"
[1372] Speech-to-text: "Please generate text from this audio. The audio file says, 'This is an important message to my family. This photo is from a trip we took together.'"
[1373] In this way, the system helps users effectively communicate their intentions by accurately understanding them and converting them into an appropriate format.
[1374] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1375] Step 1:
[1376] Displaying the input form
[1377] The user launches the application on their device and accesses the input form displayed on the main screen. The input form includes a text box, a voice recording button, and a photo upload button. Specifically, the device renders the input form and prompts the user for input. Input can be in the form of text, audio, or an image file.
[1378] Step 2:
[1379] Data entry, recording, and uploading.
[1380] The user enters text into a form, taps the voice recording button to record a voice message, and uses the photo upload button to select and upload a photo. Specifically, the device temporarily saves the entered text, the recorded voice data, and the uploaded photo data. As output, the device retains the saved data.
[1381] Step 3:
[1382] Sending data
[1383] After the user completes input, recording, and uploading, they tap the send button. The device then sends this data to the server. Specifically, the device sends the temporarily stored data to the server via an HTTP request. The input consists of temporarily stored text, audio, and image data, while the output is the transmission of data to the server.
[1384] Step 4:
[1385] Receiving and storing data
[1386] The server receives data sent from the terminal and stores it in storage. Specifically, the server stores the received data in a database and saves audio and images to file storage. The input is the data sent from the terminal, and the output is the data stored in storage.
[1387] Step 5:
[1388] Data preprocessing
[1389] The server performs preprocessing on the stored data. Specifically, the server uses an NLP (Natural Language Processing) library to clean text, a speech recognition engine (e.g., Google Speech-to-Text API) to convert speech to text, and an image analysis tool (e.g., OpenCV) to extract metadata from photos. The input consists of stored text, audio, and image data, and the output is preprocessed data.
[1390] Step 6:
[1391] Sending data to the generative AI model
[1392] The server passes pre-processed data to the generative AI model and instructs it to summarize and transform it into a specified format. Specifically, the server sends API requests to the generative AI model, inputting prompts for summarization and transformation. The input is pre-processed data, and the output is the output generated by the generative AI model.
[1393] Step 7:
[1394] Receiving and saving the generated output
[1395] The server receives the output from the generative AI model and saves it for each user. Specifically, the server saves the generated output, such as text and video, to a database. The input is the output from the generative AI model, and the output is the saved output.
[1396] Step 8:
[1397] Output format conversion
[1398] The server converts the generated output into a format that users can view, download, or share. Specifically, the server uses a file conversion tool to generate PDF or MP4 files. The input is the saved output, and the output is a file in a format usable by the user.
[1399] Step 9:
[1400] Notification to the user
[1401] The server notifies the user when the output is ready. Specifically, the server calls a notification API and sends an email or push notification to the user. The input is the output after format conversion, and the output is the notification to the user.
[1402] Step 10:
[1403] Receiving and displaying notifications
[1404] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output. Specifically, the terminal checks the notification and renders a dedicated confirmation screen. The input is the notification from the server, and the output is the confirmation screen displayed to the user.
[1405] Step 11:
[1406] Check the output and take action
[1407] The user reviews the generated output and downloads or shares it as needed. Specifically, the device displays an output viewing screen and provides download links and share buttons. The input is the output displayed on the confirmation screen, and the output is the download or sharing action performed by the user.
[1408] (Application Example 1)
[1409] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1410] Currently, there is no system that allows elderly users or those with advanced dementia, who may have difficulty communicating their wishes effectively, to easily and accurately order food delivery. In particular, there is a need for systems that understand the user's ordering intent using voice and photo data and provide appropriate suggestions based on that understanding. However, conventional systems struggle to meet these requirements. This creates a problem where elderly users and those with dementia find it difficult to easily use food delivery services.
[1411] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1412] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI to summarize and convert it into a specified format; means for providing the generated output to the user; means for converting voice data into text and extracting metadata from image data to generate food suggestions; and means for presenting the generated food suggestions to the user. This makes it possible for elderly users or users with advanced dementia who have difficulty communicating their wishes to use food delivery services easily and accurately.
[1413] "Users" refer to elderly people or individuals with dementia who have difficulty communicating their wishes appropriately and who use this system.
[1414] "Text data" refers to information in text format that is entered into the device used by the user.
[1415] "Audio data" refers to information in audio format provided by the user through voice input.
[1416] "Photo data" refers to image data uploaded by users via their devices.
[1417] "Means of receiving" refers to devices or programs that have the function of receiving text, audio, and photographic data provided by the user.
[1418] "Generative AI" refers to artificial intelligence models used to summarize received data and convert it into a specified format.
[1419] "Means for summarizing and converting to a specified format" refers to a processing mechanism that passes the received data to a generating AI, summarizes the data concisely, and converts it to the required format.
[1420] "Means of providing output" refers to devices or programs that have the function of presenting output data generated by the generation AI to the user in an appropriate format.
[1421] "Converting audio data to text" refers to the process of analyzing audio data and converting it into corresponding text data.
[1422] "Metadata" refers to information associated with photographic data, such as the date and time of shooting, location, and the type of object depicted.
[1423] "Food suggestions" refers to dishes and foods recommended by the generating AI based on the user's input data.
[1424] "Means for generating suggestions" refers to a technical system for creating suggestions tailored to the user based on the received data.
[1425] "Means of presenting proposals" refers to devices or programs that have the function of visually or audibly displaying the generated proposals to the user.
[1426] The system for realizing this invention is designed to allow users to easily utilize food delivery even if they are elderly or have difficulty communicating their wishes due to the progression of dementia. The system consists of several main components, including a terminal, a server, and a generative AI model.
[1427] Overview of program processing
[1428] 1. Processing on the terminal side
[1429] The terminal receives text, voice, and photo data entered by the user. Through an application installed on the terminal, the user can voice-input descriptions of desired dishes or favorite dishes, and upload photos of previously ordered dishes. The input data is temporarily stored on the terminal and then sent to the server.
[1430] 2. Server-side processing
[1431] The server receives data sent from the terminal and performs data preprocessing. Specifically, it converts audio data into text format (speech recognition) and extracts metadata from photo data (image analysis). After this, the data is passed to a generative AI model to generate a data summary and food suggestions. The generative AI model uses natural language processing algorithms and image recognition algorithms.
[1432] 3. Output generation and provision
[1433] The system receives suggestions generated by a generative AI model and generates output tailored to each user's requirements. A list of suggested foods is presented to the user through the application. The user can then select from this list and complete their order.
[1434] Specific example
[1435] If a user voice-inputs "I want to eat spaghetti today" and uploads a photo of spaghetti they previously ordered, the system will process it as follows:
[1436] 1. The device records the user's voice saying "I want to eat spaghetti today," retrieves a photo of spaghetti ordered previously, and sends it to the server.
[1437] 2. The server converts the audio data into text format and extracts the metadata from the photos.
[1438] 3. The generative AI model analyzes text and photo data to suggest food suitable for the user.
[1439] 4. The user selects spaghetti from the suggested food list and completes the order.
[1440] Hardware and software to be used
[1441] 1. Microphone: Used to input the user's voice.
[1442] 2. Camera: Used to take photos that users upload.
[1443] 3. Terminal (smartphone / tablet): A device used by the user to input data and communicate with the server.
[1444] 4. Server: Responsible for storing data and running the generated AI model.
[1445] 5. Generative AI Models: Algorithms that combine speech recognition, natural language processing, and image recognition.
[1446] Example of a prompt
[1447] In the user's own voice: "What I want to eat right now is, for example, curry rice or ramen, or any of my favorite dishes."
[1448] Example photo: "Please upload a photo of the curry rice you ordered previously."
[1449] This allows the system to provide an easy-to-use food delivery ordering system for elderly and dementia users, making access to meals easier.
[1450] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1451] Step 1:
[1452] Users use their devices to voice-input the characteristics of the dishes they want to eat or their favorite dishes, and upload photos of dishes they have ordered in the past.
[1453] Input: Audio data, photo data
[1454] Operation: The user inputs audio using the application's voice recording function and uploads photos using the camera function.
[1455] Output: The input audio and photo data are temporarily stored on the device.
[1456] Step 2:
[1457] The device sends the voice and photo data entered by the user to the server.
[1458] Input: Temporarily stored audio data, photo data
[1459] Operation: The terminal sends data to the server via the network.
[1460] Output: Audio and photo data transferred to the server
[1461] Step 3:
[1462] The server converts the audio data to text and extracts metadata from the photo data.
[1463] Input: Audio data, photo data
[1464] Data processing: Convert audio data into text data using speech recognition technology. Extract metadata from photo data using image analysis technology.
[1465] Operation: The server uses a specified speech recognition algorithm (e.g., Google Speech Recognition). Image recognition APIs (e.g., AWS Rekognition) are used to extract metadata from photos.
[1466] Output: Text data, photo metadata
[1467] Step 4:
[1468] The server passes text data and photo metadata to a generating AI model to generate food suggestions.
[1469] Input: Text data, photo metadata
[1470] Data processing: A generative AI model analyzes text data and photo metadata to generate food suggestions.
[1471] Operation: The server uses a generative AI model (e.g., GPT-3 or BERT) to generate optimal food suggestions for the user.
[1472] Output: List of suggested foods generated
[1473] Step 5:
[1474] The server sends the generated food suggestions to the user's terminal.
[1475] Input: List of suggested foods generated
[1476] Operation: The server sends the suggestion list to the user's terminal via the network.
[1477] Output: List of food suggestions sent to the user's terminal
[1478] Step 6:
[1479] The terminal displays a list of suggested foods to the user. The user selects from the list and completes the order.
[1480] Input: Food suggestion list
[1481] Operation: The terminal displays a list of suggestions through the application's user interface, and the user selects and places an order.
[1482] Output: Order information for food selected by the user.
[1483] Step 7:
[1484] The server sends the user's order information to the food delivery service and confirms the order.
[1485] Input: Order information for food selected by the user
[1486] Operation: The server sends order information to the food delivery service's API.
[1487] Output: Confirmed order information for food delivery service
[1488] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1489] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides it in a summarized and specified format. By combining this system with an emotion engine, it can recognize the user's emotions and adjust the content and tone of the generated output. The specific operation method and program processing of this system are described below.
[1490] 1. Data Input
[1491] Terminal-side processing
[1492] The user accesses a provided input form using their device. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user might enter the contents of a will in text format, record a voice message about memories, and upload related photos.
[1493] After a user enters or uploads data, the device temporarily stores this data and sends it to the server. This makes the entered data available within the system.
[1494] 2. Summarizing and Transformation
[1495] Server-side processing
[1496] The server receives data sent from the terminal and stores it in storage. The stored data is then prepared to be passed on to the generating AI later. Specifically, this involves verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[1497] The server passes the stored data to the generating AI, instructing it to summarize the data and convert it into a specified format. The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video message for a funeral eulogy.
[1498] 3. Emotion Recognition and Regulation
[1499] Emotional Engine Processing
[1500] Before the data is passed to the generative AI, the server passes the audio and photo data to the emotion engine for emotion recognition. For example, the emotion engine analyzes the user's emotional state from the audio data and provides the results to the generative AI. Similarly, for photo data, the emotion engine recognizes the user's emotions when analyzing the metadata.
[1501] The emotional information recognized by the emotion engine is used as a guideline for the generating AI to adjust the output content and tone. For example, if the content of a will is very moving, the tone of the text and the atmosphere of the video will be adjusted to emphasize that.
[1502] 4. Output generation and delivery
[1503] Server-side processing
[1504] The server receives the summary and output in the specified format created by the generation AI and saves it to storage. The saved output is managed on a per-user basis. Next, this output is converted into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file.
[1505] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[1506] Terminal-side processing
[1507] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly communicated.
[1508] As a concrete example, consider a scenario where a user uploads a memorable photo and records an emotionally charged audio recording of the memories associated with that photo. The device sends the audio and photo data to a server, which then passes this data to an emotion engine and a generative AI. The generative AI adjusts the output based on the photo's metadata and emotional information, ultimately generating an emotionally resonant output that combines summarized text and the photo, which is then delivered to the user.
[1509] In this way, this system accurately records the user's thoughts and provides them at the appropriate time in a form that reflects their emotions, effectively solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[1510] The following describes the processing flow.
[1511] Step 1:
[1512] The device displays an input form to the user. The form includes a text box, a voice recording button, and a photo upload button.
[1513] Step 2:
[1514] Users input text data, record voice messages, and upload photos. For example, they might enter the contents of their will in text format, record an emotionally charged voice message about a memory, and upload related photos.
[1515] Step 3:
[1516] The device temporarily stores data entered by the user. This temporary storage uses the device's memory.
[1517] Step 4:
[1518] The device sends the temporarily stored data to the server. If the transmission is successful, the device displays a confirmation message.
[1519] Step 5:
[1520] The server receives data sent from the terminal and saves it to storage. The storage location is managed individually for each user.
[1521] Step 6:
[1522] The server preprocesses the stored data. Specifically, this includes verifying text data, transcribing audio data into text, and extracting metadata from photo data.
[1523] Step 7:
[1524] The server passes the pre-processed data to the emotion engine and instructs it to analyze the user's emotions.
[1525] Step 8:
[1526] The emotion engine recognizes the user's emotions from voice and photo data and provides the results to the generating AI. For example, the emotion engine analyzes the user's emotional state from voice data and identifies emotions such as "compassion," "sadness," and "joy."
[1527] Step 9:
[1528] The server passes the emotional information obtained from the emotion engine and pre-processed data to the generating AI, instructing it to summarize and convert it into a specified format.
[1529] Step 10:
[1530] The generative AI analyzes data and generates summaries and output in a specified format that reflect emotional information. For example, if the content of a will is very moving, it will adjust the tone of the text and the atmosphere of the video to emphasize that.
[1531] Step 11:
[1532] The server receives the generated output and saves it to storage. The saved output is managed on a per-user basis.
[1533] Step 12:
[1534] The server converts the output into a format that users can view, download, or share. For example, a will video reflecting emotions might be saved as an MP4 file.
[1535] Step 13:
[1536] The server notifies the user when the output is ready. Notification methods include email and in-app notifications.
[1537] Step 14:
[1538] The terminal receives a notification from the server and displays a screen prompting the user to confirm the output.
[1539] Step 15:
[1540] Users can view the output through their device and download or share it as needed. For example, they can play the generated will video and share it with their family.
[1541] Through the steps described above, this system can analyze user-input data using an emotion engine and generative AI, record it in a way that reflects emotions, and provide it at the appropriate time.
[1542] (Example 2)
[1543] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1544] For elderly users or those with advanced dementia who have difficulty communicating their wishes, there is a need to efficiently record data such as memories and wills in a way that reflects their emotions, and provide it as appropriate output. However, conventional systems cannot recognize and reflect the user's emotions, and the content and tone of the generated output may not match the user's wishes. To solve this problem, a system is needed that can recognize the user's emotions and adjust the content and tone of the generated output.
[1545] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1546] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI to summarize and convert it into a specified format; means for providing the generated output to the user; means for recognizing emotions from the voice and photo data using an emotion engine; means for adjusting the content and tone of the output based on the recognized emotion information; and means for saving the final output to storage and notifying the user. This enables the generation of output that reflects the user's emotions, allowing them to appropriately convey their intentions and memories.
[1547] "Text data" refers to text information entered by users through input forms or similar means.
[1548] "Audio data" refers to voice messages entered by the user using the recording function.
[1549] "Photo data" refers to still image files uploaded by users.
[1550] "Means of receiving" refers to the function of acquiring text data, audio data, and photo data entered by the user from the terminal to the server.
[1551] "Generative AI" refers to artificial intelligence models used to summarize and convert data received from users into a specified format.
[1552] A "summarizing method" is a function that processes received data to summarize it and concisely organize the information.
[1553] "Means of converting to a specified format" refers to a function that converts summarized data into a format (text, video, etc.) that is easy for users to view.
[1554] "Means of provision" refers to functions that notify users of the output created by the generative AI, and enable them to view, download, and share it.
[1555] The "emotion engine" is a processing system that analyzes a user's emotions from audio and photo data and provides the results to the generating AI.
[1556] "Means of recognition" refers to a function that uses an emotion engine to identify the user's emotions from audio and photo data.
[1557] "Means of adjustment" refers to a function that allows the generating AI to change the content and tone of its output based on emotional information.
[1558] "Storage" refers to a location where generated output is saved and made accessible to users as needed.
[1559] "Notification means" refers to a notification function that informs the user when the server has finished preparing the generated output.
[1560] Modes for carrying out the invention
[1561] This invention provides a system for users who have difficulty communicating their wishes appropriately due to old age or the progression of dementia. It efficiently records data such as text, audio, and photographs, and provides them in a summarized and specified format. By combining this system with an emotion engine, it can recognize the user's emotions and adjust the content and tone of the generated output.
[1562] 1. Data Input
[1563] Terminal-side processing
[1564] The user accesses an input form provided by the system using their device. This input form includes a text box, a voice recording button, and a photo upload button. Through this form, the user can, for example, enter the contents of a will in text format, record a voice message about memories, and upload related photos. This allows the user to express their wishes in a variety of ways.
[1565] After data entry, the terminal temporarily stores this data and sends it to the server. This transmission process makes the entered data available within the system.
[1566] 2. Summarizing and Transformation
[1567] Server-side processing
[1568] The server receives data sent from the terminal and stores it in storage. The received data is processed as follows:
[1569] The text data undergoes grammatical and consistency checks.
[1570] The audio data is converted to text using speech recognition software (e.g., Google Cloud Speech-to-Text).
[1571] Metadata (such as the date and time of shooting, and location) is extracted from the photo data.
[1572] The data, after the above processing is complete, is passed to a generating AI (e.g., a GPT model) for summarization and conversion to a specified format. The generating AI uses natural language processing algorithms to perform tasks such as summarizing the contents of a will into a concise text format or generating a video message for a funeral eulogy.
[1573] 3. Emotion Recognition and Regulation
[1574] Emotional Engine Processing
[1575] Before the data is passed to the generative AI, the server passes the audio and photo data to an emotion engine (e.g., Microsoft Azure Cognitive Services) to recognize emotions. For example, the emotion engine analyzes the user's emotional state (joy, sadness, anger, etc.) from the audio data and provides the result to the generative AI. Similarly, for photo data, the emotion engine recognizes the user's emotions by analyzing the metadata.
[1576] The emotional information recognized by the emotion engine is used as a guideline for the generating AI to adjust the output content and tone. For example, if the content of a will is emotional, the tone of the text and the atmosphere of the video will be adjusted to emphasize that emotion.
[1577] 4. Output generation and provision
[1578] Server-side processing
[1579] The server receives the summary and output in the specified format created by the generation AI and saves it to storage. The saved output is managed on a per-user basis. Next, the server converts the output into a format that the user can view, download, or share. For example, a will video is saved as an MP4 file. The server then notifies the user when the output is ready. Notification methods include email notifications and in-app notifications.
[1580] Terminal-side processing
[1581] The device receives a notification from the server and displays a screen prompting the user to review the output. Through this screen, the user can review the generated output and download or share it as needed. For example, by playing the generated will video and sharing it with family, one can ensure their wishes are properly conveyed.
[1582] Specific example
[1583] As a concrete example, consider a scenario where a user uploads photos of travel memories and records an emotionally moving story related to those photos in audio. The device sends the photos and audio data to a server, which then passes them to an emotion engine and a generative AI. The emotion engine extracts emotional information from the audio, and the generative AI uses this to create an emotionally moving output. Finally, the user is provided with an output that combines summarized text and photos.
[1584] Example of a prompt
[1585] An example of a prompt to input into a generative AI model is, "Based on the audio message and photos provided by the user about past memories, please generate a summary text that evokes emotion." Using such prompts, the generative AI can produce output that accurately reflects the user's intentions.
[1586] As described above, this system accurately records the user's thoughts and provides output that reflects their emotions, thereby effectively solving the problem of elderly people and those with progressing dementia who have difficulty communicating their wishes.
[1587] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1588] Step 1:
[1589] User enters data
[1590] The user uses a device to access the system's input form. The input form includes a text box, a voice recording button, and a photo upload button. For example, the user enters the contents of a will in text format, records an audio message about memories, and uploads related photos (input: text, audio, and photo data). The device temporarily stores this data (output: data with temporarily stored content).
[1591] Step 2:
[1592] The device temporarily stores and transmits data.
[1593] After the user completes data entry, the device sends the temporarily stored data to the server (Input: Temporarily stored text, audio, and photo data). This process ensures that the data is sent to the server securely and efficiently (Output: Data sent to the server).
[1594] Step 3:
[1595] The server saves the data to the receiving storage.
[1596] The server receives data sent from the terminal and stores it in storage (Input: Data sent from the terminal). The stored data is then ready for processing and transformation (Output: Data stored in storage).
[1597] Step 4:
[1598] The server performs proofreading and preprocessing.
[1599] The server classifies the stored data and performs the following actions on each:
[1600] Text data validation (checking grammar and consistency)
[1601] Converting audio data to text (using speech recognition software to convert speech to text)
[1602] Metadata extraction from photo data (extracting information such as date, time, and location of shooting)
[1603] (Input: Saved text, audio, and photo data). The data is then ready to proceed to the next processing step (Output: Preprocessed data).
[1604] Step 5:
[1605] The server passes the data to the generating AI and instructs it on summarization and transformation.
[1606] The server passes the pre-processed data to the generating AI, requesting it to summarize and convert it to a specified format (Input: Pre-processed data). The generating AI uses natural language processing algorithms to, for example, summarize the contents of a will into a concise text format or generate a video for a funeral eulogy (Output: Summarized and converted output).
[1607] Step 6:
[1608] The server passes the data to the emotion engine to perform emotion recognition.
[1609] Before the data is passed to the generating AI, the server passes the audio and photo data to the emotion engine to recognize emotions (input: audio data, photo data). The emotion engine analyzes the user's emotional state from the audio data and also analyzes the metadata of the photo data to recognize emotions (output: recognized emotion information).
[1610] Step 7:
[1611] The AI generates the output based on emotional information, adjusting the content and tone.
[1612] The server provides the recognized emotional information to the generating AI, which then adjusts the content and tone of the output based on that information (input: recognized emotional information). For example, it sets the tone for the text or video of an emotional will (output: output reflecting the emotion).
[1613] Step 8:
[1614] The server saves the final output and notifies the user.
[1615] The server receives the final output created by the generation AI and saves it to storage (Input: Final Output). Then, it notifies the user that the output is ready (Output: Saved Output, Notification to User).
[1616] Step 9:
[1617] The device prompts for output confirmation and supports sharing.
[1618] The device receives a notification from the server and displays a screen prompting the user to review the output (Input: Notification). The user reviews the generated output through this screen and downloads or shares it as needed (Output: Reviewed output, downloaded or shared).
[1619] In this way, the system can effectively manage user data and provide output that reflects emotions.
[1620] (Application Example 2)
[1621] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1622] With an increasing number of elderly users and those with advanced dementia who have difficulty communicating their wishes effectively, there is a problem in that these users have difficulty communicating their needs regarding the products and services they are looking for in physical stores. This also hinders smooth communication with store staff, resulting in a decline in the quality of the user experience. To solve this, a system is needed that can accurately convey the user's wishes in a way that reflects their emotions.
[1623] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1624] In this invention, the server includes means for receiving text, voice, and photo data entered by the user; means for passing the received data to a generating AI for summarization and conversion into a specified format; means for recognizing the user's emotions from the received data; means for adjusting the content and tone of the output based on the emotion recognition results; and means for providing the generated output to the user. This enables users to communicate their intentions accurately and emotionally in a physical store setting.
[1625] A "user" is a person who uses the system to input data, and in particular, in physical stores, this includes elderly people who have difficulty communicating their intentions and users whose dementia is progressing.
[1626] "Text data" refers to information in text format entered by users, including requests and questions about products and services.
[1627] "Audio data" refers to audio information recorded by users, used to convey opinions and requests regarding products and services.
[1628] "Photo data" refers to image information uploaded by users to the system, including visual data related to products and services.
[1629] "Generative AI" refers to artificial intelligence (AI) technology used to summarize and convert received data into a specified format, employing natural language processing and data generation models.
[1630] An "emotion engine" refers to a technology that recognizes a user's emotions from audio and photo data, analyzing emotional information and adjusting the output accordingly.
[1631] "Output" refers to summaries or results in a specified format created by the generating AI based on data entered by the user, and is provided in formats such as text, video, and audio.
[1632] "Emotion recognition" is the process of using an emotion engine to analyze a user's emotional state from audio or photos and providing that information to a generating AI.
[1633] "Tone" refers to the emotion and atmosphere in the content and format of the output, and is adjusted according to the user's emotional state.
[1634] System Configuration
[1635] This system consists of a terminal that receives text data, voice data, and photo data entered by the user, a server that processes and analyzes the data, and a means of providing the generated output. The main components are as follows:
[1636] 1. Device: Users input text data, audio data, and photo data through a device such as a smartphone or tablet. This device is connected to the internet and has network communication capabilities to send data to the server.
[1637] 2. Server: The server stores and analyzes the received data. Using a cloud server as a hosting service is recommended here. Programming languages such as Python and related libraries are used for data analysis.
[1638] 3. Generative AI: Uses natural language processing and data generation models (e.g., Hugging Face's transformers library) to summarize user input and convert it to a specified format.
[1639] 4. Emotion Engine: Uses technologies (e.g., DeepFace library, Google Speech Recognition API) to recognize the user's emotional state from photo and audio data.
[1640] Processing flow
[1641] The server primarily processes data in the following steps:
[1642] 1. Receiving and storing data:
[1643] This function temporarily stores text data, audio data, and photo data entered by the device.
[1644] The data will be stored in cloud storage with security in mind.
[1645] 2. Speech recognition:
[1646] The server uses the Google Speech Recognition API to convert the audio data into text data.
[1647] This text data will be used for subsequent processing in the generative AI and emotion engine.
[1648] 3. Emotion recognition:
[1649] Using the DeepFace library, we recognize the user's emotions from photographic data and extract emotional information.
[1650] Emotional information is passed to the generating AI and used to adjust the tone of the output.
[1651] 4. Text summarization and conversion:
[1652] Use models from the transformers library to perform summarization and conversion to a specified format.
[1653] It incorporates emotional information into tone adjustments to generate appropriate text and video messages.
[1654] 5. Output generation and delivery:
[1655] The generated output is provided to the user. Specifically, notifications are sent via in-app notifications or email.
[1656] Users can view, download, or share the output generated through their device.
[1657] Specific example
[1658] Users record a voice message on their device to describe the details of the products they want in a physical store and upload photos of the store. The voice data is converted to text using the Google Speech Recognition API, and then the photo data undergoes emotion recognition using the DeepFace library. Based on this information, a generative AI model summarizes and provides detailed product information in a friendly tone.
[1659] Example of a prompt:
[1660] Convert the audio file "customer_request.wav" recorded by the customer in the store into text and summarize the text. Also, recognize the customer's emotions from the photo "customer_photo.jpg" and generate a response in a tone appropriate to those emotions.
[1661] This will enable elderly users and those with dementia to communicate their wishes accurately and emotionally to store staff.
[1662] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1663] Step 1:
[1664] Data entry and reception
[1665] The device receives text data entered by the user, audio data recorded, and photo data uploaded by the user. Users input this data using a smartphone or tablet via a dedicated application. The device temporarily stores the received data in local storage and sends it to a cloud server via the internet. Input is performed by the user using input forms, recording buttons, and photo upload functions, and the output in these cases is the data sent to the cloud server.
[1666] Step 2:
[1667] Data storage and preprocessing
[1668] The server receives text, audio, and photo data sent from the terminal and saves it to cloud storage. Simultaneously with saving, it begins data preprocessing. Specifically, audio data is converted to text format using the Google Speech Recognition API. This process converts the audio data into text data, which is then passed on to the next processing step. For photo data, identifiable metadata is extracted.
[1669] Step 3:
[1670] emotion recognition
[1671] The server performs emotion recognition using pre-processed data. It uses the DeepFace library to recognize the user's emotions from photographic data. Given photographic data as input, the emotion recognition engine extracts the user's emotional state (e.g., joy, sadness, surprise) as depicted in the photograph. Emotional information is obtained as output. This prepares the user's emotional information for delivery to the generating AI model.
[1672] Step 4:
[1673] Text summarization and tone adjustment
[1674] The server uses a generative AI model based on emotional information to summarize and tone-adjust text data. In particular, it uses the Hugging Face transformers library to summarize text. The input is pre-processed text and emotional information, and the output text is generated with the summary result and tone adjusted according to the emotion. Here, the tone is adjusted to be friendly when the emotion is joyful, and gentle when the emotion is sad.
[1675] Step 5:
[1676] Output generation and delivery
[1677] The generated text or other specified output formats (e.g., video messages) are delivered to the user from the server. The user is notified when the output is ready via in-app notifications or email. The user can then view, download, or share the generated output through their device. For example, the AI generates detailed information about the product the user wants to purchase, summarized in a friendly tone, making it easy for the user to communicate their intentions.
[1678] Through all these processing steps, the system enables elderly users and those with advanced dementia who have difficulty communicating to smoothly convey their intentions in physical stores.
[1679] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1680] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1681] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1682] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1683] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1684] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1685] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1686] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1687] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1688] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1689] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1690] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1691] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1692] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1693] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1694] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1695] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1696] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1697] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1698] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1699] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1700] The following is further disclosed regarding the embodiments described above.
[1701] (Claim 1)
[1702] A means of receiving text, audio, and photo data entered by the user,
[1703] A means of passing the received data to a generating AI to summarize and convert it into a specified format,
[1704] A means of providing the generated output to the user,
[1705] A system that includes this.
[1706] (Claim 2)
[1707] The system according to claim 1, wherein the generating AI converts the received audio data into text format.
[1708] (Claim 3)
[1709] The system according to claim 1, wherein the generating AI extracts metadata from received photo data, summarizes its contents, and includes them.
[1710] "Example 1"
[1711] (Claim 1)
[1712] A means of receiving text, audio, and photo data entered by the user,
[1713] A means of temporarily storing received data and sending it to a server,
[1714] A means for preprocessing data stored on a server,
[1715] A means of passing preprocessed data to a generating AI model to summarize and convert it into a specified format,
[1716] A means of saving the generated output for each user,
[1717] A means of notifying the user of the saved output and providing it in a format that can be viewed, downloaded, or shared,
[1718] A system that includes this.
[1719] (Claim 2)
[1720] The system according to claim 1, wherein a generative AI model converts received audio data into text format.
[1721] (Claim 3)
[1722] The system according to claim 1, wherein a generative AI model extracts metadata from received photo data, summarizes its contents, and includes them.
[1723] "Application Example 1"
[1724] (Claim 1)
[1725] A means of receiving text, audio, and photo data entered by the user,
[1726] A means of passing the received data to a generating AI to summarize and convert it into a specified format,
[1727] A means of providing the generated output to the user,
[1728] A means of converting audio data to text, extracting metadata from image data, and generating food suggestions,
[1729] A means of presenting the generated food suggestions to the user,
[1730] A system that includes this.
[1731] (Claim 2)
[1732] The system according to claim 1, wherein the generating AI converts the received audio data into text format.
[1733] (Claim 3)
[1734] The system according to claim 1, wherein the generating AI extracts metadata from received photo data, summarizes its contents, and includes them.
[1735] "Example 2 of combining an emotion engine"
[1736] (Claim 1)
[1737] A means of receiving text, audio, and photo data entered by the user,
[1738] A means of passing the received data to a generating AI to summarize and convert it into a specified format,
[1739] A means of providing the generated output to the user,
[1740] A means of recognizing emotions from audio and photographic data using an emotion engine,
[1741] A means of adjusting the content and tone of the output based on recognized emotional information,
[1742] A means of saving the final output to storage and notifying the user,
[1743] A system that includes this.
[1744] (Claim 2)
[1745] The system according to claim 1, wherein the generating AI converts the received audio data into text format.
[1746] (Claim 3)
[1747] The system according to claim 1, wherein the generating AI extracts metadata from received photo data, summarizes its contents, and includes them.
[1748] "Application example 2 when combining with an emotional engine"
[1749] (Claim 1)
[1750] A means of receiving text, audio, and photo data entered by the user,
[1751] A means of passing the received data to a generating AI to summarize and convert it into a specified format,
[1752] A means of recognizing the user's emotions from the received data,
[1753] A means of adjusting the content and tone of the output based on the results of emotion recognition,
[1754] A means of providing the generated output to the user,
[1755] A system that includes this.
[1756] (Claim 2)
[1757] The system according to claim 1, wherein the generating AI converts the received audio data into text format.
[1758] (Claim 3)
[1759] The system according to claim 1, wherein the generating AI extracts metadata from received photo data, summarizes its contents, and includes them. [Explanation of symbols]
[1760] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of receiving text, audio, and photo data entered by the user, A means of passing the received data to a generating AI to summarize and convert it into a specified format, A means of providing the generated output to the user, A system that includes this.
2. The system according to claim 1, wherein the generating AI converts the received audio data into text format.
3. The system according to claim 1, wherein the generating AI extracts metadata from the received photo data, summarizes its contents, and includes them.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A