System
A system using generative AI to analyze user images and facilitate conversations helps lonely individuals reminisce, addressing the lack of brain stimulation and social connections by providing enjoyable memory recall.
Patent Information
- Application Number
- JP2024130434
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
In modern society, there is an increasing number of lonely elderly people and individuals with limited opportunities to reminisce about their memories, leading to a lack of brain stimulation and social connections, and a decline in self-esteem.
A system that allows users to upload images from their devices, where a server analyzes them using generative AI to generate questions and comments, facilitating a conversation with the user, and logs the interaction for later analysis.
The system stimulates the user's brain and maintains their energy by enabling enjoyable conversations through photos, thereby activating past memories and maintaining social connections.
Smart Images

Figure 2026028136000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, the number of lonely elderly people and those with few opportunities to reminisce about their memories in their daily lives is increasing. Even if these people have photos of their memories, they often leave them alone because looking at them alone is boring. Lonely elderly people in particular lack brain stimulation and energy, and tend to lose their social connections and self-esteem. Furthermore, people who take photos of food or travels have limited opportunities to fully enjoy the memories they have recorded. To solve these issues, a system is needed that allows users to activate past memories through photos and rediscover happy memories. [Means for solving the problem]
[0005] The present invention provides a system in which users upload images from their devices, and a server receives and analyzes the images. Based on the results of the image analysis, the server uses a generation AI to generate questions and comments for the user. The device then presents the generated questions and comments to the user, and when the user responds, the response is sent from the device to the server. The server analyzes the user's responses, and the generation AI generates the next question or comment, thereby continuing the conversation with the user. In addition, all conversation logs are saved and analyzed at a later date, making it possible to effectively analyze user usage and reactions. Through this system, users can enjoy conversations through photos, stimulating their brains and maintaining their energy.
[0006] "Users" refer to people who use the system to upload images and reminisce through interactions with the generative AI.
[0007] "Terminal" refers to a device operated by a user, which uploads images, receives questions and comments from the generation AI, and inputs the user's responses.
[0008] "Server" refers to a device that receives images from users, analyzes them, works with the generation AI to generate questions and comments, and performs processing to maintain a dialogue with users.
[0009] "Image" refers to photographs or other visual data that a User uploads to the System using a Terminal.
[0010] "Image analysis" refers to the process by which the server identifies the content of the image received and extracts information to generate appropriate questions or comments based on that content.
[0011] "Generative AI" refers to an artificial intelligence system that runs inside a server and generates appropriate questions and comments based on the results of image analysis, interacting with users.
[0012] "Questions and comments" refers to questions to the user generated by the generative AI and feedback on the responses.
[0013] "Conversation log" refers to data that records the content of the conversation between the user and the generating AI.
[0014] "Post-analysis" refers to the process of later analyzing saved conversation logs to evaluate user usage and reactions. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] This invention provides a system that stimulates the user's brain and maintains their energy by interacting with a generating AI using photos of the user's memories. This system involves a multi-stage process in which the user uploads photos and looks back on their memories through conversations with the generating AI.
[0037] System Overview
[0038] The process by which a user uploads an image from their device
[0039] Users upload photos to the system using their own devices (e.g., smartphones or tablets). Users can easily select and upload photos by opening the application and going to the upload screen. Uploaded photos are first transferred to the server.
[0040] The process by which the server receives and analyzes the images
[0041] The uploaded photos are received by the server and stored in a database, where they are then analyzed using image analysis algorithms to automatically identify people, objects, and background information in the photos.
[0042] The process by which generative AI generates questions and comments for users
[0043] Based on the results of image analysis, the server sends the necessary data to the generation AI. The generation AI then generates appropriate questions and comments for the user based on the content of the photo. For example, it creates questions such as "Where was this photo taken?" or "Who is this person?"
[0044] The process by which the device presents the generated questions and comments to the user
[0045] The questions and comments generated by the AI are sent back from the server to the device, which then displays them to the user. The user can then confirm the questions and comments and enter a response.
[0046] The process of sending user responses from the terminal to the server
[0047] The response entered by the user is sent from the device to the server, and may include, for example, "These are my daughter and grandchildren."
[0048] The server analyzes the response, and the AI generates the next question or comment.
[0049] The server receives and analyzes the user's response, and the generation AI generates the next question or comment. This continues the conversation with the user and advances the dialogue process. For example, the next question might be, "Where was this photo taken?"
[0050] Continuing conversations and logging processes
[0051] This conversation process is repeated until the user is satisfied. All conversation logs are stored on the server and used later to analyze user usage and reactions.
[0052] Specific examples
[0053] Example 1: Lonely elderly people
[0054] User: Selects and uploads an old family photo.
[0055] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[0056] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[0057] Terminal: Display the question to the user.
[0058] User: Responds to the question by saying, "These are my daughter and grandchildren."
[0059] Terminal: Sends the response to the server.
[0060] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[0061] Terminal: Display the next question to the user.
[0062] Example 2: Travel photos
[0063] User: Select and upload photos of travel destinations.
[0064] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[0065] Generative AI: Generates questions like "Where is this place?" based on a photo.
[0066] Terminal: Display the question to the user.
[0067] User: Responds to the question with "This is a photo from Paris, where I went last year."
[0068] Terminal: Sends the response to the server.
[0069] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[0070] Terminal: Display the next question to the user.
[0071] As described above, through this system, users can enjoy conversations through photos, which helps stimulate the brain and maintain energy.
[0072] The processing flow will be explained below.
[0073] Step 1:
[0074] The user opens the application on their device and proceeds to the photo upload screen.
[0075] Step 2:
[0076] The user selects a photo on their device and clicks the upload button.
[0077] Step 3:
[0078] The terminal generates a request to transfer the selected photo to the server.
[0079] Step 4:
[0080] The server receives the request from the terminal and receives the photo data.
[0081] Step 5:
[0082] The server stores the received photos in a database.
[0083] Step 6:
[0084] The server passes the stored photos to an image analysis algorithm, which begins the analysis process.
[0085] Step 7:
[0086] The server performs image analysis to determine the content of the photo, such as identifying people, objects, and the background.
[0087] Step 8:
[0088] The server sends the image analysis results to the generation AI.
[0089] Step 9:
[0090] Based on the results of image analysis, the generative AI generates appropriate questions and comments for the user, such as "Where was this photo taken?" or "Who is this person?"
[0091] Step 10:
[0092] The server sends the generated questions and comments back to the device.
[0093] Step 11:
[0094] The terminal displays the questions and comments received from the server to the user.
[0095] Step 12:
[0096] The user enters a response to the question or comment on the terminal.
[0097] Step 13:
[0098] The terminal generates a request to send the user's response to the server.
[0099] Step 14:
[0100] The server receives the response from the terminal.
[0101] Step 15:
[0102] The server analyzes the received response and sends it to the generating AI.
[0103] Step 16:
[0104] Based on the user's response, the generative AI generates the next appropriate question or comment, such as "Where was this photo taken?"
[0105] Step 17:
[0106] The server sends the next generated question or comment back to the device.
[0107] Step 18:
[0108] The device will display the next question or comment to the user.
[0109] Step 19:
[0110] The user again enters a response, and this process is repeated until the user is satisfied.
[0111] Step 20:
[0112] The server stores all conversation logs.
[0113] Step 21:
[0114] The server will later analyze the saved conversation logs to evaluate the user's usage and reactions.
[0115] Example 1
[0116] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0117] Today's elderly and lonely people lack emotional communication and brain activation. As a result, they face problems such as a decline in cognitive function and loss of energy. Conventional technologies have not provided effective solutions to these issues, particularly lacking methods for enjoying conversations while reminiscing about memories. Therefore, there is a need for a system that allows users to engage in emotional communication through their own memories, activating their brains, and maintaining their energy.
[0118] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0119] In this invention, the server includes means for receiving and storing images, means for analyzing images, and means for generating questions and comments for the user, thereby generating dialogue based on images uploaded by the user from the terminal, and enabling continuous conversation with the user.
[0120] "Users" refer to people who use the system to upload their own photos and interact with the generative AI.
[0121] "Device" refers to the electronic device, such as a smartphone or tablet, that a user uses to access the system, upload photos, and respond to questions and comments from the Generative AI.
[0122] "Server" refers to the central device of the system, which receives and stores images uploaded by users, analyzes the images, generates questions and comments using generative AI, and sends them to the terminal.
[0123] "Generative AI" refers to an artificial intelligence model used by the server to generate new questions or comments based on image analysis results and user responses, such as an AI model that performs natural language processing.
[0124] "Image analysis" refers to the process of analyzing uploaded images using machine learning algorithms to automatically recognize information such as people, objects, and backgrounds in the photos.
[0125] "Question and comment generation" refers to the process in which the generative AI creates appropriate questions and comments for the user based on the results of image analysis and the user's responses.
[0126] "Conversation logs" refer to records of interactions with users on the system, and are data stored on the server. They include the content of the conversations and are used for later analysis.
[0127] "Database" refers to a system that systematically stores and manages data such as images and conversation logs received by a server.
[0128] "Image upload" refers to the process by which a user sends a photo file from their device to a server.
[0129] "Continuous dialogue" refers to a process in which the generative AI generates questions and comments one after another until the user is satisfied, maintaining an uninterrupted conversation with the user.
[0130] "Analysis results" refers to information about the content of a photo obtained through image analysis, including the recognition of people, objects, backgrounds, etc.
[0131] "Storage" refers to the process of storing data such as photos and conversation logs in a database on the server.
[0132] The present invention provides a system that allows users to upload memorable photos from their own devices and review the photos through dialogue with a generation AI, thereby stimulating the user's brain and maintaining their energy. An embodiment of this system is described in detail below.
[0133] The process by which a user uploads an image from their device
[0134] Users access the application using a device such as a smartphone or tablet. They open the photo upload screen within the application and select their memorable photos. When the user presses the "Upload" button, the photo file is sent to the server via the Internet.
[0135] The process by which the server receives and stores images
[0136] The server receives the photos sent by the user and stores them in a database. When saved, the photos are assigned a unique ID, which makes it easier to manage the photo data.
[0137] The process by which the server analyzes the image
[0138] The server runs an image analysis algorithm (such as YOLO or OpenCV) on the received photo to obtain information about people, objects, and backgrounds in the photo. The results of this analysis are stored in a database for later use.
[0139] The process by which generative AI generates questions and comments
[0140] The server sends data based on the image analysis results to a generative AI model (such as GPT-4). The generative AI uses this data to generate appropriate questions and comments for the user. For example, prompts such as "Where was this photo taken?" or "Who is this person?" are generated.
[0141] The process by which the device presents the generated questions and comments to the user
[0142] The questions and comments generated by the AI are sent from the server to the user's device, where they are displayed. The user can then confirm the questions and comments displayed on the device and enter their responses.
[0143] The process of sending the user's response to the server
[0144] The user's response is sent from the device to the server, and may contain text such as "This is my daughter and grandchildren."
[0145] The server analyzes the response, and the AI generates the next question or comment.
[0146] The server analyzes the user's response, and the AI generates the next question or comment based on the results, creating a continuous dialogue with the user.
[0147] Continuing conversations and logging processes
[0148] The dialogue process is repeated until the user is satisfied. All conversation logs are stored on the server and used for later analysis. This data provides valuable information for analyzing user usage and reactions.
[0149] Specific examples
[0150] Example 1: Lonely elderly people
[0151] User: Selects and uploads an old family photo.
[0152] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[0153] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[0154] Terminal: Display the question to the user.
[0155] User: Responds to the question by saying, "These are my daughter and grandchildren."
[0156] Terminal: Sends the response to the server.
[0157] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[0158] Terminal: Display the next question to the user.
[0159] Example 2: Travel photos
[0160] User: Select and upload photos of travel destinations.
[0161] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[0162] Generative AI: Generates questions like "Where is this place?" based on a photo.
[0163] Terminal: Display the question to the user.
[0164] User: Responds to the question with "This is a photo from Paris, where I went last year."
[0165] Terminal: Sends the response to the server.
[0166] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[0167] Terminal: Display the next question to the user.
[0168] Through this process, users can enjoy conversation while reminiscing on their memories, which helps to stimulate the brain and maintain energy.
[0169] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0170] Step 1:
[0171] The user accesses the application using a terminal and opens the photo upload screen. Then, the user selects the photo to upload and presses the "Upload" button. This action sends the photo file to the server over the Internet.
[0172] Input: A photo file selected from the user's device.
[0173] Output: The photo file sent to the server.
[0174] Specific operation: The user launches the application, logs in, selects a photo, and presses the "Upload" button.
[0175] Step 2:
[0176] The server saves the received photo files and assigns a unique ID to each photo, which is then stored in a database.
[0177] Input: Photo files sent over the internet to the server.
[0178] Output: Photo file stored in the database with a unique ID.
[0179] Specific operation: The server receives the photo file, saves it as binary data in its internal memory, and stores it in a database.
[0180] Step 3:
[0181] The server runs image analysis algorithms (such as YOLO or OpenCV) on the stored photos to extract information about people, objects, and backgrounds in the photos, and stores the results in a database.
[0182] Input: Photo files stored in the database.
[0183] Output: Data containing image analysis results (e.g. people, objects, and background in a photo).
[0184] What happens: The server runs an image analysis algorithm to identify the content of the photo.
[0185] Step 4:
[0186] The server sends data based on the image analysis results to a generation AI (for example, GPT-4), which then uses this data to generate questions and comments for the user.
[0187] Input: Image analysis results.
[0188] Output: Questions and comments generated by the generative AI.
[0189] Specific operation: The server converts the analysis result into a prompt sentence and sends it to the generation AI, and the AI model generates a response.
[0190] Step 5:
[0191] The server sends questions and comments from the generated AI to the user's device, which then displays the received questions and comments to the user.
[0192] Input: Questions and comments generated by the generative AI.
[0193] Output: Questions and comments displayed on the user's device.
[0194] Specific operation: The server sends the generated questions and comments to the terminal, which displays them to the user.
[0195] Step 6:
[0196] The user inputs their response to the questions and comments displayed on the terminal, and when they are finished, they press the "Send" button to send the response to the server.
[0197] Input: Questions and comments from the generative AI, and user responses.
[0198] Output: The user's response sent from the terminal to the server.
[0199] Specific behavior: The user reads the question or comment, enters their response, and presses the "Submit" button.
[0200] Step 7:
[0201] The server analyzes the user's response, and the AI generates the next question or comment based on the results. This process is repeated until the user is satisfied.
[0202] Input: The user's response.
[0203] Output: The next question or comment generated by the generative AI.
[0204] Specific operation: The server analyzes the user's response and sends the analysis results to the generation AI, and the AI model generates the next question or comment.
[0205] Step 8:
[0206] All conversation logs are stored on the server and used for later analysis. This data provides information for analyzing user usage and reactions.
[0207] Input: A log of the conversation between the user and the generated AI.
[0208] Output: Saved conversation logs.
[0209] Specific operation: The server stores the conversation log in a database and uses it for later analysis.
[0210] The above are the specific processing steps in this system.
[0211] (Application example 1)
[0212] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0213] To help elderly people and users who feel lonely to activate their brains and maintain their morale, we provide a system that provides emotional and cognitive stimulation by having users interact with a generative AI using photos of their memories. This system allows users to enjoy dialogue through photos, and we also create an environment where the generative AI can provide personalized content.
[0214] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0215] In this invention, the server includes a means for uploading images from the user's device, a means for receiving and analyzing the images, a means for transmitting the user's responses as voice or text from the device to the server, and a means for saving all conversation logs and using them to improve the AI generation and provide personalized content based on the user's interests. This helps to stimulate the user's brain and maintain their energy, and enables the AI generation to provide personalized dialogue based on the user's interests.
[0216] A "user" is someone who uses the system, uploads memorable photos, and interacts with the generative AI.
[0217] "Device" means a device used by a User to access the System, upload photos, and view and respond to generated questions and comments, including, for example, a smartphone, tablet, smart glasses, or head-mounted display.
[0218] "Means for uploading images" refers to the process by which a user selects a photo using a terminal and sends it to the server.
[0219] The "server" is a central system that receives and analyzes uploaded photos, generates questions and comments using generative AI, and stores conversation logs.
[0220] "Means for receiving and analyzing images" refers to the process by which the server captures the photos sent by the user and identifies the content of the photos using image analysis algorithms.
[0221] "Generative AI" is an artificial intelligence system that is connected to a server and generates appropriate questions and comments for users based on the results of image analysis.
[0222] "Means for generating questions and comments for users" refers to the process by which the generation AI creates questions and comments for users based on the results of image analysis.
[0223] The "means for presenting the generated questions and comments to the user" is a function by which the terminal displays the questions and comments sent from the server to the user.
[0224] "Means for transmitting a user's response by voice or text from the terminal to the server" refers to a process in which a user responds to a generated question or comment by voice or text and transmits it from the terminal to the server.
[0225] "Means for generating the next question or comment" refers to the process in which the server analyzes the user's response and uses a generation AI to create a question or comment to proceed with the next dialogue.
[0226] "Means of storing all conversation logs and using them to improve the generation AI and provide personalized content based on the user's interests" refers to the process by which the server records and stores all conversation history with the user, and based on that, the generation AI is further improved and personalized content is provided based on the user's interests.
[0227] The present invention provides a system that stimulates the user's brain and maintains their energy by interacting with a generating AI using memorable photos of the user. Specific embodiments of the system are described below.
[0228] System configuration
[0229] This system consists of a user's terminal, a server, a generation AI, and various means for sending and receiving data between them.
[0230] Image upload method
[0231] Users can upload memorable photos to the system using devices such as their smartphones, tablets, smart glasses, and head-mounted displays. Users select images and upload them using the device's application.
[0232] Image analysis methods
[0233] The server receives the uploaded photos and uses image analysis algorithms to analyze the content of the photos, using machine learning frameworks such as TensorFlow and Keras to identify people, objects, and backgrounds in the photos.
[0234] Question and comment generator
[0235] Based on the analysis results, generative AI operates to generate appropriate questions and comments for the user, using advanced generative models such as OpenAI's GPT-3.
[0236] User Presentation Method
[0237] The device presents the generated questions and comments to the user, who responds to the generated questions and comments using text input or voice input.
[0238] Response sending method
[0239] The user's response is sent from the device to the server, and may include, for example, a response such as "These are my daughter and grandchildren."
[0240] Next question generation method
[0241] The server analyzes the user's response, and the AI generates the next question or comment, allowing for a continuous dialogue with the user.
[0242] Conversation log storage method
[0243] All conversation logs are stored on the server and are used to improve the AI generation and provide personalized content based on the user's interests.
[0244] Specific examples
[0245] Example 1: Elderly people
[0246] An elderly person uploads an old family photo using their smartphone. The server receives the photo and uses TensorFlow to analyze the image and recognize that a family member is in it. A generative AI (e.g., GPT-3) generates the question "Who is in this photo?" and the device displays the question to the elderly person. The elderly person responds, "This is my daughter and grandchildren," and sends it to the server. The server analyzes the response and generates the next question: "What a lovely family. Where was this photo taken?"
[0247] Prompt Sentence Examples
[0248] "The Eiffel Tower is in the photo. Generate appropriate questions for the user that will remind them of their travel experiences."
[0249] This method allows users to reminisce and stimulate their brains and maintain their energy through generative AI dialogue. All dialogue logs are saved and used for later analysis. Specifically, it allows for the provision of personalized content based on the user's interests and usage.
[0250] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0251] Step 1:
[0252] The user opens the application on their device, selects a memorable photo, and presses the upload button, which sends the photo from the device to the server.
[0253] Input: A memorable photo
[0254] Process: Upload a file from the device to the server
[0255] Output: Photo files uploaded to the server
[0256] Step 2:
[0257] The server receives the uploaded photos and stores the image files in a database.
[0258] Input: Photo file sent from the device
[0259] Processing: Receiving the file and saving it to the database
[0260] Output: Photo data stored in a database
[0261] Step 3:
[0262] The server analyzes the content of the photo using image analysis algorithms, such as TensorFlow and Keras, to identify people, objects, and the background in the photo.
[0263] Input: Photo data stored in the database
[0264] Processing: Image analysis using TensorFlow and Keras
[0265] Output: Analysis results for the content of the photo (e.g. people's names, places, objects, etc.)
[0266] Step 4:
[0267] The server sends the analysis results to the generation AI, which then uses them to generate appropriate questions and comments for the user. In this example, OpenAI's GPT-3 is used.
[0268] Input: Image analysis results
[0269] Processing: Generative AI generates questions and comments
[0270] Output: Generated questions and comments
[0271] Step 5:
[0272] The server sends the generated questions and comments to the terminal, which displays them to the user.
[0273] Input: Generated questions and comments
[0274] Processing: Transmission from server to terminal, screen display
[0275] Output: Questions and comments displayed to the user
[0276] Step 6:
[0277] The user enters a response to a question or comment. The response can be entered by voice or text. The response is sent from the device to the server.
[0278] Input: User response (voice or text)
[0279] Processing: speech recognition or text input, sending data to server
[0280] Output: The user's response sent to the server
[0281] Step 7:
[0282] The server receives and analyzes the user's response. Based on the analysis results, the AI generates the next question or comment, thus continuing the dialogue with the user.
[0283] Input: User response
[0284] Processing: Analyzing responses and generating the next question or comment using AI
[0285] Output: Next question or comment
[0286] Step 8:
[0287] The server sends the next question or comment to the terminal, which displays it to the user, who responds again, and the process from step 6 to step 8 is repeated.
[0288] Input: Next question or comment
[0289] Processing: Transmission from server to terminal, screen display, user response input
[0290] Output: Repeated user interactions
[0291] Step 9:
[0292] All conversation logs are stored on the server. These logs include user usage and reactions and are used for later analysis. They serve as the basis for improving the AI generation and providing personalized content based on the user's interests.
[0293] Input: User interaction log
[0294] Processing: Data storage and later analysis
[0295] Output: Data used to improve generative AI and provide personalized content
[0296] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0297] This invention provides a system that stimulates the user's brain and maintains their morale by interacting with a generating AI using photos of the user's memories. This system also incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction.
[0298] System Overview
[0299] The process by which a user uploads an image from their device
[0300] Users upload photos to the system using their own devices (e.g., smartphones or tablets). Users can easily select and upload photos by opening the application and going to the upload screen. Uploaded photos are first transferred to the server.
[0301] The process by which the server receives and analyzes the images
[0302] The uploaded photos are received by the server and stored in a database, where they are then analyzed using image analysis algorithms to automatically identify people, objects, and background information in the photos.
[0303] The process by which generative AI generates questions and comments for users
[0304] Based on the results of image analysis, the server sends the necessary data to the generation AI. The generation AI then generates appropriate questions and comments for the user based on the content of the photo. For example, it creates questions such as "Where was this photo taken?" or "Who is this person?"
[0305] The process by which the device presents the generated questions and comments to the user
[0306] The questions and comments generated by the AI are sent back from the server to the device, which then displays them to the user. The user can then confirm the questions and comments and enter a response.
[0307] The process of sending user responses from the terminal to the server
[0308] The response entered by the user is sent from the device to the server, and may include, for example, "These are my daughter and grandchildren."
[0309] The server analyzes the response, and the AI generates the next question or comment.
[0310] The server receives and analyzes the user's response, and the generation AI generates the next question or comment. This continues the conversation with the user and advances the dialogue process. For example, the next question might be, "Where was this photo taken?"
[0311] Emotion recognition process by emotion engine
[0312] The emotion engine analyzes the user's emotions from images uploaded by the user and responses during conversations. The emotion engine uses facial expression analysis and voice analysis to determine whether the user is happy or anxious.
[0313] Tuning process for emotion-based generative AI
[0314] Based on the analysis results of the emotion engine, the generation AI will adjust the questions and comments. For example, if the user is determined to be sad, the generation AI will generate comforting comments. If the user is having fun, the generation AI will generate questions that bring out that emotion.
[0315] Continuing conversations and logging processes
[0316] This conversation process is repeated until the user is satisfied. All conversation logs are stored on the server and used to analyze user usage and reactions at a later date. The results of this analysis are used to improve the system and enhance the user experience.
[0317] Specific examples
[0318] Example 1: Lonely elderly people
[0319] User: Selects and uploads an old family photo.
[0320] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[0321] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[0322] Terminal: Display the question to the user.
[0323] User: Responds to the question by saying, "These are my daughter and grandchildren."
[0324] Terminal: Sends the response to the server.
[0325] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[0326] Emotion engine: Analyzes the user's emotions from images and responses and recognizes that the user is feeling nostalgic.
[0327] Generative AI: Based on the user's emotions, the system further adjusts the questions and comments, such as "Tell me more about your memories of that time."
[0328] Terminal: Display the next question to the user.
[0329] Example 2: Travel photos
[0330] User: Select and upload photos of travel destinations.
[0331] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[0332] Generative AI: Generates questions like "Where is this place?" based on a photo.
[0333] Terminal: Display the question to the user.
[0334] User: Responds to the question with "This is a photo from Paris, where I went last year."
[0335] Terminal: Sends the response to the server.
[0336] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[0337] Emotion engine: Analyzes the user's emotions from images and responses and recognizes what the user is enjoying.
[0338] Generative AI: Further tailor questions and comments based on user sentiment, such as "What other places have you been?"
[0339] Terminal: Display the next question to the user.
[0340] In this way, this system allows users to enjoy conversations through photos, stimulating their brains and maintaining their energy. Furthermore, the introduction of an emotion engine enables communication that responds to the user's emotions, resulting in more effective dialogue.
[0341] The processing flow will be explained below.
[0342] Step 1:
[0343] The user opens the application on their device and proceeds to the photo upload screen.
[0344] Step 2:
[0345] The user selects a photo on their device and clicks the upload button.
[0346] Step 3:
[0347] The terminal generates a request to transfer the selected photo to the server.
[0348] Step 4:
[0349] The server receives the request from the terminal and receives the photo data.
[0350] Step 5:
[0351] The server stores the received photos in a database.
[0352] Step 6:
[0353] The server passes the stored photos to an image analysis algorithm, which begins the analysis process.
[0354] Step 7:
[0355] The server performs image analysis to determine the content of the photo, such as identifying people, objects, and the background.
[0356] Step 8:
[0357] The server sends the analysis results to the generation AI.
[0358] Step 9:
[0359] Based on the results of image analysis, the generative AI generates appropriate questions and comments for the user, such as "Where was this photo taken?" or "Who is this person?"
[0360] Step 10:
[0361] The server sends the generated questions and comments back to the device.
[0362] Step 11:
[0363] The terminal displays the questions and comments received from the server to the user.
[0364] Step 12:
[0365] The user enters a response to the question or comment on the terminal.
[0366] Step 13:
[0367] The terminal generates a request to send the user's response to the server.
[0368] Step 14:
[0369] The server receives the response from the terminal.
[0370] Step 15:
[0371] The server passes the received response to the emotion engine, which analyzes the user's emotion.
[0372] Step 16:
[0373] The emotion engine recognizes emotions from user responses and images and sends the results to the generation AI.
[0374] Step 17:
[0375] The generative AI generates the next appropriate question or comment based on the emotion recognition results from the emotion engine and the user's response. For example, if the user is feeling nostalgic, it generates a question such as, "Tell me more about your memories from that time."
[0376] Step 18:
[0377] The server sends the next generated question or comment back to the device.
[0378] Step 19:
[0379] The device will display the next question or comment to the user.
[0380] Step 20:
[0381] The user again enters a response, and this process is repeated until the user is satisfied.
[0382] Step 21:
[0383] The server stores all conversation logs.
[0384] Step 22:
[0385] The server will later analyze the saved conversation logs to evaluate the user's usage and reactions.
[0386] Example 2
[0387] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0388] In modern society, the number of elderly people and individuals who feel lonely is increasing, making it important to improve their mental health and increase communication opportunities for these people. In particular, there is a demand for systems that can provide mental support by promoting interaction through past memories and photos and engaging in emotionally appropriate dialogue. Conventional systems struggle to improve the quality of emotion recognition and dialogue, and there is a lack of methods for achieving effective communication.
[0389] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0390] In this invention, the server includes means for receiving images and storing them in a database, means for generating image analysis results using an image analysis algorithm, means for analyzing the user's responses and generating the next question or comment using a generation AI, means for an emotion engine to recognize the user's emotions and for the generation AI to adjust the questions or comments based on the results, and means for saving all conversation logs and analyzing them later. This enables natural conversation based on the user's memories, provides appropriate communication according to emotions, and helps maintain and improve the user's mental health.
[0391] "Means for uploading images" refers to a series of operations and functions that allow a user to select digital images from their terminal and send them to the server.
[0392] A "database" is a system for systematically storing and managing digital information, where images and associated metadata are stored.
[0393] An "image analysis algorithm" is a computational method that analyzes the content of a digital image and extracts information about the people, objects, background, and other aspects of the photograph.
[0394] "Generative AI" is artificial intelligence that automatically generates questions and comments based on user responses and image analysis results.
[0395] An "emotion engine" is a technology that analyzes a user's emotions and adjusts the output of the generative AI based on their emotional state.
[0396] A "terminal" is a digital device used by a user, such as a smartphone or tablet.
[0397] A "conversation log" is a record of all interactions between a user and a system, data that is stored for later analysis.
[0398] A "prompt sentence" is the input text given to the generation AI, which is used to generate the next question or comment based on the user's response and the results of image analysis.
[0399] "Means for generating questions and comments" refers to the functions and operations by which the generation AI creates appropriate questions and comments for the user based on the prompt text.
[0400] This invention is a system that aims to stimulate the user's brain and maintain their energy by having them interact with a generative AI using photos of their memories. This system incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction.
[0401] System Overview
[0402] User-uploaded images
[0403] Users upload photos to the system using devices such as smartphones or tablets. Users launch a dedicated application, select memorable photos on the upload screen, and send them to the server. The uploaded photos are then transferred from the device to the server.
[0404] Receiving and analyzing images by the server
[0405] The server receives the uploaded photos and stores them in a database. The server then analyzes the content of the photos using image analysis algorithms (such as Google Cloud Vision API or Microsoft Azure's Computer Vision API). The analysis results include information about the people, objects, and background in the photos.
[0406] Question generation using generative AI
[0407] Based on the results of image analysis, the server sends the necessary data to a generative AI (such as OpenAI's GPT-4). Based on the received data, the generative AI generates appropriate questions and comments for the user. For example, it creates specific questions such as "Where was this photo taken?" or "Who is this person?"
[0408] Display of questions on the device
[0409] The generated AI's questions and comments are sent from the server to the user's device and displayed on the device. The user reads the displayed questions and comments and enters a response.
[0410] Parsing the user's response and generating the next question
[0411] The response entered by the user is sent from the device to the server. The server analyzes the received response and sends the results to the generation AI. The generation AI then generates the next most appropriate question or comment. This allows for a continuous dialogue with the user.
[0412] Emotion recognition and output adjustment by emotion engine
[0413] The emotion engine analyzes the user's emotions from images uploaded by the user and responses during conversations. The emotion engine evaluates whether the user is happy or anxious through facial expression and voice analysis. Based on the results of the emotion engine, the server adjusts the prompts sent to the generation AI and generates appropriate questions and comments.
[0414] Specific examples
[0415] Example 1: Lonely elderly people
[0416] The user selects and uploads an old family photo. The server receives the photo and uses an image analysis algorithm to recognize that a family member is in it. The generation AI generates the question, "Who is in this photo?" and displays it on the device. The user responds, "This is my daughter and grandchildren," and the device sends the response to the server. The server analyzes the response, and the generation AI generates the next question, "What a lovely family. Where was this photo taken?" Furthermore, if the emotion engine senses the user's nostalgia, it generates a question such as, "Tell me more about your memories from that time."
[0417] Example 2: Travel photos
[0418] The user selects and uploads photos of their travel destinations. The server receives the photos and uses an image analysis algorithm to recognize the travel destinations. The generation AI generates the question "Where is this place?" and displays it on the device. The user responds, "This is a photo of Paris, where I went last year," and the device sends the response to the server. The server analyzes the response, and the generation AI generates the next question, "How was Paris? Were there any places that particularly impressed you?" The emotion engine analyzes the user's emotions and, if it identifies that the user is enjoying themselves, generates questions such as "What other places did you go?"
[0419] Prompt Sentence Examples
[0420] As an example of a prompt for a generative AI model, input the following text to the generative AI:
[0421] Analyzing a photo submitted by a user reveals that it contains family members. Based on that content, generate questions or comments to naturally continue the conversation with the user. For example, "Who is in this photo?"
[0422] This system provides natural dialogue through users' memorable photos and realizes appropriate communication according to their emotions, which is expected to contribute to maintaining and improving users' mental health.
[0423] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0424] Step 1:
[0425] The user initiates the image upload process from their device.
[0426] Specific operation: The user launches the application on their smartphone or tablet, selects "Upload photos" from the menu, selects a memorable photo from their photo library, and presses the "Upload" button.
[0427] Input: A photo file selected by the user.
[0428] Output: The photo files are saved on your device and ready to be sent to the server.
[0429] Step 2:
[0430] The device sends the photo to the server.
[0431] Specific operation: The device makes an HTTP request to send the user-selected photo file to the server. The request includes the photo file and associated metadata (e.g., upload date and time, user ID).
[0432] Input: A user-selected photo file and its metadata.
[0433] Output: The photo files and metadata are sent to the server.
[0434] Step 3:
[0435] The server receives the photos and stores them in a database.
[0436] What it does: The server parses the incoming HTTP request and stores the photo file and its metadata in a database, along with information about the photo's identifier and storage location.
[0437] Input: The submitted photo file and its metadata.
[0438] Output: Photo files and associated metadata stored in a database.
[0439] Step 4:
[0440] The server analyzes the photo using image analysis algorithms.
[0441] How it works: The server sends the stored photo file to an image analysis algorithm (e.g., Google Cloud Vision API) to analyze the content of the photo, including people, objects, and background information.
[0442] Input: Photo files stored in the database.
[0443] Output: The analysis data generated as a result of the image analysis algorithm.
[0444] Step 5:
[0445] The generative AI creates prompts to generate questions and comments for the user.
[0446] Specific operation: The server creates a prompt sentence to send to the generation AI based on the image analysis results. This prompt sentence contains detailed information to facilitate the dialogue with the user. For example, "A family photo has been uploaded. Please generate a question about this photo."
[0447] Input: Image analysis results.
[0448] Output: The prompt sent to the generation AI.
[0449] Step 6:
[0450] Generative AI generates questions and comments.
[0451] How it works: A generative AI model (e.g., OpenAI's GPT-4) generates questions or comments to present to the user based on prompts received from the server, such as "Who is in this photo?"
[0452] Input: The server-generated prompt text.
[0453] Output: Generated questions and comments.
[0454] Step 7:
[0455] The server sends the generated questions and comments to the device.
[0456] Specific operation: The server creates an HTTP response to send the generated question or comment to the terminal, and sends it to the terminal.
[0457] Input: Generated questions and comments.
[0458] Output: Questions and comments sent to the user's device.
[0459] Step 8:
[0460] The device displays the question or comment to the user.
[0461] Specific operation: The device displays the received questions and comments on the screen and presents them to the user. An interface is provided so that the user can review them.
[0462] Input: Questions and comments sent by the server.
[0463] Output: Questions and comments displayed on the terminal screen.
[0464] Step 9:
[0465] The user enters a response to the question.
[0466] Specific behavior: The user reads the questions and comments displayed on the device and enters a response. For example, "These are my daughter and grandchildren."
[0467] Input: Questions or comments displayed on the device.
[0468] Output: The response text entered by the user.
[0469] Step 10:
[0470] The terminal sends the user's response to the server.
[0471] Specific operation: The terminal creates an HTTP request to send the response text entered by the user to the server.
[0472] Input: The response text entered by the user.
[0473] Output: The response text sent to the server.
[0474] Step 11:
[0475] The server analyzes the user's response and creates and sends the next prompt to the generation AI.
[0476] Specific operation: The server analyzes the user's response text, creates a new prompt based on its content, and sends the created prompt to the generation AI to generate the next question or comment.
[0477] Input: The user's response text.
[0478] Output: The next prompt sent to the generation AI.
[0479] Step 12:
[0480] Generative AI generates the next question or comment.
[0481] Specific behavior: The generative AI generates the next question or comment based on the newly received prompt, for example, "Where was this photo taken?"
[0482] Input: The newly created prompt statement.
[0483] Output: The next question or comment.
[0484] Step 13:
[0485] The emotion engine recognizes the user's emotions and adjusts the output of the generative AI.
[0486] Specific operation: The emotion engine analyzes the user's facial expressions and responses to evaluate the user's emotional state. Based on the results, it adjusts the generation AI to generate appropriate questions and comments.
[0487] Input: User's facial expression data and response content.
[0488] Output: Adjusted prompt and output of the generation AI.
[0489] Step 14:
[0490] The server sends the tailored questions and comments to the device.
[0491] Specific operation: The server sends questions and comments adjusted by the emotion engine to the terminal.
[0492] Input: Moderated questions and comments.
[0493] Output: The tailored questions and comments sent to your device.
[0494] Step 15:
[0495] The device displays the tailored question or comment to the user.
[0496] Specific operation: The device displays the adjusted questions and comments on the screen and presents them to the user.
[0497] Input: Moderated questions and comments.
[0498] Output: The adjusted questions and comments displayed on the screen.
[0499] Step 16:
[0500] The server stores all conversation logs.
[0501] What it does: The server records all interactions and stores them in a database. This information is later analyzed and used to improve the system and enhance the user experience.
[0502] Input: All dialogue content (questions, responses, and adjustment results).
[0503] Output: Conversation logs stored in a database.
[0504] (Application example 2)
[0505] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0506] The problem that this invention aims to solve is to realize more effective communication that responds to the user's emotions while stimulating the user's brain and maintaining their energy through dialogue using photos related to the user's memories. Conventional systems simply continue dialogue mechanically without taking the user's emotions into consideration, which has the problem of degrading the quality of the user experience.
[0507] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0508] In this invention, the server includes means for a user to upload images from a terminal, means for the server to receive and analyze the images, means for the generation AI to generate questions and comments for the user based on the generated analysis results, means for the terminal to present the generated questions and comments to the user, means for the terminal to send the user's response to the server, means for the server to analyze the response and for the generation AI to generate the next question or comment, means for analyzing emotions from the user's facial expressions and voice, means for the generation AI to adjust the content of the dialogue based on the emotion analysis results, means for continuously conversing with the user, and means for saving all conversation logs and analyzing them later. This allows the user to interact with the generation AI through memorable photos, and the content of the dialogue is appropriately adjusted through emotion analysis, thereby stimulating the user's brain and maintaining their energy.
[0509] "Means for users to upload images from their devices" refers to the function that allows users to send image data to the system using devices such as smartphones and tablets.
[0510] "Means for the server to receive and analyze images" refers to the function of the cloud server receiving uploaded image data and analyzing its contents using a specified image analysis algorithm.
[0511] "Means for the generation AI to generate questions and comments for the user based on the generated analysis results" refers to the function of using natural language processing technology to generate appropriate questions and comments based on the content of the image analyzed by the server.
[0512] "Means for the device to present the generated questions and comments to the user" refers to a function that displays the generated questions and comments on the user's smartphone or tablet.
[0513] "Means for transmitting a user's response from the terminal to the server" refers to a function for transmitting the response data entered by the user again to the cloud server.
[0514] "The server analyzes the response, and the generation AI generates the next question or comment" refers to the function in which the cloud server analyzes the user's response and generates the next conversation content based on the analysis results.
[0515] "Means for analyzing emotions from the user's facial expressions and voice" refers to a function that analyzes the facial expressions and voice data shown by the user during a conversation and determines their emotional state (joy, sadness, surprise, etc.).
[0516] "Means for the generation AI to adjust the content of the dialogue based on the results of emotion analysis" refers to the function by which the generation AI takes into account the user's emotional state obtained through emotion analysis and appropriately adjusts the next question or comment.
[0517] "Means for continuing conversation with the user" refers to a function that allows a dialogue between the user and the system to continue.
[0518] "A means of saving all conversation logs and analyzing them later" refers to a function that records and saves all generated conversation content and later analyzes changes in user behavior patterns and emotions based on that data.
[0519] This invention provides a system that stimulates the user's brain and maintains their morale by interacting with a generating AI using memorable photos of the user. This system also incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction. Each element of this system is explained below.
[0520] System configuration
[0521] 1. A way for users to upload images from their devices
[0522] Users use their smartphones, tablets, or other devices to select and upload images within the application, a process designed to make it easy and intuitive for users to submit photos to the system.
[0523] 2. How the server receives and analyzes images
[0524] Uploaded photos are transferred to a cloud server, which uses image analysis libraries such as the Google Vision API to automatically analyze the content of the photo and extract information about people, places, objects, and more.
[0525] 3. A means for the generative AI to generate questions and comments for users based on the generated analysis results
[0526] The results of the image analysis are sent to a generative AI (e.g., GPT-3 or GPT-4), which uses this data to generate appropriate questions and comments for the user.
[0527] For example: "Where was this photo taken?"
[0528] 4. A means for the device to present generated questions and comments to the user
[0529] The questions and comments generated by the AI are sent from the server to the user's device, where they are displayed to the user, and a dialogue begins.
[0530] 5. A means of transmitting user responses from the terminal to the server
[0531] When a user responds to a question or comment through their device, the response is sent to the cloud server.
[0532] 6. The server analyzes the response and the AI generates the next question or comment.
[0533] The server analyzes the user's responses and sends the results to the generation AI, which then generates new questions and comments to continue the dialogue with the user.
[0534] 7. Means of analyzing emotions from user facial expressions and voice
[0535] Using an emotion engine (for example, Microsoft Azure Emotion API), emotions are analyzed from the user's facial expressions and voice. The analysis results indicate the user's emotional state, such as whether they are happy or sad.
[0536] 8. How generative AI can adjust dialogue content based on emotion analysis results
[0537] Based on the analysis results of the emotion engine, the generative AI will adjust the dialogue appropriately. For example, if the user is sad, it will generate comforting comments, and if the user is happy, it will generate questions that will bring out that emotion.
[0538] For example: "Tell me more about your memories of that time."
[0539] 9. A way to continue the conversation with your users
[0540] The dialogue continues until the user is satisfied. The entire conversation process is controlled by the system.
[0541] 10. A way to store all conversation logs and analyze them later
[0542] All generated conversation logs are stored on a cloud server, allowing users' usage patterns and reactions to be analyzed at a later date, helping to improve the system and enhance the user experience.
[0543] This system allows users to enjoy conversations through memorable photos, stimulating their brains and maintaining their energy. In addition, the introduction of an emotion engine enables communication based on the user's emotions, resulting in more effective conversations.
[0544] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0545] Step 1:
[0546] A way for users to upload images from their devices
[0547] A user opens the application on a smartphone or tablet, selects a photo from the upload screen, and sends it to the system. The input is the photo file selected by the user, and the output is that the photo file is sent to the cloud server. The photo file is transferred between the device and the cloud server.
[0548] Step 2:
[0549] The server receives and analyzes the images
[0550] The server receives the photos sent from the device and stores them in a database. The stored photos are then passed to the Google Vision API for image analysis. The input is the uploaded photo file, and the output is the analyzed image content (people, places, objects, etc.) data. Image analysis uses machine learning algorithms to analyze the content of the photo.
[0551] Step 3:
[0552] A means for the AI to generate questions and comments for users based on the generated analysis results
[0553] The server sends the image analysis results to a generation AI (GPT-3 or GPT-4), which generates appropriate questions and comments. The input is the image analysis results, and the output is questions and comments to present to the user. Specifically, the analysis results are input as prompts to the generation AI, which then uses natural language processing technology to generate corresponding questions and comments.
[0554] Step 4:
[0555] A means by which the device presents generated questions and comments to the user
[0556] The server sends the generated questions and comments to the terminal and displays them to the user. The input is the generated questions and comments, and the output is what is displayed to the user. The terminal receives this and displays it on the screen.
[0557] Step 5:
[0558] A means for transmitting user responses from the terminal to the server
[0559] Users use their devices to respond to questions and comments, and then send the response data to the cloud server. The input is the user's response, and the output is the response data sent to the cloud server. Specifically, the user enters text and presses the send button, which sends the data to the server.
[0560] Step 6:
[0561] The server analyzes the response and the generation AI generates the next question or comment.
[0562] The server analyzes the user's response data and sends the analysis results to the generation AI to generate the next question or comment. The input is the user's response data, and the output is the newly generated question or comment. The response data is analyzed for text, and its content is provided to the generation AI as a prompt.
[0563] Step 7:
[0564] A means of analyzing emotions from the user's facial expressions and voice
[0565] The server receives the user's facial expression and voice data sent from the device and performs emotion analysis using an emotion engine (Microsoft Azure Emotion API). The input is facial expression and voice data, and the output is data on the user's emotional state (e.g., joy, sadness, surprise, etc.). Specifically, image and voice data is passed to the emotion engine for analysis.
[0566] Step 8:
[0567] A means for generative AI to adjust dialogue content based on emotion analysis results
[0568] The server provides the emotion analysis results to the generation AI, which then adjusts the dialogue appropriately based on the results. The input is the emotion analysis results, and the output is adjusted questions or comments. The generation AI receives prompts containing the emotion analysis results and generates appropriate dialogue content based on them. For example, if it determines that the user is sad, it generates a question such as, "Tell me more about your memories of that time."
[0569] Step 9:
[0570] A way to maintain an ongoing conversation with users
[0571] The server and the generation AI work together to continue the dialogue until the user is satisfied. The input is continuous responses from the user, and the output is newly generated questions and comments. The dialogue content is continuously generated, and the interaction with the user progresses without interruption.
[0572] Step 10:
[0573] A means to store all conversation logs and analyze them later
[0574] The server stores all dialogue logs and later analyzes changes in user behavior patterns and emotions based on them. The input is all the dialogue logs generated, and the output is the stored data and the analysis results. The log data is stored in a database and later analyzed using analysis tools.
[0575] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0576] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0577] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0578] [Second embodiment]
[0579] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0580] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0581] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0582] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0583] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0584] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0585] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0586] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0587] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0588] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0589] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0590] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0591] This invention provides a system that stimulates the user's brain and maintains their energy by interacting with a generating AI using photos of the user's memories. This system involves a multi-stage process in which the user uploads photos and looks back on their memories through conversations with the generating AI.
[0592] System Overview
[0593] The process by which a user uploads an image from their device
[0594] Users upload photos to the system using their own devices (e.g., smartphones or tablets). Users can easily select and upload photos by opening the application and going to the upload screen. Uploaded photos are first transferred to the server.
[0595] The process by which the server receives and analyzes the images
[0596] The uploaded photos are received by the server and stored in a database, where they are then analyzed using image analysis algorithms to automatically identify people, objects, and background information in the photos.
[0597] The process by which generative AI generates questions and comments for users
[0598] Based on the results of image analysis, the server sends the necessary data to the generation AI. The generation AI then generates appropriate questions and comments for the user based on the content of the photo. For example, it creates questions such as "Where was this photo taken?" or "Who is this person?"
[0599] The process by which the device presents the generated questions and comments to the user
[0600] The questions and comments generated by the AI are sent back from the server to the device, which then displays them to the user. The user can then confirm the questions and comments and enter a response.
[0601] The process of sending user responses from the terminal to the server
[0602] The response entered by the user is sent from the device to the server, and may include, for example, "These are my daughter and grandchildren."
[0603] The server analyzes the response, and the AI generates the next question or comment.
[0604] The server receives and analyzes the user's response, and the generation AI generates the next question or comment. This continues the conversation with the user and advances the dialogue process. For example, the next question might be, "Where was this photo taken?"
[0605] Continuing conversations and logging processes
[0606] This conversation process is repeated until the user is satisfied. All conversation logs are stored on the server and used later to analyze user usage and reactions.
[0607] Specific examples
[0608] Example 1: Lonely elderly people
[0609] User: Selects and uploads an old family photo.
[0610] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[0611] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[0612] Terminal: Display the question to the user.
[0613] User: Responds to the question by saying, "These are my daughter and grandchildren."
[0614] Terminal: Sends the response to the server.
[0615] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[0616] Terminal: Display the next question to the user.
[0617] Example 2: Travel photos
[0618] User: Select and upload photos of travel destinations.
[0619] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[0620] Generative AI: Generates questions like "Where is this place?" based on a photo.
[0621] Terminal: Display the question to the user.
[0622] User: Responds to the question with "This is a photo from Paris, where I went last year."
[0623] Terminal: Sends the response to the server.
[0624] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[0625] Terminal: Display the next question to the user.
[0626] As described above, through this system, users can enjoy conversations through photos, which helps stimulate the brain and maintain energy.
[0627] The processing flow will be explained below.
[0628] Step 1:
[0629] The user opens the application on their device and proceeds to the photo upload screen.
[0630] Step 2:
[0631] The user selects a photo on their device and clicks the upload button.
[0632] Step 3:
[0633] The terminal generates a request to transfer the selected photo to the server.
[0634] Step 4:
[0635] The server receives the request from the terminal and receives the photo data.
[0636] Step 5:
[0637] The server stores the received photos in a database.
[0638] Step 6:
[0639] The server passes the stored photos to an image analysis algorithm, which begins the analysis process.
[0640] Step 7:
[0641] The server performs image analysis to determine the content of the photo, such as identifying people, objects, and the background.
[0642] Step 8:
[0643] The server sends the image analysis results to the generation AI.
[0644] Step 9:
[0645] Based on the results of image analysis, the generative AI generates appropriate questions and comments for the user, such as "Where was this photo taken?" or "Who is this person?"
[0646] Step 10:
[0647] The server sends the generated questions and comments back to the device.
[0648] Step 11:
[0649] The terminal displays the questions and comments received from the server to the user.
[0650] Step 12:
[0651] The user enters a response to the question or comment on the terminal.
[0652] Step 13:
[0653] The terminal generates a request to send the user's response to the server.
[0654] Step 14:
[0655] The server receives the response from the terminal.
[0656] Step 15:
[0657] The server analyzes the received response and sends it to the generating AI.
[0658] Step 16:
[0659] Based on the user's response, the generative AI generates the next appropriate question or comment, such as "Where was this photo taken?"
[0660] Step 17:
[0661] The server sends the next generated question or comment back to the device.
[0662] Step 18:
[0663] The device will display the next question or comment to the user.
[0664] Step 19:
[0665] The user again enters a response, and this process is repeated until the user is satisfied.
[0666] Step 20:
[0667] The server stores all conversation logs.
[0668] Step 21:
[0669] The server will later analyze the saved conversation logs to evaluate the user's usage and reactions.
[0670] Example 1
[0671] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0672] Today's elderly and lonely people lack emotional communication and brain activation. As a result, they face problems such as a decline in cognitive function and loss of energy. Conventional technologies have not provided effective solutions to these issues, particularly lacking methods for enjoying conversations while reminiscing about memories. Therefore, there is a need for a system that allows users to engage in emotional communication through their own memories, activating their brains, and maintaining their energy.
[0673] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0674] In this invention, the server includes means for receiving and storing images, means for analyzing images, and means for generating questions and comments for the user, thereby generating dialogue based on images uploaded by the user from the terminal, and enabling continuous conversation with the user.
[0675] "Users" refer to people who use the system to upload their own photos and interact with the generative AI.
[0676] "Device" refers to the electronic device, such as a smartphone or tablet, that a user uses to access the system, upload photos, and respond to questions and comments from the Generative AI.
[0677] "Server" refers to the central device of the system, which receives and stores images uploaded by users, analyzes the images, generates questions and comments using generative AI, and sends them to the terminal.
[0678] "Generative AI" refers to an artificial intelligence model used by the server to generate new questions or comments based on image analysis results and user responses, such as an AI model that performs natural language processing.
[0679] "Image analysis" refers to the process of analyzing uploaded images using machine learning algorithms to automatically recognize information such as people, objects, and backgrounds in the photos.
[0680] "Question and comment generation" refers to the process in which the generative AI creates appropriate questions and comments for the user based on the results of image analysis and the user's responses.
[0681] "Conversation logs" refer to records of interactions with users on the system, and are data stored on the server. They include the content of the conversations and are used for later analysis.
[0682] "Database" refers to a system that systematically stores and manages data such as images and conversation logs received by a server.
[0683] "Image upload" refers to the process by which a user sends a photo file from their device to a server.
[0684] "Continuous dialogue" refers to a process in which the generative AI generates questions and comments one after another until the user is satisfied, maintaining an uninterrupted conversation with the user.
[0685] "Analysis results" refers to information about the content of a photo obtained through image analysis, including the recognition of people, objects, backgrounds, etc.
[0686] "Storage" refers to the process of storing data such as photos and conversation logs in a database on the server.
[0687] The present invention provides a system that allows users to upload memorable photos from their own devices and review the photos through dialogue with a generation AI, thereby stimulating the user's brain and maintaining their energy. An embodiment of this system is described in detail below.
[0688] The process by which a user uploads an image from their device
[0689] Users access the application using a device such as a smartphone or tablet. They open the photo upload screen within the application and select their memorable photos. When the user presses the "Upload" button, the photo file is sent to the server via the Internet.
[0690] The process by which the server receives and stores images
[0691] The server receives the photos sent by the user and stores them in a database. When saved, the photos are assigned a unique ID, which makes it easier to manage the photo data.
[0692] The process by which the server analyzes the image
[0693] The server runs an image analysis algorithm (such as YOLO or OpenCV) on the received photo to obtain information about people, objects, and backgrounds in the photo. The results of this analysis are stored in a database for later use.
[0694] The process by which generative AI generates questions and comments
[0695] The server sends data based on the image analysis results to a generative AI model (such as GPT-4). The generative AI uses this data to generate appropriate questions and comments for the user. For example, prompts such as "Where was this photo taken?" or "Who is this person?" are generated.
[0696] The process by which the device presents the generated questions and comments to the user
[0697] The questions and comments generated by the AI are sent from the server to the user's device, where they are displayed. The user can then confirm the questions and comments displayed on the device and enter their responses.
[0698] The process of sending the user's response to the server
[0699] The user's response is sent from the device to the server, and may contain text such as "This is my daughter and grandchildren."
[0700] The server analyzes the response, and the AI generates the next question or comment.
[0701] The server analyzes the user's response, and the AI generates the next question or comment based on the results, creating a continuous dialogue with the user.
[0702] Continuing conversations and logging processes
[0703] The dialogue process is repeated until the user is satisfied. All conversation logs are stored on the server and used for later analysis. This data provides valuable information for analyzing user usage and reactions.
[0704] Specific examples
[0705] Example 1: Lonely elderly people
[0706] User: Selects and uploads an old family photo.
[0707] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[0708] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[0709] Terminal: displays the question to the user.
[0710] User: Responds to the question by saying, "These are my daughter and grandchildren."
[0711] Terminal: Sends the response to the server.
[0712] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[0713] Terminal: Display the next question to the user.
[0714] Example 2: Travel photos
[0715] User: Select and upload photos of travel destinations.
[0716] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[0717] Generative AI: Generates questions like "Where is this place?" based on a photo.
[0718] Terminal: displays the question to the user.
[0719] User: Responds to the question with "This is a photo from Paris, where I went last year."
[0720] Terminal: Sends the response to the server.
[0721] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[0722] Terminal: Display the next question to the user.
[0723] Through this process, users can enjoy conversation while reminiscing on their memories, which helps to stimulate the brain and maintain energy.
[0724] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0725] Step 1:
[0726] The user accesses the application using a terminal and opens the photo upload screen. Then, the user selects the photo to upload and presses the "Upload" button. This action sends the photo file to the server over the Internet.
[0727] Input: A photo file selected from the user's device.
[0728] Output: The photo file sent to the server.
[0729] Specific operation: The user launches the application, logs in, selects a photo, and presses the "Upload" button.
[0730] Step 2:
[0731] The server saves the received photo files and assigns a unique ID to each photo, which is then stored in a database.
[0732] Input: Photo files sent over the internet to the server.
[0733] Output: Photo file stored in the database with a unique ID.
[0734] Specific operation: The server receives the photo file, saves it as binary data in its internal memory, and stores it in a database.
[0735] Step 3:
[0736] The server runs image analysis algorithms (such as YOLO or OpenCV) on the stored photos to extract information about people, objects, and backgrounds in the photos, and stores the results in a database.
[0737] Input: Photo files stored in the database.
[0738] Output: Data containing image analysis results (e.g. people, objects, and background in a photo).
[0739] What happens: The server runs an image analysis algorithm to identify the content of the photo.
[0740] Step 4:
[0741] The server sends data based on the image analysis results to a generation AI (for example, GPT-4), which then uses this data to generate questions and comments for the user.
[0742] Input: Image analysis results.
[0743] Output: Questions and comments generated by the generative AI.
[0744] Specific operation: The server converts the analysis result into a prompt sentence and sends it to the generation AI, and the AI model generates a response.
[0745] Step 5:
[0746] The server sends questions and comments from the generated AI to the user's device, which then displays the received questions and comments to the user.
[0747] Input: Questions and comments generated by the generative AI.
[0748] Output: Questions and comments displayed on the user's device.
[0749] Specific operation: The server sends the generated questions and comments to the terminal, which displays them to the user.
[0750] Step 6:
[0751] The user inputs their response to the questions and comments displayed on the terminal, and when they are finished, they press the "Send" button to send the response to the server.
[0752] Input: Questions and comments from the generative AI, and user responses.
[0753] Output: The user's response sent from the terminal to the server.
[0754] Specific behavior: The user reads the question or comment, enters their response, and presses the "Submit" button.
[0755] Step 7:
[0756] The server analyzes the user's response, and the AI generates the next question or comment based on the results. This process is repeated until the user is satisfied.
[0757] Input: The user's response.
[0758] Output: The next question or comment generated by the generative AI.
[0759] Specific operation: The server analyzes the user's response and sends the analysis results to the generation AI, and the AI model generates the next question or comment.
[0760] Step 8:
[0761] All conversation logs are stored on the server and used for later analysis. This data provides information for analyzing user usage and reactions.
[0762] Input: A log of the conversation between the user and the generated AI.
[0763] Output: Saved conversation logs.
[0764] Specific operation: The server stores the conversation log in a database and uses it for later analysis.
[0765] The above are the specific processing steps in this system.
[0766] (Application example 1)
[0767] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0768] To help elderly people and users who feel lonely to activate their brains and maintain their morale, we provide a system that provides emotional and cognitive stimulation by having users interact with a generative AI using photos of their memories. This system allows users to enjoy dialogue through photos, and we also create an environment where the generative AI can provide personalized content.
[0769] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0770] In this invention, the server includes a means for uploading images from the user's device, a means for receiving and analyzing the images, a means for transmitting the user's responses as voice or text from the device to the server, and a means for saving all conversation logs and using them to improve the AI generation and provide personalized content based on the user's interests. This helps to stimulate the user's brain and maintain their energy, and enables the AI generation to provide personalized dialogue based on the user's interests.
[0771] A "user" is someone who uses the system, uploads memorable photos, and interacts with the generative AI.
[0772] "Device" means a device used by a User to access the System, upload photos, and view and respond to generated questions and comments, including, for example, a smartphone, tablet, smart glasses, or head-mounted display.
[0773] "Means for uploading images" refers to the process by which a user selects a photo using a terminal and sends it to the server.
[0774] The "server" is a central system that receives and analyzes uploaded photos, generates questions and comments using generative AI, and stores conversation logs.
[0775] "Means for receiving and analyzing images" refers to the process by which the server captures the photos sent by the user and identifies the content of the photos using image analysis algorithms.
[0776] "Generative AI" is an artificial intelligence system that is connected to a server and generates appropriate questions and comments for users based on the results of image analysis.
[0777] "Means for generating questions and comments for users" refers to the process by which the generation AI creates questions and comments for users based on the results of image analysis.
[0778] The "means for presenting the generated questions and comments to the user" is a function by which the terminal displays the questions and comments sent from the server to the user.
[0779] "Means for transmitting a user's response by voice or text from the terminal to the server" refers to a process in which a user responds to a generated question or comment by voice or text and transmits it from the terminal to the server.
[0780] "Means for generating the next question or comment" refers to the process in which the server analyzes the user's response and uses a generation AI to create a question or comment to proceed with the next dialogue.
[0781] "Means of storing all conversation logs and using them to improve the generation AI and provide personalized content based on the user's interests" refers to the process by which the server records and stores all conversation history with the user, and based on that, the generation AI is further improved and personalized content is provided based on the user's interests.
[0782] The present invention provides a system that stimulates the user's brain and maintains their energy by interacting with a generating AI using memorable photos of the user. Specific embodiments of the system are described below.
[0783] System configuration
[0784] This system consists of a user's terminal, a server, a generation AI, and various means for sending and receiving data between them.
[0785] Image upload method
[0786] Users can upload memorable photos to the system using devices such as their smartphones, tablets, smart glasses, and head-mounted displays. Users select images and upload them using the device's application.
[0787] Image analysis methods
[0788] The server receives the uploaded photos and uses image analysis algorithms to analyze the content of the photos, using machine learning frameworks such as TensorFlow and Keras to identify people, objects, and backgrounds in the photos.
[0789] Question and comment generator
[0790] Based on the analysis results, generative AI operates to generate appropriate questions and comments for the user, using advanced generative models such as OpenAI's GPT-3.
[0791] User Presentation Method
[0792] The device presents the generated questions and comments to the user, who responds to the generated questions and comments using text input or voice input.
[0793] Response sending method
[0794] The user's response is sent from the device to the server, and may include, for example, a response such as "These are my daughter and grandchildren."
[0795] Next question generation method
[0796] The server analyzes the user's response, and the AI generates the next question or comment, allowing for a continuous dialogue with the user.
[0797] Conversation log storage method
[0798] All conversation logs are stored on the server and are used to improve the AI generation and provide personalized content based on the user's interests.
[0799] Specific examples
[0800] Example 1: Elderly people
[0801] An elderly person uploads an old family photo using their smartphone. The server receives the photo and uses TensorFlow to analyze the image and recognize that a family member is in it. A generative AI (e.g., GPT-3) generates the question "Who is in this photo?" and the device displays the question to the elderly person. The elderly person responds, "This is my daughter and grandchildren," and sends it to the server. The server analyzes the response and generates the next question: "What a lovely family. Where was this photo taken?"
[0802] Prompt Sentence Examples
[0803] "The Eiffel Tower is in the photo. Generate appropriate questions for the user that will remind them of their travel experiences."
[0804] This method allows users to reminisce and stimulate their brains and maintain their energy through generative AI dialogue. All dialogue logs are saved and used for later analysis. Specifically, it allows for the provision of personalized content based on the user's interests and usage.
[0805] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0806] Step 1:
[0807] The user opens the application on their device, selects a memorable photo, and presses the upload button, which sends the photo from the device to the server.
[0808] Input: A memorable photo
[0809] Process: Upload a file from the device to the server
[0810] Output: Photo files uploaded to the server
[0811] Step 2:
[0812] The server receives the uploaded photos and stores the image files in a database.
[0813] Input: Photo file sent from the device
[0814] Processing: Receiving the file and saving it to the database
[0815] Output: Photo data stored in a database
[0816] Step 3:
[0817] The server analyzes the content of the photo using image analysis algorithms, such as TensorFlow and Keras, to identify people, objects, and the background in the photo.
[0818] Input: Photo data stored in the database
[0819] Processing: Image analysis using TensorFlow and Keras
[0820] Output: Analysis results for the content of the photo (e.g. people's names, places, objects, etc.)
[0821] Step 4:
[0822] The server sends the analysis results to the generation AI, which then uses them to generate appropriate questions and comments for the user. In this example, OpenAI's GPT-3 is used.
[0823] Input: Image analysis results
[0824] Processing: Generative AI generates questions and comments
[0825] Output: Generated questions and comments
[0826] Step 5:
[0827] The server sends the generated questions and comments to the terminal, which displays them to the user.
[0828] Input: Generated questions and comments
[0829] Processing: Transmission from server to terminal, screen display
[0830] Output: Questions and comments displayed to the user
[0831] Step 6:
[0832] The user enters a response to a question or comment. The response can be entered by voice or text. The response is sent from the device to the server.
[0833] Input: User response (voice or text)
[0834] Processing: speech recognition or text input, sending data to server
[0835] Output: The user's response sent to the server
[0836] Step 7:
[0837] The server receives and analyzes the user's response. Based on the analysis results, the AI generates the next question or comment, thus continuing the dialogue with the user.
[0838] Input: User response
[0839] Processing: Analyzing responses and generating the next question or comment using AI
[0840] Output: Next question or comment
[0841] Step 8:
[0842] The server sends the next question or comment to the terminal, which displays it to the user, who responds again, and the process from step 6 to step 8 is repeated.
[0843] Input: Next question or comment
[0844] Processing: Transmission from server to terminal, screen display, user response input
[0845] Output: Repeated user interactions
[0846] Step 9:
[0847] All conversation logs are stored on the server. These logs include user usage and reactions and are used for later analysis. They serve as the basis for improving the AI generation and providing personalized content based on the user's interests.
[0848] Input: User interaction log
[0849] Processing: Data storage and later analysis
[0850] Output: Data used to improve generative AI and provide personalized content
[0851] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0852] This invention provides a system that stimulates the user's brain and maintains their morale by interacting with a generating AI using photos of the user's memories. This system also incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction.
[0853] System Overview
[0854] The process by which a user uploads an image from their device
[0855] Users upload photos to the system using their own devices (e.g., smartphones or tablets). Users can easily select and upload photos by opening the application and going to the upload screen. Uploaded photos are first transferred to the server.
[0856] The process by which the server receives and analyzes the images
[0857] The uploaded photos are received by the server and stored in a database, where they are then analyzed using image analysis algorithms to automatically identify people, objects, and background information in the photos.
[0858] The process by which generative AI generates questions and comments for users
[0859] Based on the results of image analysis, the server sends the necessary data to the generation AI. The generation AI then generates appropriate questions and comments for the user based on the content of the photo. For example, it creates questions such as "Where was this photo taken?" or "Who is this person?"
[0860] The process by which the device presents the generated questions and comments to the user
[0861] The questions and comments generated by the AI are sent back from the server to the device, which then displays them to the user. The user can then confirm the questions and comments and enter a response.
[0862] The process of sending user responses from the terminal to the server
[0863] The response entered by the user is sent from the device to the server, and may include, for example, "These are my daughter and grandchildren."
[0864] The server analyzes the response, and the AI generates the next question or comment.
[0865] The server receives and analyzes the user's response, and the generation AI generates the next question or comment. This continues the conversation with the user and advances the dialogue process. For example, the next question might be, "Where was this photo taken?"
[0866] Emotion recognition process by emotion engine
[0867] The emotion engine analyzes the user's emotions from images uploaded by the user and responses during conversations. The emotion engine uses facial expression analysis and voice analysis to determine whether the user is happy or anxious.
[0868] Tuning process for emotion-based generative AI
[0869] Based on the analysis results of the emotion engine, the generation AI will adjust the questions and comments. For example, if the user is determined to be sad, the generation AI will generate comforting comments. If the user is having fun, the generation AI will generate questions that bring out that emotion.
[0870] Continuing conversations and logging processes
[0871] This conversation process is repeated until the user is satisfied. All conversation logs are stored on the server and used to analyze user usage and reactions at a later date. The results of this analysis are used to improve the system and enhance the user experience.
[0872] Specific examples
[0873] Example 1: Lonely elderly people
[0874] User: Selects and uploads an old family photo.
[0875] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[0876] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[0877] Terminal: displays the question to the user.
[0878] User: Responds to the question by saying, "These are my daughter and grandchildren."
[0879] Terminal: Sends the response to the server.
[0880] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[0881] Emotion engine: Analyzes the user's emotions from images and responses and recognizes that the user is feeling nostalgic.
[0882] Generative AI: Based on the user's emotions, the system further adjusts the questions and comments, such as "Tell me more about your memories of that time."
[0883] Terminal: Display the next question to the user.
[0884] Example 2: Travel photos
[0885] User: Select and upload photos of travel destinations.
[0886] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[0887] Generative AI: Generates questions like "Where is this place?" based on a photo.
[0888] Terminal: displays the question to the user.
[0889] User: Responds to the question with "This is a photo from Paris, where I went last year."
[0890] Terminal: Sends the response to the server.
[0891] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[0892] Emotion engine: Analyzes the user's emotions from images and responses and recognizes what the user is enjoying.
[0893] Generative AI: Further tailor questions and comments based on user sentiment, such as "What other places have you been?"
[0894] Terminal: Display the next question to the user.
[0895] In this way, this system allows users to enjoy conversations through photos, stimulating their brains and maintaining their energy. Furthermore, the introduction of an emotion engine enables communication that responds to the user's emotions, resulting in more effective dialogue.
[0896] The processing flow will be explained below.
[0897] Step 1:
[0898] The user opens the application on their device and proceeds to the photo upload screen.
[0899] Step 2:
[0900] The user selects a photo on their device and clicks the upload button.
[0901] Step 3:
[0902] The terminal generates a request to transfer the selected photo to the server.
[0903] Step 4:
[0904] The server receives the request from the terminal and receives the photo data.
[0905] Step 5:
[0906] The server stores the received photos in a database.
[0907] Step 6:
[0908] The server passes the stored photos to an image analysis algorithm, which begins the analysis process.
[0909] Step 7:
[0910] The server performs image analysis to determine the content of the photo, such as identifying people, objects, and the background.
[0911] Step 8:
[0912] The server sends the analysis results to the generation AI.
[0913] Step 9:
[0914] Based on the results of image analysis, the generative AI generates appropriate questions and comments for the user, such as "Where was this photo taken?" or "Who is this person?"
[0915] Step 10:
[0916] The server sends the generated questions and comments back to the device.
[0917] Step 11:
[0918] The terminal displays the questions and comments received from the server to the user.
[0919] Step 12:
[0920] The user enters a response to the question or comment on the terminal.
[0921] Step 13:
[0922] The terminal generates a request to send the user's response to the server.
[0923] Step 14:
[0924] The server receives the response from the terminal.
[0925] Step 15:
[0926] The server passes the received response to the emotion engine, which analyzes the user's emotion.
[0927] Step 16:
[0928] The emotion engine recognizes emotions from user responses and images and sends the results to the generation AI.
[0929] Step 17:
[0930] The generative AI generates the next appropriate question or comment based on the emotion recognition results from the emotion engine and the user's response. For example, if the user is feeling nostalgic, it generates a question such as, "Tell me more about your memories from that time."
[0931] Step 18:
[0932] The server sends the next generated question or comment back to the device.
[0933] Step 19:
[0934] The device will display the next question or comment to the user.
[0935] Step 20:
[0936] The user again enters a response, and this process is repeated until the user is satisfied.
[0937] Step 21:
[0938] The server stores all conversation logs.
[0939] Step 22:
[0940] The server will later analyze the saved conversation logs to evaluate the user's usage and reactions.
[0941] Example 2
[0942] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0943] In modern society, the number of elderly people and individuals who feel lonely is increasing, making it important to improve their mental health and increase communication opportunities for these people. In particular, there is a demand for systems that can provide mental support by promoting interaction through past memories and photos and engaging in emotionally appropriate dialogue. Conventional systems struggle to improve the quality of emotion recognition and dialogue, and there is a lack of methods for achieving effective communication.
[0944] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0945] In this invention, the server includes means for receiving images and storing them in a database, means for generating image analysis results using an image analysis algorithm, means for analyzing the user's responses and generating the next question or comment using a generation AI, means for an emotion engine to recognize the user's emotions and for the generation AI to adjust the questions or comments based on the results, and means for saving all conversation logs and analyzing them later. This enables natural conversation based on the user's memories, provides appropriate communication according to emotions, and helps maintain and improve the user's mental health.
[0946] "Means for uploading images" refers to a series of operations and functions that allow a user to select digital images from their terminal and send them to the server.
[0947] A "database" is a system for systematically storing and managing digital information, where images and associated metadata are stored.
[0948] An "image analysis algorithm" is a computational method that analyzes the content of a digital image and extracts information about the people, objects, background, and other aspects of the photograph.
[0949] "Generative AI" is artificial intelligence that automatically generates questions and comments based on user responses and image analysis results.
[0950] An "emotion engine" is a technology that analyzes a user's emotions and adjusts the output of the generative AI based on their emotional state.
[0951] A "terminal" is a digital device used by a user, such as a smartphone or tablet.
[0952] A "conversation log" is a record of all interactions between a user and a system, data that is stored for later analysis.
[0953] A "prompt sentence" is the input text given to the generation AI, which is used to generate the next question or comment based on the user's response and the results of image analysis.
[0954] "Means for generating questions and comments" refers to the functions and operations by which the generation AI creates appropriate questions and comments for the user based on the prompt text.
[0955] This invention is a system that aims to stimulate the user's brain and maintain their energy by having them interact with a generative AI using photos of their memories. This system incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction.
[0956] System Overview
[0957] User-uploaded images
[0958] Users upload photos to the system using devices such as smartphones or tablets. Users launch a dedicated application, select memorable photos on the upload screen, and send them to the server. The uploaded photos are then transferred from the device to the server.
[0959] Receiving and analyzing images by the server
[0960] The server receives the uploaded photos and stores them in a database. The server then analyzes the content of the photos using image analysis algorithms (such as Google Cloud Vision API or Microsoft Azure's Computer Vision API). The analysis results include information about the people, objects, and background in the photos.
[0961] Question generation using generative AI
[0962] Based on the results of image analysis, the server sends the necessary data to a generative AI (such as OpenAI's GPT-4). Based on the received data, the generative AI generates appropriate questions and comments for the user. For example, it creates specific questions such as "Where was this photo taken?" or "Who is this person?"
[0963] Display of questions on the device
[0964] The generated AI's questions and comments are sent from the server to the user's device and displayed on the device. The user reads the displayed questions and comments and enters a response.
[0965] Parsing the user's response and generating the next question
[0966] The response entered by the user is sent from the device to the server. The server analyzes the received response and sends the results to the generation AI. The generation AI then generates the next most appropriate question or comment. This allows for a continuous dialogue with the user.
[0967] Emotion recognition and output adjustment by emotion engine
[0968] The emotion engine analyzes the user's emotions from images uploaded by the user and responses during conversations. The emotion engine evaluates whether the user is happy or anxious through facial expression and voice analysis. Based on the results of the emotion engine, the server adjusts the prompts sent to the generation AI and generates appropriate questions and comments.
[0969] Specific examples
[0970] Example 1: Lonely elderly people
[0971] The user selects and uploads an old family photo. The server receives the photo and uses an image analysis algorithm to recognize that a family member is in it. The generation AI generates the question, "Who is in this photo?" and displays it on the device. The user responds, "This is my daughter and grandchildren," and the device sends the response to the server. The server analyzes the response, and the generation AI generates the next question, "What a lovely family. Where was this photo taken?" Furthermore, if the emotion engine senses the user's nostalgia, it generates a question such as, "Tell me more about your memories from that time."
[0972] Example 2: Travel photos
[0973] The user selects and uploads photos of their travel destinations. The server receives the photos and uses an image analysis algorithm to recognize the travel destinations. The generation AI generates the question "Where is this place?" and displays it on the device. The user responds, "This is a photo of Paris, where I went last year," and the device sends the response to the server. The server analyzes the response, and the generation AI generates the next question, "How was Paris? Were there any places that particularly impressed you?" The emotion engine analyzes the user's emotions and, if it identifies that the user is enjoying themselves, generates questions such as "What other places did you go?"
[0974] Prompt Sentence Examples
[0975] As an example of a prompt for a generative AI model, input the following text to the generative AI:
[0976] Analyzing a photo submitted by a user reveals that it contains family members. Based on that content, generate questions or comments to naturally continue the conversation with the user. For example, "Who is in this photo?"
[0977] This system provides natural dialogue through users' memorable photos and realizes appropriate communication according to their emotions, which is expected to contribute to maintaining and improving users' mental health.
[0978] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0979] Step 1:
[0980] The user initiates the image upload process from their device.
[0981] Specific operation: The user launches the application on their smartphone or tablet, selects "Upload photos" from the menu, selects a memorable photo from their photo library, and presses the "Upload" button.
[0982] Input: A photo file selected by the user.
[0983] Output: The photo files are saved on your device and ready to be sent to the server.
[0984] Step 2:
[0985] The device sends the photo to the server.
[0986] Specific operation: The device makes an HTTP request to send the user-selected photo file to the server. The request includes the photo file and associated metadata (e.g., upload date and time, user ID).
[0987] Input: A user-selected photo file and its metadata.
[0988] Output: The photo files and metadata are sent to the server.
[0989] Step 3:
[0990] The server receives the photos and stores them in a database.
[0991] What it does: The server parses the incoming HTTP request and stores the photo file and its metadata in a database, along with information about the photo's identifier and storage location.
[0992] Input: The submitted photo file and its metadata.
[0993] Output: Photo files and associated metadata stored in a database.
[0994] Step 4:
[0995] The server analyzes the photo using image analysis algorithms.
[0996] How it works: The server sends the stored photo file to an image analysis algorithm (e.g., Google Cloud Vision API) to analyze the content of the photo, including people, objects, and background information.
[0997] Input: Photo files stored in the database.
[0998] Output: The analysis data generated as a result of the image analysis algorithm.
[0999] Step 5:
[1000] The generative AI creates prompts to generate questions and comments for the user.
[1001] Specific operation: The server creates a prompt sentence to send to the generation AI based on the image analysis results. This prompt sentence contains detailed information to facilitate the dialogue with the user. For example, "A family photo has been uploaded. Please generate a question about this photo."
[1002] Input: Image analysis results.
[1003] Output: The prompt sent to the generation AI.
[1004] Step 6:
[1005] Generative AI generates questions and comments.
[1006] How it works: A generative AI model (e.g., OpenAI's GPT-4) generates questions or comments to present to the user based on prompts received from the server, such as "Who is in this photo?"
[1007] Input: The server-generated prompt text.
[1008] Output: Generated questions and comments.
[1009] Step 7:
[1010] The server sends the generated questions and comments to the device.
[1011] Specific operation: The server creates an HTTP response to send the generated question or comment to the terminal, and sends it to the terminal.
[1012] Input: Generated questions and comments.
[1013] Output: Questions and comments sent to the user's device.
[1014] Step 8:
[1015] The device displays the question or comment to the user.
[1016] Specific operation: The device displays the received questions and comments on the screen and presents them to the user. An interface is provided so that the user can review them.
[1017] Input: Questions and comments sent by the server.
[1018] Output: Questions and comments displayed on the terminal screen.
[1019] Step 9:
[1020] The user enters a response to the question.
[1021] Specific behavior: The user reads the questions and comments displayed on the device and enters a response. For example, "These are my daughter and grandchildren."
[1022] Input: Questions or comments displayed on the device.
[1023] Output: The response text entered by the user.
[1024] Step 10:
[1025] The terminal sends the user's response to the server.
[1026] Specific operation: The terminal creates an HTTP request to send the response text entered by the user to the server.
[1027] Input: The response text entered by the user.
[1028] Output: The response text sent to the server.
[1029] Step 11:
[1030] The server analyzes the user's response and creates and sends the next prompt to the generation AI.
[1031] Specific operation: The server analyzes the user's response text, creates a new prompt based on its content, and sends the created prompt to the generation AI to generate the next question or comment.
[1032] Input: The user's response text.
[1033] Output: The next prompt sent to the generation AI.
[1034] Step 12:
[1035] Generative AI generates the next question or comment.
[1036] Specific behavior: The generative AI generates the next question or comment based on the newly received prompt, for example, "Where was this photo taken?"
[1037] Input: The newly created prompt statement.
[1038] Output: The next question or comment.
[1039] Step 13:
[1040] The emotion engine recognizes the user's emotions and adjusts the output of the generative AI.
[1041] Specific operation: The emotion engine analyzes the user's facial expressions and responses to evaluate the user's emotional state. Based on the results, it adjusts the generation AI to generate appropriate questions and comments.
[1042] Input: User's facial expression data and response content.
[1043] Output: Adjusted prompt and output of the generation AI.
[1044] Step 14:
[1045] The server sends the tailored questions and comments to the device.
[1046] Specific operation: The server sends questions and comments adjusted by the emotion engine to the terminal.
[1047] Input: Moderated questions and comments.
[1048] Output: The tailored questions and comments sent to your device.
[1049] Step 15:
[1050] The device displays the tailored question or comment to the user.
[1051] Specific operation: The device displays the adjusted questions and comments on the screen and presents them to the user.
[1052] Input: Moderated questions and comments.
[1053] Output: The adjusted questions and comments displayed on the screen.
[1054] Step 16:
[1055] The server stores all conversation logs.
[1056] What it does: The server records all interactions and stores them in a database. This information is later analyzed and used to improve the system and enhance the user experience.
[1057] Input: All dialogue content (questions, responses, and adjustment results).
[1058] Output: Conversation logs stored in a database.
[1059] (Application example 2)
[1060] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1061] The problem that this invention aims to solve is to realize more effective communication that responds to the user's emotions while stimulating the user's brain and maintaining their energy through dialogue using photos related to the user's memories. Conventional systems simply continue dialogue mechanically without taking the user's emotions into consideration, which has the problem of degrading the quality of the user experience.
[1062] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1063] In this invention, the server includes means for a user to upload images from a terminal, means for the server to receive and analyze the images, means for the generation AI to generate questions and comments for the user based on the generated analysis results, means for the terminal to present the generated questions and comments to the user, means for the terminal to send the user's response to the server, means for the server to analyze the response and for the generation AI to generate the next question or comment, means for analyzing emotions from the user's facial expressions and voice, means for the generation AI to adjust the content of the dialogue based on the emotion analysis results, means for continuously conversing with the user, and means for saving all conversation logs and analyzing them later. This allows the user to interact with the generation AI through memorable photos, and the content of the dialogue is appropriately adjusted through emotion analysis, thereby stimulating the user's brain and maintaining their energy.
[1064] "Means for users to upload images from their devices" refers to the function that allows users to send image data to the system using devices such as smartphones and tablets.
[1065] "Means for the server to receive and analyze images" refers to the function of the cloud server receiving uploaded image data and analyzing its contents using a specified image analysis algorithm.
[1066] "Means for the generation AI to generate questions and comments for the user based on the generated analysis results" refers to the function of using natural language processing technology to generate appropriate questions and comments based on the content of the image analyzed by the server.
[1067] "Means for the device to present the generated questions and comments to the user" refers to a function that displays the generated questions and comments on the user's smartphone or tablet.
[1068] "Means for transmitting a user's response from the terminal to the server" refers to a function for transmitting the response data entered by the user again to the cloud server.
[1069] "The server analyzes the response, and the generation AI generates the next question or comment" refers to the function in which the cloud server analyzes the user's response and generates the next conversation content based on the analysis results.
[1070] "Means for analyzing emotions from the user's facial expressions and voice" refers to a function that analyzes the facial expressions and voice data shown by the user during a conversation and determines their emotional state (joy, sadness, surprise, etc.).
[1071] "Means for the generation AI to adjust the content of the dialogue based on the results of emotion analysis" refers to the function by which the generation AI takes into account the user's emotional state obtained through emotion analysis and appropriately adjusts the next question or comment.
[1072] "Means for continuing conversation with the user" refers to a function that allows a dialogue between the user and the system to continue.
[1073] "A means of saving all conversation logs and analyzing them later" refers to a function that records and saves all generated conversation content and later analyzes changes in user behavior patterns and emotions based on that data.
[1074] This invention provides a system that stimulates the user's brain and maintains their morale by interacting with a generating AI using memorable photos of the user. This system also incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction. Each element of this system is explained below.
[1075] System configuration
[1076] 1. A way for users to upload images from their devices
[1077] Users use their smartphones, tablets, or other devices to select and upload images within the application, a process designed to make it easy and intuitive for users to submit photos to the system.
[1078] 2. How the server receives and analyzes images
[1079] Uploaded photos are transferred to a cloud server, which uses image analysis libraries such as the Google Vision API to automatically analyze the content of the photo and extract information about people, places, objects, and more.
[1080] 3. A means for the generative AI to generate questions and comments for users based on the generated analysis results
[1081] The results of the image analysis are sent to a generative AI (e.g., GPT-3 or GPT-4), which uses this data to generate appropriate questions and comments for the user.
[1082] For example: "Where was this photo taken?"
[1083] 4. A means for the device to present generated questions and comments to the user
[1084] The questions and comments generated by the AI are sent from the server to the user's device, where they are displayed to the user, and a dialogue begins.
[1085] 5. A means of transmitting user responses from the terminal to the server
[1086] When a user responds to a question or comment through their device, the response is sent to the cloud server.
[1087] 6. The server analyzes the response and the AI generates the next question or comment.
[1088] The server analyzes the user's responses and sends the results to the generation AI, which then generates new questions and comments to continue the dialogue with the user.
[1089] 7. Means of analyzing emotions from user facial expressions and voice
[1090] Using an emotion engine (for example, Microsoft Azure Emotion API), emotions are analyzed from the user's facial expressions and voice. The analysis results indicate the user's emotional state, such as whether they are happy or sad.
[1091] 8. How generative AI can adjust dialogue content based on emotion analysis results
[1092] Based on the analysis results of the emotion engine, the generative AI will adjust the dialogue appropriately. For example, if the user is sad, it will generate comforting comments, and if the user is happy, it will generate questions that will bring out that emotion.
[1093] For example: "Tell me more about your memories of that time."
[1094] 9. A way to continue the conversation with your users
[1095] The dialogue continues until the user is satisfied. The entire conversation process is controlled by the system.
[1096] 10. A way to store all conversation logs and analyze them later
[1097] All generated conversation logs are stored on a cloud server, allowing users' usage patterns and reactions to be analyzed at a later date, helping to improve the system and enhance the user experience.
[1098] This system allows users to enjoy conversations through memorable photos, stimulating their brains and maintaining their energy. In addition, the introduction of an emotion engine enables communication based on the user's emotions, resulting in more effective conversations.
[1099] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1100] Step 1:
[1101] A way for users to upload images from their devices
[1102] A user opens the application on a smartphone or tablet, selects a photo from the upload screen, and sends it to the system. The input is the photo file selected by the user, and the output is that the photo file is sent to the cloud server. The photo file is transferred between the device and the cloud server.
[1103] Step 2:
[1104] The server receives and analyzes the images
[1105] The server receives the photos sent from the device and stores them in a database. The stored photos are then passed to the Google Vision API for image analysis. The input is the uploaded photo file, and the output is the analyzed image content (people, places, objects, etc.) data. Image analysis uses machine learning algorithms to analyze the content of the photo.
[1106] Step 3:
[1107] A means for the AI to generate questions and comments for users based on the generated analysis results
[1108] The server sends the image analysis results to a generation AI (GPT-3 or GPT-4), which generates appropriate questions and comments. The input is the image analysis results, and the output is questions and comments to present to the user. Specifically, the analysis results are input as prompts to the generation AI, which then uses natural language processing technology to generate corresponding questions and comments.
[1109] Step 4:
[1110] A means by which the device presents generated questions and comments to the user
[1111] The server sends the generated questions and comments to the terminal and displays them to the user. The input is the generated questions and comments, and the output is what is displayed to the user. The terminal receives this and displays it on the screen.
[1112] Step 5:
[1113] A means for transmitting user responses from the terminal to the server
[1114] Users use their devices to respond to questions and comments, and then send the response data to the cloud server. The input is the user's response, and the output is the response data sent to the cloud server. Specifically, the user enters text and presses the send button, which sends the data to the server.
[1115] Step 6:
[1116] The server analyzes the response and the generation AI generates the next question or comment.
[1117] The server analyzes the user's response data and sends the analysis results to the generation AI to generate the next question or comment. The input is the user's response data, and the output is the newly generated question or comment. The response data is analyzed for text, and its content is provided to the generation AI as a prompt.
[1118] Step 7:
[1119] A means of analyzing emotions from the user's facial expressions and voice
[1120] The server receives the user's facial expression and voice data sent from the device and performs emotion analysis using an emotion engine (Microsoft Azure Emotion API). The input is facial expression and voice data, and the output is data on the user's emotional state (e.g., joy, sadness, surprise, etc.). Specifically, image and voice data is passed to the emotion engine for analysis.
[1121] Step 8:
[1122] A means for generative AI to adjust dialogue content based on emotion analysis results
[1123] The server provides the emotion analysis results to the generation AI, which then adjusts the dialogue appropriately based on the results. The input is the emotion analysis results, and the output is adjusted questions or comments. The generation AI receives prompts containing the emotion analysis results and generates appropriate dialogue content based on them. For example, if it determines that the user is sad, it generates a question such as, "Tell me more about your memories of that time."
[1124] Step 9:
[1125] A way to maintain an ongoing conversation with users
[1126] The server and the generation AI work together to continue the dialogue until the user is satisfied. The input is continuous responses from the user, and the output is newly generated questions and comments. The dialogue content is continuously generated, and the interaction with the user progresses without interruption.
[1127] Step 10:
[1128] A means to store all conversation logs and analyze them later
[1129] The server stores all dialogue logs and later analyzes changes in user behavior patterns and emotions based on them. The input is all the dialogue logs generated, and the output is the stored data and the analysis results. The log data is stored in a database and later analyzed using analysis tools.
[1130] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1131] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1132] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1133] [Third embodiment]
[1134] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1135] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1136] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1137] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1138] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1139] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1140] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1141] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1142] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1143] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1144] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1145] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1146] This invention provides a system that stimulates the user's brain and maintains their energy by interacting with a generating AI using photos of the user's memories. This system involves a multi-stage process in which the user uploads photos and looks back on their memories through conversations with the generating AI.
[1147] System Overview
[1148] The process by which a user uploads an image from their device
[1149] Users upload photos to the system using their own devices (e.g., smartphones or tablets). Users can easily select and upload photos by opening the application and going to the upload screen. Uploaded photos are first transferred to the server.
[1150] The process by which the server receives and analyzes the images
[1151] The uploaded photos are received by the server and stored in a database, where they are then analyzed using image analysis algorithms to automatically identify people, objects, and background information in the photos.
[1152] The process by which generative AI generates questions and comments for users
[1153] Based on the results of image analysis, the server sends the necessary data to the generation AI. The generation AI then generates appropriate questions and comments for the user based on the content of the photo. For example, it creates questions such as "Where was this photo taken?" or "Who is this person?"
[1154] The process by which the device presents the generated questions and comments to the user
[1155] The questions and comments generated by the AI are sent back from the server to the device, which then displays them to the user. The user can then confirm the questions and comments and enter a response.
[1156] The process of sending user responses from the terminal to the server
[1157] The response entered by the user is sent from the device to the server, and may include, for example, "These are my daughter and grandchildren."
[1158] The server analyzes the response, and the AI generates the next question or comment.
[1159] The server receives and analyzes the user's response, and the generation AI generates the next question or comment. This continues the conversation with the user and advances the dialogue process. For example, the next question might be, "Where was this photo taken?"
[1160] Continuing conversations and logging processes
[1161] This conversation process is repeated until the user is satisfied. All conversation logs are stored on the server and used later to analyze user usage and reactions.
[1162] Specific examples
[1163] Example 1: Lonely elderly people
[1164] User: Selects and uploads an old family photo.
[1165] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[1166] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[1167] Terminal: displays the question to the user.
[1168] User: Responds to the question by saying, "These are my daughter and grandchildren."
[1169] Terminal: Sends the response to the server.
[1170] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[1171] Terminal: Display the next question to the user.
[1172] Example 2: Travel photos
[1173] User: Select and upload photos of travel destinations.
[1174] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[1175] Generative AI: Generates questions like "Where is this place?" based on a photo.
[1176] Terminal: Display the question to the user.
[1177] User: Responds to the question with "This is a photo from Paris, where I went last year."
[1178] Terminal: Sends the response to the server.
[1179] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[1180] Terminal: Display the next question to the user.
[1181] As described above, through this system, users can enjoy conversations through photos, which helps stimulate the brain and maintain energy.
[1182] The processing flow will be explained below.
[1183] Step 1:
[1184] The user opens the application on their device and proceeds to the photo upload screen.
[1185] Step 2:
[1186] The user selects a photo on their device and clicks the upload button.
[1187] Step 3:
[1188] The terminal generates a request to transfer the selected photo to the server.
[1189] Step 4:
[1190] The server receives the request from the terminal and receives the photo data.
[1191] Step 5:
[1192] The server stores the received photos in a database.
[1193] Step 6:
[1194] The server passes the stored photos to an image analysis algorithm, which begins the analysis process.
[1195] Step 7:
[1196] The server performs image analysis to determine the content of the photo, such as identifying people, objects, and the background.
[1197] Step 8:
[1198] The server sends the image analysis results to the generation AI.
[1199] Step 9:
[1200] Based on the results of image analysis, the generative AI generates appropriate questions and comments for the user, such as "Where was this photo taken?" or "Who is this person?"
[1201] Step 10:
[1202] The server sends the generated questions and comments back to the device.
[1203] Step 11:
[1204] The terminal displays the questions and comments received from the server to the user.
[1205] Step 12:
[1206] The user enters a response to the question or comment on the terminal.
[1207] Step 13:
[1208] The terminal generates a request to send the user's response to the server.
[1209] Step 14:
[1210] The server receives the response from the terminal.
[1211] Step 15:
[1212] The server analyzes the received response and sends it to the generating AI.
[1213] Step 16:
[1214] Based on the user's response, the generative AI generates the next appropriate question or comment, such as "Where was this photo taken?"
[1215] Step 17:
[1216] The server sends the next generated question or comment back to the device.
[1217] Step 18:
[1218] The device will display the next question or comment to the user.
[1219] Step 19:
[1220] The user again enters a response, and this process is repeated until the user is satisfied.
[1221] Step 20:
[1222] The server stores all conversation logs.
[1223] Step 21:
[1224] The server will later analyze the saved conversation logs to evaluate the user's usage and reactions.
[1225] Example 1
[1226] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1227] Today's elderly and lonely people lack emotional communication and brain activation. As a result, they face problems such as a decline in cognitive function and loss of energy. Conventional technologies have not provided effective solutions to these issues, particularly lacking methods for enjoying conversations while reminiscing about memories. Therefore, there is a need for a system that allows users to engage in emotional communication through their own memories, activating their brains, and maintaining their energy.
[1228] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1229] In this invention, the server includes means for receiving and storing images, means for analyzing images, and means for generating questions and comments for the user, thereby generating dialogue based on images uploaded by the user from the terminal, and enabling continuous conversation with the user.
[1230] "Users" refer to people who use the system to upload their own photos and interact with the generative AI.
[1231] "Device" refers to the electronic device, such as a smartphone or tablet, that a user uses to access the system, upload photos, and respond to questions and comments from the Generative AI.
[1232] "Server" refers to the central device of the system, which receives and stores images uploaded by users, analyzes the images, generates questions and comments using generative AI, and sends them to the terminal.
[1233] "Generative AI" refers to an artificial intelligence model used by the server to generate new questions or comments based on image analysis results and user responses, such as an AI model that performs natural language processing.
[1234] "Image analysis" refers to the process of analyzing uploaded images using machine learning algorithms to automatically recognize information such as people, objects, and backgrounds in the photos.
[1235] "Question and comment generation" refers to the process in which the generative AI creates appropriate questions and comments for the user based on the results of image analysis and the user's responses.
[1236] "Conversation logs" refer to records of interactions with users on the system, and are data stored on the server. They include the content of the conversations and are used for later analysis.
[1237] "Database" refers to a system that systematically stores and manages data such as images and conversation logs received by a server.
[1238] "Image upload" refers to the process by which a user sends a photo file from their device to a server.
[1239] "Continuous dialogue" refers to a process in which the generative AI generates questions and comments one after another until the user is satisfied, maintaining an uninterrupted conversation with the user.
[1240] "Analysis results" refers to information about the content of a photo obtained through image analysis, including the recognition of people, objects, backgrounds, etc.
[1241] "Storage" refers to the process of storing data such as photos and conversation logs in a database on the server.
[1242] The present invention provides a system that allows users to upload memorable photos from their own devices and review the photos through dialogue with a generation AI, thereby stimulating the user's brain and maintaining their energy. An embodiment of this system is described in detail below.
[1243] The process by which a user uploads an image from their device
[1244] Users access the application using a device such as a smartphone or tablet. They open the photo upload screen within the application and select their memorable photos. When the user presses the "Upload" button, the photo file is sent to the server via the Internet.
[1245] The process by which the server receives and stores images
[1246] The server receives the photos sent by the user and stores them in a database. When saved, the photos are assigned a unique ID, which makes it easier to manage the photo data.
[1247] The process by which the server analyzes the image
[1248] The server runs an image analysis algorithm (such as YOLO or OpenCV) on the received photo to obtain information about people, objects, and backgrounds in the photo. The results of this analysis are stored in a database for later use.
[1249] The process by which generative AI generates questions and comments
[1250] The server sends data based on the image analysis results to a generative AI model (such as GPT-4). The generative AI uses this data to generate appropriate questions and comments for the user. For example, prompts such as "Where was this photo taken?" or "Who is this person?" are generated.
[1251] The process by which the device presents the generated questions and comments to the user
[1252] The questions and comments generated by the AI are sent from the server to the user's device, where they are displayed. The user can then confirm the questions and comments displayed on the device and enter their responses.
[1253] The process of sending the user's response to the server
[1254] The user's response is sent from the device to the server, and may contain text such as "This is my daughter and grandchildren."
[1255] The server analyzes the response, and the AI generates the next question or comment.
[1256] The server analyzes the user's response, and the AI generates the next question or comment based on the results, creating a continuous dialogue with the user.
[1257] Continuing conversations and logging processes
[1258] The dialogue process is repeated until the user is satisfied. All conversation logs are stored on the server and used for later analysis. This data provides valuable information for analyzing user usage and reactions.
[1259] Specific examples
[1260] Example 1: Lonely elderly people
[1261] User: Selects and uploads an old family photo.
[1262] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[1263] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[1264] Terminal: Display the question to the user.
[1265] User: Responds to the question by saying, "These are my daughter and grandchildren."
[1266] Terminal: Sends the response to the server.
[1267] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[1268] Terminal: Display the next question to the user.
[1269] Example 2: Travel photos
[1270] User: Select and upload photos of travel destinations.
[1271] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[1272] Generative AI: Generates questions like "Where is this place?" based on a photo.
[1273] Terminal: Display the question to the user.
[1274] User: Responds to the question with "This is a photo from Paris, where I went last year."
[1275] Terminal: Sends the response to the server.
[1276] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[1277] Terminal: Display the next question to the user.
[1278] Through this process, users can enjoy conversation while reminiscing on their memories, which helps to stimulate the brain and maintain energy.
[1279] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1280] Step 1:
[1281] The user accesses the application using a terminal and opens the photo upload screen. Then, the user selects the photo to upload and presses the "Upload" button. This action sends the photo file to the server over the Internet.
[1282] Input: A photo file selected from the user's device.
[1283] Output: The photo file sent to the server.
[1284] Specific operation: The user launches the application, logs in, selects a photo, and presses the "Upload" button.
[1285] Step 2:
[1286] The server saves the received photo files and assigns a unique ID to each photo, which is then stored in a database.
[1287] Input: Photo files sent over the internet to the server.
[1288] Output: Photo file stored in the database with a unique ID.
[1289] Specific operation: The server receives the photo file, saves it as binary data in its internal memory, and stores it in a database.
[1290] Step 3:
[1291] The server runs image analysis algorithms (such as YOLO or OpenCV) on the stored photos to extract information about people, objects, and backgrounds in the photos, and stores the results in a database.
[1292] Input: Photo files stored in the database.
[1293] Output: Data containing image analysis results (e.g. people, objects, and background in a photo).
[1294] What happens: The server runs an image analysis algorithm to identify the content of the photo.
[1295] Step 4:
[1296] The server sends data based on the image analysis results to a generation AI (for example, GPT-4), which then uses this data to generate questions and comments for the user.
[1297] Input: Image analysis results.
[1298] Output: Questions and comments generated by the generative AI.
[1299] Specific operation: The server converts the analysis result into a prompt sentence and sends it to the generation AI, and the AI model generates a response.
[1300] Step 5:
[1301] The server sends questions and comments from the generated AI to the user's device, which then displays the received questions and comments to the user.
[1302] Input: Questions and comments generated by the generative AI.
[1303] Output: Questions and comments displayed on the user's device.
[1304] Specific operation: The server sends the generated questions and comments to the terminal, which displays them to the user.
[1305] Step 6:
[1306] The user inputs their response to the questions and comments displayed on the terminal, and when they are finished, they press the "Send" button to send the response to the server.
[1307] Input: Questions and comments from the generative AI, and user responses.
[1308] Output: The user's response sent from the terminal to the server.
[1309] Specific behavior: The user reads the question or comment, enters their response, and presses the "Submit" button.
[1310] Step 7:
[1311] The server analyzes the user's response, and the AI generates the next question or comment based on the results. This process is repeated until the user is satisfied.
[1312] Input: The user's response.
[1313] Output: The next question or comment generated by the generative AI.
[1314] Specific operation: The server analyzes the user's response and sends the analysis results to the generation AI, and the AI model generates the next question or comment.
[1315] Step 8:
[1316] All conversation logs are stored on the server and used for later analysis. This data provides information for analyzing user usage and reactions.
[1317] Input: A log of the conversation between the user and the generated AI.
[1318] Output: Saved conversation logs.
[1319] Specific operation: The server stores the conversation log in a database and uses it for later analysis.
[1320] The above are the specific processing steps in this system.
[1321] (Application example 1)
[1322] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1323] To help elderly people and users who feel lonely to activate their brains and maintain their morale, we provide a system that provides emotional and cognitive stimulation by having users interact with a generative AI using photos of their memories. This system allows users to enjoy dialogue through photos, and we also create an environment where the generative AI can provide personalized content.
[1324] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1325] In this invention, the server includes a means for uploading images from the user's device, a means for receiving and analyzing the images, a means for transmitting the user's responses as voice or text from the device to the server, and a means for saving all conversation logs and using them to improve the AI generation and provide personalized content based on the user's interests. This helps to stimulate the user's brain and maintain their energy, and enables the AI generation to provide personalized dialogue based on the user's interests.
[1326] A "user" is someone who uses the system, uploads memorable photos, and interacts with the generative AI.
[1327] "Device" means a device used by a User to access the System, upload photos, and view and respond to generated questions and comments, including, for example, a smartphone, tablet, smart glasses, or head-mounted display.
[1328] "Means for uploading images" refers to the process by which a user selects a photo using a terminal and sends it to the server.
[1329] The "server" is a central system that receives and analyzes uploaded photos, generates questions and comments using generative AI, and stores conversation logs.
[1330] "Means for receiving and analyzing images" refers to the process by which the server captures the photos sent by the user and identifies the content of the photos using image analysis algorithms.
[1331] "Generative AI" is an artificial intelligence system that is connected to a server and generates appropriate questions and comments for users based on the results of image analysis.
[1332] "Means for generating questions and comments for users" refers to the process by which the generation AI creates questions and comments for users based on the results of image analysis.
[1333] The "means for presenting the generated questions and comments to the user" is a function by which the terminal displays the questions and comments sent from the server to the user.
[1334] "Means for transmitting a user's response by voice or text from the terminal to the server" refers to a process in which a user responds to a generated question or comment by voice or text and transmits it from the terminal to the server.
[1335] "Means for generating the next question or comment" refers to the process in which the server analyzes the user's response and uses a generation AI to create a question or comment to proceed with the next dialogue.
[1336] "Means of storing all conversation logs and using them to improve the generation AI and provide personalized content based on the user's interests" refers to the process by which the server records and stores all conversation history with the user, and based on that, the generation AI is further improved and personalized content is provided based on the user's interests.
[1337] The present invention provides a system that stimulates the user's brain and maintains their energy by interacting with a generating AI using memorable photos of the user. Specific embodiments of the system are described below.
[1338] System configuration
[1339] This system consists of a user's terminal, a server, a generation AI, and various means for sending and receiving data between them.
[1340] Image upload method
[1341] Users can upload memorable photos to the system using devices such as their smartphones, tablets, smart glasses, and head-mounted displays. Users select images and upload them using the device's application.
[1342] Image analysis methods
[1343] The server receives the uploaded photos and uses image analysis algorithms to analyze the content of the photos, using machine learning frameworks such as TensorFlow and Keras to identify people, objects, and backgrounds in the photos.
[1344] Question and comment generator
[1345] Based on the analysis results, generative AI operates to generate appropriate questions and comments for the user, using advanced generative models such as OpenAI's GPT-3.
[1346] User Presentation Method
[1347] The device presents the generated questions and comments to the user, who responds to the generated questions and comments using text input or voice input.
[1348] Response sending method
[1349] The user's response is sent from the device to the server, and may include, for example, a response such as "These are my daughter and grandchildren."
[1350] Next question generation method
[1351] The server analyzes the user's response, and the AI generates the next question or comment, allowing for a continuous dialogue with the user.
[1352] Conversation log storage method
[1353] All conversation logs are stored on the server and are used to improve the AI generation and provide personalized content based on the user's interests.
[1354] Specific examples
[1355] Example 1: Elderly people
[1356] An elderly person uploads an old family photo using their smartphone. The server receives the photo and uses TensorFlow to analyze the image and recognize that a family member is in it. A generative AI (e.g., GPT-3) generates the question "Who is in this photo?" and the device displays the question to the elderly person. The elderly person responds, "This is my daughter and grandchildren," and sends it to the server. The server analyzes the response and generates the next question: "What a lovely family. Where was this photo taken?"
[1357] Prompt Sentence Examples
[1358] "The Eiffel Tower is in the photo. Generate appropriate questions for the user that will remind them of their travel experiences."
[1359] This method allows users to reminisce and stimulate their brains and maintain their energy through generative AI dialogue. All dialogue logs are saved and used for later analysis. Specifically, it allows for the provision of personalized content based on the user's interests and usage.
[1360] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1361] Step 1:
[1362] The user opens the application on their device, selects a memorable photo, and presses the upload button, which sends the photo from the device to the server.
[1363] Input: A memorable photo
[1364] Process: Upload a file from the device to the server
[1365] Output: Photo files uploaded to the server
[1366] Step 2:
[1367] The server receives the uploaded photos and stores the image files in a database.
[1368] Input: Photo file sent from the device
[1369] Processing: Receiving the file and saving it to the database
[1370] Output: Photo data stored in a database
[1371] Step 3:
[1372] The server analyzes the content of the photo using image analysis algorithms, such as TensorFlow and Keras, to identify people, objects, and the background in the photo.
[1373] Input: Photo data stored in the database
[1374] Processing: Image analysis using TensorFlow and Keras
[1375] Output: Analysis results for the content of the photo (e.g. people's names, places, objects, etc.)
[1376] Step 4:
[1377] The server sends the analysis results to the generation AI, which then uses them to generate appropriate questions and comments for the user. In this example, OpenAI's GPT-3 is used.
[1378] Input: Image analysis results
[1379] Processing: Generative AI generates questions and comments
[1380] Output: Generated questions and comments
[1381] Step 5:
[1382] The server sends the generated questions and comments to the terminal, which displays them to the user.
[1383] Input: Generated questions and comments
[1384] Processing: Transmission from server to terminal, screen display
[1385] Output: Questions and comments displayed to the user
[1386] Step 6:
[1387] The user enters a response to a question or comment. The response can be entered by voice or text. The response is sent from the device to the server.
[1388] Input: User response (voice or text)
[1389] Processing: speech recognition or text input, sending data to server
[1390] Output: The user's response sent to the server
[1391] Step 7:
[1392] The server receives and analyzes the user's response. Based on the analysis results, the AI generates the next question or comment, thus continuing the dialogue with the user.
[1393] Input: User response
[1394] Processing: Analyzing responses and generating the next question or comment using AI
[1395] Output: Next question or comment
[1396] Step 8:
[1397] The server sends the next question or comment to the terminal, which displays it to the user, who responds again, and the process from step 6 to step 8 is repeated.
[1398] Input: Next question or comment
[1399] Processing: Transmission from server to terminal, screen display, user response input
[1400] Output: Repeated user interactions
[1401] Step 9:
[1402] All conversation logs are stored on the server. These logs include user usage and reactions and are used for later analysis. They serve as the basis for improving the AI generation and providing personalized content based on the user's interests.
[1403] Input: User interaction log
[1404] Processing: Data storage and later analysis
[1405] Output: Data used to improve generative AI and provide personalized content
[1406] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1407] This invention provides a system that stimulates the user's brain and maintains their morale by interacting with a generating AI using photos of the user's memories. This system also incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction.
[1408] System Overview
[1409] The process by which a user uploads an image from their device
[1410] Users upload photos to the system using their own devices (e.g., smartphones or tablets). Users can easily select and upload photos by opening the application and going to the upload screen. Uploaded photos are first transferred to the server.
[1411] The process by which the server receives and analyzes the images
[1412] The uploaded photos are received by the server and stored in a database, where they are then analyzed using image analysis algorithms to automatically identify people, objects, and background information in the photos.
[1413] The process by which generative AI generates questions and comments for users
[1414] Based on the results of image analysis, the server sends the necessary data to the generation AI. The generation AI then generates appropriate questions and comments for the user based on the content of the photo. For example, it creates questions such as "Where was this photo taken?" or "Who is this person?"
[1415] The process by which the device presents the generated questions and comments to the user
[1416] The questions and comments generated by the AI are sent back from the server to the device, which then displays them to the user. The user can then confirm the questions and comments and enter a response.
[1417] The process of sending user responses from the terminal to the server
[1418] The response entered by the user is sent from the device to the server, and may include, for example, "These are my daughter and grandchildren."
[1419] The server analyzes the response, and the AI generates the next question or comment.
[1420] The server receives and analyzes the user's response, and the generation AI generates the next question or comment. This continues the conversation with the user and advances the dialogue process. For example, the next question might be, "Where was this photo taken?"
[1421] Emotion recognition process by emotion engine
[1422] The emotion engine analyzes the user's emotions from images uploaded by the user and responses during conversations. The emotion engine uses facial expression analysis and voice analysis to determine whether the user is happy or anxious.
[1423] Tuning process for emotion-based generative AI
[1424] Based on the analysis results of the emotion engine, the generation AI will adjust the questions and comments. For example, if the user is determined to be sad, the generation AI will generate comforting comments. If the user is having fun, the generation AI will generate questions that bring out that emotion.
[1425] Continuing conversations and logging processes
[1426] This conversation process is repeated until the user is satisfied. All conversation logs are stored on the server and used to analyze user usage and reactions at a later date. The results of this analysis are used to improve the system and enhance the user experience.
[1427] Specific examples
[1428] Example 1: Lonely elderly people
[1429] User: Selects and uploads an old family photo.
[1430] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[1431] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[1432] Terminal: Display the question to the user.
[1433] User: Responds to the question by saying, "These are my daughter and grandchildren."
[1434] Terminal: Sends the response to the server.
[1435] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[1436] Emotion engine: Analyzes the user's emotions from images and responses and recognizes that the user is feeling nostalgic.
[1437] Generative AI: Based on the user's emotions, the system further adjusts the questions and comments, such as "Tell me more about your memories of that time."
[1438] Terminal: Display the next question to the user.
[1439] Example 2: Travel photos
[1440] User: Select and upload photos of travel destinations.
[1441] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[1442] Generative AI: Generates questions like "Where is this place?" based on a photo.
[1443] Terminal: Display the question to the user.
[1444] User: Responds to the question with "This is a photo from Paris, where I went last year."
[1445] Terminal: Sends the response to the server.
[1446] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[1447] Emotion engine: Analyzes the user's emotions from images and responses and recognizes what the user is enjoying.
[1448] Generative AI: Further tailor questions and comments based on user sentiment, such as "What other places have you been?"
[1449] Terminal: Display the next question to the user.
[1450] In this way, this system allows users to enjoy conversations through photos, stimulating their brains and maintaining their energy. Furthermore, the introduction of an emotion engine enables communication that responds to the user's emotions, resulting in more effective dialogue.
[1451] The processing flow will be explained below.
[1452] Step 1:
[1453] The user opens the application on their device and proceeds to the photo upload screen.
[1454] Step 2:
[1455] The user selects a photo on their device and clicks the upload button.
[1456] Step 3:
[1457] The terminal generates a request to transfer the selected photo to the server.
[1458] Step 4:
[1459] The server receives the request from the terminal and receives the photo data.
[1460] Step 5:
[1461] The server stores the received photos in a database.
[1462] Step 6:
[1463] The server passes the stored photos to an image analysis algorithm, which begins the analysis process.
[1464] Step 7:
[1465] The server performs image analysis to determine the content of the photo, such as identifying people, objects, and the background.
[1466] Step 8:
[1467] The server sends the analysis results to the generation AI.
[1468] Step 9:
[1469] Based on the results of image analysis, the generative AI generates appropriate questions and comments for the user, such as "Where was this photo taken?" or "Who is this person?"
[1470] Step 10:
[1471] The server sends the generated questions and comments back to the device.
[1472] Step 11:
[1473] The terminal displays the questions and comments received from the server to the user.
[1474] Step 12:
[1475] The user enters a response to the question or comment on the terminal.
[1476] Step 13:
[1477] The terminal generates a request to send the user's response to the server.
[1478] Step 14:
[1479] The server receives the response from the terminal.
[1480] Step 15:
[1481] The server passes the received response to the emotion engine, which analyzes the user's emotion.
[1482] Step 16:
[1483] The emotion engine recognizes emotions from user responses and images and sends the results to the generation AI.
[1484] Step 17:
[1485] The generative AI generates the next appropriate question or comment based on the emotion recognition results from the emotion engine and the user's response. For example, if the user is feeling nostalgic, it generates a question such as, "Tell me more about your memories from that time."
[1486] Step 18:
[1487] The server sends the next generated question or comment back to the device.
[1488] Step 19:
[1489] The device will display the next question or comment to the user.
[1490] Step 20:
[1491] The user again enters a response, and this process is repeated until the user is satisfied.
[1492] Step 21:
[1493] The server stores all conversation logs.
[1494] Step 22:
[1495] The server will later analyze the saved conversation logs to evaluate the user's usage and reactions.
[1496] Example 2
[1497] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1498] In modern society, the number of elderly people and individuals who feel lonely is increasing, making it important to improve their mental health and increase communication opportunities for these people. In particular, there is a demand for systems that can provide mental support by promoting interaction through past memories and photos and engaging in emotionally appropriate dialogue. Conventional systems struggle to improve the quality of emotion recognition and dialogue, and there is a lack of methods for achieving effective communication.
[1499] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1500] In this invention, the server includes means for receiving images and storing them in a database, means for generating image analysis results using an image analysis algorithm, means for analyzing the user's responses and generating the next question or comment using a generation AI, means for an emotion engine to recognize the user's emotions and for the generation AI to adjust the questions or comments based on the results, and means for saving all conversation logs and analyzing them later. This enables natural conversation based on the user's memories, provides appropriate communication according to emotions, and helps maintain and improve the user's mental health.
[1501] "Means for uploading images" refers to a series of operations and functions that allow a user to select digital images from their terminal and send them to the server.
[1502] A "database" is a system for systematically storing and managing digital information, where images and associated metadata are stored.
[1503] An "image analysis algorithm" is a computational method that analyzes the content of a digital image and extracts information about the people, objects, background, and other aspects of the photograph.
[1504] "Generative AI" is artificial intelligence that automatically generates questions and comments based on user responses and image analysis results.
[1505] An "emotion engine" is a technology that analyzes a user's emotions and adjusts the output of the generative AI based on their emotional state.
[1506] A "terminal" is a digital device used by a user, such as a smartphone or tablet.
[1507] A "conversation log" is a record of all interactions between a user and a system, data that is stored for later analysis.
[1508] A "prompt sentence" is the input text given to the generation AI, which is used to generate the next question or comment based on the user's response and the results of image analysis.
[1509] "Means for generating questions and comments" refers to the functions and operations by which the generation AI creates appropriate questions and comments for the user based on the prompt text.
[1510] This invention is a system that aims to stimulate the user's brain and maintain their energy by having them interact with a generative AI using photos of their memories. This system incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction.
[1511] System Overview
[1512] User-uploaded images
[1513] Users upload photos to the system using devices such as smartphones or tablets. They launch a dedicated application, select memorable photos on the upload screen, and send them to the server. The uploaded photos are then transferred from the device to the server.
[1514] Receiving and analyzing images by the server
[1515] The server receives the uploaded photos and stores them in a database. The server then analyzes the content of the photos using image analysis algorithms (such as Google Cloud Vision API or Microsoft Azure's Computer Vision API). The analysis results include information about the people, objects, and background in the photos.
[1516] Question generation using generative AI
[1517] Based on the results of image analysis, the server sends the necessary data to a generative AI (such as OpenAI's GPT-4). Based on the received data, the generative AI generates appropriate questions and comments for the user. For example, it creates specific questions such as "Where was this photo taken?" or "Who is this person?"
[1518] Display of questions on the device
[1519] The generated AI's questions and comments are sent from the server to the user's device and displayed on the device. The user reads the displayed questions and comments and enters a response.
[1520] Parsing the user's response and generating the next question
[1521] The response entered by the user is sent from the device to the server. The server analyzes the received response and sends the results to the generation AI. The generation AI then generates the next most appropriate question or comment. This allows for a continuous dialogue with the user.
[1522] Emotion recognition and output adjustment by emotion engine
[1523] The emotion engine analyzes the user's emotions from images uploaded by the user and responses during conversations. The emotion engine evaluates whether the user is happy or anxious through facial expression and voice analysis. Based on the results of the emotion engine, the server adjusts the prompts sent to the generation AI and generates appropriate questions and comments.
[1524] Specific examples
[1525] Example 1: Lonely elderly people
[1526] The user selects and uploads an old family photo. The server receives the photo and uses an image analysis algorithm to recognize that a family member is in it. The generation AI generates the question, "Who is in this photo?" and displays it on the device. The user responds, "This is my daughter and grandchildren," and the device sends the response to the server. The server analyzes the response, and the generation AI generates the next question, "What a lovely family. Where was this photo taken?" Furthermore, if the emotion engine senses the user's nostalgia, it generates a question such as, "Tell me more about your memories from that time."
[1527] Example 2: Travel photos
[1528] The user selects and uploads photos of their travel destinations. The server receives the photos and uses an image analysis algorithm to recognize the travel destinations. The generation AI generates the question "Where is this place?" and displays it on the device. The user responds, "This is a photo of Paris, where I went last year," and the device sends the response to the server. The server analyzes the response, and the generation AI generates the next question, "How was Paris? Were there any places that particularly impressed you?" The emotion engine analyzes the user's emotions and, if it identifies that the user is enjoying themselves, generates questions such as "What other places did you go?"
[1529] Prompt Sentence Examples
[1530] As an example of a prompt for a generative AI model, input the following text to the generative AI:
[1531] Analyzing a photo submitted by a user reveals that it contains family members. Based on that content, generate questions or comments to naturally continue the conversation with the user. For example, "Who is in this photo?"
[1532] This system provides natural dialogue through users' memorable photos and realizes appropriate communication according to their emotions, which is expected to contribute to maintaining and improving users' mental health.
[1533] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1534] Step 1:
[1535] The user initiates the image upload process from their device.
[1536] Specific operation: The user launches the application on their smartphone or tablet, selects "Upload photos" from the menu, selects a memorable photo from their photo library, and presses the "Upload" button.
[1537] Input: A photo file selected by the user.
[1538] Output: The photo files are saved on your device and ready to be sent to the server.
[1539] Step 2:
[1540] The device sends the photo to the server.
[1541] Specific operation: The device makes an HTTP request to send the user-selected photo file to the server. The request includes the photo file and associated metadata (e.g., upload date and time, user ID).
[1542] Input: A user-selected photo file and its metadata.
[1543] Output: The photo files and metadata are sent to the server.
[1544] Step 3:
[1545] The server receives the photos and stores them in a database.
[1546] What it does: The server parses the incoming HTTP request and stores the photo file and its metadata in a database, along with information about the photo's identifier and storage location.
[1547] Input: The submitted photo file and its metadata.
[1548] Output: Photo files and associated metadata stored in a database.
[1549] Step 4:
[1550] The server analyzes the photo using image analysis algorithms.
[1551] How it works: The server sends the stored photo file to an image analysis algorithm (e.g., Google Cloud Vision API) to analyze the content of the photo, including people, objects, and background information.
[1552] Input: Photo files stored in the database.
[1553] Output: The analysis data generated as a result of the image analysis algorithm.
[1554] Step 5:
[1555] The generative AI creates prompts to generate questions and comments for the user.
[1556] Specific operation: The server creates a prompt sentence to send to the generation AI based on the image analysis results. This prompt sentence contains detailed information to facilitate the dialogue with the user. For example, "A family photo has been uploaded. Please generate a question about this photo."
[1557] Input: Image analysis results.
[1558] Output: The prompt sent to the generation AI.
[1559] Step 6:
[1560] Generative AI generates questions and comments.
[1561] How it works: A generative AI model (e.g., OpenAI's GPT-4) generates questions or comments to present to the user based on prompts received from the server, such as "Who is in this photo?"
[1562] Input: The server-generated prompt text.
[1563] Output: Generated questions and comments.
[1564] Step 7:
[1565] The server sends the generated questions and comments to the device.
[1566] Specific operation: The server creates an HTTP response to send the generated question or comment to the terminal, and sends it to the terminal.
[1567] Input: Generated questions and comments.
[1568] Output: Questions and comments sent to the user's device.
[1569] Step 8:
[1570] The device displays the question or comment to the user.
[1571] Specific operation: The device displays the received questions and comments on the screen and presents them to the user. An interface is provided so that the user can review them.
[1572] Input: Questions and comments sent by the server.
[1573] Output: Questions and comments displayed on the terminal screen.
[1574] Step 9:
[1575] The user enters a response to the question.
[1576] Specific behavior: The user reads the questions and comments displayed on the device and enters a response. For example, "These are my daughter and grandchildren."
[1577] Input: Questions or comments displayed on the device.
[1578] Output: The response text entered by the user.
[1579] Step 10:
[1580] The terminal sends the user's response to the server.
[1581] Specific operation: The terminal creates an HTTP request to send the response text entered by the user to the server.
[1582] Input: The response text entered by the user.
[1583] Output: The response text sent to the server.
[1584] Step 11:
[1585] The server analyzes the user's response and creates and sends the next prompt to the generation AI.
[1586] Specific operation: The server analyzes the user's response text, creates a new prompt based on its content, and sends the created prompt to the generation AI to generate the next question or comment.
[1587] Input: The user's response text.
[1588] Output: The next prompt sent to the generation AI.
[1589] Step 12:
[1590] Generative AI generates the next question or comment.
[1591] Specific behavior: The generative AI generates the next question or comment based on the newly received prompt, for example, "Where was this photo taken?"
[1592] Input: The newly created prompt statement.
[1593] Output: The next question or comment.
[1594] Step 13:
[1595] The emotion engine recognizes the user's emotions and adjusts the output of the generative AI.
[1596] Specific operation: The emotion engine analyzes the user's facial expressions and responses to evaluate the user's emotional state. Based on the results, it adjusts the generation AI to generate appropriate questions and comments.
[1597] Input: User's facial expression data and response content.
[1598] Output: Adjusted prompt and output of the generation AI.
[1599] Step 14:
[1600] The server sends the tailored questions and comments to the device.
[1601] Specific operation: The server sends questions and comments adjusted by the emotion engine to the terminal.
[1602] Input: Moderated questions and comments.
[1603] Output: The tailored questions and comments sent to your device.
[1604] Step 15:
[1605] The device displays the tailored question or comment to the user.
[1606] Specific operation: The device displays the adjusted questions and comments on the screen and presents them to the user.
[1607] Input: Moderated questions and comments.
[1608] Output: The adjusted questions and comments displayed on the screen.
[1609] Step 16:
[1610] The server stores all conversation logs.
[1611] What it does: The server records all interactions and stores them in a database. This information is later analyzed and used to improve the system and enhance the user experience.
[1612] Input: All dialogue content (questions, responses, and adjustment results).
[1613] Output: Conversation logs stored in a database.
[1614] (Application example 2)
[1615] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1616] The problem that this invention aims to solve is to realize more effective communication that responds to the user's emotions while stimulating the user's brain and maintaining their energy through dialogue using photos related to the user's memories. Conventional systems simply continue dialogue mechanically without taking the user's emotions into consideration, which has the problem of degrading the quality of the user experience.
[1617] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1618] In this invention, the server includes means for a user to upload images from a terminal, means for the server to receive and analyze the images, means for the generation AI to generate questions and comments for the user based on the generated analysis results, means for the terminal to present the generated questions and comments to the user, means for the terminal to send the user's response to the server, means for the server to analyze the response and for the generation AI to generate the next question or comment, means for analyzing emotions from the user's facial expressions and voice, means for the generation AI to adjust the content of the dialogue based on the emotion analysis results, means for continuously conversing with the user, and means for saving all conversation logs and analyzing them later. This allows the user to interact with the generation AI through memorable photos, and the content of the dialogue is appropriately adjusted through emotion analysis, thereby stimulating the user's brain and maintaining their energy.
[1619] "Means for users to upload images from their devices" refers to the function that allows users to send image data to the system using devices such as smartphones and tablets.
[1620] "Means for the server to receive and analyze images" refers to the function of the cloud server receiving uploaded image data and analyzing its contents using a specified image analysis algorithm.
[1621] "Means for the generation AI to generate questions and comments for the user based on the generated analysis results" refers to the function of using natural language processing technology to generate appropriate questions and comments based on the content of the image analyzed by the server.
[1622] "Means for the device to present the generated questions and comments to the user" refers to a function that displays the generated questions and comments on the user's smartphone or tablet.
[1623] "Means for transmitting a user's response from the terminal to the server" refers to a function for transmitting the response data entered by the user again to the cloud server.
[1624] "The server analyzes the response, and the generation AI generates the next question or comment" refers to the function in which the cloud server analyzes the user's response and generates the next conversation content based on the analysis results.
[1625] "Means for analyzing emotions from the user's facial expressions and voice" refers to a function that analyzes the facial expressions and voice data shown by the user during a conversation and determines their emotional state (joy, sadness, surprise, etc.).
[1626] "Means for the generation AI to adjust the content of the dialogue based on the results of emotion analysis" refers to the function by which the generation AI takes into account the user's emotional state obtained through emotion analysis and appropriately adjusts the next question or comment.
[1627] "Means for continuing conversation with the user" refers to a function that allows a dialogue between the user and the system to continue.
[1628] "A means of saving all conversation logs and analyzing them later" refers to a function that records and saves all generated conversation content and later analyzes changes in user behavior patterns and emotions based on that data.
[1629] This invention provides a system that stimulates the user's brain and maintains their morale by interacting with a generating AI using memorable photos of the user. This system also incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction. Each element of this system is explained below.
[1630] System configuration
[1631] 1. A way for users to upload images from their devices
[1632] Users use their smartphones, tablets, or other devices to select and upload images within the application, a process designed to make it easy and intuitive for users to submit photos to the system.
[1633] 2. How the server receives and analyzes images
[1634] Uploaded photos are transferred to a cloud server, which uses image analysis libraries such as the Google Vision API to automatically analyze the content of the photo and extract information about people, places, objects, and more.
[1635] 3. A means for the generative AI to generate questions and comments for users based on the generated analysis results
[1636] The results of the image analysis are sent to a generative AI (e.g., GPT-3 or GPT-4), which uses this data to generate appropriate questions and comments for the user.
[1637] For example: "Where was this photo taken?"
[1638] 4. A means for the device to present generated questions and comments to the user
[1639] The questions and comments generated by the AI are sent from the server to the user's device, where they are displayed to the user, and a dialogue begins.
[1640] 5. A means of transmitting user responses from the terminal to the server
[1641] When a user responds to a question or comment through their device, the response is sent to the cloud server.
[1642] 6. The server analyzes the response and the AI generates the next question or comment.
[1643] The server analyzes the user's responses and sends the results to the generation AI, which then generates new questions and comments to continue the dialogue with the user.
[1644] 7. Means of analyzing emotions from user facial expressions and voice
[1645] Using an emotion engine (for example, Microsoft Azure Emotion API), emotions are analyzed from the user's facial expressions and voice. The analysis results indicate the user's emotional state, such as whether they are happy or sad.
[1646] 8. How generative AI can adjust dialogue content based on emotion analysis results
[1647] Based on the analysis results of the emotion engine, the generative AI will adjust the dialogue appropriately. For example, if the user is sad, it will generate comforting comments, and if the user is happy, it will generate questions that will bring out that emotion.
[1648] For example: "Tell me more about your memories of that time."
[1649] 9. A way to continue the conversation with your users
[1650] The dialogue continues until the user is satisfied. The entire conversation process is controlled by the system.
[1651] 10. A way to store all conversation logs and analyze them later
[1652] All generated conversation logs are stored on a cloud server, allowing users' usage patterns and reactions to be analyzed at a later date, helping to improve the system and enhance the user experience.
[1653] This system allows users to enjoy conversations through memorable photos, stimulating their brains and maintaining their energy. In addition, the introduction of an emotion engine enables communication based on the user's emotions, resulting in more effective conversations.
[1654] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1655] Step 1:
[1656] A way for users to upload images from their devices
[1657] A user opens the application on a smartphone or tablet, selects a photo from the upload screen, and sends it to the system. The input is the photo file selected by the user, and the output is that the photo file is sent to the cloud server. The photo file is transferred between the device and the cloud server.
[1658] Step 2:
[1659] The server receives and analyzes the images
[1660] The server receives the photos sent from the device and stores them in a database. The stored photos are then passed to the Google Vision API for image analysis. The input is the uploaded photo file, and the output is the analyzed image content (people, places, objects, etc.) data. Image analysis uses machine learning algorithms to analyze the content of the photo.
[1661] Step 3:
[1662] A means for the AI to generate questions and comments for users based on the generated analysis results
[1663] The server sends the image analysis results to a generation AI (GPT-3 or GPT-4), which generates appropriate questions and comments. The input is the image analysis results, and the output is questions and comments to present to the user. Specifically, the analysis results are input as prompts to the generation AI, which then uses natural language processing technology to generate corresponding questions and comments.
[1664] Step 4:
[1665] A means by which the device presents generated questions and comments to the user
[1666] The server sends the generated questions and comments to the terminal and displays them to the user. The input is the generated questions and comments, and the output is what is displayed to the user. The terminal receives this and displays it on the screen.
[1667] Step 5:
[1668] A means for transmitting user responses from the terminal to the server
[1669] Users use their devices to respond to questions and comments, and then send the response data to the cloud server. The input is the user's response, and the output is the response data sent to the cloud server. Specifically, the user enters text and presses the send button, which sends the data to the server.
[1670] Step 6:
[1671] The server analyzes the response and the generation AI generates the next question or comment.
[1672] The server analyzes the user's response data and sends the analysis results to the generation AI to generate the next question or comment. The input is the user's response data, and the output is the newly generated question or comment. The response data is analyzed for text, and its content is provided to the generation AI as a prompt.
[1673] Step 7:
[1674] A means of analyzing emotions from the user's facial expressions and voice
[1675] The server receives the user's facial expression and voice data sent from the device and performs emotion analysis using an emotion engine (Microsoft Azure Emotion API). The input is facial expression and voice data, and the output is data on the user's emotional state (e.g., joy, sadness, surprise, etc.). Specifically, image and voice data is passed to the emotion engine for analysis.
[1676] Step 8:
[1677] A means for generative AI to adjust dialogue content based on emotion analysis results
[1678] The server provides the emotion analysis results to the generation AI, which then adjusts the dialogue appropriately based on the results. The input is the emotion analysis results, and the output is adjusted questions or comments. The generation AI receives prompts containing the emotion analysis results and generates appropriate dialogue content based on them. For example, if it determines that the user is sad, it generates a question such as, "Tell me more about your memories of that time."
[1679] Step 9:
[1680] A way to maintain an ongoing conversation with users
[1681] The server and the generation AI work together to continue the dialogue until the user is satisfied. The input is continuous responses from the user, and the output is newly generated questions and comments. The dialogue content is continuously generated, and the interaction with the user progresses without interruption.
[1682] Step 10:
[1683] A means to store all conversation logs and analyze them later
[1684] The server stores all dialogue logs and later analyzes changes in user behavior patterns and emotions based on them. The input is all the dialogue logs generated, and the output is the stored data and the analysis results. The log data is stored in a database and later analyzed using analysis tools.
[1685] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1686] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1687] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1688] [Fourth embodiment]
[1689] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1690] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1691] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1692] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1693] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1694] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1695] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1696] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1697] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1698] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1699] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1700] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1701] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1702] This invention provides a system that stimulates the user's brain and maintains their energy by interacting with a generating AI using photos of the user's memories. This system involves a multi-stage process in which the user uploads photos and looks back on their memories through conversations with the generating AI.
[1703] System Overview
[1704] The process by which a user uploads an image from their device
[1705] Users upload photos to the system using their own devices (e.g., smartphones or tablets). Users can easily select and upload photos by opening the application and going to the upload screen. Uploaded photos are first transferred to the server.
[1706] The process by which the server receives and analyzes the images
[1707] The uploaded photos are received by the server and stored in a database, where they are then analyzed using image analysis algorithms to automatically identify people, objects, and background information in the photos.
[1708] The process by which generative AI generates questions and comments for users
[1709] Based on the results of image analysis, the server sends the necessary data to the generation AI. The generation AI then generates appropriate questions and comments for the user based on the content of the photo. For example, it creates questions such as "Where was this photo taken?" or "Who is this person?"
[1710] The process by which the device presents the generated questions and comments to the user
[1711] The questions and comments generated by the AI are sent back from the server to the device, which then displays them to the user. The user can then confirm the questions and comments and enter a response.
[1712] The process of sending user responses from the terminal to the server
[1713] The response entered by the user is sent from the device to the server, and may include, for example, "These are my daughter and grandchildren."
[1714] The server analyzes the response, and the AI generates the next question or comment.
[1715] The server receives and analyzes the user's response, and the generation AI generates the next question or comment. This continues the conversation with the user and advances the dialogue process. For example, the next question might be, "Where was this photo taken?"
[1716] Continuing conversations and logging processes
[1717] This conversation process is repeated until the user is satisfied. All conversation logs are stored on the server and used later to analyze user usage and reactions.
[1718] Specific examples
[1719] Example 1: Lonely elderly people
[1720] User: Selects and uploads an old family photo.
[1721] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[1722] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[1723] Terminal: Display the question to the user.
[1724] User: Responds to the question by saying, "These are my daughter and grandchildren."
[1725] Terminal: Sends the response to the server.
[1726] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[1727] Terminal: Display the next question to the user.
[1728] Example 2: Travel photos
[1729] User: Select and upload photos of travel destinations.
[1730] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[1731] Generative AI: Generates questions like "Where is this place?" based on a photo.
[1732] Terminal: Display the question to the user.
[1733] User: Responds to the question with "This is a photo from Paris, where I went last year."
[1734] Terminal: Sends the response to the server.
[1735] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[1736] Terminal: Display the next question to the user.
[1737] As described above, through this system, users can enjoy conversations through photos, which helps stimulate the brain and maintain energy.
[1738] The processing flow will be explained below.
[1739] Step 1:
[1740] The user opens the application on their device and proceeds to the photo upload screen.
[1741] Step 2:
[1742] The user selects a photo on their device and clicks the upload button.
[1743] Step 3:
[1744] The terminal generates a request to transfer the selected photo to the server.
[1745] Step 4:
[1746] The server receives the request from the terminal and receives the photo data.
[1747] Step 5:
[1748] The server stores the received photos in a database.
[1749] Step 6:
[1750] The server passes the stored photos to an image analysis algorithm, which begins the analysis process.
[1751] Step 7:
[1752] The server performs image analysis to determine the content of the photo, such as identifying people, objects, and the background.
[1753] Step 8:
[1754] The server sends the image analysis results to the generation AI.
[1755] Step 9:
[1756] Based on the results of image analysis, the generative AI generates appropriate questions and comments for the user, such as "Where was this photo taken?" or "Who is this person?"
[1757] Step 10:
[1758] The server sends the generated questions and comments back to the device.
[1759] Step 11:
[1760] The terminal displays the questions and comments received from the server to the user.
[1761] Step 12:
[1762] The user enters a response to the question or comment on the terminal.
[1763] Step 13:
[1764] The terminal generates a request to send the user's response to the server.
[1765] Step 14:
[1766] The server receives the response from the terminal.
[1767] Step 15:
[1768] The server analyzes the received response and sends it to the generating AI.
[1769] Step 16:
[1770] Based on the user's response, the generative AI generates the next appropriate question or comment, such as "Where was this photo taken?"
[1771] Step 17:
[1772] The server sends the next generated question or comment back to the device.
[1773] Step 18:
[1774] The device will display the next question or comment to the user.
[1775] Step 19:
[1776] The user again enters a response, and this process is repeated until the user is satisfied.
[1777] Step 20:
[1778] The server stores all conversation logs.
[1779] Step 21:
[1780] The server will later analyze the saved conversation logs to evaluate the user's usage and reactions.
[1781] Example 1
[1782] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1783] Today's elderly and lonely people lack emotional communication and brain activation. As a result, they face problems such as a decline in cognitive function and loss of energy. Conventional technologies have not provided effective solutions to these issues, particularly lacking methods for enjoying conversations while reminiscing about memories. Therefore, there is a need for a system that allows users to engage in emotional communication through their own memories, activating their brains, and maintaining their energy.
[1784] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1785] In this invention, the server includes means for receiving and storing images, means for analyzing images, and means for generating questions and comments for the user, thereby generating dialogue based on images uploaded by the user from the terminal, and enabling continuous conversation with the user.
[1786] "Users" refer to people who use the system to upload their own photos and interact with the generative AI.
[1787] "Device" refers to the electronic device, such as a smartphone or tablet, that a user uses to access the system, upload photos, and respond to questions and comments from the Generative AI.
[1788] "Server" refers to the central device of the system, which receives and stores images uploaded by users, analyzes the images, generates questions and comments using generative AI, and sends them to the terminal.
[1789] "Generative AI" refers to an artificial intelligence model used by the server to generate new questions or comments based on image analysis results and user responses, such as an AI model that performs natural language processing.
[1790] "Image analysis" refers to the process of analyzing uploaded images using machine learning algorithms to automatically recognize information such as people, objects, and backgrounds in the photos.
[1791] "Question and comment generation" refers to the process in which the generative AI creates appropriate questions and comments for the user based on the results of image analysis and the user's responses.
[1792] "Conversation logs" refer to records of interactions with users on the system, and are data stored on the server. They include the content of the conversations and are used for later analysis.
[1793] "Database" refers to a system that systematically stores and manages data such as images and conversation logs received by a server.
[1794] "Image upload" refers to the process by which a user sends a photo file from their device to a server.
[1795] "Continuous dialogue" refers to a process in which the generative AI generates questions and comments one after another until the user is satisfied, maintaining an uninterrupted conversation with the user.
[1796] "Analysis results" refers to information about the content of a photo obtained through image analysis, including the recognition of people, objects, backgrounds, etc.
[1797] "Storage" refers to the process of storing data such as photos and conversation logs in a database on the server.
[1798] The present invention provides a system that allows users to upload memorable photos from their own devices and review the photos through dialogue with a generation AI, thereby stimulating the user's brain and maintaining their energy. An embodiment of this system is described in detail below.
[1799] The process by which a user uploads an image from their device
[1800] Users access the application using a device such as a smartphone or tablet. They open the photo upload screen within the application and select their memorable photos. When the user presses the "Upload" button, the photo file is sent to the server via the Internet.
[1801] The process by which the server receives and stores images
[1802] The server receives the photos sent by the user and stores them in a database. When saved, the photos are assigned a unique ID, which makes it easier to manage the photo data.
[1803] The process by which the server analyzes the image
[1804] The server runs an image analysis algorithm (such as YOLO or OpenCV) on the received photo to obtain information about people, objects, and backgrounds in the photo. The results of this analysis are stored in a database for later use.
[1805] The process by which generative AI generates questions and comments
[1806] The server sends data based on the image analysis results to a generative AI model (such as GPT-4). The generative AI uses this data to generate appropriate questions and comments for the user. For example, prompts such as "Where was this photo taken?" or "Who is this person?" are generated.
[1807] The process by which the device presents the generated questions and comments to the user
[1808] The questions and comments generated by the AI are sent from the server to the user's device, where they are displayed. The user can then confirm the questions and comments displayed on the device and enter their responses.
[1809] The process of sending the user's response to the server
[1810] The user's response is sent from the device to the server, and may contain text such as "This is my daughter and grandchildren."
[1811] The server analyzes the response, and the AI generates the next question or comment.
[1812] The server analyzes the user's response, and the AI generates the next question or comment based on the results, creating a continuous dialogue with the user.
[1813] Continuing conversations and logging processes
[1814] The dialogue process is repeated until the user is satisfied. All conversation logs are stored on the server and used for later analysis. This data provides valuable information for analyzing user usage and reactions.
[1815] Specific examples
[1816] Example 1: Lonely elderly people
[1817] User: Selects and uploads an old family photo.
[1818] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[1819] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[1820] Terminal: Display the question to the user.
[1821] User: Responds to the question by saying, "These are my daughter and grandchildren."
[1822] Terminal: Sends the response to the server.
[1823] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[1824] Terminal: Display the next question to the user.
[1825] Example 2: Travel photos
[1826] User: Select and upload photos of travel destinations.
[1827] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[1828] Generative AI: Generates questions like "Where is this place?" based on a photo.
[1829] Terminal: displays the question to the user.
[1830] User: Responds to the question with "This is a photo from Paris, where I went last year."
[1831] Terminal: Sends the response to the server.
[1832] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[1833] Terminal: Display the next question to the user.
[1834] Through this process, users can enjoy conversation while reminiscing on their memories, which helps to stimulate the brain and maintain energy.
[1835] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1836] Step 1:
[1837] The user accesses the application using a terminal and opens the photo upload screen. Then, the user selects the photo to upload and presses the "Upload" button. This action sends the photo file to the server over the Internet.
[1838] Input: A photo file selected from the user's device.
[1839] Output: The photo file sent to the server.
[1840] Specific operation: The user launches the application, logs in, selects a photo, and presses the "Upload" button.
[1841] Step 2:
[1842] The server saves the received photo files and assigns a unique ID to each photo, which is then stored in a database.
[1843] Input: Photo files sent over the internet to the server.
[1844] Output: Photo file stored in the database with a unique ID.
[1845] Specific operation: The server receives the photo file, saves it as binary data in its internal memory, and stores it in a database.
[1846] Step 3:
[1847] The server runs image analysis algorithms (such as YOLO or OpenCV) on the stored photos to extract information about people, objects, and backgrounds in the photos, and stores the results in a database.
[1848] Input: Photo files stored in the database.
[1849] Output: Data containing image analysis results (e.g. people, objects, and background in a photo).
[1850] What happens: The server runs an image analysis algorithm to identify the content of the photo.
[1851] Step 4:
[1852] The server sends data based on the image analysis results to a generation AI (for example, GPT-4), which then uses this data to generate questions and comments for the user.
[1853] Input: Image analysis results.
[1854] Output: Questions and comments generated by the generative AI.
[1855] Specific operation: The server converts the analysis result into a prompt sentence and sends it to the generation AI, and the AI model generates a response.
[1856] Step 5:
[1857] The server sends questions and comments from the generated AI to the user's device, which then displays the received questions and comments to the user.
[1858] Input: Questions and comments generated by the generative AI.
[1859] Output: Questions and comments displayed on the user's device.
[1860] Specific operation: The server sends the generated questions and comments to the terminal, which displays them to the user.
[1861] Step 6:
[1862] The user inputs their response to the questions and comments displayed on the terminal, and when they are finished, they press the "Send" button to send the response to the server.
[1863] Input: Questions and comments from the generative AI, and user responses.
[1864] Output: The user's response sent from the terminal to the server.
[1865] Specific behavior: The user reads the question or comment, enters their response, and presses the "Submit" button.
[1866] Step 7:
[1867] The server analyzes the user's response, and the AI generates the next question or comment based on the results. This process is repeated until the user is satisfied.
[1868] Input: The user's response.
[1869] Output: The next question or comment generated by the generative AI.
[1870] Specific operation: The server analyzes the user's response and sends the analysis results to the generation AI, and the AI model generates the next question or comment.
[1871] Step 8:
[1872] All conversation logs are stored on the server and used for later analysis. This data provides information for analyzing user usage and reactions.
[1873] Input: A log of the conversation between the user and the generated AI.
[1874] Output: Saved conversation logs.
[1875] Specific operation: The server stores the conversation log in a database and uses it for later analysis.
[1876] The above are the specific processing steps in this system.
[1877] (Application example 1)
[1878] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1879] To help elderly people and users who feel lonely to activate their brains and maintain their morale, we provide a system that provides emotional and cognitive stimulation by having users interact with a generative AI using photos of their memories. This system allows users to enjoy dialogue through photos, and we also create an environment where the generative AI can provide personalized content.
[1880] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1881] In this invention, the server includes a means for uploading images from the user's device, a means for receiving and analyzing the images, a means for transmitting the user's responses as voice or text from the device to the server, and a means for saving all conversation logs and using them to improve the AI generation and provide personalized content based on the user's interests. This helps to stimulate the user's brain and maintain their energy, and enables the AI generation to provide personalized dialogue based on the user's interests.
[1882] A "user" is someone who uses the system, uploads memorable photos, and interacts with the generative AI.
[1883] "Device" means a device used by a User to access the System, upload photos, and view and respond to generated questions and comments, including, for example, a smartphone, tablet, smart glasses, or head-mounted display.
[1884] "Means for uploading images" refers to the process by which a user selects a photo using a terminal and sends it to the server.
[1885] The "server" is a central system that receives and analyzes uploaded photos, generates questions and comments using generative AI, and stores conversation logs.
[1886] "Means for receiving and analyzing images" refers to the process by which the server captures the photos sent by the user and identifies the content of the photos using image analysis algorithms.
[1887] "Generative AI" is an artificial intelligence system that is connected to a server and generates appropriate questions and comments for users based on the results of image analysis.
[1888] "Means for generating questions and comments for users" refers to the process by which the generation AI creates questions and comments for users based on the results of image analysis.
[1889] The "means for presenting the generated questions and comments to the user" is a function by which the terminal displays the questions and comments sent from the server to the user.
[1890] "Means for transmitting a user's response by voice or text from the terminal to the server" refers to a process in which a user responds to a generated question or comment by voice or text and transmits it from the terminal to the server.
[1891] "Means for generating the next question or comment" refers to the process in which the server analyzes the user's response and uses a generation AI to create a question or comment to proceed with the next dialogue.
[1892] "Means of storing all conversation logs and using them to improve the generation AI and provide personalized content based on the user's interests" refers to the process by which the server records and stores all conversation history with the user, and based on that, the generation AI is further improved and personalized content is provided based on the user's interests.
[1893] The present invention provides a system that stimulates the user's brain and maintains their energy by interacting with a generating AI using memorable photos of the user. Specific embodiments of the system are described below.
[1894] System configuration
[1895] This system consists of a user's terminal, a server, a generation AI, and various means for sending and receiving data between them.
[1896] Image upload method
[1897] Users can upload memorable photos to the system using devices such as their smartphones, tablets, smart glasses, and head-mounted displays. Users select images and upload them using the device's application.
[1898] Image analysis methods
[1899] The server receives the uploaded photos and uses image analysis algorithms to analyze the content of the photos, using machine learning frameworks such as TensorFlow and Keras to identify people, objects, and backgrounds in the photos.
[1900] Question and comment generator
[1901] Based on the analysis results, generative AI operates to generate appropriate questions and comments for the user, using advanced generative models such as OpenAI's GPT-3.
[1902] User Presentation Method
[1903] The device presents the generated questions and comments to the user, who responds to the generated questions and comments using text input or voice input.
[1904] Response sending method
[1905] The user's response is sent from the device to the server, and may include, for example, a response such as "These are my daughter and grandchildren."
[1906] Next question generation method
[1907] The server analyzes the user's response, and the AI generates the next question or comment, allowing for a continuous dialogue with the user.
[1908] Conversation log storage method
[1909] All conversation logs are stored on the server and are used to improve the AI generation and provide personalized content based on the user's interests.
[1910] Specific examples
[1911] Example 1: Elderly people
[1912] An elderly person uploads an old family photo using their smartphone. The server receives the photo and uses TensorFlow to analyze the image and recognize that a family member is in it. A generative AI (e.g., GPT-3) generates the question "Who is in this photo?" and the device displays the question to the elderly person. The elderly person responds, "This is my daughter and grandchildren," and sends it to the server. The server analyzes the response and generates the next question: "What a lovely family. Where was this photo taken?"
[1913] Prompt Sentence Examples
[1914] "The Eiffel Tower is in the photo. Generate appropriate questions for the user that will remind them of their travel experiences."
[1915] This method allows users to reminisce and stimulate their brains and maintain their energy through generative AI dialogue. All dialogue logs are saved and used for later analysis. Specifically, it allows for the provision of personalized content based on the user's interests and usage.
[1916] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1917] Step 1:
[1918] The user opens the application on their device, selects a memorable photo, and presses the upload button, which sends the photo from the device to the server.
[1919] Input: A memorable photo
[1920] Process: Upload a file from the device to the server
[1921] Output: Photo files uploaded to the server
[1922] Step 2:
[1923] The server receives the uploaded photos and stores the image files in a database.
[1924] Input: Photo file sent from the device
[1925] Processing: Receiving the file and saving it to the database
[1926] Output: Photo data stored in a database
[1927] Step 3:
[1928] The server analyzes the content of the photo using image analysis algorithms, such as TensorFlow and Keras, to identify people, objects, and the background in the photo.
[1929] Input: Photo data stored in the database
[1930] Processing: Image analysis using TensorFlow and Keras
[1931] Output: Analysis results for the content of the photo (e.g. people's names, places, objects, etc.)
[1932] Step 4:
[1933] The server sends the analysis results to the generation AI, which then uses them to generate appropriate questions and comments for the user. In this example, OpenAI's GPT-3 is used.
[1934] Input: Image analysis results
[1935] Processing: Generative AI generates questions and comments
[1936] Output: Generated questions and comments
[1937] Step 5:
[1938] The server sends the generated questions and comments to the terminal, which displays them to the user.
[1939] Input: Generated questions and comments
[1940] Processing: Transmission from server to terminal, screen display
[1941] Output: Questions and comments displayed to the user
[1942] Step 6:
[1943] The user enters a response to a question or comment. The response can be entered by voice or text. The response is sent from the device to the server.
[1944] Input: User response (voice or text)
[1945] Processing: speech recognition or text input, sending data to server
[1946] Output: The user's response sent to the server
[1947] Step 7:
[1948] The server receives and analyzes the user's response. Based on the analysis results, the AI generates the next question or comment, thus continuing the dialogue with the user.
[1949] Input: User response
[1950] Processing: Analyzing responses and generating the next question or comment using AI
[1951] Output: Next question or comment
[1952] Step 8:
[1953] The server sends the next question or comment to the terminal, which displays it to the user, who responds again, and the process from step 6 to step 8 is repeated.
[1954] Input: Next question or comment
[1955] Processing: Transmission from server to terminal, screen display, user response input
[1956] Output: Repeated user interactions
[1957] Step 9:
[1958] All conversation logs are stored on the server. These logs include user usage and reactions and are used for later analysis. They serve as the basis for improving the AI generation and providing personalized content based on the user's interests.
[1959] Input: User interaction log
[1960] Processing: Data storage and later analysis
[1961] Output: Data used to improve generative AI and provide personalized content
[1962] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1963] This invention provides a system that stimulates the user's brain and maintains their morale by interacting with a generating AI using photos of the user's memories. This system also incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction.
[1964] System Overview
[1965] The process by which a user uploads an image from their device
[1966] Users upload photos to the system using their own devices (e.g., smartphones or tablets). Users can easily select and upload photos by opening the application and going to the upload screen. Uploaded photos are first transferred to the server.
[1967] The process by which the server receives and analyzes the images
[1968] The uploaded photos are received by the server and stored in a database, where they are then analyzed using image analysis algorithms to automatically identify people, objects, and background information in the photos.
[1969] The process by which generative AI generates questions and comments for users
[1970] Based on the results of image analysis, the server sends the necessary data to the generation AI. The generation AI then generates appropriate questions and comments for the user based on the content of the photo. For example, it creates questions such as "Where was this photo taken?" or "Who is this person?"
[1971] The process by which the device presents the generated questions and comments to the user
[1972] The questions and comments generated by the AI are sent back from the server to the device, which then displays them to the user. The user can then confirm the questions and comments and enter a response.
[1973] The process of sending user responses from the terminal to the server
[1974] The response entered by the user is sent from the device to the server, and may include, for example, "These are my daughter and grandchildren."
[1975] The server analyzes the response, and the AI generates the next question or comment.
[1976] The server receives and analyzes the user's response, and the generation AI generates the next question or comment. This continues the conversation with the user and advances the dialogue process. For example, the next question might be, "Where was this photo taken?"
[1977] Emotion recognition process by emotion engine
[1978] The emotion engine analyzes the user's emotions from images uploaded by the user and responses during conversations. The emotion engine uses facial expression analysis and voice analysis to determine whether the user is happy or anxious.
[1979] Tuning process for emotion-based generative AI
[1980] Based on the analysis results of the emotion engine, the generation AI will adjust the questions and comments. For example, if the user is determined to be sad, the generation AI will generate comforting comments. If the user is having fun, the generation AI will generate questions that bring out that emotion.
[1981] Continuing conversations and logging processes
[1982] This conversation process is repeated until the user is satisfied. All conversation logs are stored on the server and used to analyze user usage and reactions at a later date. The results of this analysis are used to improve the system and enhance the user experience.
[1983] Specific examples
[1984] Example 1: Lonely elderly people
[1985] User: Selects and uploads an old family photo.
[1986] Server: Receives the photo and performs image analysis. It recognizes that family members are in the photo.
[1987] Generative AI: Generates questions based on a photo, such as "Who is in this photo?"
[1988] Terminal: Display the question to the user.
[1989] User: Responds to the question by saying, "These are my daughter and grandchildren."
[1990] Terminal: Sends the response to the server.
[1991] Server: The response is analyzed and the generation AI generates the next question: "What a lovely family. Where was this photo taken?"
[1992] Emotion engine: Analyzes the user's emotions from images and responses and recognizes that the user is feeling nostalgic.
[1993] Generative AI: Based on the user's emotions, the system further adjusts the questions and comments, such as "Tell me more about your memories of that time."
[1994] Terminal: Display the next question to the user.
[1995] Example 2: Travel photos
[1996] User: Select and upload photos of travel destinations.
[1997] Server: Receives the photo and performs image analysis to determine whether it contains a travel destination.
[1998] Generative AI: Generates questions like "Where is this place?" based on a photo.
[1999] Terminal: Display the question to the user.
[2000] User: Responds to the question with "This is a photo from Paris, where I went last year."
[2001] Terminal: Sends the response to the server.
[2002] Server: The response is analyzed and the generation AI generates the next question: "How was Paris? Were there any places that particularly impressed you?"
[2003] Emotion engine: Analyzes the user's emotions from images and responses and recognizes what the user is enjoying.
[2004] Generative AI: Further tailor questions and comments based on user sentiment, such as "What other places have you been?"
[2005] Terminal: Display the next question to the user.
[2006] In this way, this system allows users to enjoy conversations through photos, stimulating their brains and maintaining their energy. Furthermore, the introduction of an emotion engine enables communication that responds to the user's emotions, resulting in more effective dialogue.
[2007] The processing flow will be explained below.
[2008] Step 1:
[2009] The user opens the application on their device and proceeds to the photo upload screen.
[2010] Step 2:
[2011] The user selects a photo on their device and clicks the upload button.
[2012] Step 3:
[2013] The terminal generates a request to transfer the selected photo to the server.
[2014] Step 4:
[2015] The server receives the request from the terminal and receives the photo data.
[2016] Step 5:
[2017] The server stores the received photos in a database.
[2018] Step 6:
[2019] The server passes the stored photos to an image analysis algorithm, which begins the analysis process.
[2020] Step 7:
[2021] The server performs image analysis to determine the content of the photo, such as identifying people, objects, and the background.
[2022] Step 8:
[2023] The server sends the analysis results to the generation AI.
[2024] Step 9:
[2025] Based on the results of image analysis, the generative AI generates appropriate questions and comments for the user, such as "Where was this photo taken?" or "Who is this person?"
[2026] Step 10:
[2027] The server sends the generated questions and comments back to the device.
[2028] Step 11:
[2029] The terminal displays the questions and comments received from the server to the user.
[2030] Step 12:
[2031] The user enters a response to the question or comment on the terminal.
[2032] Step 13:
[2033] The terminal generates a request to send the user's response to the server.
[2034] Step 14:
[2035] The server receives the response from the terminal.
[2036] Step 15:
[2037] The server passes the received response to the emotion engine, which analyzes the user's emotion.
[2038] Step 16:
[2039] The emotion engine recognizes emotions from user responses and images and sends the results to the generation AI.
[2040] Step 17:
[2041] The generative AI generates the next appropriate question or comment based on the emotion recognition results from the emotion engine and the user's response. For example, if the user is feeling nostalgic, it generates a question such as, "Tell me more about your memories from that time."
[2042] Step 18:
[2043] The server sends the next generated question or comment back to the device.
[2044] Step 19:
[2045] The device will display the next question or comment to the user.
[2046] Step 20:
[2047] The user again enters a response, and this process is repeated until the user is satisfied.
[2048] Step 21:
[2049] The server stores all conversation logs.
[2050] Step 22:
[2051] The server will later analyze the saved conversation logs to evaluate the user's usage and reactions.
[2052] Example 2
[2053] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2054] In modern society, the number of elderly people and individuals who feel lonely is increasing, making it important to improve their mental health and increase communication opportunities for these people. In particular, there is a demand for systems that can provide mental support by promoting interaction through past memories and photos and engaging in emotionally appropriate dialogue. Conventional systems struggle to improve the quality of emotion recognition and dialogue, and there is a lack of methods for achieving effective communication.
[2055] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[2056] In this invention, the server includes means for receiving images and storing them in a database, means for generating image analysis results using an image analysis algorithm, means for analyzing the user's responses and generating the next question or comment using a generation AI, means for an emotion engine to recognize the user's emotions and for the generation AI to adjust the questions or comments based on the results, and means for saving all conversation logs and analyzing them later. This enables natural conversation based on the user's memories, provides appropriate communication according to emotions, and helps maintain and improve the user's mental health.
[2057] "Means for uploading images" refers to a series of operations and functions that allow a user to select digital images from their terminal and send them to the server.
[2058] A "database" is a system for systematically storing and managing digital information, where images and associated metadata are stored.
[2059] An "image analysis algorithm" is a computational method that analyzes the content of a digital image and extracts information about the people, objects, background, and other aspects of the photograph.
[2060] "Generative AI" is artificial intelligence that automatically generates questions and comments based on user responses and image analysis results.
[2061] An "emotion engine" is a technology that analyzes a user's emotions and adjusts the output of the generative AI based on their emotional state.
[2062] A "terminal" is a digital device used by a user, such as a smartphone or tablet.
[2063] A "conversation log" is a record of all interactions between a user and a system, data that is stored for later analysis.
[2064] A "prompt sentence" is the input text given to the generation AI, which is used to generate the next question or comment based on the user's response and the results of image analysis.
[2065] "Means for generating questions and comments" refers to the functions and operations by which the generation AI creates appropriate questions and comments for the user based on the prompt text.
[2066] This invention is a system that aims to stimulate the user's brain and maintain their energy by having them interact with a generative AI using photos of their memories. This system incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction.
[2067] System Overview
[2068] User-uploaded images
[2069] Users upload photos to the system using devices such as smartphones or tablets. Users launch a dedicated application, select memorable photos on the upload screen, and send them to the server. The uploaded photos are then transferred from the device to the server.
[2070] Receiving and analyzing images by the server
[2071] The server receives the uploaded photos and stores them in a database. The server then analyzes the content of the photos using image analysis algorithms (such as Google Cloud Vision API or Microsoft Azure's Computer Vision API). The analysis results include information about the people, objects, and background in the photos.
[2072] Question generation using generative AI
[2073] Based on the results of image analysis, the server sends the necessary data to a generative AI (such as OpenAI's GPT-4). Based on the received data, the generative AI generates appropriate questions and comments for the user. For example, it creates specific questions such as "Where was this photo taken?" or "Who is this person?"
[2074] Display of questions on the device
[2075] The generated AI's questions and comments are sent from the server to the user's device and displayed on the device. The user reads the displayed questions and comments and enters a response.
[2076] Parsing the user's response and generating the next question
[2077] The response entered by the user is sent from the device to the server. The server analyzes the received response and sends the results to the generation AI. The generation AI then generates the next most appropriate question or comment. This allows for a continuous dialogue with the user.
[2078] Emotion recognition and output adjustment by emotion engine
[2079] The emotion engine analyzes the user's emotions from images uploaded by the user and responses during conversations. The emotion engine evaluates whether the user is happy or anxious through facial expression and voice analysis. Based on the results of the emotion engine, the server adjusts the prompts sent to the generation AI and generates appropriate questions and comments.
[2080] Specific examples
[2081] Example 1: Lonely elderly people
[2082] The user selects and uploads an old family photo. The server receives the photo and uses an image analysis algorithm to recognize that a family member is in it. The generation AI generates the question, "Who is in this photo?" and displays it on the device. The user responds, "This is my daughter and grandchildren," and the device sends the response to the server. The server analyzes the response, and the generation AI generates the next question, "What a lovely family. Where was this photo taken?" Furthermore, if the emotion engine senses the user's nostalgia, it generates a question such as, "Tell me more about your memories from that time."
[2083] Example 2: Travel photos
[2084] The user selects and uploads photos of their travel destinations. The server receives the photos and uses an image analysis algorithm to recognize the travel destinations. The generation AI generates the question "Where is this place?" and displays it on the device. The user responds, "This is a photo of Paris, where I went last year," and the device sends the response to the server. The server analyzes the response, and the generation AI generates the next question, "How was Paris? Were there any places that particularly impressed you?" The emotion engine analyzes the user's emotions and, if it identifies that the user is enjoying themselves, generates questions such as "What other places did you go?"
[2085] Prompt Sentence Examples
[2086] As an example of a prompt for a generative AI model, input the following text to the generative AI:
[2087] Analyzing a photo submitted by a user reveals that it contains family members. Based on that content, generate questions or comments to naturally continue the conversation with the user. For example, "Who is in this photo?"
[2088] This system provides natural dialogue through users' memorable photos and realizes appropriate communication according to their emotions, which is expected to contribute to maintaining and improving users' mental health.
[2089] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2090] Step 1:
[2091] The user initiates the image upload process from their device.
[2092] Specific operation: The user launches the application on their smartphone or tablet, selects "Upload photos" from the menu, selects a memorable photo from their photo library, and presses the "Upload" button.
[2093] Input: A photo file selected by the user.
[2094] Output: The photo files are saved on your device and ready to be sent to the server.
[2095] Step 2:
[2096] The device sends the photo to the server.
[2097] Specific operation: The device makes an HTTP request to send the user-selected photo file to the server. The request includes the photo file and associated metadata (e.g., upload date and time, user ID).
[2098] Input: A user-selected photo file and its metadata.
[2099] Output: The photo files and metadata are sent to the server.
[2100] Step 3:
[2101] The server receives the photos and stores them in a database.
[2102] What it does: The server parses the incoming HTTP request and stores the photo file and its metadata in a database, along with information about the photo's identifier and storage location.
[2103] Input: The submitted photo file and its metadata.
[2104] Output: Photo files and associated metadata stored in a database.
[2105] Step 4:
[2106] The server analyzes the photo using image analysis algorithms.
[2107] How it works: The server sends the stored photo file to an image analysis algorithm (e.g., Google Cloud Vision API) to analyze the content of the photo, including people, objects, and background information.
[2108] Input: Photo files stored in the database.
[2109] Output: The analysis data generated as a result of the image analysis algorithm.
[2110] Step 5:
[2111] The generative AI creates prompts to generate questions and comments for the user.
[2112] Specific operation: The server creates a prompt sentence to send to the generation AI based on the image analysis results. This prompt sentence contains detailed information to facilitate the dialogue with the user. For example, "A family photo has been uploaded. Please generate a question about this photo."
[2113] Input: Image analysis results.
[2114] Output: The prompt sent to the generation AI.
[2115] Step 6:
[2116] Generative AI generates questions and comments.
[2117] How it works: A generative AI model (e.g., OpenAI's GPT-4) generates questions or comments to present to the user based on prompts received from the server, such as "Who is in this photo?"
[2118] Input: The server-generated prompt text.
[2119] Output: Generated questions and comments.
[2120] Step 7:
[2121] The server sends the generated questions and comments to the device.
[2122] Specific operation: The server creates an HTTP response to send the generated question or comment to the terminal, and sends it to the terminal.
[2123] Input: Generated questions and comments.
[2124] Output: Questions and comments sent to the user's device.
[2125] Step 8:
[2126] The device displays the question or comment to the user.
[2127] Specific operation: The device displays the received questions and comments on the screen and presents them to the user. An interface is provided so that the user can review them.
[2128] Input: Questions and comments sent by the server.
[2129] Output: Questions and comments displayed on the terminal screen.
[2130] Step 9:
[2131] The user enters a response to the question.
[2132] Specific behavior: The user reads the questions and comments displayed on the device and enters a response. For example, "These are my daughter and grandchildren."
[2133] Input: Questions or comments displayed on the device.
[2134] Output: The response text entered by the user.
[2135] Step 10:
[2136] The terminal sends the user's response to the server.
[2137] Specific operation: The terminal creates an HTTP request to send the response text entered by the user to the server.
[2138] Input: The response text entered by the user.
[2139] Output: The response text sent to the server.
[2140] Step 11:
[2141] The server analyzes the user's response and creates and sends the next prompt to the generation AI.
[2142] Specific operation: The server analyzes the user's response text, creates a new prompt based on its content, and sends the created prompt to the generation AI to generate the next question or comment.
[2143] Input: The user's response text.
[2144] Output: The next prompt sent to the generation AI.
[2145] Step 12:
[2146] Generative AI generates the next question or comment.
[2147] Specific behavior: The generative AI generates the next question or comment based on the newly received prompt, for example, "Where was this photo taken?"
[2148] Input: The newly created prompt statement.
[2149] Output: The next question or comment.
[2150] Step 13:
[2151] The emotion engine recognizes the user's emotions and adjusts the output of the generative AI.
[2152] Specific operation: The emotion engine analyzes the user's facial expressions and responses to evaluate the user's emotional state. Based on the results, it adjusts the generation AI to generate appropriate questions and comments.
[2153] Input: User's facial expression data and response content.
[2154] Output: Adjusted prompt and output of the generation AI.
[2155] Step 14:
[2156] The server sends the tailored questions and comments to the device.
[2157] Specific operation: The server sends questions and comments adjusted by the emotion engine to the terminal.
[2158] Input: Moderated questions and comments.
[2159] Output: The tailored questions and comments sent to your device.
[2160] Step 15:
[2161] The device displays the tailored question or comment to the user.
[2162] Specific operation: The device displays the adjusted questions and comments on the screen and presents them to the user.
[2163] Input: Moderated questions and comments.
[2164] Output: The adjusted questions and comments displayed on the screen.
[2165] Step 16:
[2166] The server stores all conversation logs.
[2167] What it does: The server records all interactions and stores them in a database. This information is later analyzed and used to improve the system and enhance the user experience.
[2168] Input: All dialogue content (questions, responses, and adjustment results).
[2169] Output: Conversation logs stored in a database.
[2170] (Application example 2)
[2171] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2172] The problem that this invention aims to solve is to realize more effective communication that responds to the user's emotions while stimulating the user's brain and maintaining their energy through dialogue using photos related to the user's memories. Conventional systems simply continue dialogue mechanically without taking the user's emotions into consideration, which has the problem of degrading the quality of the user experience.
[2173] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2174] In this invention, the server includes means for a user to upload images from a terminal, means for the server to receive and analyze the images, means for the generation AI to generate questions and comments for the user based on the generated analysis results, means for the terminal to present the generated questions and comments to the user, means for the terminal to send the user's response to the server, means for the server to analyze the response and for the generation AI to generate the next question or comment, means for analyzing emotions from the user's facial expressions and voice, means for the generation AI to adjust the content of the dialogue based on the emotion analysis results, means for continuously conversing with the user, and means for saving all conversation logs and analyzing them later. This allows the user to interact with the generation AI through memorable photos, and the content of the dialogue is appropriately adjusted through emotion analysis, thereby stimulating the user's brain and maintaining their energy.
[2175] "Means for users to upload images from their devices" refers to the function that allows users to send image data to the system using devices such as smartphones and tablets.
[2176] "Means for the server to receive and analyze images" refers to the function of the cloud server receiving uploaded image data and analyzing its contents using a specified image analysis algorithm.
[2177] "Means for the generation AI to generate questions and comments for the user based on the generated analysis results" refers to the function of using natural language processing technology to generate appropriate questions and comments based on the content of the image analyzed by the server.
[2178] "Means for the device to present the generated questions and comments to the user" refers to a function that displays the generated questions and comments on the user's smartphone or tablet.
[2179] "Means for transmitting a user's response from the terminal to the server" refers to a function for transmitting the response data entered by the user again to the cloud server.
[2180] "The server analyzes the response, and the generation AI generates the next question or comment" refers to the function in which the cloud server analyzes the user's response and generates the next conversation content based on the analysis results.
[2181] "Means for analyzing emotions from the user's facial expressions and voice" refers to a function that analyzes the facial expressions and voice data shown by the user during a conversation and determines their emotional state (joy, sadness, surprise, etc.).
[2182] "Means for the generation AI to adjust the content of the dialogue based on the results of emotion analysis" refers to the function by which the generation AI takes into account the user's emotional state obtained through emotion analysis and appropriately adjusts the next question or comment.
[2183] "Means for continuing conversation with the user" refers to a function that allows a dialogue between the user and the system to continue.
[2184] "A means of saving all conversation logs and analyzing them later" refers to a function that records and saves all generated conversation content and later analyzes changes in user behavior patterns and emotions based on that data.
[2185] This invention provides a system that stimulates the user's brain and maintains their morale by interacting with a generating AI using memorable photos of the user. This system also incorporates an emotion engine that recognizes the user's emotions, improving the quality of the interaction. Each element of this system is explained below.
[2186] System configuration
[2187] 1. A way for users to upload images from their devices
[2188] Users use their smartphones, tablets, or other devices to select and upload images within the application, a process designed to make it easy and intuitive for users to submit photos to the system.
[2189] 2. How the server receives and analyzes images
[2190] Uploaded photos are transferred to a cloud server, which uses image analysis libraries such as the Google Vision API to automatically analyze the content of the photo and extract information about people, places, objects, and more.
[2191] 3. A means for the generative AI to generate questions and comments for users based on the generated analysis results
[2192] The results of the image analysis are sent to a generative AI (e.g., GPT-3 or GPT-4), which uses this data to generate appropriate questions and comments for the user.
[2193] For example: "Where was this photo taken?"
[2194] 4. A means for the device to present generated questions and comments to the user
[2195] The questions and comments generated by the AI are sent from the server to the user's device, where they are displayed to the user, and a dialogue begins.
[2196] 5. A means of transmitting user responses from the terminal to the server
[2197] When a user responds to a question or comment through their device, the response is sent to the cloud server.
[2198] 6. The server analyzes the response and the AI generates the next question or comment.
[2199] The server analyzes the user's responses and sends the results to the generation AI, which then generates new questions and comments to continue the dialogue with the user.
[2200] 7. Means of analyzing emotions from user facial expressions and voice
[2201] Using an emotion engine (for example, Microsoft Azure Emotion API), emotions are analyzed from the user's facial expressions and voice. The analysis results indicate the user's emotional state, such as whether they are happy or sad.
[2202] 8. How generative AI can adjust dialogue content based on emotion analysis results
[2203] Based on the analysis results of the emotion engine, the generative AI will adjust the dialogue appropriately. For example, if the user is sad, it will generate comforting comments, and if the user is happy, it will generate questions that will bring out that emotion.
[2204] For example: "Tell me more about your memories of that time."
[2205] 9. A way to continue the conversation with your users
[2206] The dialogue continues until the user is satisfied. The entire conversation process is controlled by the system.
[2207] 10. A way to store all conversation logs and analyze them later
[2208] All generated conversation logs are stored on a cloud server, allowing users' usage patterns and reactions to be analyzed at a later date, helping to improve the system and enhance the user experience.
[2209] This system allows users to enjoy conversations through memorable photos, stimulating their brains and maintaining their energy. In addition, the introduction of an emotion engine enables communication based on the user's emotions, resulting in more effective conversations.
[2210] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2211] Step 1:
[2212] A way for users to upload images from their devices
[2213] A user opens the application on a smartphone or tablet, selects a photo from the upload screen, and sends it to the system. The input is the photo file selected by the user, and the output is that the photo file is sent to the cloud server. The photo file is transferred between the device and the cloud server.
[2214] Step 2:
[2215] The server receives and analyzes the images
[2216] The server receives the photos sent from the device and stores them in a database. The stored photos are then passed to the Google Vision API for image analysis. The input is the uploaded photo file, and the output is the analyzed image content (people, places, objects, etc.) data. Image analysis uses machine learning algorithms to analyze the content of the photo.
[2217] Step 3:
[2218] A means for the AI to generate questions and comments for users based on the generated analysis results
[2219] The server sends the image analysis results to a generation AI (GPT-3 or GPT-4), which generates appropriate questions and comments. The input is the image analysis results, and the output is questions and comments to present to the user. Specifically, the analysis results are input as prompts to the generation AI, which then uses natural language processing technology to generate corresponding questions and comments.
[2220] Step 4:
[2221] A means by which the device presents generated questions and comments to the user
[2222] The server sends the generated questions and comments to the terminal and displays them to the user. The input is the generated questions and comments, and the output is what is displayed to the user. The terminal receives this and displays it on the screen.
[2223] Step 5:
[2224] A means for transmitting user responses from the terminal to the server
[2225] Users use their devices to respond to questions and comments, and then send the response data to the cloud server. The input is the user's response, and the output is the response data sent to the cloud server. Specifically, the user enters text and presses the send button, which sends the data to the server.
[2226] Step 6:
[2227] The server analyzes the response and the generation AI generates the next question or comment.
[2228] The server analyzes the user's response data and sends the analysis results to the generation AI to generate the next question or comment. The input is the user's response data, and the output is the newly generated question or comment. The response data is analyzed for text, and its content is provided to the generation AI as a prompt.
[2229] Step 7:
[2230] A means of analyzing emotions from the user's facial expressions and voice
[2231] The server receives the user's facial expression and voice data sent from the device and performs emotion analysis using an emotion engine (Microsoft Azure Emotion API). The input is facial expression and voice data, and the output is data on the user's emotional state (e.g., joy, sadness, surprise, etc.). Specifically, image and voice data is passed to the emotion engine for analysis.
[2232] Step 8:
[2233] A means for generative AI to adjust dialogue content based on emotion analysis results
[2234] The server provides the emotion analysis results to the generation AI, which then adjusts the dialogue appropriately based on the results. The input is the emotion analysis results, and the output is adjusted questions or comments. The generation AI receives prompts containing the emotion analysis results and generates appropriate dialogue content based on them. For example, if it determines that the user is sad, it generates a question such as, "Tell me more about your memories of that time."
[2235] Step 9:
[2236] A way to maintain an ongoing conversation with users
[2237] The server and the generation AI work together to continue the dialogue until the user is satisfied. The input is continuous responses from the user, and the output is newly generated questions and comments. The dialogue content is continuously generated, and the interaction with the user progresses without interruption.
[2238] Step 10:
[2239] A means to store all conversation logs and analyze them later
[2240] The server stores all dialogue logs and later analyzes changes in user behavior patterns and emotions based on them. The input is all the dialogue logs generated, and the output is the stored data and the analysis results. The log data is stored in a database and later analyzed using analysis tools.
[2241] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2242] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2243] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2244] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2245] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric...
Claims
1. a means for a user to upload an image from a terminal; means for the server to receive and analyze the images; A means for the generation AI to generate questions and comments for the user based on the generated analysis results; means for the terminal to present the generated questions and comments to the user; means for transmitting a user's response from the terminal to a server; The server analyzes the response and the AI generates the next question or comment. a means for carrying out an ongoing conversation with the user; A means to store all conversation logs and analyze them later, A system including:
2. 2. The system according to claim 1, wherein all conversation logs are saved and analyzed at a later date to analyze the usage status and reactions of users.
3. The system according to claim 1, wherein the generating AI has a function of generating new questions and comments based on the user's responses.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A