system

A system that uses natural language processing to generate questions from past photographs, facilitating dialogues to improve cognitive function and emotional engagement in elderly individuals.

JP2026085764APending Publication Date: 2026-05-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-11-13
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

Elderly individuals living alone often lack social interaction, leading to insufficient brain activation and accelerated dementia due to reduced memory recall opportunities, necessitating a system to engage them in conversations based on past memories.

Method used

A system that retrieves image data from photographic storage, generates questions using natural language processing, and facilitates a dialogue process through a terminal device to recall past memories, thereby maintaining and improving cognitive function.

Benefits of technology

The system effectively activates the brain by prompting users to recall past memories through interactive dialogues, enhancing cognitive function and emotional engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026085764000001_ABST
    Figure 2026085764000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for acquiring image data stored on a photographic storage medium, A means for generating a question using natural language processing based on the aforementioned image data, The means of transmitting the generated question to a terminal device and displaying it to the user, A means for obtaining a response from a user and generating a new question based on the response, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the case of the elderly, especially those living alone, when there is no one to talk to on a daily basis, the activation of the brain becomes insufficient, which may cause the progression of dementia to accelerate due to this. In particular, due to the decrease in the opportunity to recall past memories, there is concern about the decline of important brain functions such as memory and thinking ability. The present invention aims to provide conversations based on old photos stored in photo storage to activate the brains of the elderly in order to address such a situation.

Means for Solving the Problems

[0005] This invention provides means for acquiring image data stored on a photographic storage medium and means for generating questions using natural language processing based on this image data. By transmitting the generated questions to a terminal device and displaying them to the user, the user has the opportunity to recall past memories. Furthermore, by acquiring responses from the user and generating new questions based on those responses, a series of dialogues is realized, promoting recollection in the elderly. Through this dialogue process, it is possible to maintain and improve cognitive function.

[0006] A "photographic storage medium" is a digital storage system for saving image data.

[0007] "Image data" refers to visual information such as photographs and paintings that are stored in digital format.

[0008] "Natural language processing" is a technology that enables computers to understand, generate, and process human language.

[0009] A "means of generating questions" is a method of creating questions that ask for relevant information based on image data.

[0010] A "terminal device" refers to an electronic device that a user operates to display and process digital data.

[0011] "Means of displaying to the user" refers to methods of presenting information or messages visually or audibly via a terminal device.

[0012] "Means of obtaining responses" refers to methods of collecting verbal or written responses from users.

[0013] "Means of enabling dialogue" refers to methods of maintaining communication with the user through a series of questions and answers. [Brief explanation of the drawing]

[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

MODE FOR CARRYING OUT THE INVENTION

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] This invention is a system that provides conversational content using past photographs, with the aim of maintaining and improving the cognitive function of elderly people. The individual components of this system are described in detail below.

[0036] First, the server accesses the photo storage medium and retrieves the stored image data. The photo storage medium is digital storage containing photos that the user has saved in the past, and the server accesses it via an API to identify what kind of memories the user has.

[0037] The server performs natural language processing on the acquired image data to generate questions related to the photos. For example, if the photos are from a specific trip, a question like, "Please tell us about your fondest memories of this place," might be generated. These questions are customized using image metadata and AI-powered image recognition.

[0038] Next, the generated question and its corresponding photo are sent from the server to the terminal. The terminal receives this information and displays the photo and question to the user. The interface is designed to be intuitive and visually easy to understand, allowing the user to recall past experiences.

[0039] The user responds to questions displayed on the device using spoken or written language. This response is then sent to the server via speech recognition and text analysis technologies.

[0040] The server then reviews the user's response and generates further questions. This is to allow the user to delve deeper into their memory and continue the conversation by analyzing their response. For example, if the user talked about travel, the next question might be, "What was the most memorable event from that trip?"

[0041] In this way, by repeating the flow of dialogue through photographs and questions, the user's past memories are recalled and activated. This helps to maintain the user's cognitive function and, consequently, contribute to improving their health. This series of steps constitutes a specific embodiment of the present invention.

[0042] The following describes the processing flow.

[0043] Step 1:

[0044] The user logs into the photo storage service and grants permission to access their photo data. This process begins with secure access to the account through an authentication process.

[0045] Step 2:

[0046] The server receives the user's authentication information and retrieves image data stored on the photo storage medium from that account. Here, the photos are organized based on date and tag information.

[0047] Step 3:

[0048] The server references the metadata of the image data and applies a selection algorithm to determine which photos are most relevant to the user's memories.

[0049] Step 4:

[0050] The server analyzes the content related to the selected photo and uses natural language processing techniques to generate questions about that photo. For example, if it's a travel photo, it might generate a question like, "What is your most memorable experience at this place?"

[0051] Step 5:

[0052] The server sends the generated question and photo to the terminal. This communication is conducted using a secure data transmission protocol.

[0053] Step 6:

[0054] The device configures and displays an interface on the screen to visually present the received question and photo to the user. Here, the photo is displayed prominently, with the question shown below in text format.

[0055] Step 7:

[0056] The user answers the questions displayed on the screen using voice or text. This response is collected using the device's input device.

[0057] Step 8:

[0058] The device sends the response data obtained from the user to the server. In this process, the voice data may be converted to text.

[0059] Step 9:

[0060] The server analyzes the user's responses and generates new questions. This analysis uses an AI model to construct new questions that are relevant to past responses.

[0061] Step 10:

[0062] The server sends a newly generated question to the terminal and resumes the interaction with the user. This process is repeated as long as the user remains interested.

[0063] (Example 1)

[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0065] Currently, many elderly people suffer from cognitive decline, and there is a need for effective means to maintain and improve their cognitive function. Traditional methods often rely on monotonous memory training and limited interaction with others, which is inefficient. Furthermore, there are currently few interactive tools that promote memory activation.

[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0067] In this invention, the server includes means for acquiring image information stored in a data storage device, means for generating question and answer using natural language processing based on the image information, and communication means for securely transferring information. This makes it easier for users to recall memories through images, and enables effective maintenance and improvement of cognitive function through interactive dialogue.

[0068] A "data storage device" is a physical or virtual container for storing and managing information in digital format.

[0069] "Image information" refers to a collection of visual data that has been digitized as still images or videos and stored electronically.

[0070] "Natural language processing" is a field of technology that enables computer systems to understand, analyze, and generate human language.

[0071] "Question and answer" is a two-way communication method consisting of questions generated based on specific information and the answers provided to them.

[0072] A "display device" is hardware used to visually present digital information to a user, and includes screens, monitors, and other similar devices.

[0073] A "response" is an audio or written response given by a user to the information presented.

[0074] "Communication methods" refer to methods and technologies for safely and efficiently sending and receiving data.

[0075] "Memory activation" refers to the cognitive process of recalling past events and re-recognizing information related to them.

[0076] "Dialogue" is the process or means by which two or more people exchange information.

[0077] This invention is a dialogue system aimed at enhancing the user's cognitive functions using digitized past memories. The specific implementation details are shown below.

[0078] The server first accesses the data storage device to retrieve image information saved by the user. This image information is stored in the user's cloud storage service (for example, a common online storage platform). The server also uses an API to access and identify the relevant photos through specific metadata.

[0079] The server applies natural language processing to the acquired image information and uses a generative AI model to generate questions and answers. Specifically, it analyzes the content of the photograph using an image recognition API and creates questions based on the results. This Q&A can be obtained by inputting the generated text as a prompt. For example, a possible prompt sentence is "What memories are associated with this photograph?"

[0080] The generated question-and-answer format is transmitted from the server to the terminal via a secure communication method. The terminal consists of smartphones and tablet devices, and uses a dedicated application to display images and questions to the user through a user interface. This allows the user to interact intuitively.

[0081] Users can respond to displayed questions with voice or text. This allows users to recall and deepen their memories. Voice information is converted into text data using speech recognition software and sent to the server.

[0082] Through this process, the server generates new questions and continues the dialogue with the user. This makes it easier for the user to recall past events, resulting in the maintenance and improvement of their cognitive function.

[0083] The system of this invention aims to provide a novel method for enhancing users' cognitive abilities through interaction, thereby creating a comfortable and effective user experience.

[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0085] Step 1:

[0086] The server retrieves image information from a data storage device. The input to this process is photo data stored in a cloud storage service. The server searches for image information via an API and retrieves a specific photo previously saved by the user. The output is an image file containing the user's metadata.

[0087] Step 2:

[0088] The server applies natural language processing to the acquired image information. The input for this step is the image information acquired in step 1. The server uses an image recognition API to analyze the objects and location information of the photograph. The result of the data processing is metadata that can be used by the generative AI model, and this is used to generate a question-and-answer session.

[0089] Step 3:

[0090] The server generates question-and-answer responses using a generative AI model. The input for this step is the metadata and prompt template generated in step 2. The server inputs the prompt into the generative AI model, and the AI ​​creates customized questions. The output is the customized questions presented to the user.

[0091] Step 4:

[0092] The server sends the generated question and answer to the terminal. The input is the question and image information generated in step 3. The data is transmitted via a secure communication method and received by the terminal. The output is the question and photo displayed on the terminal.

[0093] Step 5:

[0094] The terminal displays images and questions to the user through a user interface. This step receives the questions and images, which are the output from step 4, as input. Specifically, it displays a photograph on the display device and shows an interactive question box. It provides a visually easy-to-understand interface for the user.

[0095] Step 6:

[0096] The user responds to the displayed question. The input is the question presented in step 5. If the user responds verbally, the terminal uses speech recognition technology to convert it into text data. The output is the user's response text sent to the server.

[0097] Step 7:

[0098] The server generates additional questions based on the user's responses. The input for this step is the text data received from step 6. The server analyzes the responses and generates the next question by inputting new prompt sentences into the generating AI model. The output is a new question to facilitate further dialogue.

[0099] (Application Example 1)

[0100] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0101] Traditional virtual stores typically offer suggestions to customers, but lack personalized experiences that reflect individual purchase history and preferences. Furthermore, it's difficult for customers to intuitively select products that reflect their preferences within the virtual space. This results in insufficient improvement in customer satisfaction and stimulation of purchasing intent.

[0102] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0103] In this invention, the server includes means for acquiring image information stored on a photographic storage medium, means for generating questions using natural language processing based on the image information, and means for providing a personalized customer service experience based on past transaction history and object information. This enables suggestions tailored to the individual preferences of the customer and effectively stimulates the customer's purchasing intent through a personalized experience.

[0104] A "photo storage medium" is a storage device that digitally holds image information previously taken or saved by a user.

[0105] "Image information" refers to visual data stored on a photographic storage medium, and includes metadata related to specific objects or scenes.

[0106] Natural language processing is a technology that enables computers to understand, interpret, and generate human language, and is used to generate relevant questions from image information.

[0107] A "terminal device" is a device used by users for interaction, and its role is to display questions and suggestions and receive responses.

[0108] A "response" is the verbal or written feedback that a user provides in response to a question, which forms the basis for generating further dialogue and suggestions.

[0109] A "virtual space" is a visual, three-dimensional environment created using digital technology, providing a space for users to interact.

[0110] "Target information" refers to data about products and services related to past transactions and customer interests.

[0111] "Customer service experience" refers to a personalized shopping environment provided through user interaction during the shopping process.

[0112] This invention provides an interactive system for improving cognitive function in the elderly, and specific embodiments thereof are shown below.

[0113] The server accesses the photo storage medium and retrieves image information previously saved by the user in digital format. This image information includes metadata about specific objects and scenes. The server uses natural language processing techniques to generate relevant questions based on the retrieved image information. These questions are customized based on the image content and metadata. For example, a question such as "What are your fondest memories of this place?" might be generated from a photograph of a specific trip.

[0114] The generated questions are sent to a terminal and presented to the user. The terminal is a device that interacts with the user, displaying questions and corresponding images, and receiving responses from the customer verbally or in writing. The user's responses are sent to a server, which analyzes their content and uses it as a basis for generating new questions and suggestions.

[0115] Furthermore, the server can provide a personalized customer service experience based on past transaction history and product information. It presents products to the user in a virtual space, engages in real-time interaction, and makes optimal suggestions to increase the user's interest.

[0116] The system is implemented using Python on the Django framework, and as mentioned above, it uses the A-Frame framework for VR space construction and the Google® Speech-to-Text API for voice input. This allows users to explore products three-dimensionally in a virtual space and enjoy a personalized purchasing experience based on their interests.

[0117] As a concrete example, a conversation might take place about sneakers the user has purchased in the past. If the user responds to a prompt such as, "Please tell me about your experience when you first wore these sneakers," with "They were very comfortable and I could wear them all day without getting tired," the next question might be, "In what situations do you usually wear them since then?" In this way, the system helps maintain cognitive function by retrieving and activating the user's memories.

[0118] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0119] Step 1:

[0120] The server accesses the photo storage medium and retrieves image information previously saved by the user. This image information is visual data stored in digital format and also includes metadata. The server accesses the database via an API and efficiently reads this data.

[0121] Step 2:

[0122] The server uses natural language processing techniques to generate relevant questions based on the acquired image information. Using the acquired image information and its metadata as input, and employing image recognition algorithms, it identifies specific objects and scenes, forming prompt sentences. For example, it might generate a question in natural language such as, "Please tell me about your fondest memories of this place."

[0123] Step 3:

[0124] The server sends the generated question to the terminal. The terminal displays the presented question and associated image to the user. Here, the user can intuitively view the photo and question on the interface to facilitate recalling past memories. The terminal receives the question and image sent from the server as input and displays them visually to the user.

[0125] Step 4:

[0126] The user responds to the displayed question verbally or in text. The device converts the user's voice response into text data using the Google Speech-to-Text API. This converted text is sent to the server as input for use in the next step.

[0127] Step 5:

[0128] The server analyzes user responses and generates new questions and suggestions. Using natural language processing technology, it analyzes keywords and sentiments within the text data to generate questions that further engage the user's interest in the next step. For example, if a user responds that they were "very comfortable," the server can generate a secondary question such as, "In what situations do you usually wear them now?"

[0129] Step 6:

[0130] The server uses past transaction history and related information to construct optimal recommendations within the virtual space. These recommendations are designed to increase user interest and provide a personalized purchasing experience. The recommendations are sent to the user's device and displayed in real time within the virtual space.

[0131] Step 7:

[0132] The terminal plays the role of constructing a virtual space and provides the user with a three-dimensional environment using the A-Frame framework. This allows the user to intuitively and effectively consider products based on personalized information. It receives suggestions from the server as input and generates a virtual environment that visually represents them.

[0133] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0134] This invention provides a system that allows elderly individuals to engage in dialogue based on past photographs, promoting the maintenance of cognitive function while simultaneously offering interactions that take their emotional state into consideration. Embodiments of this invention are described below from the perspectives of the terminal, server, and user.

[0135] First, the user logs into the photo storage service and grants the system permission to access their photo archive. Once the user begins using the system, the server retrieves image data stored on the user's photo storage media. This image data is organized based on date and tags, and photos that are emotionally meaningful to the user are extracted.

[0136] For the acquired image data, the server uses natural language processing technology to generate questions related to the photograph. At this time, the server is equipped with an emotion engine that analyzes the user's past conversation history and emotional state data to adjust the content and tone of the questions.

[0137] Next, the server sends the generated question and its corresponding photo to the terminal. The terminal then displays these to the user visually and audibly, setting up an intuitive and user-friendly interface to present the question along with the photo.

[0138] The user responds to questions displayed on the device using voice or text. When the device receives the response data, it analyzes the user's voice tone and the emotions contained in the text, and sends this information to the server in real time.

[0139] As soon as the server receives a user's response, it uses an emotion engine to analyze the user's emotional state and adjusts the next question based on this information. For example, it may continue asking in-depth questions on topics the user found particularly interesting, or move to a different topic if the user indicates discomfort.

[0140] This allows users to experience dialogue that takes their current emotional state into consideration and enjoy richer interactions. This invention enhances user psychological satisfaction while simultaneously contributing to the maintenance of healthy cognitive function.

[0141] The following describes the processing flow.

[0142] Step 1:

[0143] The user logs into the photo storage service and grants the system permission to access image data. During this process, the user's account information is authenticated.

[0144] Step 2:

[0145] The server retrieves image data from the user's photo storage. This includes a retrieval process using an API. The server organizes the retrieved images based on date and tag information.

[0146] Step 3:

[0147] The server uses an algorithm to select photos that are emotionally important to the user from the organized image data. This selection is based on the photo's metadata and the user's past preferences.

[0148] Step 4:

[0149] The server generates questions using natural language processing techniques related to the selected photo. Simultaneously, it uses an emotion engine to adjust the tone and content of the questions to match the user's past conversation history and emotional state.

[0150] Step 5:

[0151] The server sends the generated question and associated photo to the terminal. The terminal receives this and prepares to display the photo and question in the user interface.

[0152] Step 6:

[0153] The device displays photos and questions to the user through a visually appealing and intuitive interface. Information is presented in a way that makes it easy for the user to recall past events.

[0154] Step 7:

[0155] The user responds to the questions displayed on the device using voice or text. The user's response is analyzed by an emotion engine, which considers the tone of their voice and the content of their sentences.

[0156] Step 8:

[0157] The device transmits emotional data contained in the user's voice and text responses to the server in real time.

[0158] Step 9:

[0159] The server generates new questions based on the user's responses and sentiment analysis results. In this process, questions that take into account the user's current emotional state are selected.

[0160] Step 10:

[0161] The server sends the newly generated question back to the terminal and then resumes the conversation with the user. This process is repeated as long as the user wishes to continue the conversation.

[0162] (Example 2)

[0163] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0164] In the elderly, there is a need for interactive systems that maintain cognitive function and improve emotional state based on past memories. However, existing technologies struggle to generate questions that fully consider emotional meaning and to adjust dialogue according to the user's real-time emotional state. Therefore, the challenge is to provide an interactive dialogue system that improves user psychological satisfaction and supports healthy cognitive function.

[0165] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0166] In this invention, the server includes means for acquiring data from an image storage device, means for analyzing information related to the image using natural language processing technology and generating questions, and means for analyzing the user's response with an emotion analysis algorithm and adjusting the content of the next dialogue. This makes it possible to generate questions and adjust the dialogue while taking into account the user's emotional state, thereby effectively improving the maintenance of cognitive function and emotional satisfaction in the elderly.

[0167] An "image storage device" is a storage medium for recording and storing digital image data.

[0168] "Data" refers to a series of recorded events and figures that include information about a user.

[0169] "Natural language processing technology" is a technology that enables computers to understand and generate human language.

[0170] "Question generation" is the process of constructing questions based on a specific purpose or context in order to elicit answers.

[0171] An "output device" is a device used to display or transmit digital information in a way that is perceptible to the user.

[0172] "Analysis" is the process of breaking down input information or data and understanding its meaning and relationships.

[0173] An "emotion analysis algorithm" is a computational method used to identify emotional states from text or audio data.

[0174] "Dialogue adjustment" is a process that dynamically changes the content and progress of a dialogue based on the user's responses and status.

[0175] This system is an interactive platform designed to help elderly people maintain cognitive function and improve emotional well-being by allowing them to engage in conversations based on past photographs.

[0176] First, the user logs into the photo storage service and grants the system permission to access their photo archive. This process begins with the use of a user ID and password.

[0177] Once login is complete, the server retrieves the user's image data from the image storage device. This includes photos uploaded from digital cameras and smartphones. Based on this data, the server uses dates and tags to select images that appear to be emotionally important.

[0178] The server uses natural language processing techniques to generate relevant questions for selected images. This involves a generative AI model, and an emotion engine takes into account the user's emotional state and past conversation history. This method provides the user with highly relevant content.

[0179] Next, the server sends the generated question and associated photos to the terminal. The terminal provides the user with an intuitive and easy-to-use interface, visually through the display and audibly through the speaker. The user can then interact with the associated question along with their favorite photos.

[0180] When a user responds to a question, the device captures the response as voice or text and sends it to the server in real time. The device uses a speech recognition system to convert the response to text and also analyzes the tone of voice.

[0181] For example, when a user selects a photo of autumn foliage, the server might generate a question such as, "Who did you go to this place with?". An example of a prompt might be, "Create a question to elicit memories related to the user's photo. The photo shows a landscape with autumn leaves."

[0182] Based on the above, this system aims to analyze the user's emotional state, improve psychological satisfaction through appropriate dialogue, and maintain healthy cognitive function in the elderly.

[0183] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0184] Step 1:

[0185] The user logs in to the photo storage service. In this step, the user enters their user ID and password, and their access rights to the system are verified. After successful login, the system grants the user access to their photo archive. As output, the user's identification information is sent to the server, and the user session begins.

[0186] Step 2:

[0187] The server retrieves user image data from an image storage device. Inputs include user identification information and associated data storage information. Image data is extracted from the database, and emotionally significant images are selected based on date and tags. The output is a list of images meaningful to the user, serving as intermediate data for processing by the server.

[0188] Step 3:

[0189] The server uses selected image data and natural language processing techniques to generate questions. The input consists of selected images and tag information, and contextual analysis is performed by a generative AI model. This analysis outputs questions related to the content of the photos, which are then stored on the server. The data processing used here includes text analysis and question output.

[0190] Step 4:

[0191] The server sends the generated question and its corresponding image to the terminal. The input is the generated question and image data. The output is the visual and auditory content displayed on the user interface. The terminal adjusts the interface to display the question on the screen and to present it audibly using its speaker.

[0192] Step 5:

[0193] The user responds to questions displayed on the device using voice or text. The input is the user's voice or text response, and the output is recognized text data. The device uses speech recognition to convert the response to text and sends it to the server in real time. The response data is used for sentiment analysis in the next step.

[0194] Step 6:

[0195] The server receives the user's response and analyzes it using an emotion analysis algorithm. The input is converted text data, and the analysis identifies the user's emotional state. The output is an emotion statement, which is used as input data when constructing the next dialogue. This allows for the adjustment of dialogue content to reflect the user's emotions appropriately.

[0196] (Application Example 2)

[0197] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0198] There is a need to provide a system that supports the maintenance of cognitive function in the elderly while enabling interaction that leverages past memories. However, conventional systems have struggled to achieve effective interaction that fully reflects the emotions and past memories of individual users. Furthermore, there is a lack of means to enhance affinity for the elderly by personalizing the shopping experience in a virtual environment.

[0199] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0200] In this invention, the server includes means for acquiring image data stored on a photographic storage medium, means for generating questions using natural language processing based on the image data, and means for cross-referencing the acquired image data with product information to identify products associated with the user's past memories. This allows the user to relive past memories while receiving product information linked to those memories. Furthermore, by using this information to provide a friendly purchasing experience through heartwarming dialogue, it is possible to support the maintenance of cognitive function in the elderly and realize interactions that are sensitive to their feelings.

[0201] A "photo storage medium" is a digital storage device that stores a user's past photographic data and enables the retrieval of image data.

[0202] "Image data" refers to digital information of still images stored on a photographic storage medium, and it serves as the basic information for generating questions and dialogues based on its content.

[0203] "Natural language processing" is a technology that enables computers to understand and generate human language, and in this system, it is used for generating questions and analyzing user responses.

[0204] "Means for generating questions" refers to a function that creates questions to ask the user based on image data, providing a foundation for interactions related to the user's past memories.

[0205] A "terminal device" is hardware that a user uses as an interface, and it is a device that presents generated questions and identified product information to the user visually and audibly.

[0206] A "means of obtaining a response" refers to a function that receives a response from the user, either in voice or text, and uses it as the basis for the next interaction.

[0207] "Cross-referencing" is a technique that compares acquired image data and information derived from it with another dataset, such as product information, to find relationships between them.

[0208] "Means of displaying related information" refers to a function that visually or audibly presents product information related to past photographs to the user, thereby linking memories with the current purchasing experience.

[0209] "Means for generating heartwarming dialogue" refers to technologies that create interactions that take into account the user's emotions and promote emotional exchange that goes beyond mere information provision.

[0210] The system of this invention operates via the user's smart device, a cloud-based server, and an internet connection. The user first grants the system access rights to a photo storage service. This allows the server to retrieve image data from the photo storage medium.

[0211] The server uses natural language processing techniques to generate relevant questions based on the acquired image data. In this process, it leverages an emotion engine (e.g., Google Cloud Natural Language API) to analyze the user's past conversation history and emotional state data. The generated questions and related product information are then sent to the user's smart glasses interface.

[0212] The device presents information to the user visually and audibly, and accepts responses in voice or text. These responses are analyzed in real time and sent to a server to assess the user's emotional state. The server then uses this information to tailor the next interaction, generating in-depth discussions about topics or products of particular interest to the user.

[0213] For example, if a product related to a "birthday" photo taken by a user in the past is presented in a virtual store, the user might interact with the product by saying, "This is the cake you photographed with your child in {year}. Do you remember?" This information is provided to the user through an intuitive interface, resulting in a familiar and engaging shopping experience.

[0214] An example of a prompt statement is set as follows:

[0215] "To suggest products related to past memories, use photos and user stories related to 'birthdays' to create heartwarming conversations."

[0216] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0217] Step 1:

[0218] The server retrieves image data stored on the photo storage medium after the user logs into the photo storage service. In this step, all image files associated with the specific user are retrieved from the accessed database, and the metadata contained in each file (e.g., date, tags) is organized. The input is the account information for the photo storage service, and the output is the organized image data set.

[0219] Step 2:

[0220] The server uses the acquired image data to generate questions related to the user's past photos through natural language processing. The input here is the image data acquired in the previous step, and by utilizing the emotion engine, questions are output that take into account the user's past conversation history and emotional state.

[0221] Step 3:

[0222] The server cross-references the generated questions and retrieved image data with the product information database to identify products related to the user's past memories. The input for this step is image data and question information based on dates and tags, and the output is related product information.

[0223] Step 4:

[0224] The terminal visually and audibly presents the user with questions and related product information sent from the server. The input here consists of questions and product information from the server, which are output to the user through an intuitive interface. Specifically, the information is displayed on the smart glasses' screen using video and audio.

[0225] Step 5:

[0226] The user responds to questions displayed on the device using voice or text. The input consists of the presented question and the user's voice / text response, while the output is response data that includes the user's intent.

[0227] Step 6:

[0228] The server analyzes the user's response and uses an emotion engine to evaluate the user's emotional state. The input for this step is the user's response data, and the output is the user's emotional state and tailored question information to generate the next dialogue.

[0229] Step 7:

[0230] The server adjusts the next interaction based on the user's emotional state, generates new questions and relevant product information as needed, and sends them back to the terminal. The input is emotional information and dialogue data as analysis results, and the output is interaction information as needed. This cycle allows the user to experience emotionally sensitive interactions.

[0231] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0232] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0233] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0234] [Second Embodiment]

[0235] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0236] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0237] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0238] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0239] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0240] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0241] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0242] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0243] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0244] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0245] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0246] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0247] This invention is a system that provides conversational content using past photographs, with the aim of maintaining and improving the cognitive function of elderly people. The individual components of this system are described in detail below.

[0248] First, the server accesses the photo storage medium and retrieves the stored image data. The photo storage medium is digital storage containing photos that the user has saved in the past, and the server accesses it via an API to identify what kind of memories the user has.

[0249] The server performs natural language processing on the acquired image data to generate questions related to the photos. For example, if the photos are from a specific trip, a question like, "Please tell us about your fondest memories of this place," might be generated. These questions are customized using image metadata and AI-powered image recognition.

[0250] Next, the generated question and its corresponding photo are sent from the server to the terminal. The terminal receives this information and displays the photo and question to the user. The interface is designed to be intuitive and visually easy to understand, allowing the user to recall past experiences.

[0251] The user responds to questions displayed on the device using spoken or written language. This response is then sent to the server via speech recognition and text analysis technologies.

[0252] The server then reviews the user's response and generates further questions. This is to allow the user to delve deeper into their memory and continue the conversation by analyzing their response. For example, if the user talked about travel, the next question might be, "What was the most memorable event from that trip?"

[0253] In this way, by repeating the flow of dialogue through photographs and questions, the user's past memories are recalled and activated. This helps to maintain the user's cognitive function and, consequently, contribute to improving their health. This series of steps constitutes a specific embodiment of the present invention.

[0254] The following describes the processing flow.

[0255] Step 1:

[0256] The user logs into the photo storage service and grants permission to access their photo data. This process begins with secure access to the account through an authentication process.

[0257] Step 2:

[0258] The server receives the user's authentication information and retrieves image data stored on the photo storage medium from that account. Here, the photos are organized based on date and tag information.

[0259] Step 3:

[0260] The server references the metadata of the image data and applies a selection algorithm to determine which photos are most relevant to the user's memories.

[0261] Step 4:

[0262] The server analyzes the content related to the selected photo and uses natural language processing techniques to generate questions about that photo. For example, if it's a travel photo, it might generate a question like, "What is your most memorable experience at this place?"

[0263] Step 5:

[0264] The server sends the generated question and photo to the terminal. This communication is conducted using a secure data transmission protocol.

[0265] Step 6:

[0266] The device configures and displays an interface on the screen to visually present the received question and photo to the user. Here, the photo is displayed prominently, with the question shown below in text format.

[0267] Step 7:

[0268] The user answers the questions displayed on the screen using voice or text. This response is collected using the device's input device.

[0269] Step 8:

[0270] The device sends the response data obtained from the user to the server. In this process, the voice data may be converted to text.

[0271] Step 9:

[0272] The server analyzes the user's responses and generates new questions. This analysis uses an AI model to construct new questions that are relevant to past responses.

[0273] Step 10:

[0274] The server sends a newly generated question to the terminal and resumes the interaction with the user. This process is repeated as long as the user remains interested.

[0275] (Example 1)

[0276] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0277] Currently, many elderly people suffer from cognitive decline, and there is a need for effective means to maintain and improve their cognitive function. Traditional methods often rely on monotonous memory training and limited interaction with others, which is inefficient. Furthermore, there are currently few interactive tools that promote memory activation.

[0278] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0279] In this invention, the server includes means for acquiring image information stored in a data storage device, means for generating question and answer using natural language processing based on the image information, and communication means for securely transferring information. This makes it easier for users to recall memories through images, and enables effective maintenance and improvement of cognitive function through interactive dialogue.

[0280] A "data storage device" is a physical or virtual container for storing and managing information in digital form.

[0281] "Image information" refers to a group of visual data digitized as still images or videos and stored electronically.

[0282] "Natural language processing" is a technical field for a computer system to understand, analyze, and generate human language.

[0283] "Question and answer" is a two-way communication means consisting of questions generated based on specific information and answers to them.

[0284] A "display device" is hardware for visually presenting digital information to a user, including a screen, monitor, etc. <​​​​​​​​​​​​​​​​​​​​The server first accesses the data storage device to retrieve image information saved by the user. This image information is stored in the user's cloud storage service (for example, a common online storage platform). The server also uses an API to access and identify the relevant photos through specific metadata.

[0291] The server applies natural language processing to the acquired image information and uses a generative AI model to generate questions and answers. Specifically, it analyzes the content of the photograph using an image recognition API and creates questions based on the results. This Q&A can be obtained by inputting the generated text as a prompt. For example, a possible prompt sentence is "What memories are associated with this photograph?"

[0292] The generated question-and-answer format is transmitted from the server to the terminal via a secure communication method. The terminal consists of smartphones and tablet devices, and uses a dedicated application to display images and questions to the user through a user interface. This allows the user to interact intuitively.

[0293] Users can respond to displayed questions with voice or text. This allows users to recall and deepen their memories. Voice information is converted into text data using speech recognition software and sent to the server.

[0294] Through this process, the server generates new questions and continues the dialogue with the user. This makes it easier for the user to recall past events, resulting in the maintenance and improvement of their cognitive function.

[0295] The system of this invention aims to provide a novel method for enhancing users' cognitive abilities through interaction, thereby creating a comfortable and effective user experience.

[0296] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0297] Step 1:

[0298] The server retrieves image information from a data storage device. The input to this process is photo data stored in a cloud storage service. The server searches for image information via an API and retrieves a specific photo previously saved by the user. The output is an image file containing the user's metadata.

[0299] Step 2:

[0300] The server applies natural language processing to the acquired image information. The input for this step is the image information acquired in step 1. The server uses an image recognition API to analyze the objects and location information of the photograph. The result of the data processing is metadata that can be used by the generative AI model, and this is used to generate a question-and-answer session.

[0301] Step 3:

[0302] The server generates question-and-answer responses using a generative AI model. The input for this step is the metadata and prompt template generated in step 2. The server inputs the prompt into the generative AI model, and the AI ​​creates customized questions. The output is the customized questions presented to the user.

[0303] Step 4:

[0304] The server sends the generated question and answer to the terminal. The input is the question and image information generated in step 3. The data is transmitted via a secure communication method and received by the terminal. The output is the question and photo displayed on the terminal.

[0305] Step 5:

[0306] The terminal displays images and questions to the user through the user interface. This step receives as input the questions and images that are the output from Step 4. As a specific operation, it displays a photo on the display device and shows an interactive question box, providing a visually easy interface for the user.

[0307] Step 6:

[0308] The user returns a response to the displayed question. The input is the question presented in Step 5. If the user responds verbally, the terminal uses speech recognition technology to convert this into text data. The output is the user's response text that is sent to the server.

[0309] Step 7:

[0310] The server generates additional questions based on the user's response. The input for this step is the text data received from Step 6. The server analyzes the response content and generates the next question by inputting a new prompt sentence into the generation AI model. The output is a new question to facilitate further dialogue.

[0311] (Application Example 1)

[0312] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0313] In a conventional virtual store, there is an issue that proposals for customers are general and lack a personalized experience that reflects individual purchase histories and preferences. There is also a problem that it is difficult for customers to intuitively make product selections that reflect their own preferences within the virtual space. As a result, the situation is such that customer satisfaction cannot be sufficiently improved and the purchasing motivation cannot be effectively stimulated.

[0314] The specific processing by the specific processing unit 290 of the data processing device in Application Example 1 is realized by the following means.

[0315] In this invention, the server includes means for acquiring image information stored on a photographic storage medium, means for generating questions using natural language processing based on the image information, and means for providing a personalized customer service experience based on past transaction history and object information. This enables suggestions tailored to the individual preferences of the customer and effectively stimulates the customer's purchasing intent through a personalized experience.

[0316] A "photo storage medium" is a storage device that digitally holds image information previously taken or saved by a user.

[0317] "Image information" refers to visual data stored on a photographic storage medium, and includes metadata related to specific objects or scenes.

[0318] Natural language processing is a technology that enables computers to understand, interpret, and generate human language, and is used to generate relevant questions from image information.

[0319] A "terminal device" is a device used by users for interaction, and its role is to display questions and suggestions and receive responses.

[0320] A "response" is the verbal or written feedback that a user provides in response to a question, which forms the basis for generating further dialogue and suggestions.

[0321] A "virtual space" is a visual, three-dimensional environment created using digital technology, providing a space for users to interact.

[0322] "Target information" refers to data about products and services related to past transactions and customer interests.

[0323] "Customer service experience" refers to a personalized shopping environment provided through user interaction during the shopping process.

[0324] This invention provides an interactive system for improving cognitive function in the elderly, and specific embodiments thereof are shown below.

[0325] The server accesses the photo storage medium and retrieves image information previously saved by the user in digital format. This image information includes metadata about specific objects and scenes. The server uses natural language processing techniques to generate relevant questions based on the retrieved image information. These questions are customized based on the image content and metadata. For example, a question such as "What are your fondest memories of this place?" might be generated from a photograph of a specific trip.

[0326] The generated questions are sent to a terminal and presented to the user. The terminal is a device that interacts with the user, displaying questions and corresponding images, and receiving responses from the customer verbally or in writing. The user's responses are sent to a server, which analyzes their content and uses it as a basis for generating new questions and suggestions.

[0327] Furthermore, the server can provide a personalized customer service experience based on past transaction history and product information. It presents products to the user in a virtual space, engages in real-time interaction, and makes optimal suggestions to increase the user's interest.

[0328] The system is implemented using Python on the Django framework, and as mentioned earlier, it uses the A-Frame framework for VR space construction and the Google Speech-to-Text API for voice input. This allows users to explore products three-dimensionally in a virtual space and enjoy a personalized purchasing experience based on their interests.

[0329] As a concrete example, a conversation might take place about sneakers the user has purchased in the past. If the user responds to a prompt such as, "Please tell me about your experience when you first wore these sneakers," with "They were very comfortable and I could wear them all day without getting tired," the next question might be, "In what situations do you usually wear them since then?" In this way, the system helps maintain cognitive function by retrieving and activating the user's memories.

[0330] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0331] Step 1:

[0332] The server accesses the photo storage medium and retrieves image information previously saved by the user. This image information is visual data stored in digital format and also includes metadata. The server accesses the database via an API and efficiently reads this data.

[0333] Step 2:

[0334] The server uses natural language processing techniques to generate relevant questions based on the acquired image information. Using the acquired image information and its metadata as input, and employing image recognition algorithms, it identifies specific objects and scenes, forming prompt sentences. For example, it might generate a question in natural language such as, "Please tell me about your fondest memories of this place."

[0335] Step 3:

[0336] The server sends the generated question to the terminal. The terminal displays the presented question and associated image to the user. Here, the user can intuitively view the photo and question on the interface to facilitate recalling past memories. The terminal receives the question and image sent from the server as input and displays them visually to the user.

[0337] Step 4:

[0338] The user responds to the displayed question verbally or in text. The device converts the user's voice response into text data using the Google Speech-to-Text API. This converted text is sent to the server as input for use in the next step.

[0339] Step 5:

[0340] The server analyzes user responses and generates new questions and suggestions. Using natural language processing technology, it analyzes keywords and sentiments within the text data to generate questions that further engage the user's interest in the next step. For example, if a user responds that they were "very comfortable," the server can generate a secondary question such as, "In what situations do you usually wear them now?"

[0341] Step 6:

[0342] The server uses past transaction history and related information to construct optimal recommendations within the virtual space. These recommendations are designed to increase user interest and provide a personalized purchasing experience. The recommendations are sent to the user's device and displayed in real time within the virtual space.

[0343] Step 7:

[0344] The terminal plays the role of constructing a virtual space and provides the user with a three-dimensional environment using the A-Frame framework. This allows the user to intuitively and effectively consider products based on personalized information. It receives suggestions from the server as input and generates a virtual environment that visually represents them.

[0345] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0346] This invention provides a system that allows elderly individuals to engage in dialogue based on past photographs, promoting the maintenance of cognitive function while simultaneously offering interactions that take their emotional state into consideration. Embodiments of this invention are described below from the perspectives of the terminal, server, and user.

[0347] First, the user logs into the photo storage service and grants the system permission to access their photo archive. Once the user begins using the system, the server retrieves image data stored on the user's photo storage media. This image data is organized based on date and tags, and photos that are emotionally meaningful to the user are extracted.

[0348] For the acquired image data, the server uses natural language processing technology to generate questions related to the photograph. At this time, the server is equipped with an emotion engine that analyzes the user's past conversation history and emotional state data to adjust the content and tone of the questions.

[0349] Next, the server sends the generated question and its corresponding photo to the terminal. The terminal then displays these to the user visually and audibly, setting up an intuitive and user-friendly interface to present the question along with the photo.

[0350] The user responds to questions displayed on the device using voice or text. When the device receives the response data, it analyzes the user's voice tone and the emotions contained in the text, and sends this information to the server in real time.

[0351] As soon as the server receives a user's response, it uses an emotion engine to analyze the user's emotional state and adjusts the next question based on this information. For example, it may continue asking in-depth questions on topics the user found particularly interesting, or move to a different topic if the user indicates discomfort.

[0352] This allows users to experience dialogue that takes their current emotional state into consideration and enjoy richer interactions. This invention enhances user psychological satisfaction while simultaneously contributing to the maintenance of healthy cognitive function.

[0353] The following describes the processing flow.

[0354] Step 1:

[0355] The user logs into the photo storage service and grants the system permission to access image data. During this process, the user's account information is authenticated.

[0356] Step 2:

[0357] The server retrieves image data from the user's photo storage. This includes a retrieval process using an API. The server organizes the retrieved images based on date and tag information.

[0358] Step 3:

[0359] The server uses an algorithm to select photos that are emotionally important to the user from the organized image data. This selection is based on the photo's metadata and the user's past preferences.

[0360] Step 4:

[0361] The server generates questions using natural language processing techniques related to the selected photo. Simultaneously, it uses an emotion engine to adjust the tone and content of the questions to match the user's past conversation history and emotional state.

[0362] Step 5:

[0363] The server sends the generated question and associated photo to the terminal. The terminal receives this and prepares to display the photo and question in the user interface.

[0364] Step 6:

[0365] The device displays photos and questions to the user through a visually appealing and intuitive interface. Information is presented in a way that makes it easy for the user to recall past events.

[0366] Step 7:

[0367] The user responds to the questions displayed on the device using voice or text. The user's response is analyzed by an emotion engine, which considers the tone of their voice and the content of their sentences.

[0368] Step 8:

[0369] The device transmits emotional data contained in the user's voice and text responses to the server in real time.

[0370] Step 9:

[0371] The server generates new questions based on the user's responses and sentiment analysis results. In this process, questions that take into account the user's current emotional state are selected.

[0372] Step 10:

[0373] The server sends the newly generated question back to the terminal and then resumes the conversation with the user. This process is repeated as long as the user wishes to continue the conversation.

[0374] (Example 2)

[0375] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0376] In the elderly, there is a need for interactive systems that maintain cognitive function and improve emotional state based on past memories. However, existing technologies struggle to generate questions that fully consider emotional meaning and to adjust dialogue according to the user's real-time emotional state. Therefore, the challenge is to provide an interactive dialogue system that improves user psychological satisfaction and supports healthy cognitive function.

[0377] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0378] In this invention, the server includes means for acquiring data from an image storage device, means for analyzing information related to the image using natural language processing technology and generating questions, and means for analyzing the user's response with an emotion analysis algorithm and adjusting the content of the next dialogue. This makes it possible to generate questions and adjust the dialogue while taking into account the user's emotional state, thereby effectively improving the maintenance of cognitive function and emotional satisfaction in the elderly.

[0379] An "image storage device" is a storage medium for recording and storing digital image data.

[0380] "Data" refers to a series of recorded events and figures that include information about a user.

[0381] "Natural language processing technology" is a technology that enables computers to understand and generate human language.

[0382] "Question generation" is the process of constructing questions based on a specific purpose or context in order to elicit answers.

[0383] An "output device" is a device used to display or transmit digital information in a way that is perceptible to the user.

[0384] "Analysis" is the process of breaking down input information or data and understanding its meaning and relationships.

[0385] An "emotion analysis algorithm" is a computational method used to identify emotional states from text or audio data.

[0386] "Dialogue adjustment" is a process that dynamically changes the content and progress of a dialogue based on the user's responses and status.

[0387] This system is an interactive platform designed to help elderly people maintain cognitive function and improve emotional well-being by allowing them to engage in conversations based on past photographs.

[0388] First, the user logs into the photo storage service and grants the system permission to access their photo archive. This process begins with the use of a user ID and password.

[0389] Once login is complete, the server retrieves the user's image data from the image storage device. This includes photos uploaded from digital cameras and smartphones. Based on this data, the server uses dates and tags to select images that appear to be emotionally important.

[0390] The server uses natural language processing techniques to generate relevant questions for selected images. This involves a generative AI model, and an emotion engine takes into account the user's emotional state and past conversation history. This method provides the user with highly relevant content.

[0391] Next, the server sends the generated question and associated photos to the terminal. The terminal provides the user with an intuitive and easy-to-use interface, visually through the display and audibly through the speaker. The user can then interact with the associated question along with their favorite photos.

[0392] When a user responds to a question, the device captures the response as voice or text and sends it to the server in real time. The device uses a speech recognition system to convert the response to text and also analyzes the tone of voice.

[0393] For example, when a user selects a photo of autumn foliage, the server might generate a question such as, "Who did you go to this place with?". An example of a prompt might be, "Create a question to elicit memories related to the user's photo. The photo shows a landscape with autumn leaves."

[0394] Based on the above, this system aims to analyze the user's emotional state, improve psychological satisfaction through appropriate dialogue, and maintain healthy cognitive function in the elderly.

[0395] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0396] Step 1:

[0397] The user logs in to the photo storage service. In this step, the user enters their user ID and password, and their access rights to the system are verified. After successful login, the system grants the user access to their photo archive. As output, the user's identification information is sent to the server, and the user session begins.

[0398] Step 2:

[0399] The server retrieves user image data from an image storage device. Inputs include user identification information and associated data storage information. Image data is extracted from the database, and emotionally significant images are selected based on date and tags. The output is a list of images meaningful to the user, serving as intermediate data for processing by the server.

[0400] Step 3:

[0401] The server uses selected image data and natural language processing techniques to generate questions. The input consists of selected images and tag information, and contextual analysis is performed by a generative AI model. This analysis outputs questions related to the content of the photos, which are then stored on the server. The data processing used here includes text analysis and question output.

[0402] Step 4:

[0403] The server sends the generated question and its corresponding image to the terminal. The input is the generated question and image data. The output is the visual and auditory content displayed on the user interface. The terminal adjusts the interface to display the question on the screen and to present it audibly using its speaker.

[0404] Step 5:

[0405] The user responds to questions displayed on the device using voice or text. The input is the user's voice or text response, and the output is recognized text data. The device uses speech recognition to convert the response to text and sends it to the server in real time. The response data is used for sentiment analysis in the next step.

[0406] Step 6:

[0407] The server receives the user's response and analyzes it using an emotion analysis algorithm. The input is converted text data, and the analysis identifies the user's emotional state. The output is an emotion statement, which is used as input data when constructing the next dialogue. This allows for the adjustment of dialogue content to reflect the user's emotions appropriately.

[0408] (Application Example 2)

[0409] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0410] There is a need to provide a system that supports the maintenance of cognitive function in the elderly while enabling interaction that leverages past memories. However, conventional systems have struggled to achieve effective interaction that fully reflects the emotions and past memories of individual users. Furthermore, there is a lack of means to enhance affinity for the elderly by personalizing the shopping experience in a virtual environment.

[0411] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0412] In this invention, the server includes means for acquiring image data stored on a photographic storage medium, means for generating questions using natural language processing based on the image data, and means for cross-referencing the acquired image data with product information to identify products associated with the user's past memories. This allows the user to relive past memories while receiving product information linked to those memories. Furthermore, by using this information to provide a friendly purchasing experience through heartwarming dialogue, it is possible to support the maintenance of cognitive function in the elderly and realize interactions that are sensitive to their feelings.

[0413] A "photo storage medium" is a digital storage device that stores a user's past photographic data and enables the retrieval of image data.

[0414] "Image data" refers to digital information of still images stored on a photographic storage medium, and it serves as the basic information for generating questions and dialogues based on its content.

[0415] "Natural language processing" is a technology that enables computers to understand and generate human language, and in this system, it is used for generating questions and analyzing user responses.

[0416] "Means for generating questions" refers to a function that creates questions to ask the user based on image data, providing a foundation for interactions related to the user's past memories.

[0417] A "terminal device" is hardware that a user uses as an interface, and it is a device that presents generated questions and identified product information to the user visually and audibly.

[0418] A "means of obtaining a response" refers to a function that receives a response from the user, either in voice or text, and uses it as the basis for the next interaction.

[0419] "Cross-referencing" is a technique that compares acquired image data and information derived from it with another dataset, such as product information, to find relationships between them.

[0420] "Means of displaying related information" refers to a function that visually or audibly presents product information related to past photographs to the user, thereby linking memories with the current purchasing experience.

[0421] "Means for generating heartwarming dialogue" refers to technologies that create interactions that take into account the user's emotions and promote emotional exchange that goes beyond mere information provision.

[0422] The system of this invention operates via the user's smart device, a cloud-based server, and an internet connection. The user first grants the system access rights to a photo storage service. This allows the server to retrieve image data from the photo storage medium.

[0423] The server uses natural language processing techniques to generate relevant questions based on the acquired image data. In this process, it leverages an emotion engine (e.g., Google Cloud Natural Language API) to analyze the user's past conversation history and emotional state data. The generated questions and related product information are then sent to the user's smart glasses interface.

[0424] The device presents information to the user visually and audibly, and accepts responses in voice or text. These responses are analyzed in real time and sent to a server to assess the user's emotional state. The server then uses this information to tailor the next interaction, generating in-depth discussions about topics or products of particular interest to the user.

[0425] For example, if a product related to a "birthday" photo taken by a user in the past is presented in a virtual store, the user might interact with the product by saying, "This is the cake you photographed with your child in {year}. Do you remember?" This information is provided to the user through an intuitive interface, resulting in a familiar and engaging shopping experience.

[0426] An example of a prompt statement is set as follows:

[0427] "To suggest products related to past memories, use photos and user stories related to 'birthdays' to create heartwarming conversations."

[0428] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0429] Step 1:

[0430] The server retrieves image data stored on the photo storage medium after the user logs into the photo storage service. In this step, all image files associated with the specific user are retrieved from the accessed database, and the metadata contained in each file (e.g., date, tags) is organized. The input is the account information for the photo storage service, and the output is the organized image data set.

[0431] Step 2:

[0432] The server uses the acquired image data to generate questions related to the user's past photos through natural language processing. The input here is the image data acquired in the previous step, and by utilizing the emotion engine, questions are output that take into account the user's past conversation history and emotional state.

[0433] Step 3:

[0434] The server cross-references the generated questions and retrieved image data with the product information database to identify products related to the user's past memories. The input for this step is image data and question information based on dates and tags, and the output is related product information.

[0435] Step 4:

[0436] The terminal visually and audibly presents the user with questions and related product information sent from the server. The input here consists of questions and product information from the server, which are output to the user through an intuitive interface. Specifically, the information is displayed on the smart glasses' screen using video and audio.

[0437] Step 5:

[0438] The user responds to questions displayed on the device using voice or text. The input consists of the presented question and the user's voice / text response, while the output is response data that includes the user's intent.

[0439] Step 6:

[0440] The server analyzes the user's response and uses an emotion engine to evaluate the user's emotional state. The input for this step is the user's response data, and the output is the user's emotional state and tailored question information to generate the next dialogue.

[0441] Step 7:

[0442] The server adjusts the next interaction based on the user's emotional state, generates new questions and relevant product information as needed, and sends them back to the terminal. The input is emotional information and dialogue data as analysis results, and the output is interaction information as needed. This cycle allows the user to experience emotionally sensitive interactions.

[0443] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0444] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0445] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0446] [Third Embodiment]

[0447] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0448] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0449] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0450] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0451] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0453] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0454] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0455] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0456] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0457] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0458] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0459] This invention is a system that provides conversational content using past photographs, with the aim of maintaining and improving the cognitive function of elderly people. The individual components of this system are described in detail below.

[0460] First, the server accesses the photo storage medium and retrieves the stored image data. The photo storage medium is digital storage containing photos that the user has saved in the past, and the server accesses it via an API to identify what kind of memories the user has.

[0461] The server performs natural language processing on the acquired image data to generate questions related to the photos. For example, if the photos are from a specific trip, a question like, "Please tell us about your fondest memories of this place," might be generated. These questions are customized using image metadata and AI-powered image recognition.

[0462] Next, the generated question and its corresponding photo are sent from the server to the terminal. The terminal receives this information and displays the photo and question to the user. The interface is designed to be intuitive and visually easy to understand, allowing the user to recall past experiences.

[0463] The user responds to questions displayed on the device using spoken or written language. This response is then sent to the server via speech recognition and text analysis technologies.

[0464] The server then reviews the user's response and generates further questions. This is to allow the user to delve deeper into their memory and continue the conversation by analyzing their response. For example, if the user talked about travel, the next question might be, "What was the most memorable event from that trip?"

[0465] In this way, by repeating the flow of dialogue through photographs and questions, the user's past memories are recalled and activated. This helps to maintain the user's cognitive function and, consequently, contribute to improving their health. This series of steps constitutes a specific embodiment of the present invention.

[0466] The following describes the processing flow.

[0467] Step 1:

[0468] The user logs into the photo storage service and grants permission to access their photo data. This process begins with secure access to the account through an authentication process.

[0469] Step 2:

[0470] The server receives the user's authentication information and retrieves image data stored on the photo storage medium from that account. Here, the photos are organized based on date and tag information.

[0471] Step 3:

[0472] The server references the metadata of the image data and applies a selection algorithm to determine which photos are most relevant to the user's memories.

[0473] Step 4:

[0474] The server analyzes the content related to the selected photo and uses natural language processing techniques to generate questions about that photo. For example, if it's a travel photo, it might generate a question like, "What is your most memorable experience at this place?"

[0475] Step 5:

[0476] The server sends the generated question and photo to the terminal. This communication is conducted using a secure data transmission protocol.

[0477] Step 6:

[0478] The device configures and displays an interface on the screen to visually present the received question and photo to the user. Here, the photo is displayed prominently, with the question shown below in text format.

[0479] Step 7:

[0480] The user answers the questions displayed on the screen using voice or text. This response is collected using the device's input device.

[0481] Step 8:

[0482] The device sends the response data obtained from the user to the server. In this process, the voice data may be converted to text.

[0483] Step 9:

[0484] The server analyzes the user's responses and generates new questions. This analysis uses an AI model to construct new questions that are relevant to past responses.

[0485] Step 10:

[0486] The server sends a newly generated question to the terminal and resumes the interaction with the user. This process is repeated as long as the user remains interested.

[0487] (Example 1)

[0488] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0489] Currently, many elderly people suffer from cognitive decline, and there is a need for effective means to maintain and improve their cognitive function. Traditional methods often rely on monotonous memory training and limited interaction with others, which is inefficient. Furthermore, there are currently few interactive tools that promote memory activation.

[0490] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0491] In this invention, the server includes means for acquiring image information stored in a data storage device, means for generating question and answer using natural language processing based on the image information, and communication means for securely transferring information. This makes it easier for users to recall memories through images, and enables effective maintenance and improvement of cognitive function through interactive dialogue.

[0492] A "data storage device" is a physical or virtual container for storing and managing information in digital format.

[0493] "Image information" refers to a collection of visual data that has been digitized as still images or videos and stored electronically.

[0494] "Natural language processing" is a field of technology that enables computer systems to understand, analyze, and generate human language.

[0495] "Question and answer" is a two-way communication method consisting of questions generated based on specific information and the answers provided to them.

[0496] A "display device" is hardware used to visually present digital information to a user, and includes screens, monitors, and other similar devices.

[0497] A "response" is an audio or written response given by a user to the information presented.

[0498] "Communication methods" refer to methods and technologies for safely and efficiently sending and receiving data.

[0499] "Memory activation" refers to the cognitive process of recalling past events and re-recognizing information related to them.

[0500] "Dialogue" is the process or means by which two or more people exchange information.

[0501] This invention is a dialogue system aimed at enhancing the user's cognitive functions using digitized past memories. The specific implementation details are shown below.

[0502] The server first accesses the data storage device to retrieve image information saved by the user. This image information is stored in the user's cloud storage service (for example, a common online storage platform). The server also uses an API to access and identify the relevant photos through specific metadata.

[0503] The server applies natural language processing to the acquired image information and uses a generative AI model to generate questions and answers. Specifically, it analyzes the content of the photograph using an image recognition API and creates questions based on the results. This Q&A can be obtained by inputting the generated text as a prompt. For example, a possible prompt sentence is "What memories are associated with this photograph?"

[0504] The generated question-and-answer format is transmitted from the server to the terminal via a secure communication method. The terminal consists of smartphones and tablet devices, and uses a dedicated application to display images and questions to the user through a user interface. This allows the user to interact intuitively.

[0505] Users can respond to displayed questions with voice or text. This allows users to recall and deepen their memories. Voice information is converted into text data using speech recognition software and sent to the server.

[0506] Through this process, the server generates new questions and continues the dialogue with the user. This makes it easier for the user to recall past events, resulting in the maintenance and improvement of their cognitive function.

[0507] The system of this invention aims to provide a novel method for enhancing users' cognitive abilities through interaction, thereby creating a comfortable and effective user experience.

[0508] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0509] Step 1:

[0510] The server retrieves image information from a data storage device. The input to this process is photo data stored in a cloud storage service. The server searches for image information via an API and retrieves a specific photo previously saved by the user. The output is an image file containing the user's metadata.

[0511] Step 2:

[0512] The server applies natural language processing to the acquired image information. The input for this step is the image information acquired in step 1. The server uses an image recognition API to analyze the objects and location information of the photograph. The result of the data processing is metadata that can be used by the generative AI model, and this is used to generate a question-and-answer session.

[0513] Step 3:

[0514] The server generates question-and-answer responses using a generative AI model. The input for this step is the metadata and prompt template generated in step 2. The server inputs the prompt into the generative AI model, and the AI ​​creates customized questions. The output is the customized questions presented to the user.

[0515] Step 4:

[0516] The server sends the generated question and answer to the terminal. The input is the question and image information generated in step 3. The data is transmitted via a secure communication method and received by the terminal. The output is the question and photo displayed on the terminal.

[0517] Step 5:

[0518] The terminal displays images and questions to the user through a user interface. This step receives the questions and images, which are the output from step 4, as input. Specifically, it displays a photograph on the display device and shows an interactive question box. It provides a visually easy-to-understand interface for the user.

[0519] Step 6:

[0520] The user responds to the displayed question. The input is the question presented in step 5. If the user responds verbally, the terminal uses speech recognition technology to convert it into text data. The output is the user's response text sent to the server.

[0521] Step 7:

[0522] The server generates additional questions based on the user's responses. The input for this step is the text data received from step 6. The server analyzes the responses and generates the next question by inputting new prompt sentences into the generating AI model. The output is a new question to facilitate further dialogue.

[0523] (Application Example 1)

[0524] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0525] Traditional virtual stores typically offer suggestions to customers, but lack personalized experiences that reflect individual purchase history and preferences. Furthermore, it's difficult for customers to intuitively select products that reflect their preferences within the virtual space. This results in insufficient improvement in customer satisfaction and stimulation of purchasing intent.

[0526] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0527] In this invention, the server includes means for acquiring image information stored on a photographic storage medium, means for generating questions using natural language processing based on the image information, and means for providing a personalized customer service experience based on past transaction history and object information. This enables suggestions tailored to the individual preferences of the customer and effectively stimulates the customer's purchasing intent through a personalized experience.

[0528] A "photo storage medium" is a storage device that digitally holds image information previously taken or saved by a user.

[0529] "Image information" refers to visual data stored on a photographic storage medium, and includes metadata related to specific objects or scenes.

[0530] Natural language processing is a technology that enables computers to understand, interpret, and generate human language, and is used to generate relevant questions from image information.

[0531] A "terminal device" is a device used by users for interaction, and its role is to display questions and suggestions and receive responses.

[0532] A "response" is the verbal or written feedback that a user provides in response to a question, which forms the basis for generating further dialogue and suggestions.

[0533] A "virtual space" is a visual, three-dimensional environment created using digital technology, providing a space for users to interact.

[0534] "Target information" refers to data about products and services related to past transactions and customer interests.

[0535] "Customer service experience" refers to a personalized shopping environment provided through user interaction during the shopping process.

[0536] This invention provides an interactive system for improving cognitive function in the elderly, and specific embodiments thereof are shown below.

[0537] The server accesses the photo storage medium and retrieves image information previously saved by the user in digital format. This image information includes metadata about specific objects and scenes. The server uses natural language processing techniques to generate relevant questions based on the retrieved image information. These questions are customized based on the image content and metadata. For example, a question such as "What are your fondest memories of this place?" might be generated from a photograph of a specific trip.

[0538] The generated questions are sent to a terminal and presented to the user. The terminal is a device that interacts with the user, displaying questions and corresponding images, and receiving responses from the customer verbally or in writing. The user's responses are sent to a server, which analyzes their content and uses it as a basis for generating new questions and suggestions.

[0539] Furthermore, the server can provide a personalized customer service experience based on past transaction history and product information. It presents products to the user in a virtual space, engages in real-time interaction, and makes optimal suggestions to increase the user's interest.

[0540] The system is implemented using Python on the Django framework, and as mentioned earlier, it uses the A-Frame framework for VR space construction and the Google Speech-to-Text API for voice input. This allows users to explore products three-dimensionally in a virtual space and enjoy a personalized purchasing experience based on their interests.

[0541] As a concrete example, a conversation might take place about sneakers the user has purchased in the past. If the user responds to a prompt such as, "Please tell me about your experience when you first wore these sneakers," with "They were very comfortable and I could wear them all day without getting tired," the next question might be, "In what situations do you usually wear them since then?" In this way, the system helps maintain cognitive function by retrieving and activating the user's memories.

[0542] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0543] Step 1:

[0544] The server accesses the photo storage medium and retrieves image information previously saved by the user. This image information is visual data stored in digital format and also includes metadata. The server accesses the database via an API and efficiently reads this data.

[0545] Step 2:

[0546] The server uses natural language processing techniques to generate relevant questions based on the acquired image information. Using the acquired image information and its metadata as input, and employing image recognition algorithms, it identifies specific objects and scenes, forming prompt sentences. For example, it might generate a question in natural language such as, "Please tell me about your fondest memories of this place."

[0547] Step 3:

[0548] The server sends the generated question to the terminal. The terminal displays the presented question and associated image to the user. Here, the user can intuitively view the photo and question on the interface to facilitate recalling past memories. The terminal receives the question and image sent from the server as input and displays them visually to the user.

[0549] Step 4:

[0550] The user responds to the displayed question verbally or in text. The device converts the user's voice response into text data using the Google Speech-to-Text API. This converted text is sent to the server as input for use in the next step.

[0551] Step 5:

[0552] The server analyzes user responses and generates new questions and suggestions. Using natural language processing technology, it analyzes keywords and sentiments within the text data to generate questions that further engage the user's interest in the next step. For example, if a user responds that they were "very comfortable," the server can generate a secondary question such as, "In what situations do you usually wear them now?"

[0553] Step 6:

[0554] The server uses past transaction history and related information to construct optimal recommendations within the virtual space. These recommendations are designed to increase user interest and provide a personalized purchasing experience. The recommendations are sent to the user's device and displayed in real time within the virtual space.

[0555] Step 7:

[0556] The terminal plays the role of constructing a virtual space and provides the user with a three-dimensional environment using the A-Frame framework. This allows the user to intuitively and effectively consider products based on personalized information. It receives suggestions from the server as input and generates a virtual environment that visually represents them.

[0557] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0558] This invention provides a system that allows elderly individuals to engage in dialogue based on past photographs, promoting the maintenance of cognitive function while simultaneously offering interactions that take their emotional state into consideration. Embodiments of this invention are described below from the perspectives of the terminal, server, and user.

[0559] First, the user logs into the photo storage service and grants the system permission to access their photo archive. Once the user begins using the system, the server retrieves image data stored on the user's photo storage media. This image data is organized based on date and tags, and photos that are emotionally meaningful to the user are extracted.

[0560] For the acquired image data, the server uses natural language processing technology to generate questions related to the photograph. At this time, the server is equipped with an emotion engine that analyzes the user's past conversation history and emotional state data to adjust the content and tone of the questions.

[0561] Next, the server sends the generated question and its corresponding photo to the terminal. The terminal then displays these to the user visually and audibly, setting up an intuitive and user-friendly interface to present the question along with the photo.

[0562] The user responds to questions displayed on the device using voice or text. When the device receives the response data, it analyzes the user's voice tone and the emotions contained in the text, and sends this information to the server in real time.

[0563] As soon as the server receives a user's response, it uses an emotion engine to analyze the user's emotional state and adjusts the next question based on this information. For example, it may continue asking in-depth questions on topics the user found particularly interesting, or move to a different topic if the user indicates discomfort.

[0564] This allows users to experience dialogue that takes their current emotional state into consideration and enjoy richer interactions. This invention enhances user psychological satisfaction while simultaneously contributing to the maintenance of healthy cognitive function.

[0565] The following describes the processing flow.

[0566] Step 1:

[0567] The user logs into the photo storage service and grants the system permission to access image data. During this process, the user's account information is authenticated.

[0568] Step 2:

[0569] The server retrieves image data from the user's photo storage. This includes a retrieval process using an API. The server organizes the retrieved images based on date and tag information.

[0570] Step 3:

[0571] The server uses an algorithm to select photos that are emotionally important to the user from the organized image data. This selection is based on the photo's metadata and the user's past preferences.

[0572] Step 4:

[0573] The server generates questions using natural language processing techniques related to the selected photo. Simultaneously, it uses an emotion engine to adjust the tone and content of the questions to match the user's past conversation history and emotional state.

[0574] Step 5:

[0575] The server sends the generated question and associated photo to the terminal. The terminal receives this and prepares to display the photo and question in the user interface.

[0576] Step 6:

[0577] The device displays photos and questions to the user through a visually appealing and intuitive interface. Information is presented in a way that makes it easy for the user to recall past events.

[0578] Step 7:

[0579] The user responds to the questions displayed on the device using voice or text. The user's response is analyzed by an emotion engine, which considers the tone of their voice and the content of their sentences.

[0580] Step 8:

[0581] The device transmits emotional data contained in the user's voice and text responses to the server in real time.

[0582] Step 9:

[0583] The server generates new questions based on the user's responses and sentiment analysis results. In this process, questions that take into account the user's current emotional state are selected.

[0584] Step 10:

[0585] The server sends the newly generated question back to the terminal and then resumes the conversation with the user. This process is repeated as long as the user wishes to continue the conversation.

[0586] (Example 2)

[0587] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0588] In the elderly, there is a need for interactive systems that maintain cognitive function and improve emotional state based on past memories. However, existing technologies struggle to generate questions that fully consider emotional meaning and to adjust dialogue according to the user's real-time emotional state. Therefore, the challenge is to provide an interactive dialogue system that improves user psychological satisfaction and supports healthy cognitive function.

[0589] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0590] In this invention, the server includes means for acquiring data from an image storage device, means for analyzing information related to the image using natural language processing technology and generating questions, and means for analyzing the user's response with an emotion analysis algorithm and adjusting the content of the next dialogue. This makes it possible to generate questions and adjust the dialogue while taking into account the user's emotional state, thereby effectively improving the maintenance of cognitive function and emotional satisfaction in the elderly.

[0591] An "image storage device" is a storage medium for recording and storing digital image data.

[0592] "Data" refers to a series of recorded events and figures that include information about a user.

[0593] "Natural language processing technology" is a technology that enables computers to understand and generate human language.

[0594] "Question generation" is the process of constructing questions based on a specific purpose or context in order to elicit answers.

[0595] An "output device" is a device used to display or transmit digital information in a way that is perceptible to the user.

[0596] "Analysis" is the process of breaking down input information or data and understanding its meaning and relationships.

[0597] An "emotion analysis algorithm" is a computational method used to identify emotional states from text or audio data.

[0598] "Dialogue adjustment" is a process that dynamically changes the content and progress of a dialogue based on the user's responses and status.

[0599] This system is an interactive platform designed to help elderly people maintain cognitive function and improve emotional well-being by allowing them to engage in conversations based on past photographs.

[0600] First, the user logs into the photo storage service and grants the system permission to access their photo archive. This process begins with the use of a user ID and password.

[0601] Once login is complete, the server retrieves the user's image data from the image storage device. This includes photos uploaded from digital cameras and smartphones. Based on this data, the server uses dates and tags to select images that appear to be emotionally important.

[0602] The server uses natural language processing techniques to generate relevant questions for selected images. This involves a generative AI model, and an emotion engine takes into account the user's emotional state and past conversation history. This method provides the user with highly relevant content.

[0603] Next, the server sends the generated question and associated photos to the terminal. The terminal provides the user with an intuitive and easy-to-use interface, visually through the display and audibly through the speaker. The user can then interact with the associated question along with their favorite photos.

[0604] When a user responds to a question, the device captures the response as voice or text and sends it to the server in real time. The device uses a speech recognition system to convert the response to text and also analyzes the tone of voice.

[0605] For example, when a user selects a photo of autumn foliage, the server might generate a question such as, "Who did you go to this place with?". An example of a prompt might be, "Create a question to elicit memories related to the user's photo. The photo shows a landscape with autumn leaves."

[0606] Based on the above, this system aims to analyze the user's emotional state, improve psychological satisfaction through appropriate dialogue, and maintain healthy cognitive function in the elderly.

[0607] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0608] Step 1:

[0609] The user logs in to the photo storage service. In this step, the user enters their user ID and password, and their access rights to the system are verified. After successful login, the system grants the user access to their photo archive. As output, the user's identification information is sent to the server, and the user session begins.

[0610] Step 2:

[0611] The server retrieves user image data from an image storage device. Inputs include user identification information and associated data storage information. Image data is extracted from the database, and emotionally significant images are selected based on date and tags. The output is a list of images meaningful to the user, serving as intermediate data for processing by the server.

[0612] Step 3:

[0613] The server uses selected image data and natural language processing techniques to generate questions. The input consists of selected images and tag information, and contextual analysis is performed by a generative AI model. This analysis outputs questions related to the content of the photos, which are then stored on the server. The data processing used here includes text analysis and question output.

[0614] Step 4:

[0615] The server sends the generated question and its corresponding image to the terminal. The input is the generated question and image data. The output is the visual and auditory content displayed on the user interface. The terminal adjusts the interface to display the question on the screen and to present it audibly using its speaker.

[0616] Step 5:

[0617] The user responds to questions displayed on the device using voice or text. The input is the user's voice or text response, and the output is recognized text data. The device uses speech recognition to convert the response to text and sends it to the server in real time. The response data is used for sentiment analysis in the next step.

[0618] Step 6:

[0619] The server receives the user's response and analyzes it using an emotion analysis algorithm. The input is converted text data, and the analysis identifies the user's emotional state. The output is an emotion statement, which is used as input data when constructing the next dialogue. This allows for the adjustment of dialogue content to reflect the user's emotions appropriately.

[0620] (Application Example 2)

[0621] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0622] There is a need to provide a system that supports the maintenance of cognitive function in the elderly while enabling interaction that leverages past memories. However, conventional systems have struggled to achieve effective interaction that fully reflects the emotions and past memories of individual users. Furthermore, there is a lack of means to enhance affinity for the elderly by personalizing the shopping experience in a virtual environment.

[0623] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0624] In this invention, the server includes means for acquiring image data stored on a photographic storage medium, means for generating questions using natural language processing based on the image data, and means for cross-referencing the acquired image data with product information to identify products associated with the user's past memories. This allows the user to relive past memories while receiving product information linked to those memories. Furthermore, by using this information to provide a friendly purchasing experience through heartwarming dialogue, it is possible to support the maintenance of cognitive function in the elderly and realize interactions that are sensitive to their feelings.

[0625] A "photo storage medium" is a digital storage device that stores a user's past photographic data and enables the retrieval of image data.

[0626] "Image data" refers to digital information of still images stored on a photographic storage medium, and it serves as the basic information for generating questions and dialogues based on its content.

[0627] "Natural language processing" is a technology that enables computers to understand and generate human language, and in this system, it is used for generating questions and analyzing user responses.

[0628] "Means for generating questions" refers to a function that creates questions to ask the user based on image data, providing a foundation for interactions related to the user's past memories.

[0629] A "terminal device" is hardware that a user uses as an interface, and it is a device that presents generated questions and identified product information to the user visually and audibly.

[0630] A "means of obtaining a response" refers to a function that receives a response from the user, either in voice or text, and uses it as the basis for the next interaction.

[0631] "Cross-referencing" is a technique that compares acquired image data and information derived from it with another dataset, such as product information, to find relationships between them.

[0632] "Means of displaying related information" refers to a function that visually or audibly presents product information related to past photographs to the user, thereby linking memories with the current purchasing experience.

[0633] "Means for generating heartwarming dialogue" refers to technologies that create interactions that take into account the user's emotions and promote emotional exchange that goes beyond mere information provision.

[0634] The system of this invention operates via the user's smart device, a cloud-based server, and an internet connection. The user first grants the system access rights to a photo storage service. This allows the server to retrieve image data from the photo storage medium.

[0635] The server uses natural language processing techniques to generate relevant questions based on the acquired image data. In this process, it leverages an emotion engine (e.g., Google Cloud Natural Language API) to analyze the user's past conversation history and emotional state data. The generated questions and related product information are then sent to the user's smart glasses interface.

[0636] The device presents information to the user visually and audibly, and accepts responses in voice or text. These responses are analyzed in real time and sent to a server to assess the user's emotional state. The server then uses this information to tailor the next interaction, generating in-depth discussions about topics or products of particular interest to the user.

[0637] For example, if a product related to a "birthday" photo taken by a user in the past is presented in a virtual store, the user might interact with the product by saying, "This is the cake you photographed with your child in {year}. Do you remember?" This information is provided to the user through an intuitive interface, resulting in a familiar and engaging shopping experience.

[0638] An example of a prompt statement is set as follows:

[0639] "To suggest products related to past memories, use photos and user stories related to 'birthdays' to create heartwarming conversations."

[0640] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0641] Step 1:

[0642] The server retrieves image data stored on the photo storage medium after the user logs into the photo storage service. In this step, all image files associated with the specific user are retrieved from the accessed database, and the metadata contained in each file (e.g., date, tags) is organized. The input is the account information for the photo storage service, and the output is the organized image data set.

[0643] Step 2:

[0644] The server uses the acquired image data to generate questions related to the user's past photos through natural language processing. The input here is the image data acquired in the previous step, and by utilizing the emotion engine, questions are output that take into account the user's past conversation history and emotional state.

[0645] Step 3:

[0646] The server cross-references the generated questions and retrieved image data with the product information database to identify products related to the user's past memories. The input for this step is image data and question information based on dates and tags, and the output is related product information.

[0647] Step 4:

[0648] The terminal visually and audibly presents the user with questions and related product information sent from the server. The input here consists of questions and product information from the server, which are output to the user through an intuitive interface. Specifically, the information is displayed on the smart glasses' screen using video and audio.

[0649] Step 5:

[0650] The user responds to questions displayed on the device using voice or text. The input consists of the presented question and the user's voice / text response, while the output is response data that includes the user's intent.

[0651] Step 6:

[0652] The server analyzes the user's response and uses an emotion engine to evaluate the user's emotional state. The input for this step is the user's response data, and the output is the user's emotional state and tailored question information to generate the next dialogue.

[0653] Step 7:

[0654] The server adjusts the next interaction based on the user's emotional state, generates new questions and relevant product information as needed, and sends them back to the terminal. The input is emotional information and dialogue data as analysis results, and the output is interaction information as needed. This cycle allows the user to experience emotionally sensitive interactions.

[0655] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0656] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0657] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0658] [Fourth Embodiment]

[0659] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0660] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0661] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0662] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0663] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0664] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0665] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0666] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0667] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0668] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0669] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0670] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0671] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0672] This invention is a system that provides conversational content using past photographs, with the aim of maintaining and improving the cognitive function of elderly people. The individual components of this system are described in detail below.

[0673] First, the server accesses the photo storage medium and retrieves the stored image data. The photo storage medium is digital storage containing photos that the user has saved in the past, and the server accesses it via an API to identify what kind of memories the user has.

[0674] The server performs natural language processing on the acquired image data to generate questions related to the photos. For example, if the photos are from a specific trip, a question like, "Please tell us about your fondest memories of this place," might be generated. These questions are customized using image metadata and AI-powered image recognition.

[0675] Next, the generated question and its corresponding photo are sent from the server to the terminal. The terminal receives this information and displays the photo and question to the user. The interface is designed to be intuitive and visually easy to understand, allowing the user to recall past experiences.

[0676] The user responds to questions displayed on the device using spoken or written language. This response is then sent to the server via speech recognition and text analysis technologies.

[0677] The server then reviews the user's response and generates further questions. This is to allow the user to delve deeper into their memory and continue the conversation by analyzing their response. For example, if the user talked about travel, the next question might be, "What was the most memorable event from that trip?"

[0678] In this way, by repeating the flow of dialogue through photographs and questions, the user's past memories are recalled and activated. This helps to maintain the user's cognitive function and, consequently, contribute to improving their health. This series of steps constitutes a specific embodiment of the present invention.

[0679] The following describes the processing flow.

[0680] Step 1:

[0681] The user logs into the photo storage service and grants permission to access their photo data. This process begins with secure access to the account through an authentication process.

[0682] Step 2:

[0683] The server receives the user's authentication information and retrieves image data stored on the photo storage medium from that account. Here, the photos are organized based on date and tag information.

[0684] Step 3:

[0685] The server references the metadata of the image data and applies a selection algorithm to determine which photos are most relevant to the user's memories.

[0686] Step 4:

[0687] The server analyzes the content related to the selected photo and uses natural language processing techniques to generate questions about that photo. For example, if it's a travel photo, it might generate a question like, "What is your most memorable experience at this place?"

[0688] Step 5:

[0689] The server sends the generated question and photo to the terminal. This communication is conducted using a secure data transmission protocol.

[0690] Step 6:

[0691] The device configures and displays an interface on the screen to visually present the received question and photo to the user. Here, the photo is displayed prominently, with the question shown below in text format.

[0692] Step 7:

[0693] The user answers the questions displayed on the screen using voice or text. This response is collected using the device's input device.

[0694] Step 8:

[0695] The device sends the response data obtained from the user to the server. In this process, the voice data may be converted to text.

[0696] Step 9:

[0697] The server analyzes the user's responses and generates new questions. This analysis uses an AI model to construct new questions that are relevant to past responses.

[0698] Step 10:

[0699] The server sends a newly generated question to the terminal and resumes the interaction with the user. This process is repeated as long as the user remains interested.

[0700] (Example 1)

[0701] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0702] Currently, many elderly people suffer from cognitive decline, and there is a need for effective means to maintain and improve their cognitive function. Traditional methods often rely on monotonous memory training and limited interaction with others, which is inefficient. Furthermore, there are currently few interactive tools that promote memory activation.

[0703] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0704] In this invention, the server includes means for acquiring image information stored in a data storage device, means for generating question and answer using natural language processing based on the image information, and communication means for securely transferring information. This makes it easier for users to recall memories through images, and enables effective maintenance and improvement of cognitive function through interactive dialogue.

[0705] A "data storage device" is a physical or virtual container for storing and managing information in digital format.

[0706] "Image information" refers to a collection of visual data that has been digitized as still images or videos and stored electronically.

[0707] "Natural language processing" is a field of technology that enables computer systems to understand, analyze, and generate human language.

[0708] "Question and answer" is a two-way communication method consisting of questions generated based on specific information and the answers provided to them.

[0709] A "display device" is hardware used to visually present digital information to a user, and includes screens, monitors, and other similar devices.

[0710] A "response" is an audio or written response given by a user to the information presented.

[0711] "Communication methods" refer to methods and technologies for safely and efficiently sending and receiving data.

[0712] "Memory activation" refers to the cognitive process of recalling past events and re-recognizing information related to them.

[0713] "Dialogue" is the process or means by which two or more people exchange information.

[0714] This invention is a dialogue system aimed at enhancing the user's cognitive functions using digitized past memories. The specific implementation details are shown below.

[0715] The server first accesses the data storage device to retrieve image information saved by the user. This image information is stored in the user's cloud storage service (for example, a common online storage platform). The server also uses an API to access and identify the relevant photos through specific metadata.

[0716] The server applies natural language processing to the acquired image information and uses a generative AI model to generate questions and answers. Specifically, it analyzes the content of the photograph using an image recognition API and creates questions based on the results. This Q&A can be obtained by inputting the generated text as a prompt. For example, a possible prompt sentence is "What memories are associated with this photograph?"

[0717] The generated question-and-answer format is transmitted from the server to the terminal via a secure communication method. The terminal consists of smartphones and tablet devices, and uses a dedicated application to display images and questions to the user through a user interface. This allows the user to interact intuitively.

[0718] Users can respond to displayed questions with voice or text. This allows users to recall and deepen their memories. Voice information is converted into text data using speech recognition software and sent to the server.

[0719] Through this process, the server generates new questions and continues the dialogue with the user. This makes it easier for the user to recall past events, resulting in the maintenance and improvement of their cognitive function.

[0720] The system of this invention aims to provide a novel method for enhancing users' cognitive abilities through interaction, thereby creating a comfortable and effective user experience.

[0721] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0722] Step 1:

[0723] The server retrieves image information from a data storage device. The input to this process is photo data stored in a cloud storage service. The server searches for image information via an API and retrieves a specific photo previously saved by the user. The output is an image file containing the user's metadata.

[0724] Step 2:

[0725] The server applies natural language processing to the acquired image information. The input for this step is the image information acquired in step 1. The server uses an image recognition API to analyze the objects and location information of the photograph. The result of the data processing is metadata that can be used by the generative AI model, and this is used to generate a question-and-answer session.

[0726] Step 3:

[0727] The server generates question-and-answer responses using a generative AI model. The input for this step is the metadata and prompt template generated in step 2. The server inputs the prompt into the generative AI model, and the AI ​​creates customized questions. The output is the customized questions presented to the user.

[0728] Step 4:

[0729] The server sends the generated question and answer to the terminal. The input is the question and image information generated in step 3. The data is transmitted via a secure communication method and received by the terminal. The output is the question and photo displayed on the terminal.

[0730] Step 5:

[0731] The terminal displays images and questions to the user through a user interface. This step receives the questions and images, which are the output from step 4, as input. Specifically, it displays a photograph on the display device and shows an interactive question box. It provides a visually easy-to-understand interface for the user.

[0732] Step 6:

[0733] The user responds to the displayed question. The input is the question presented in step 5. If the user responds verbally, the terminal uses speech recognition technology to convert it into text data. The output is the user's response text sent to the server.

[0734] Step 7:

[0735] The server generates additional questions based on the user's responses. The input for this step is the text data received from step 6. The server analyzes the responses and generates the next question by inputting new prompt sentences into the generating AI model. The output is a new question to facilitate further dialogue.

[0736] (Application Example 1)

[0737] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0738] Traditional virtual stores typically offer suggestions to customers, but lack personalized experiences that reflect individual purchase history and preferences. Furthermore, it's difficult for customers to intuitively select products that reflect their preferences within the virtual space. This results in insufficient improvement in customer satisfaction and stimulation of purchasing intent.

[0739] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0740] In this invention, the server includes means for acquiring image information stored on a photographic storage medium, means for generating questions using natural language processing based on the image information, and means for providing a personalized customer service experience based on past transaction history and object information. This enables suggestions tailored to the individual preferences of the customer and effectively stimulates the customer's purchasing intent through a personalized experience.

[0741] A "photo storage medium" is a storage device that digitally holds image information previously taken or saved by a user.

[0742] "Image information" refers to visual data stored on a photographic storage medium, and includes metadata related to specific objects or scenes.

[0743] Natural language processing is a technology that enables computers to understand, interpret, and generate human language, and is used to generate relevant questions from image information.

[0744] A "terminal device" is a device used by users for interaction, and its role is to display questions and suggestions and receive responses.

[0745] A "response" is the verbal or written feedback that a user provides in response to a question, which forms the basis for generating further dialogue and suggestions.

[0746] A "virtual space" is a visual, three-dimensional environment created using digital technology, providing a space for users to interact.

[0747] "Target information" refers to data about products and services related to past transactions and customer interests.

[0748] "Customer service experience" refers to a personalized shopping environment provided through user interaction during the shopping process.

[0749] This invention provides an interactive system for improving cognitive function in the elderly, and specific embodiments thereof are shown below.

[0750] The server accesses the photo storage medium and retrieves image information previously saved by the user in digital format. This image information includes metadata about specific objects and scenes. The server uses natural language processing techniques to generate relevant questions based on the retrieved image information. These questions are customized based on the image content and metadata. For example, a question such as "What are your fondest memories of this place?" might be generated from a photograph of a specific trip.

[0751] The generated questions are sent to a terminal and presented to the user. The terminal is a device that interacts with the user, displaying questions and corresponding images, and receiving responses from the customer verbally or in writing. The user's responses are sent to a server, which analyzes their content and uses it as a basis for generating new questions and suggestions.

[0752] Furthermore, the server can provide a personalized customer service experience based on past transaction history and product information. It presents products to the user in a virtual space, engages in real-time interaction, and makes optimal suggestions to increase the user's interest.

[0753] The system is implemented using Python on the Django framework, and as mentioned earlier, it uses the A-Frame framework for VR space construction and the Google Speech-to-Text API for voice input. This allows users to explore products three-dimensionally in a virtual space and enjoy a personalized purchasing experience based on their interests.

[0754] As a concrete example, a conversation might take place about sneakers the user has purchased in the past. If the user responds to a prompt such as, "Please tell me about your experience when you first wore these sneakers," with "They were very comfortable and I could wear them all day without getting tired," the next question might be, "In what situations do you usually wear them since then?" In this way, the system helps maintain cognitive function by retrieving and activating the user's memories.

[0755] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0756] Step 1:

[0757] The server accesses the photo storage medium and retrieves image information previously saved by the user. This image information is visual data stored in digital format and also includes metadata. The server accesses the database via an API and efficiently reads this data.

[0758] Step 2:

[0759] The server uses natural language processing techniques to generate relevant questions based on the acquired image information. Using the acquired image information and its metadata as input, and employing image recognition algorithms, it identifies specific objects and scenes, forming prompt sentences. For example, it might generate a question in natural language such as, "Please tell me about your fondest memories of this place."

[0760] Step 3:

[0761] The server sends the generated question to the terminal. The terminal displays the presented question and associated image to the user. Here, the user can intuitively view the photo and question on the interface to facilitate recalling past memories. The terminal receives the question and image sent from the server as input and displays them visually to the user.

[0762] Step 4:

[0763] The user responds to the displayed question verbally or in text. The device converts the user's voice response into text data using the Google Speech-to-Text API. This converted text is sent to the server as input for use in the next step.

[0764] Step 5:

[0765] The server analyzes user responses and generates new questions and suggestions. Using natural language processing technology, it analyzes keywords and sentiments within the text data to generate questions that further engage the user's interest in the next step. For example, if a user responds that they were "very comfortable," the server can generate a secondary question such as, "In what situations do you usually wear them now?"

[0766] Step 6:

[0767] The server uses past transaction history and related information to construct optimal recommendations within the virtual space. These recommendations are designed to increase user interest and provide a personalized purchasing experience. The recommendations are sent to the user's device and displayed in real time within the virtual space.

[0768] Step 7:

[0769] The terminal plays the role of constructing a virtual space and provides the user with a three-dimensional environment using the A-Frame framework. This allows the user to intuitively and effectively consider products based on personalized information. It receives suggestions from the server as input and generates a virtual environment that visually represents them.

[0770] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0771] This invention provides a system that allows elderly individuals to engage in dialogue based on past photographs, promoting the maintenance of cognitive function while simultaneously offering interactions that take their emotional state into consideration. Embodiments of this invention are described below from the perspectives of the terminal, server, and user.

[0772] First, the user logs into the photo storage service and grants the system permission to access their photo archive. Once the user begins using the system, the server retrieves image data stored on the user's photo storage media. This image data is organized based on date and tags, and photos that are emotionally meaningful to the user are extracted.

[0773] For the acquired image data, the server uses natural language processing technology to generate questions related to the photograph. At this time, the server is equipped with an emotion engine that analyzes the user's past conversation history and emotional state data to adjust the content and tone of the questions.

[0774] Next, the server sends the generated question and its corresponding photo to the terminal. The terminal then displays these to the user visually and audibly, setting up an intuitive and user-friendly interface to present the question along with the photo.

[0775] The user responds to questions displayed on the device using voice or text. When the device receives the response data, it analyzes the user's voice tone and the emotions contained in the text, and sends this information to the server in real time.

[0776] As soon as the server receives a user's response, it uses an emotion engine to analyze the user's emotional state and adjusts the next question based on this information. For example, it may continue asking in-depth questions on topics the user found particularly interesting, or move to a different topic if the user indicates discomfort.

[0777] This allows users to experience dialogue that takes their current emotional state into consideration and enjoy richer interactions. This invention enhances user psychological satisfaction while simultaneously contributing to the maintenance of healthy cognitive function.

[0778] The following describes the processing flow.

[0779] Step 1:

[0780] The user logs into the photo storage service and grants the system permission to access image data. During this process, the user's account information is authenticated.

[0781] Step 2:

[0782] The server retrieves image data from the user's photo storage. This includes a retrieval process using an API. The server organizes the retrieved images based on date and tag information.

[0783] Step 3:

[0784] The server uses an algorithm to select photos that are emotionally important to the user from the organized image data. This selection is based on the photo's metadata and the user's past preferences.

[0785] Step 4:

[0786] The server generates questions using natural language processing techniques related to the selected photo. Simultaneously, it uses an emotion engine to adjust the tone and content of the questions to match the user's past conversation history and emotional state.

[0787] Step 5:

[0788] The server sends the generated question and associated photo to the terminal. The terminal receives this and prepares to display the photo and question in the user interface.

[0789] Step 6:

[0790] The device displays photos and questions to the user through a visually appealing and intuitive interface. Information is presented in a way that makes it easy for the user to recall past events.

[0791] Step 7:

[0792] The user responds to the questions displayed on the device using voice or text. The user's response is analyzed by an emotion engine, which considers the tone of their voice and the content of their sentences.

[0793] Step 8:

[0794] The device transmits emotional data contained in the user's voice and text responses to the server in real time.

[0795] Step 9:

[0796] The server generates new questions based on the user's responses and sentiment analysis results. In this process, questions that take into account the user's current emotional state are selected.

[0797] Step 10:

[0798] The server sends the newly generated question back to the terminal and then resumes the conversation with the user. This process is repeated as long as the user wishes to continue the conversation.

[0799] (Example 2)

[0800] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0801] In the elderly, there is a need for interactive systems that maintain cognitive function and improve emotional state based on past memories. However, existing technologies struggle to generate questions that fully consider emotional meaning and to adjust dialogue according to the user's real-time emotional state. Therefore, the challenge is to provide an interactive dialogue system that improves user psychological satisfaction and supports healthy cognitive function.

[0802] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0803] In this invention, the server includes means for acquiring data from an image storage device, means for analyzing information related to the image using natural language processing technology and generating questions, and means for analyzing the user's response with an emotion analysis algorithm and adjusting the content of the next dialogue. This makes it possible to generate questions and adjust the dialogue while taking into account the user's emotional state, thereby effectively improving the maintenance of cognitive function and emotional satisfaction in the elderly.

[0804] An "image storage device" is a storage medium for recording and storing digital image data.

[0805] "Data" refers to a series of recorded events and figures that include information about a user.

[0806] "Natural language processing technology" is a technology that enables computers to understand and generate human language.

[0807] "Question generation" is the process of constructing questions based on a specific purpose or context in order to elicit answers.

[0808] An "output device" is a device used to display or transmit digital information in a way that is perceptible to the user.

[0809] "Analysis" is the process of breaking down input information or data and understanding its meaning and relationships.

[0810] An "emotion analysis algorithm" is a computational method used to identify emotional states from text or audio data.

[0811] "Dialogue adjustment" is a process that dynamically changes the content and progress of a dialogue based on the user's responses and status.

[0812] This system is an interactive platform designed to help elderly people maintain cognitive function and improve emotional well-being by allowing them to engage in conversations based on past photographs.

[0813] First, the user logs into the photo storage service and grants the system permission to access their photo archive. This process begins with the use of a user ID and password.

[0814] Once login is complete, the server retrieves the user's image data from the image storage device. This includes photos uploaded from digital cameras and smartphones. Based on this data, the server uses dates and tags to select images that appear to be emotionally important.

[0815] The server uses natural language processing techniques to generate relevant questions for selected images. This involves a generative AI model, and an emotion engine takes into account the user's emotional state and past conversation history. This method provides the user with highly relevant content.

[0816] Next, the server sends the generated question and associated photos to the terminal. The terminal provides the user with an intuitive and easy-to-use interface, visually through the display and audibly through the speaker. The user can then interact with the associated question along with their favorite photos.

[0817] When a user responds to a question, the device captures the response as voice or text and sends it to the server in real time. The device uses a speech recognition system to convert the response to text and also analyzes the tone of voice.

[0818] For example, when a user selects a photo of autumn foliage, the server might generate a question such as, "Who did you go to this place with?". An example of a prompt might be, "Create a question to elicit memories related to the user's photo. The photo shows a landscape with autumn leaves."

[0819] Based on the above, this system aims to analyze the user's emotional state, improve psychological satisfaction through appropriate dialogue, and maintain healthy cognitive function in the elderly.

[0820] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0821] Step 1:

[0822] The user logs in to the photo storage service. In this step, the user enters their user ID and password, and their access rights to the system are verified. After successful login, the system grants the user access to their photo archive. As output, the user's identification information is sent to the server, and the user session begins.

[0823] Step 2:

[0824] The server retrieves user image data from an image storage device. Inputs include user identification information and associated data storage information. Image data is extracted from the database, and emotionally significant images are selected based on date and tags. The output is a list of images meaningful to the user, serving as intermediate data for processing by the server.

[0825] Step 3:

[0826] The server uses selected image data and natural language processing techniques to generate questions. The input consists of selected images and tag information, and contextual analysis is performed by a generative AI model. This analysis outputs questions related to the content of the photos, which are then stored on the server. The data processing used here includes text analysis and question output.

[0827] Step 4:

[0828] The server sends the generated question and its corresponding image to the terminal. The input is the generated question and image data. The output is the visual and auditory content displayed on the user interface. The terminal adjusts the interface to display the question on the screen and to present it audibly using its speaker.

[0829] Step 5:

[0830] The user responds to questions displayed on the device using voice or text. The input is the user's voice or text response, and the output is recognized text data. The device uses speech recognition to convert the response to text and sends it to the server in real time. The response data is used for sentiment analysis in the next step.

[0831] Step 6:

[0832] The server receives the user's response and analyzes it using an emotion analysis algorithm. The input is converted text data, and the analysis identifies the user's emotional state. The output is an emotion statement, which is used as input data when constructing the next dialogue. This allows for the adjustment of dialogue content to reflect the user's emotions appropriately.

[0833] (Application Example 2)

[0834] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0835] There is a need to provide a system that supports the maintenance of cognitive function in the elderly while enabling interaction that leverages past memories. However, conventional systems have struggled to achieve effective interaction that fully reflects the emotions and past memories of individual users. Furthermore, there is a lack of means to enhance affinity for the elderly by personalizing the shopping experience in a virtual environment.

[0836] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0837] In this invention, the server includes means for acquiring image data stored on a photographic storage medium, means for generating questions using natural language processing based on the image data, and means for cross-referencing the acquired image data with product information to identify products associated with the user's past memories. This allows the user to relive past memories while receiving product information linked to those memories. Furthermore, by using this information to provide a friendly purchasing experience through heartwarming dialogue, it is possible to support the maintenance of cognitive function in the elderly and realize interactions that are sensitive to their feelings.

[0838] A "photo storage medium" is a digital storage device that stores a user's past photographic data and enables the retrieval of image data.

[0839] "Image data" refers to digital information of still images stored on a photographic storage medium, and it serves as the basic information for generating questions and dialogues based on its content.

[0840] "Natural language processing" is a technology that enables computers to understand and generate human language, and in this system, it is used for generating questions and analyzing user responses.

[0841] "Means for generating questions" refers to a function that creates questions to ask the user based on image data, providing a foundation for interactions related to the user's past memories.

[0842] A "terminal device" is hardware that a user uses as an interface, and it is a device that presents generated questions and identified product information to the user visually and audibly.

[0843] A "means of obtaining a response" refers to a function that receives a response from the user, either in voice or text, and uses it as the basis for the next interaction.

[0844] "Cross-referencing" is a technique that compares acquired image data and information derived from it with another dataset, such as product information, to find relationships between them.

[0845] "Means of displaying related information" refers to a function that visually or audibly presents product information related to past photographs to the user, thereby linking memories with the current purchasing experience.

[0846] "Means for generating heartwarming dialogue" refers to technologies that create interactions that take into account the user's emotions and promote emotional exchange that goes beyond mere information provision.

[0847] The system of this invention operates via the user's smart device, a cloud-based server, and an internet connection. The user first grants the system access rights to a photo storage service. This allows the server to retrieve image data from the photo storage medium.

[0848] The server uses natural language processing techniques to generate relevant questions based on the acquired image data. In this process, it leverages an emotion engine (e.g., Google Cloud Natural Language API) to analyze the user's past conversation history and emotional state data. The generated questions and related product information are then sent to the user's smart glasses interface.

[0849] The device presents information to the user visually and audibly, and accepts responses in voice or text. These responses are analyzed in real time and sent to a server to assess the user's emotional state. The server then uses this information to tailor the next interaction, generating in-depth discussions about topics or products of particular interest to the user.

[0850] For example, if a product related to a "birthday" photo taken by a user in the past is presented in a virtual store, the user might interact with the product by saying, "This is the cake you photographed with your child in {year}. Do you remember?" This information is provided to the user through an intuitive interface, resulting in a familiar and engaging shopping experience.

[0851] An example of a prompt statement is set as follows:

[0852] "To suggest products related to past memories, use photos and user stories related to 'birthdays' to create heartwarming conversations."

[0853] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0854] Step 1:

[0855] The server retrieves image data stored on the photo storage medium after the user logs into the photo storage service. In this step, all image files associated with the specific user are retrieved from the accessed database, and the metadata contained in each file (e.g., date, tags) is organized. The input is the account information for the photo storage service, and the output is the organized image data set.

[0856] Step 2:

[0857] The server uses the acquired image data to generate questions related to the user's past photos through natural language processing. The input here is the image data acquired in the previous step, and by utilizing the emotion engine, questions are output that take into account the user's past conversation history and emotional state.

[0858] Step 3:

[0859] The server cross-references the generated questions and retrieved image data with the product information database to identify products related to the user's past memories. The input for this step is image data and question information based on dates and tags, and the output is related product information.

[0860] Step 4:

[0861] The terminal visually and audibly presents the user with questions and related product information sent from the server. The input here consists of questions and product information from the server, which are output to the user through an intuitive interface. Specifically, the information is displayed on the smart glasses' screen using video and audio.

[0862] Step 5:

[0863] The user responds to questions displayed on the device using voice or text. The input consists of the presented question and the user's voice / text response, while the output is response data that includes the user's intent.

[0864] Step 6:

[0865] The server analyzes the user's response and uses an emotion engine to evaluate the user's emotional state. The input for this step is the user's response data, and the output is the user's emotional state and tailored question information to generate the next dialogue.

[0866] Step 7:

[0867] The server adjusts the next interaction based on the user's emotional state, generates new questions and relevant product information as needed, and sends them back to the terminal. The input is emotional information and dialogue data as analysis results, and the output is interaction information as needed. This cycle allows the user to experience emotionally sensitive interactions.

[0868] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0869] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0870] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0871] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0872] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0873] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0874] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0875] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0876] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0877] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0878] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0879] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0880] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0881] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0882] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0883] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0884] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0885] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0886] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0887] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0888] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0889] The following is further disclosed regarding the embodiments described above.

[0890] (Claim 1)

[0891] A means for acquiring image data stored on a photographic storage medium,

[0892] A means for generating a question using natural language processing based on the aforementioned image data,

[0893] The means of transmitting the generated question to a terminal device and displaying it to the user,

[0894] A means for obtaining a response from a user and generating a new question based on the response,

[0895] A system that includes this.

[0896] (Claim 2)

[0897] The system according to claim 1, wherein the terminal device is configured to receive user responses in both voice and text.

[0898] (Claim 3)

[0899] The system according to claim 1, further comprising means for analyzing the user's response and constructing the following dialogue for the purpose of promoting the user's recollection.

[0900] "Example 1"

[0901] (Claim 1)

[0902] A means for acquiring image information stored in a data storage device,

[0903] A means for generating a question-and-answer response using natural language processing based on the aforementioned image information,

[0904] A means for transmitting the generated question and answer to a display device and presenting it to the user,

[0905] A means for obtaining responses from users and generating new questions and answers based on those responses,

[0906] A means for analyzing the user's response and constructing additional dialogue to stimulate the user's memory,

[0907] A system that includes means of communication for securely transferring information.

[0908] (Claim 2)

[0909] The system according to claim 1, wherein the display device is configured to receive user responses in both audio and text.

[0910] (Claim 3)

[0911] The system according to claim 1, wherein the generated question and answer is customized based on the results of the analysis of image information and is designed to make it easier for the user to recall memories.

[0912] "Application Example 1"

[0913] (Claim 1)

[0914] A means for acquiring image information stored on a photographic storage medium,

[0915] A means for generating a question using natural language processing based on the aforementioned image information,

[0916] The means of transmitting the generated question to a terminal device and displaying it to the user,

[0917] A means for obtaining a response from a user and generating a new question based on the response,

[0918] A means of providing personalized customer service experiences based on past transaction history and product information,

[0919] A means of presenting products to customers in a virtual space and generating subsequent suggestions based on their responses,

[0920] A system that includes this.

[0921] (Claim 2)

[0922] The system according to claim 1, wherein the terminal device is configured to receive user responses in both voice and text.

[0923] (Claim 3)

[0924] The system according to claim 1, further comprising means for analyzing the user's response, estimating the user's interests, and constructing the next dialogue.

[0925] "Example 2 of combining an emotion engine"

[0926] (Claim 1)

[0927] A means of acquiring data from an image storage device,

[0928] A means for determining whether the aforementioned data is emotionally important to the user,

[0929] A means for analyzing information related to an image using natural language processing technology and generating questions,

[0930] A means for transmitting the generated question to an output device and presenting it visually or audibly,

[0931] A means of analyzing the tone of the user's voice and the content of their responses, and transmitting them to a data processing device in real time,

[0932] A means of analyzing user responses using an emotion analysis algorithm and adjusting the content of the next dialogue,

[0933] A system that includes this.

[0934] (Claim 2)

[0935] The system according to claim 1, wherein the output device is configured to receive and analyze the user's response in both voice and text.

[0936] (Claim 3)

[0937] The system according to claim 1, further comprising means for adaptively modifying the next dialogue based on the user's psychological state and constructing the dialogue for the purpose of enhancing the user's interest and emotional comfort.

[0938] "Application example 2 when combining with an emotional engine"

[0939] (Claim 1)

[0940] A means for acquiring image data stored on a photographic storage medium,

[0941] A means for generating a question using natural language processing based on the aforementioned image data,

[0942] The means of transmitting the generated question to a terminal device and displaying it to the user,

[0943] A means for obtaining a response from a user and generating a new question based on the response,

[0944] A means of cross-referencing acquired image data with product information to identify products associated with the user's past memories,

[0945] A means for displaying the identified product information in association with the user's past photos and generating a heartwarming conversation,

[0946] A system that includes this.

[0947] (Claim 2)

[0948] The system according to claim 1, wherein the terminal device is configured to receive user responses in both voice and text, and further provides an intuitive interface via a smart device.

[0949] (Claim 3)

[0950] The system according to claim 1, further comprising means for analyzing the user's response, constructing the next dialogue for the purpose of promoting the user's recollection, and providing further interaction with products of interest. [Explanation of symbols]

[0951] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring image data stored on a photographic storage medium, A means for generating a question using natural language processing based on the aforementioned image data, The means of transmitting the generated question to a terminal device and displaying it to the user, A means for obtaining a response from a user and generating a new question based on the response, A system that includes this.

2. The system according to claim 1, wherein the terminal device is configured to receive user responses in both voice and text.

3. The system according to claim 1, further comprising means for analyzing the user's response and constructing the following dialogue for the purpose of promoting the user's recollection.