System
A system that uses AI to generate personalized videos and dialogues based on past photos and memories addresses the inefficiency of current rehabilitation methods, effectively reducing caregiver burden.
Patent Information
- Application Number
- JP2024116468
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Current rehabilitation methods for dementia patients are inefficient and labor-intensive, making it difficult to provide personalized care, which increases the burden on caregivers and medical professionals.
A system that collects past photos and memories, generates personalized videos using AI, and facilitates interactive dialogues to stimulate individual memories and provide tailored rehabilitation plans.
Efficiently provides personalized rehabilitation plans for dementia patients, reducing the burden on caregivers and medical professionals by using AI to generate videos and responses.
Smart Images

Figure 2026014994000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Currently, there are a wide variety of rehabilitation methods for dementia patients, but providing the optimal rehabilitation plan for each individual patient remains a challenge. Furthermore, as the aging society progresses, the number of dementia patients is also increasing, placing a greater burden on caregivers and medical professionals. Conventional methods rely heavily on manual labor, making it difficult to provide efficient rehabilitation. Therefore, there is a need for a system that provides rehabilitation tailored to each individual dementia patient and reduces the burden on caregivers and medical professionals. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for collecting past photos and related memories, a means for generating videos based on the collected data, a means for providing the generated videos to a dementia patient, a means for analyzing a user's voice input, a means for generating responses based on the analyzed voice data, and a means for providing the generated responses to the user. This system can stimulate the individual memories of a dementia patient and activate their brain. Furthermore, by using artificial intelligence technology to generate videos and responses, it is possible to efficiently provide an individual rehabilitation plan, reducing the burden on caregivers and medical professionals.
[0006] "Photos from the past" are photographs that record the past lives and experiences of dementia patients.
[0007] "Memories" refer to stories about past events and experiences told by people with dementia and their families.
[0008] "Means of collection" refers to hardware and software mechanisms for obtaining past photos and stories of memories from users.
[0009] "Means for generating video" refers to technology for creating visual and audio content based on collected past photographs and reminiscences.
[0010] "Means for providing" refers to a method or device for showing or letting a dementia patient listen to the generated video or response.
[0011] "Means for analyzing user voice input" refers to technology that receives speech from a dementia patient, converts it into digital data, and understands its content.
[0012] The "means for generating a response" is a technology that automatically creates an appropriate reply based on the content of the voice input.
[0013] "Artificial intelligence technology" refers to technology that uses advanced algorithms, including machine learning and natural language processing, to imitate human knowledge and abilities. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and reminiscences and provides them to patients. The system uses artificial intelligence technology to create individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[0036] The program for this system mainly consists of the following phases:
[0037] Data Collection Phase
[0038] Users (dementia patients and their families) input past photos and episodes through a dedicated device or application. Photos are collected by scanning or uploading, and episodes are input as text or voice data. The device temporarily stores the collected data and sends it to a server.
[0039] Data Processing Phase
[0040] The server receives the photos and text data sent from the device. The received data is analyzed by an artificial intelligence model to extract basic information about the episode. Based on this information, an AI video generation module operates and automatically generates videos that include emotional elements and specific locations. The generated video data is stored in cloud storage.
[0041] Content provision phase
[0042] The user requests playback of the generated video from a dedicated device or application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[0043] Dialogue Phase
[0044] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[0045] Specific examples
[0046] For example, if a family member of a dementia patient inputs a photo of the family cherry blossom viewing and an episode from that time into the system, the process would be as follows:
[0047] 1. The user (family member) logs in to the application, scans a photo, and enters the text "This photo was taken when my family went cherry blossom viewing."
[0048] 2. The device sends the photo and text data to the server.
[0049] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[0050] 4. The user (patient) requests video playback and watches the video on their device.
[0051] 5. The user (patient) says, "The cherry blossoms were beautiful at this time." The device analyzes the voice and sends it to the server.
[0052] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[0053] In this way, the system can stimulate past memories and promote brain activity in dementia patients. Furthermore, by utilizing artificial intelligence, it can efficiently provide individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[0054] The processing flow will be explained below.
[0055] Step 1:
[0056] Users (dementia patients and their families) log in to their accounts through a dedicated device or application.
[0057] Step 2:
[0058] Users scan or upload old photos and enter related memories as text or audio data.
[0059] Step 3:
[0060] The device temporarily stores the entered photo and text data in local storage.
[0061] Step 4:
[0062] The device transmits the photos and text data stored in the local storage to the server.
[0063] Step 5:
[0064] The server receives the photo and text data sent from the terminal.
[0065] Step 6:
[0066] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode from the text data.
[0067] Step 7:
[0068] The server uses an AI video generation module to generate videos based on the extracted information, including emotional elements and specific locations.
[0069] Step 8:
[0070] The server stores the generated video data in cloud storage.
[0071] Step 9:
[0072] The user requests playback of the generated video from a dedicated terminal or application.
[0073] Step 10:
[0074] The server receives the user's request and delivers the video data to the terminal in streaming or download format.
[0075] Step 11:
[0076] The terminal plays the distributed video and allows the user to view it.
[0077] Step 12:
[0078] The user (a dementia patient) can start a conversation based on the content of the video, for example, saying, "This day was really fun."
[0079] Step 13:
[0080] The device analyzes the user's voice input and converts it into text data using voice recognition technology.
[0081] Step 14:
[0082] The terminal transmits the analyzed voice data (text) to the server.
[0083] Step 15:
[0084] The server generates a response using a natural language generation (NLG) module based on the received voice data.
[0085] Step 16:
[0086] The server transmits the generated response data to the terminal.
[0087] Step 17:
[0088] The terminal uses a voice synthesis function to reproduce the received response data and convey it to the user.
[0089] Example 1
[0090] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0091] In the rehabilitation of dementia patients, it is important to stimulate individual memories and elicit emotional responses. However, rehabilitation using past photographs and anecdotes can be laborious for some caregivers and medical professionals. Furthermore, these tasks tend to rely on subjective judgment and are difficult to perform consistently. This can result in the insufficient effectiveness of dementia rehabilitation. Therefore, there is a need for an efficient and individually tailored rehabilitation system.
[0092] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0093] In this invention, the server includes: a means for a user to input past photos and episodes; a means for transmitting the input photos and episodes to the server; a means for the server to analyze the received data using an artificial intelligence model; a means for automatically generating videos based on the analysis results; a means for storing the generated video data in cloud storage; a means for a user to request video playback; a means for the server to receive the request and deliver the video data to a terminal; a means for the terminal to play the video; a means for analyzing the user's voice input and converting it into text data; a means for the server to generate a response based on the voice input; and a means for transmitting the generated response to the terminal and playing it as audio. This allows a user to input past photos and episodes and efficiently stimulate the memory of a dementia patient using the generated video. Furthermore, the use of artificial intelligence can provide a consistent rehabilitation plan, reducing the burden on caregivers and medical professionals.
[0094] "User" refers to the person who operates the system and inputs data, such as a dementia patient, their family member, or caregiver.
[0095] A "terminal" is an electronic device that allows a user to input data, receive data from a server, and execute processing.
[0096] A "server" is a central processing system that receives data sent from users and performs processes such as data analysis, video generation, and response generation.
[0097] "Artificial intelligence model" refers to machine learning algorithms and techniques used for complex information processing such as data analysis, video generation, and response generation.
[0098] "Cloud storage" refers to internet-based data storage servers that remotely store generated video data and other data and make it accessible as needed.
[0099] An "episode" refers to a specific memory or event related to a photo entered by the user.
[0100] A "request" refers to an action in which a user requests a specific operation or process from the system.
[0101] "Voice input" refers to voice data input by a user speaking to the system using a microphone or the like.
[0102] "Text data" refers to data in the form of a string of characters obtained by analyzing voice input.
[0103] "Video generation" refers to the process of creating visual and audio content based on photographs and anecdotes.
[0104] "Response generation" refers to the process of generating an appropriate reaction or reply to a user's voice input.
[0105] "Analysis" refers to the processing of collected data to extract useful information and to perform subsequent processing based on that information.
[0106] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and episodes and provides them to patients. The system uses artificial intelligence technology to analyze data entered by users (dementia patients, their families, or caregivers) and create individual rehabilitation plans.
[0107] First, a user uses a dedicated device or application to input past photos and related episodes. For example, the user scans a photo of a family cherry blossom viewing and enters a text description such as "This photo was taken when my family went cherry blossom viewing." This data is temporarily stored by the device and then sent to the server.
[0108] The server receives the photos and text data sent from the device. The received data is analyzed using artificial intelligence models such as TensorFlow and OpenAI's GPT. The analysis extracts basic information about the episode and automatically generates a video that includes emotional elements and a specific location (in this case, cherry blossom viewing). The generated video data is stored in cloud storage such as Amazon S3.
[0109] The user requests playback of the generated video from a dedicated device or application. The server receives the request, retrieves the video data stored in cloud storage, and delivers it to the device via streaming or download. The device plays the received video, and the user (dementia patient) watches it.
[0110] A dementia patient can initiate a dialogue based on the content of the video. For example, they can say, "The cherry blossoms were beautiful at this time." The device analyzes this voice input and converts it into text data. The speech is converted to text using speech recognition technology such as the Google Speech-to-Text API. The converted text data is sent to a server, which uses an artificial intelligence model such as OpenAI's GPT-3 to generate an appropriate response. This response is sent to the device and played back using a text-to-speech engine that converts text to speech.
[0111] A specific example is shown below. For example, if a user inputs "a photo of a family cherry blossom viewing party" and an episode from that event, the system operates as follows.
[0112] 1. The user uses the application to scan a photo of a family cherry blossom viewing party and enters the message, "This photo was taken when my family went cherry blossom viewing."
[0113] 2. The device temporarily stores the photo and text data and sends them to the server.
[0114] 3. The server receives the data and uses an AI model (e.g., TensorFlow) to generate a "cherry blossom viewing"-themed video, which includes footage of cherry blossom viewing and scenes of families having fun.
[0115] 4. The user (patient) requests video playback and watches the video on their device.
[0116] 5. The user (patient) says, "The cherry blossoms were beautiful at this time."
[0117] 6. The device analyzes the voice, converts it into text data, and sends it to the server.
[0118] 7. The server generates a response saying, "It was really beautiful. Your grandma was with you at the time," and plays it back through the device.
[0119] An example prompt is, "Based on photos and stories from this family's cherry blossom viewing, create a video that includes emotional elements and specific locations."
[0120] In this way, the system can stimulate past memories and promote brain activity in dementia patients, and by utilizing artificial intelligence, it can efficiently provide individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[0121] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0122] Step 1:
[0123] User data entry and saving
[0124] The user inputs past photos and episodes using a dedicated device or application. For example, the user scans a photo of a family cherry blossom viewing party and enters a text description such as "This photo is from when we went cherry blossom viewing with my family." The input data (photo and text data) is temporarily saved by the device. The input data is the photo file and text data, and the output is the temporarily saved data.
[0125] Step 2:
[0126] Sending data
[0127] The device sends the saved photos and text data to the server. Specifically, the device sends the data by making an HTTP POST request to the server. The input is the temporarily saved photos and text data, and the output is the data sent to the server.
[0128] Step 3:
[0129] Data reception and analysis
[0130] The server receives the photos and text data sent from the device. The received data is input into an artificial intelligence model such as Google's TensorFlow or OpenAI's GPT, which extracts basic information about the episode from the text data and matches it with the photos to analyze its relevance. The input is the photos and text data sent to the server, and the output is the analyzed basic information about the episode.
[0131] Step 4:
[0132] Video generation
[0133] The server runs an AI video generation module based on the analysis results. It automatically generates videos that include emotional elements or specific locations (e.g., cherry blossom viewing). Specifically, the AI model selects video material based on the analysis results, and the video is edited using an editor. The input is basic information about the analyzed episode, and the output is the generated video data.
[0134] Step 5:
[0135] Saving videos
[0136] The server saves the generated video data in cloud storage (e.g., Amazon S3). It uploads the video data to the cloud using an API. The input is the generated video data, and the output is the video data saved in the cloud storage.
[0137] Step 6:
[0138] Video Request
[0139] A user requests video playback from a dedicated device or application. Specifically, the user presses the play button on the device to send a playback request to the server. The input is the user's playback request, and the output is a playback request to the server.
[0140] Step 7:
[0141] Video data acquisition and distribution
[0142] The server receives a request from a user and retrieves video data from cloud storage. The retrieved video data is then delivered to the device via streaming or download. The input is the video data retrieved from cloud storage, and the output is the video data delivered to the device.
[0143] Step 8:
[0144] Play video
[0145] The device plays the received video data. The video is displayed using a dedicated playback application or a standard video player. The input is the video data delivered to the device, and the output is the played video.
[0146] Step 9:
[0147] Analyzing voice input
[0148] The patient begins a conversation based on the content of the video. For example, they might say, "The cherry blossoms were beautiful at this time." The device analyzes this voice input and converts it into text data using the Google Speech-to-Text API or similar. The input is the patient's voice, and the output is text data.
[0149] Step 10:
[0150] Generate and send the response
[0151] The device sends the converted text data to the server. The server generates an appropriate response based on the received text data using an artificial intelligence model such as OpenAI's GPT. The generated response is sent to the device in text format. The input is text data, and the output is the generated response.
[0152] Step 11:
[0153] Response playback
[0154] The device plays the response sent from the server as audio. It uses a text-to-speech engine to convert the response into audio and plays it back. The input is the generated response text data, and the output is the played audio response.
[0155] (Application example 1)
[0156] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0157] Rehabilitation is important for dementia patients to stimulate their memories and improve their cognitive function. However, effective rehabilitation requires the creation of individual rehabilitation plans, which places a heavy burden on caregivers and medical professionals. Furthermore, there is a problem with brick-and-mortar rehabilitation support, where there is no effective way to provide videos and dialogue based on past memories.
[0158] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0159] In this invention, the server includes a means for collecting past photos and related episodes, a means for generating videos based on the collected data, a means for providing the generated videos to the dementia patient, a means for analyzing a user's voice input, a means for generating responses based on the analyzed voice data, and a means for providing videos and interactive dialogues based on past memories in a physical store, thereby stimulating the memory of the dementia patient and effectively implementing rehabilitation, while reducing the burden on caregivers and medical professionals.
[0160] "Past photos and related episodes" is information provided by a patient with dementia or their family that indicates images taken in the past and memories or events related to the images.
[0161] "Means of collection" refers to the methods and technologies used to collect photos and stories via devices such as tablets and smartphones.
[0162] "Means for generating video" refers to methods and algorithms for creating video using artificial intelligence technology based on collected photos and episodes.
[0163] "Means of delivery to dementia patients" refers to methods of providing the generated videos and other rehabilitation content to dementia patients via devices such as smartphones and tablets.
[0164] "Means for analyzing user voice input" means techniques for collecting voice data using a microphone or other voice capture device and analyzing the voice data.
[0165] "Means for generating a response based on analyzed voice data" refers to a method or algorithm for generating an appropriate reply based on the analyzed user's voice input.
[0166] "Means for providing videos and interactive dialogues based on past memories in physical locations" means technologies and methods for providing videos based on past photos and episodes and related dialogues in physical locations such as cafes and rehabilitation centers.
[0167] This invention is a rehabilitation system for dementia patients, and aims to provide videos and interactive dialogues based on past photos and episodes in a physical store. This system consistently performs everything from collecting photos and episodes to processing the data, providing content, and generating dialogues.
[0168] The system that realizes this application example is made up of a number of hardware and software components, each of which will be described in detail below.
[0169] 1. Data Collection Phase
[0170] Users (dementia patients and their families) enter past photos and episodes through dedicated terminals located in physical stores or smartphone applications. Photos are collected by scanning or uploading, and episodes are entered as text or voice data. The terminal temporarily stores this data and later sends it to a server.
[0171] 2. Data Processing Phase
[0172] The server receives the photos and text data sent from the device. The received data is analyzed by an artificial intelligence model (using TensorFlow, for example) to extract basic information about the episode. Based on this information, an AI video generation module operates and automatically generates videos that include emotional elements and specific locations. The generated video data is stored in cloud storage.
[0173] 3. Content provision phase
[0174] The user then requests playback of the generated video again from a dedicated device or smartphone application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[0175] 4. Dialogue Phase
[0176] The patient with dementia initiates a dialogue based on the content of the video. For example, they might say, "I really enjoyed this day." The device analyzes the voice input and converts it into text data (for example, using a Transformer-based NLP model). The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[0177] Specific examples
[0178] For example, if a family member of a dementia patient inputs a photo of the family cherry blossom viewing and an episode from that time into the system, the process would be as follows:
[0179] 1. The user (family member) logs in to a terminal at a physical store, scans a photo, and enters a text story such as, "This photo was taken when the family went cherry blossom viewing."
[0180] 2. The device sends the photo and text data to the server.
[0181] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[0182] 4. The user (patient) requests video playback and watches the video on a tablet in the store.
[0183] 5. When the user (patient) says, "The cherry blossoms were beautiful at this time," the device analyzes the voice and sends it to the server.
[0184] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[0185] Prompt Sentence Examples
[0186] "If you look at a photo of a family cherry blossom viewing and say the cherry blossoms were beautiful, how will the AI respond?"
[0187] "What kind of comments would you have the AI make when shown photos of past summer festivals?"
[0188] In this way, the system can stimulate past memories in dementia patients, enabling effective rehabilitation while reducing the burden on caregivers and medical professionals.
[0189] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0190] Step 1: Data collection phase
[0191] Specific operation: Users (dementia patients and their families) input past photos and episodes using a dedicated terminal installed in a physical store or a smartphone application. Users scan or upload photos and enter an episode in text, such as "This photo was taken when my family went cherry blossom viewing."
[0192] Input: scanned or uploaded photos, text or audio episodes
[0193] Output: Photos and episode data temporarily saved on the device
[0194] Step 2: Data transmission phase
[0195] Specific operation: The device sends the temporarily stored photos and episode data to the server. During this process, the data is compressed and encrypted as necessary.
[0196] Input: Photos and episode data stored on the device
[0197] Output: Photos and episode data sent to the server
[0198] Step 3: Data analysis phase
[0199] How it works: The server analyzes the received photo and episode data. An artificial intelligence model (using, for example, TensorFlow) extracts basic information about the episode and identifies related themes (e.g., "cherry blossom viewing") based on the analysis results.
[0200] Input: Photos and episode data sent to the server
[0201] Output: Basic information about the analyzed episodes and identified themes
[0202] Step 4: Video Generation Phase
[0203] How it works: Based on the analysis results, the server uses an AI video generation module to automatically generate videos that incorporate emotional elements and specific locations, including footage of actual cherry blossom viewing and scenes of families having fun.
[0204] Input: Basic information about the analyzed episode and the identified themes
[0205] Output: Generated video data
[0206] Step 5: Video saving phase
[0207] Specific operation: The generated video data is stored in cloud storage. Before storing, the video data is compressed and encrypted as necessary.
[0208] Input: Generated video data
[0209] Output: Video data stored in cloud storage
[0210] Step 6: Playback Request Phase
[0211] Specific operation: A user requests playback of the generated video using a dedicated device or a smartphone application. The request is sent to the server.
[0212] Input: User's playback request
[0213] Output: Playback request sent to the server
[0214] Step 7: Video Distribution Phase
[0215] Specific operation: The server receives a playback request and delivers the video data stored in cloud storage to the device in streaming or download format.
[0216] Input: Playback request sent to the server, video data stored in cloud storage
[0217] Output: Video data delivered to the device
[0218] Step 8: Video viewing phase
[0219] Specific operation: The device plays the distributed video and allows the dementia patient to watch it. As the patient watches the video, past memories are stimulated.
[0220] Input: Video data delivered to the device
[0221] Output: Video currently being played by the dementia patient
[0222] Step 9: Audio Analysis Phase
[0223] Specific operation: The dementia patient begins a dialogue based on the content of the video. When a voice input such as "This day was really fun," the device analyzes the voice and converts it into text data.
[0224] Input: Voice input from dementia patients
[0225] Output: Speech input converted to text data
[0226] Step 10: Response generation phase
[0227] Specific operation: The server receives the converted text data and uses artificial intelligence techniques (e.g., a Transformer-based NLP model) to generate an appropriate response. The generated response is then sent from the server to the device.
[0228] Input: Voice input converted to text data
[0229] Output: The server-generated response
[0230] Step 11: Response delivery phase
[0231] Specific operation: The device plays back the response sent from the server and provides it to the dementia patient. Responses such as "It was really beautiful. Your grandmother was with you at the time" are played back.
[0232] Input: The response sent by the server
[0233] Output: Response voice heard by dementia patient
[0234] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0235] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and reminiscences and provides them to patients. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can provide an individual rehabilitation plan more effectively.
[0236] The program for this system mainly consists of the following phases:
[0237] Data Collection Phase
[0238] Users (people with dementia or their families) log in to their accounts through a dedicated device or application. They then scan or upload past photos and enter related memories as text or voice data. The device temporarily stores the collected photos and text data in its local storage and then sends them to the server.
[0239] Data Processing Phase
[0240] The server receives the photos and text data sent from the device. The received data is input into a natural language processing (NLP) module to extract basic information about the episode. Based on this information, an AI-based video generation module operates to automatically generate videos that include emotional elements and specific locations. The generated video data is then stored in cloud storage.
[0241] Content provision phase
[0242] The user requests playback of the generated video from a dedicated device or application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[0243] Dialogue Phase
[0244] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[0245] Emotion Recognition Phase
[0246] The emotion engine analyzes the user's voice input and facial expression data, determining their emotional state and generating emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[0247] Specific examples
[0248] For example, a family member of a dementia patient can enter "photos from cherry blossom viewing" and "anecdotes from that time" into the system, and the process will go as follows:
[0249] 1. The user (family member) logs in to the application, scans a photo, and enters the text "This photo was taken when my family went cherry blossom viewing."
[0250] 2. The device sends the photo and text data to the server.
[0251] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[0252] 4. The user (patient) requests video playback and watches the video on their device.
[0253] 5. The user (patient) says, "The cherry blossoms were beautiful at this time." The device analyzes the voice and sends it to the server.
[0254] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[0255] 7. The emotion engine analyzes the user's emotions and adjusts the next dialogue content based on the results. For example, if the user is happy, it will provide more stories related to happiness.
[0256] In this way, the combination of emotion recognition capabilities can more effectively advance the personalized rehabilitation plan for dementia patients, while providing responses that correspond to the user's emotional state makes interactions more natural and effective.
[0257] The processing flow will be explained below.
[0258] Step 1:
[0259] Users (dementia patients and their families) log in to their accounts through a dedicated device or application.
[0260] Step 2:
[0261] Users scan or upload old photos and enter related memories as text or audio data.
[0262] Step 3:
[0263] The device temporarily stores the entered photo and text data in local storage.
[0264] Step 4:
[0265] The device transmits the photos and text data stored in the local storage to the server.
[0266] Step 5:
[0267] The server receives the photo and text data sent from the terminal.
[0268] Step 6:
[0269] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode from the text data.
[0270] Step 7:
[0271] The server uses an AI video generation module to generate videos based on the extracted information, including emotional elements and specific locations.
[0272] Step 8:
[0273] The server stores the generated video data in cloud storage.
[0274] Step 9:
[0275] The user requests playback of the generated video from a dedicated terminal or application.
[0276] Step 10:
[0277] The server receives the user's request and delivers the video data to the terminal in streaming or download format.
[0278] Step 11:
[0279] The terminal plays the distributed video and allows the user to view it.
[0280] Step 12:
[0281] The user (a dementia patient) can start a conversation based on the content of the video, for example, saying, "This day was really fun."
[0282] Step 13:
[0283] The device analyzes the user's voice input and converts it into text data using voice recognition technology.
[0284] Step 14:
[0285] The terminal transmits the analyzed voice data (text) to the server.
[0286] Step 15:
[0287] The server generates a response using a natural language generation (NLG) module based on the received voice data.
[0288] Step 16:
[0289] The server transmits the generated response data to the terminal.
[0290] Step 17:
[0291] The terminal uses a voice synthesis function to reproduce the received response data and convey it to the user.
[0292] Step 18:
[0293] The device inputs the user's voice and facial expression data into an emotion engine to analyze the user's emotional state. For example, emotion tags such as "joy," "sadness," and "excitement" are generated.
[0294] Step 19:
[0295] The server automatically adjusts the content of the next video and dialogue responses based on the user's emotions recognized by the emotion engine. For example, if the user is happy, the server generates a video that includes more episodes related to happiness.
[0296] Step 20:
[0297] The user can receive continually adapted rehabilitation through updated content.
[0298] The above is the specific processing flow of the system of the invention combined with the emotion engine. In this way, by providing rehabilitation that corresponds to the user's emotions, it is possible to effectively stimulate the memory of dementia patients and realize more natural and effective dialogue.
[0299] Example 2
[0300] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0301] In rehabilitation for dementia patients, there is a need to provide effective stimulation that corresponds to the individual memories and emotions of each patient. Conventional rehabilitation methods have difficulty customizing to correspond to the memories and emotions of each patient, which often limits their effectiveness. For this reason, there is a need to develop a system that can automatically provide an individual rehabilitation plan that corresponds to the patient's memory ability and emotions.
[0302] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0303] In this invention, the server includes: means for a user to collect past photos and related memories; means for transmitting the collected data from the terminal to the server; means for the server to analyze the received data using a natural language processing module and extract basic information about the episode; means for generating a video using a generative AI model based on the extracted information; means for storing the generated video in cloud storage; means for a user to request video playback; means for the server to deliver video data to the terminal in response to the request; means for the terminal to play the video for the user; means for analyzing the user's voice input and converting it into text data; means for transmitting the converted voice data to the server and generating an appropriate response; means for transmitting the response data generated by the server to the terminal and playing it; means for analyzing the user's voice and facial expressions using an emotion engine to determine the user's emotional state; and means for adjusting the content to be provided next based on the emotional state. This makes it possible to automatically provide rehabilitation tailored to the individual memories and emotions of dementia patients.
[0304] "User" refers to people who use the system to input past photos and memories, or dementia patients and their families.
[0305] "Terminal" refers to the device through which a user accesses the system, uploads photos, and performs voice input. Specifically, this includes smartphones, tablets, and PCs.
[0306] "Server" refers to a computer system that receives, analyzes, stores, and manages data sent by users.
[0307] "Natural language processing module" refers to a software module for analyzing received text data and extracting basic information about an episode.
[0308] "Generative AI model" refers to an artificial intelligence algorithm that automatically generates videos based on extracted information.
[0309] "Cloud storage" refers to an online storage service for storing generated video data.
[0310] An "emotion engine" refers to a computer program that analyzes a user's voice and facial expressions to determine their emotional state and generate appropriate responses and content.
[0311] A "request" refers to an operation by a user to request playback of a video through a terminal.
[0312] "Video" refers to video content generated from past photos and related memories.
[0313] "Text data" refers to data that has been analyzed and converted into text information from a user's voice.
[0314] "Response" refers to an appropriate reply generated by a server in response to a user's input.
[0315] "Content" refers to videos, responses, and other information provided to users.
[0316] "Emotional state" refers to the emotional state such as joy, sadness, excitement, etc., analyzed from the user's voice and facial expressions.
[0317] "Rehabilitation" refers to treatment aimed at stimulating the memories and emotions of dementia patients and maintaining or improving their functions.
[0318] This invention is a system aimed at the rehabilitation of dementia patients, which generates videos based on past photos and reminiscences and provides them to patients. It also uses an emotion engine that recognizes the user's emotions to more effectively provide individual rehabilitation plans.
[0319] Overall flow
[0320] 1. Data Collection Phase
[0321] Users log in to the system using a dedicated device or application, then scan or upload old photos and enter related memories as text or voice data. The device temporarily saves this data in its local storage and then sends it to the server.
[0322] 2. Data Processing Phase
[0323] The server receives the photos and text data sent from the device. The received data is input into a natural language processing (NLP) module to extract basic information about the episode. Based on the extracted information, a generative AI model generates videos containing emotional elements. The generated videos are then stored in cloud storage.
[0324] 3. Content provision phase
[0325] The user requests video playback from a dedicated device or application. The server receives this request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[0326] 4. Dialogue Phase
[0327] The patient can start a dialogue based on the content of the video. For example, if the patient says, "I had a really fun day," the device analyzes the voice input and converts it into text data. The converted data is sent to the server, which then generates an appropriate response and plays it back through the device.
[0328] 5. Emotion Recognition Phase
[0329] The user's voice and facial expression data are analyzed by the emotion engine, which determines the user's emotional state and generates emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[0330] Hardware and software used
[0331] Devices: Smartphones, tablets, computers, etc.
[0332] Server: A high-performance computer that receives, analyzes, stores, and runs generative AI models.
[0333] Natural Language Processing Module: Software that performs text analysis (e.g., SpaCy, NLTK, etc.).
[0334] Generative AI models: Artificial intelligence algorithms that generate videos (e.g., GPT-3, DALL-E).
[0335] Emotion engine: A software module that performs emotion analysis (e.g., IBM Watson, Microsoft Azure Emotion API).
[0336] Cloud storage: An internet storage service for saving video data (e.g., Amazon S3, Google Cloud Storage).
[0337] Specific examples
[0338] For example, if a family member of a dementia patient enters "photos from a cherry blossom viewing party" and "an episode from that time" into the system, the process would be as follows:
[0339] 1. The user (family member) logs in to the application, enters the text "This photo was taken when my family went cherry blossom viewing," and scans the photo.
[0340] 2. The device temporarily stores the photo and text data and sends them to the server.
[0341] 3. The server receives the data and uses a natural language processing module to extract basic information about the episodes on the theme of "cherry blossom viewing." A generative AI model then generates a video based on that information and saves it in cloud storage.
[0342] 4. The user (patient) uses the application to request video playback and watches the video on their device.
[0343] 5. When the user (patient) says, "The cherry blossoms were beautiful at this time," the device captures the voice, converts it into text, and sends it to the server. The server generates an appropriate response (e.g., "They were really beautiful. Your grandmother was with you at the time") and plays it back on the device.
[0344] 6. The emotion engine analyzes the user's emotional state, and if they are happy, it will provide them with the next relevant episode.
[0345] Prompt Sentence Examples
[0346] An example of a prompt for a generative AI model is:
[0347] text
[0348] "Upload a photo of you and your family going cherry blossom viewing, enter the text 'We had a great time enjoying cherry blossom viewing in the park on this day,' and generate a two-minute video with the theme 'Cherry Blossom Park.'"
[0349] In this way, combining emotion recognition capabilities can effectively advance individual rehabilitation for dementia patients. Also, by providing responses that correspond to the user's emotional state, interactions become more natural and effective.
[0350] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0351] Step 1:
[0352] A user logs in to a dedicated terminal or application.
[0353] Specific operation: The user launches the application and logs in by entering their account information (username and password).
[0354] Input: Username, Password.
[0355] Output: The user is logged in.
[0356] Step 2:
[0357] Users scan or upload old photos and enter their memories.
[0358] What happens: The user uses the device's camera to scan a photo or upload an existing photo file, then uses the keyboard to enter text or the microphone to enter voice input.
[0359] Input: Photo files, text data, or audio data.
[0360] Output: Collected photos and text data.
[0361] Step 3:
[0362] The device temporarily stores the collected data in local storage and then transmits it to the server.
[0363] Specific operation: The device temporarily saves the photo and text data as files in local storage, and then sends the data to the server.
[0364] Input: Collected photo files and text data.
[0365] Output: The data sent to the server.
[0366] Step 4:
[0367] The server receives the data sent from the terminal.
[0368] Specific operation: The server receives the photos and text data sent from the device via the network and temporarily stores them in the server's storage.
[0369] Input: Photo files and text data sent from the device.
[0370] Output: Data stored on the server.
[0371] Step 5:
[0372] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode.
[0373] Specific operation: The server inputs the received text data into a natural language processing (NLP) module to extract basic information such as keywords, date and time, location, and names of people.
[0374] Input: Received text data.
[0375] Output: Basic information about the extracted episode.
[0376] Step 6:
[0377] The server generates a video using a generative AI model based on the extracted information.
[0378] How it works: The server inputs the extracted basic information into the generative AI model, builds a video scenario based on it, and then generates a video by combining related images and music.
[0379] Input: Basic information.
[0380] Output: The generated video data.
[0381] Step 7:
[0382] The server stores the generated video in cloud storage.
[0383] Specific operation: The server saves the generated video to a cloud storage service and generates a URL for accessing it.
[0384] Input: The generated video data.
[0385] Output: Videos stored in cloud storage, access URL.
[0386] Step 8:
[0387] The user requests video playback from a dedicated device or application.
[0388] Specific operation: The user logs in to the application again and selects the video they want to play from the generated video list page.
[0389] Input: Login information, video request.
[0390] Output: Video playback request.
[0391] Step 9:
[0392] The server distributes the video data to the terminal.
[0393] Specific operation: The server searches for video data in response to a user request and delivers it to the device in streaming or download format.
[0394] Input: A video playback request.
[0395] Output: Video data delivery to the device.
[0396] Step 10:
[0397] The terminal plays the video and allows the user to watch it.
[0398] Specific operation: The device decodes the received video data and plays it on the display. The user watches the video on the device.
[0399] Input: Received video data.
[0400] Output: The video shown on the display.
[0401] Step 11:
[0402] Users begin to interact based on the content of the video.
[0403] Specific operation: The user talks while watching the video (e.g., "The cherry blossoms were beautiful at this time."). The device captures the user's voice with a microphone.
[0404] Input: User's voice input.
[0405] Output: The captured audio data.
[0406] Step 12:
[0407] The device analyzes the voice, converts it into text data, and sends it to the server.
[0408] Specific operation: The device uses a speech recognition engine to convert the user's voice into text data and send it to the server.
[0409] Input: The captured audio data.
[0410] Output: The text data sent to the server.
[0411] Step 13:
[0412] The server generates an appropriate response based on the data it receives.
[0413] Specific operation: The server understands the context from the received text data and generates an appropriate response using a generative AI model.
[0414] Input: Received text data.
[0415] Output: The generated response data.
[0416] Step 14:
[0417] The terminal plays back the generated response.
[0418] Specific operation: The device converts the response data received from the server into speech using a speech synthesis engine and plays it back to the user.
[0419] Input: The response data received.
[0420] Output: The audio response played.
[0421] Step 15:
[0422] The emotion engine analyzes the user's voice and facial expression data to determine their emotional state.
[0423] How it works: The device captures the user's facial expressions with a camera and records their voice. The emotion engine analyzes this data to determine the user's emotional state.
[0424] Input: User's facial expression data, voice data.
[0425] Output: Emotion tag (e.g., happy, sad, excited).
[0426] Step 16:
[0427] Tailor your next piece of content based on your emotional state.
[0428] Specific operation: Based on the emotion tag, the server adjusts the next video and dialogue response according to the user's emotional state.
[0429] Input: emotion tag.
[0430] Output: A tailored content delivery plan.
[0431] (Application example 2)
[0432] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0433] Conventional rehabilitation systems provide the same content to all workers, including those with dementia and those unfamiliar with work processes, making it difficult to respond appropriately to individual needs and emotional states. Furthermore, there were insufficient means to properly reproduce and visually present past work procedures and episodes, making it difficult for workers to efficiently understand work procedures.
[0434] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting past photos and related memories, means for generating a video based on the collected data, means for providing the generated video to the target person, means for analyzing the user's voice input, means for generating a response based on the analyzed voice data, means for providing the generated response to the user, means for generating a video that reproduces a work procedure based on the collected data, and means for recognizing the user's emotional state and adjusting the content of the video and the response based on the emotional state. This enables appropriate responses according to individual needs and emotional states, and by providing a visual reproduction of past work procedures and episodes, the target person can efficiently understand the information.
[0435] "Past photographs and related reminiscences" are visual information related to the subject's past experiences or episodes and accompanying text or audio data.
[0436] "Collected data" is a collection of information provided by users, including photographs, text data, audio data, etc.
[0437] "Video generation means" refers to technology or devices that use collected data to create videos that convey information visually and audibly.
[0438] The "means for providing a video" refers to a method or device for allowing a target person to view the generated video.
[0439] "Means for analyzing user voice input" refers to technology that converts the user's voice into text data and understands and analyzes its content.
[0440] A "means for generating a response" is a technique or device that determines an appropriate response or next action based on the analyzed voice data.
[0441] "Means for generating videos that reproduce work procedures based on collected data" refers to technology or devices that visually reproduce past work procedures or episodes and provide them as videos in an easy-to-understand format.
[0442] "Means for recognizing the user's emotional state" refers to technology or devices that analyze the user's voice, facial expressions, behavior, etc. to determine and identify their emotional state.
[0443] The "means for adjusting the response content" refers to a technique or device for appropriately changing the video or response content provided depending on the user's emotional state or situation.
[0444] This invention is a rehabilitation system for dementia patients and workers, which aims to generate videos based on past photos and reminiscences and provide them to users. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides an individual rehabilitation plan.
[0445] Program processing explanation
[0446] Data Collection Phase
[0447] Terminal
[0448] Users log in to their account using a dedicated device (smartphone or head-mounted display) and upload past photos and related memories. This data is entered as text or voice data, temporarily saved in local storage, and then sent to the server.
[0449] Data Processing Phase
[0450] server
[0451] The server receives the photos and text data sent from the device. The received data is analyzed using a natural language processing (NLP) module (e.g., SpaCy or Google NLP API) to extract basic information about the episode. Based on this information, an AI-based video generation module (e.g., OpenAI or DeepMind) runs and automatically generates videos based on work procedures and reminiscences. The generated video data is stored in cloud storage (e.g., Google Cloud Storage or AWS S3).
[0452] Content provision phase
[0453] Terminal
[0454] The user requests playback of the generated video from a dedicated terminal. The server receives the request and delivers the video data to the terminal via streaming or download. The terminal then plays the received video and allows the user to watch it.
[0455] Dialogue Phase
[0456] Server and Device
[0457] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[0458] Emotion Recognition Phase
[0459] server
[0460] The user's voice input and facial expression data are analyzed by an emotion engine (e.g., Microsoft Azure Cognitive Services or Amazon Rekognition). The emotion engine determines the user's emotional state and generates emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[0461] Specific examples
[0462] For example, the following flow would be used for a rehabilitation system that supports operations within a factory.
[0463] 1. A worker uses a terminal to upload photos and procedures related to past "factory line maintenance work."
[0464] 2. The server receives the data and uses a video generation AI model to generate a video that reproduces the "factory line maintenance work procedures."
[0465] 3. The worker requests "Maintenance procedure playback" and watches the video.
[0466] 4. The worker asks, "What should I pay attention to in this part?" The device analyzes the voice and sends it to the server.
[0467] 5. The server generates a response saying, "It is important to check the safety devices in this section. In particular, make sure this switch is working," and plays it back through the terminal.
[0468] Prompt Sentence Examples
[0469] Generate a video that recreates a past maintenance procedure on a factory line based on the following data:
[0470] Photo: Maintenance work on a factory line
[0471] Procedure: 1. Turn off the power. 2. Check the safety devices. 3. Clean the conveyor belt. 4. Test the operation.
[0472] Your video should include the following elements:
[0473] 1. Detailed explanation of each step
[0474] 2. Points to pay particular attention to
[0475] 3. Clear visual and audio guidance for easy understanding by workers
[0476] Also, generate appropriate responses according to the worker's emotions.
[0477] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0478] Step 1: Data collection
[0479] The user logs into their account using a dedicated device (smartphone or head-mounted display) and uploads past photos and related memories. The input is photo data and text or voice data, which are temporarily stored in local storage. The output is data that is saved in local storage and then sent to the server. The specific actions that take place in this step are scanning photos, inputting text and voice data, saving the data in local storage, and sending it to the server.
[0480] Step 2: Data reception and analysis
[0481] The server receives the photo and text data sent from the device. The input is the data sent from the device, and the output is the data stored on the server. This data is input into a natural language processing (NLP) module, which extracts basic information about the episode. The specific operations performed here are receiving the data, storing it, and analyzing it using the NLP module.
[0482] Step 3: Video Generation
[0483] The AI video generation module in the server generates videos based on the analyzed episode information. The input is the data analyzed by the NLP module, and the output is the generated video data. The specific operations are as follows: input data into the AI model, generate video, and save it to cloud storage.
[0484] Step 4: Submit your video
[0485] A user requests playback of a generated video from a dedicated device. The input is a video playback request, and the output is video data that is streamed or downloaded to the user. Specific operations include accepting the request, obtaining the video data, and delivering it to the device.
[0486] Step 5: Dialogue analysis
[0487] The device receives and analyzes the audio emitted by the user while watching a video. The input is audio data, and the output is the analysis results converted into text data. The specific operations are to collect audio, analyze the audio, convert it into text data, and send it to the server.
[0488] Step 6: Response Generation
[0489] The server generates an appropriate response based on the received voice data. The input is the analyzed voice data, and the output is the generated response data. The specific operations are analyzing the voice data, determining the response content, and generating the response.
[0490] Step 7: Provide a response
[0491] The generated response is provided to the user through the terminal. The input is the generated response data, and the output is the response voice heard by the user. The specific operations are receiving the response data, generating the voice, and delivering it to the user.
[0492] Step 8: Emotion Recognition
[0493] The emotion engine analyzes the user's voice input and facial expression data. The input is the user's voice and facial expression data, and the output is an emotion tag. Specific operations include collecting voice and facial expression data, analyzing emotions, and generating emotion tags.
[0494] Step 9: Content Adjustment
[0495] The next video or response content is adjusted based on the emotion tag. The input is the emotion tag and the next content data to be provided, and the output is the adjusted video or response content. The specific operations are emotion tag analysis, content adjustment, and next content generation.
[0496] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0497] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0498] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0499] [Second embodiment]
[0500] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0501] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0502] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0503] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0504] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0505] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0506] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0507] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0508] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0509] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0510] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0511] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0512] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and reminiscences and provides them to patients. The system uses artificial intelligence technology to create individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[0513] The program for this system mainly consists of the following phases:
[0514] Data Collection Phase
[0515] Users (dementia patients and their families) input past photos and episodes through a dedicated device or application. Photos are collected by scanning or uploading, and episodes are input as text or voice data. The device temporarily stores the collected data and sends it to a server.
[0516] Data Processing Phase
[0517] The server receives the photos and text data sent from the device. The received data is analyzed by an artificial intelligence model to extract basic information about the episode. Based on this information, an AI video generation module operates and automatically generates videos that include emotional elements and specific locations. The generated video data is stored in cloud storage.
[0518] Content provision phase
[0519] The user requests playback of the generated video from a dedicated device or application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[0520] Dialogue Phase
[0521] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[0522] Specific examples
[0523] For example, if a family member of a dementia patient inputs a photo of the family cherry blossom viewing and an episode from that time into the system, the process would be as follows:
[0524] 1. The user (family member) logs in to the application, scans a photo, and enters the text "This photo was taken when my family went cherry blossom viewing."
[0525] 2. The device sends the photo and text data to the server.
[0526] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[0527] 4. The user (patient) requests video playback and watches the video on their device.
[0528] 5. The user (patient) says, "The cherry blossoms were beautiful at this time." The device analyzes the voice and sends it to the server.
[0529] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[0530] In this way, the system can stimulate past memories and promote brain activity in dementia patients. Furthermore, by utilizing artificial intelligence, it can efficiently provide individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[0531] The processing flow will be explained below.
[0532] Step 1:
[0533] Users (dementia patients and their families) log in to their accounts through a dedicated device or application.
[0534] Step 2:
[0535] Users scan or upload old photos and enter related memories as text or audio data.
[0536] Step 3:
[0537] The device temporarily stores the entered photo and text data in local storage.
[0538] Step 4:
[0539] The device transmits the photos and text data stored in the local storage to the server.
[0540] Step 5:
[0541] The server receives the photo and text data sent from the terminal.
[0542] Step 6:
[0543] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode from the text data.
[0544] Step 7:
[0545] The server uses an AI video generation module to generate videos based on the extracted information, including emotional elements and specific locations.
[0546] Step 8:
[0547] The server stores the generated video data in cloud storage.
[0548] Step 9:
[0549] The user requests playback of the generated video from a dedicated terminal or application.
[0550] Step 10:
[0551] The server receives the user's request and delivers the video data to the terminal in streaming or download format.
[0552] Step 11:
[0553] The terminal plays the distributed video and allows the user to view it.
[0554] Step 12:
[0555] The user (a dementia patient) can start a conversation based on the content of the video, for example, saying, "This day was really fun."
[0556] Step 13:
[0557] The device analyzes the user's voice input and converts it into text data using voice recognition technology.
[0558] Step 14:
[0559] The terminal transmits the analyzed voice data (text) to the server.
[0560] Step 15:
[0561] The server generates a response using a natural language generation (NLG) module based on the received voice data.
[0562] Step 16:
[0563] The server transmits the generated response data to the terminal.
[0564] Step 17:
[0565] The terminal uses a voice synthesis function to reproduce the received response data and convey it to the user.
[0566] Example 1
[0567] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0568] In the rehabilitation of dementia patients, it is important to stimulate individual memories and elicit emotional responses. However, rehabilitation using past photographs and anecdotes can be laborious for some caregivers and medical professionals. Furthermore, these tasks tend to rely on subjective judgment and are difficult to perform consistently. This can result in the insufficient effectiveness of dementia rehabilitation. Therefore, there is a need for an efficient and individually tailored rehabilitation system.
[0569] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0570] In this invention, the server includes: a means for a user to input past photos and episodes; a means for transmitting the input photos and episodes to the server; a means for the server to analyze the received data using an artificial intelligence model; a means for automatically generating videos based on the analysis results; a means for storing the generated video data in cloud storage; a means for a user to request video playback; a means for the server to receive the request and deliver the video data to a terminal; a means for the terminal to play the video; a means for analyzing the user's voice input and converting it into text data; a means for the server to generate a response based on the voice input; and a means for transmitting the generated response to the terminal and playing it as audio. This allows a user to input past photos and episodes and efficiently stimulate the memory of a dementia patient using the generated video. Furthermore, the use of artificial intelligence can provide a consistent rehabilitation plan, reducing the burden on caregivers and medical professionals.
[0571] "User" refers to the person who operates the system and inputs data, such as a dementia patient, their family member, or caregiver.
[0572] A "terminal" is an electronic device that allows a user to input data, receive data from a server, and execute processing.
[0573] A "server" is a central processing system that receives data sent from users and performs processes such as data analysis, video generation, and response generation.
[0574] "Artificial intelligence model" refers to machine learning algorithms and techniques used for complex information processing such as data analysis, video generation, and response generation.
[0575] "Cloud storage" refers to internet-based data storage servers that remotely store generated video data and other data and make it accessible as needed.
[0576] An "episode" refers to a specific memory or event related to a photo entered by the user.
[0577] A "request" refers to an action in which a user requests a specific operation or process from the system.
[0578] "Voice input" refers to voice data input by a user speaking to the system using a microphone or the like.
[0579] "Text data" refers to data in the form of a string of characters obtained by analyzing voice input.
[0580] "Video generation" refers to the process of creating visual and audio content based on photographs and anecdotes.
[0581] "Response generation" refers to the process of generating an appropriate reaction or reply to a user's voice input.
[0582] "Analysis" refers to the processing of collected data to extract useful information and to perform subsequent processing based on that information.
[0583] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and episodes and provides them to patients. The system uses artificial intelligence technology to analyze data entered by users (dementia patients, their families, or caregivers) and create individual rehabilitation plans.
[0584] First, a user uses a dedicated device or application to input past photos and related episodes. For example, the user scans a photo of a family cherry blossom viewing and enters a text description such as "This photo was taken when my family went cherry blossom viewing." This data is temporarily stored by the device and then sent to the server.
[0585] The server receives the photos and text data sent from the device. The received data is analyzed using artificial intelligence models such as TensorFlow and OpenAI's GPT. The analysis extracts basic information about the episode and automatically generates a video that includes emotional elements and a specific location (in this case, cherry blossom viewing). The generated video data is stored in cloud storage such as Amazon S3.
[0586] The user requests playback of the generated video from a dedicated device or application. The server receives the request, retrieves the video data stored in cloud storage, and delivers it to the device via streaming or download. The device plays the received video, and the user (dementia patient) watches it.
[0587] A dementia patient can initiate a dialogue based on the content of the video. For example, they can say, "The cherry blossoms were beautiful at this time." The device analyzes this voice input and converts it into text data. The speech is converted to text using speech recognition technology such as the Google Speech-to-Text API. The converted text data is sent to a server, which uses an artificial intelligence model such as OpenAI's GPT-3 to generate an appropriate response. This response is sent to the device and played back using a text-to-speech engine that converts text to speech.
[0588] A specific example is shown below. For example, if a user inputs "a photo of a family cherry blossom viewing party" and an episode from that event, the system operates as follows.
[0589] 1. The user uses the application to scan a photo of a family cherry blossom viewing party and enters the message, "This photo was taken when my family went cherry blossom viewing."
[0590] 2. The device temporarily stores the photo and text data and sends them to the server.
[0591] 3. The server receives the data and uses an AI model (e.g., TensorFlow) to generate a "cherry blossom viewing"-themed video, which includes footage of cherry blossom viewing and scenes of families having fun.
[0592] 4. The user (patient) requests video playback and watches the video on their device.
[0593] 5. The user (patient) says, "The cherry blossoms were beautiful at this time."
[0594] 6. The device analyzes the voice, converts it into text data, and sends it to the server.
[0595] 7. The server generates a response saying, "It was really beautiful. Your grandma was with you at the time," and plays it back through the device.
[0596] An example prompt is, "Based on photos and stories from this family's cherry blossom viewing, create a video that includes emotional elements and specific locations."
[0597] In this way, the system can stimulate past memories and promote brain activity in dementia patients, and by utilizing artificial intelligence, it can efficiently provide individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[0598] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0599] Step 1:
[0600] User data entry and saving
[0601] The user inputs past photos and episodes using a dedicated device or application. For example, the user scans a photo of a family cherry blossom viewing party and enters a text description such as "This photo is from when we went cherry blossom viewing with my family." The input data (photo and text data) is temporarily saved by the device. The input data is the photo file and text data, and the output is the temporarily saved data.
[0602] Step 2:
[0603] Sending data
[0604] The device sends the saved photos and text data to the server. Specifically, the device sends the data by making an HTTP POST request to the server. The input is the temporarily saved photos and text data, and the output is the data sent to the server.
[0605] Step 3:
[0606] Data reception and analysis
[0607] The server receives the photos and text data sent from the device. The received data is input into an artificial intelligence model such as Google's TensorFlow or OpenAI's GPT, which extracts basic information about the episode from the text data and matches it with the photos to analyze its relevance. The input is the photos and text data sent to the server, and the output is the analyzed basic information about the episode.
[0608] Step 4:
[0609] Video generation
[0610] The server runs an AI video generation module based on the analysis results. It automatically generates videos that include emotional elements or specific locations (e.g., cherry blossom viewing). Specifically, the AI model selects video material based on the analysis results, and the video is edited using an editor. The input is basic information about the analyzed episode, and the output is the generated video data.
[0611] Step 5:
[0612] Saving videos
[0613] The server saves the generated video data in cloud storage (e.g., Amazon S3). It uploads the video data to the cloud using an API. The input is the generated video data, and the output is the video data saved in the cloud storage.
[0614] Step 6:
[0615] Video Request
[0616] A user requests video playback from a dedicated device or application. Specifically, the user presses the play button on the device to send a playback request to the server. The input is the user's playback request, and the output is a playback request to the server.
[0617] Step 7:
[0618] Video data acquisition and distribution
[0619] The server receives a request from a user and retrieves video data from cloud storage. The retrieved video data is then delivered to the device via streaming or download. The input is the video data retrieved from cloud storage, and the output is the video data delivered to the device.
[0620] Step 8:
[0621] Play video
[0622] The device plays the received video data. The video is displayed using a dedicated playback application or a standard video player. The input is the video data delivered to the device, and the output is the played video.
[0623] Step 9:
[0624] Analyzing voice input
[0625] The patient begins a conversation based on the content of the video. For example, they might say, "The cherry blossoms were beautiful at this time." The device analyzes this voice input and converts it into text data using the Google Speech-to-Text API or similar. The input is the patient's voice, and the output is text data.
[0626] Step 10:
[0627] Generate and send the response
[0628] The device sends the converted text data to the server. The server generates an appropriate response based on the received text data using an artificial intelligence model such as OpenAI's GPT. The generated response is sent to the device in text format. The input is text data, and the output is the generated response.
[0629] Step 11:
[0630] Response playback
[0631] The device plays the response sent from the server as audio. It uses a text-to-speech engine to convert the response into audio and plays it back. The input is the generated response text data, and the output is the played audio response.
[0632] (Application example 1)
[0633] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0634] Rehabilitation is important for dementia patients to stimulate their memories and improve their cognitive function. However, effective rehabilitation requires the creation of individual rehabilitation plans, which places a heavy burden on caregivers and medical professionals. Furthermore, there is a problem with brick-and-mortar rehabilitation support, where there is no effective way to provide videos and dialogue based on past memories.
[0635] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0636] In this invention, the server includes a means for collecting past photos and related episodes, a means for generating videos based on the collected data, a means for providing the generated videos to the dementia patient, a means for analyzing a user's voice input, a means for generating responses based on the analyzed voice data, and a means for providing videos and interactive dialogues based on past memories in a physical store, thereby stimulating the memory of the dementia patient and effectively implementing rehabilitation, while reducing the burden on caregivers and medical professionals.
[0637] "Past photos and related episodes" is information provided by a patient with dementia or their family that indicates images taken in the past and memories or events related to the images.
[0638] "Means of collection" refers to the methods and technologies used to collect photos and stories via devices such as tablets and smartphones.
[0639] "Means for generating video" refers to methods and algorithms for creating video using artificial intelligence technology based on collected photos and episodes.
[0640] "Means of delivery to dementia patients" refers to methods of providing the generated videos and other rehabilitation content to dementia patients via devices such as smartphones and tablets.
[0641] "Means for analyzing user voice input" means techniques for collecting voice data using a microphone or other voice capture device and analyzing the voice data.
[0642] "Means for generating a response based on analyzed voice data" refers to a method or algorithm for generating an appropriate reply based on the analyzed user's voice input.
[0643] "Means for providing videos and interactive dialogues based on past memories in physical locations" means technologies and methods for providing videos based on past photos and episodes and related dialogues in physical locations such as cafes and rehabilitation centers.
[0644] This invention is a rehabilitation system for dementia patients, and aims to provide videos and interactive dialogues based on past photos and episodes in a physical store. This system consistently performs everything from collecting photos and episodes to processing the data, providing content, and generating dialogues.
[0645] The system that realizes this application example is made up of a number of hardware and software components, each of which will be described in detail below.
[0646] 1. Data Collection Phase
[0647] Users (dementia patients and their families) enter past photos and episodes through dedicated terminals located in physical stores or smartphone applications. Photos are collected by scanning or uploading, and episodes are entered as text or voice data. The terminal temporarily stores this data and later sends it to a server.
[0648] 2. Data Processing Phase
[0649] The server receives the photos and text data sent from the device. The received data is analyzed by an artificial intelligence model (using TensorFlow, for example) to extract basic information about the episode. Based on this information, an AI video generation module operates and automatically generates videos that include emotional elements and specific locations. The generated video data is stored in cloud storage.
[0650] 3. Content provision phase
[0651] The user then requests playback of the generated video again from a dedicated device or smartphone application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[0652] 4. Dialogue Phase
[0653] The patient with dementia initiates a dialogue based on the content of the video. For example, they might say, "I really enjoyed this day." The device analyzes the voice input and converts it into text data (for example, using a Transformer-based NLP model). The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[0654] Specific examples
[0655] For example, if a family member of a dementia patient inputs a photo of the family cherry blossom viewing and an episode from that time into the system, the process would be as follows:
[0656] 1. The user (family member) logs in to a terminal at a physical store, scans a photo, and enters a text story such as, "This photo was taken when the family went cherry blossom viewing."
[0657] 2. The device sends the photo and text data to the server.
[0658] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[0659] 4. The user (patient) requests video playback and watches the video on a tablet in the store.
[0660] 5. When the user (patient) says, "The cherry blossoms were beautiful at this time," the device analyzes the voice and sends it to the server.
[0661] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[0662] Prompt Sentence Examples
[0663] "If you look at a photo of a family cherry blossom viewing and say the cherry blossoms were beautiful, how will the AI respond?"
[0664] "What kind of comments would you have the AI make when shown photos of past summer festivals?"
[0665] In this way, the system can stimulate past memories in dementia patients, enabling effective rehabilitation while reducing the burden on caregivers and medical professionals.
[0666] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0667] Step 1: Data collection phase
[0668] Specific operation: Users (dementia patients and their families) input past photos and episodes using a dedicated terminal installed in a physical store or a smartphone application. Users scan or upload photos and enter an episode in text, such as "This photo was taken when my family went cherry blossom viewing."
[0669] Input: scanned or uploaded photos, text or audio episodes
[0670] Output: Photos and episode data temporarily saved on the device
[0671] Step 2: Data transmission phase
[0672] Specific operation: The device sends the temporarily stored photos and episode data to the server. During this process, the data is compressed and encrypted as necessary.
[0673] Input: Photos and episode data stored on the device
[0674] Output: Photos and episode data sent to the server
[0675] Step 3: Data analysis phase
[0676] How it works: The server analyzes the received photo and episode data. An artificial intelligence model (using, for example, TensorFlow) extracts basic information about the episode and identifies related themes (e.g., "cherry blossom viewing") based on the analysis results.
[0677] Input: Photos and episode data sent to the server
[0678] Output: Basic information about the analyzed episodes and identified themes
[0679] Step 4: Video Generation Phase
[0680] How it works: Based on the analysis results, the server uses an AI video generation module to automatically generate videos that incorporate emotional elements and specific locations, including footage of actual cherry blossom viewing and scenes of families having fun.
[0681] Input: Basic information about the analyzed episode and the identified themes
[0682] Output: Generated video data
[0683] Step 5: Video saving phase
[0684] Specific operation: The generated video data is stored in cloud storage. Before storing, the video data is compressed and encrypted as necessary.
[0685] Input: Generated video data
[0686] Output: Video data stored in cloud storage
[0687] Step 6: Playback Request Phase
[0688] Specific operation: A user requests playback of the generated video using a dedicated device or a smartphone application. The request is sent to the server.
[0689] Input: User's playback request
[0690] Output: Playback request sent to the server
[0691] Step 7: Video Distribution Phase
[0692] Specific operation: The server receives a playback request and delivers the video data stored in cloud storage to the device in streaming or download format.
[0693] Input: Playback request sent to the server, video data stored in cloud storage
[0694] Output: Video data delivered to the device
[0695] Step 8: Video viewing phase
[0696] Specific operation: The device plays the distributed video and allows the dementia patient to watch it. As the patient watches the video, past memories are stimulated.
[0697] Input: Video data delivered to the device
[0698] Output: Video currently being played by the dementia patient
[0699] Step 9: Audio Analysis Phase
[0700] Specific operation: The dementia patient begins a dialogue based on the content of the video. When a voice input such as "This day was really fun," the device analyzes the voice and converts it into text data.
[0701] Input: Voice input from dementia patients
[0702] Output: Speech input converted to text data
[0703] Step 10: Response generation phase
[0704] Specific operation: The server receives the converted text data and uses artificial intelligence techniques (e.g., a Transformer-based NLP model) to generate an appropriate response. The generated response is then sent from the server to the device.
[0705] Input: Voice input converted to text data
[0706] Output: The server-generated response
[0707] Step 11: Response delivery phase
[0708] Specific operation: The device plays back the response sent from the server and provides it to the dementia patient. Responses such as "It was really beautiful. Your grandmother was with you at the time" are played back.
[0709] Input: The response sent by the server
[0710] Output: Response voice heard by dementia patient
[0711] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0712] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and reminiscences and provides them to patients. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can provide an individual rehabilitation plan more effectively.
[0713] The program for this system mainly consists of the following phases:
[0714] Data Collection Phase
[0715] Users (people with dementia or their families) log in to their accounts through a dedicated device or application. They then scan or upload past photos and enter related memories as text or voice data. The device temporarily stores the collected photos and text data in its local storage and then sends them to the server.
[0716] Data Processing Phase
[0717] The server receives the photos and text data sent from the device. The received data is input into a natural language processing (NLP) module to extract basic information about the episode. Based on this information, an AI-based video generation module operates to automatically generate videos that include emotional elements and specific locations. The generated video data is then stored in cloud storage.
[0718] Content provision phase
[0719] The user requests playback of the generated video from a dedicated device or application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[0720] Dialogue Phase
[0721] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[0722] Emotion Recognition Phase
[0723] The emotion engine analyzes the user's voice input and facial expression data, determining their emotional state and generating emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[0724] Specific examples
[0725] For example, a family member of a dementia patient can enter "photos from cherry blossom viewing" and "anecdotes from that time" into the system, and the process will go as follows:
[0726] 1. The user (family member) logs in to the application, scans a photo, and enters the text "This photo was taken when my family went cherry blossom viewing."
[0727] 2. The device sends the photo and text data to the server.
[0728] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[0729] 4. The user (patient) requests video playback and watches the video on their device.
[0730] 5. The user (patient) says, "The cherry blossoms were beautiful at this time." The device analyzes the voice and sends it to the server.
[0731] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[0732] 7. The emotion engine analyzes the user's emotions and adjusts the next dialogue content based on the results. For example, if the user is happy, it will provide more stories related to happiness.
[0733] In this way, the combination of emotion recognition capabilities can more effectively advance the personalized rehabilitation plan for dementia patients, while providing responses that correspond to the user's emotional state makes interactions more natural and effective.
[0734] The processing flow will be explained below.
[0735] Step 1:
[0736] Users (dementia patients and their families) log in to their accounts through a dedicated device or application.
[0737] Step 2:
[0738] Users scan or upload old photos and enter related memories as text or audio data.
[0739] Step 3:
[0740] The device temporarily stores the entered photo and text data in local storage.
[0741] Step 4:
[0742] The device transmits the photos and text data stored in the local storage to the server.
[0743] Step 5:
[0744] The server receives the photo and text data sent from the terminal.
[0745] Step 6:
[0746] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode from the text data.
[0747] Step 7:
[0748] The server uses an AI video generation module to generate videos based on the extracted information, including emotional elements and specific locations.
[0749] Step 8:
[0750] The server stores the generated video data in cloud storage.
[0751] Step 9:
[0752] The user requests playback of the generated video from a dedicated terminal or application.
[0753] Step 10:
[0754] The server receives the user's request and delivers the video data to the terminal in streaming or download format.
[0755] Step 11:
[0756] The terminal plays the distributed video and allows the user to view it.
[0757] Step 12:
[0758] The user (a dementia patient) can start a conversation based on the content of the video, for example, saying, "This day was really fun."
[0759] Step 13:
[0760] The device analyzes the user's voice input and converts it into text data using voice recognition technology.
[0761] Step 14:
[0762] The terminal transmits the analyzed voice data (text) to the server.
[0763] Step 15:
[0764] The server generates a response using a natural language generation (NLG) module based on the received voice data.
[0765] Step 16:
[0766] The server transmits the generated response data to the terminal.
[0767] Step 17:
[0768] The terminal uses a voice synthesis function to reproduce the received response data and convey it to the user.
[0769] Step 18:
[0770] The device inputs the user's voice and facial expression data into an emotion engine to analyze the user's emotional state. For example, emotion tags such as "joy," "sadness," and "excitement" are generated.
[0771] Step 19:
[0772] The server automatically adjusts the content of the next video and dialogue responses based on the user's emotions recognized by the emotion engine. For example, if the user is happy, the server generates a video that includes more episodes related to happiness.
[0773] Step 20:
[0774] The user can receive continually adapted rehabilitation through updated content.
[0775] The above is the specific processing flow of the system of the invention combined with the emotion engine. In this way, by providing rehabilitation that corresponds to the user's emotions, it is possible to effectively stimulate the memory of dementia patients and realize more natural and effective dialogue.
[0776] Example 2
[0777] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0778] In rehabilitation for dementia patients, there is a need to provide effective stimulation that corresponds to the individual memories and emotions of each patient. Conventional rehabilitation methods have difficulty customizing to correspond to the memories and emotions of each patient, which often limits their effectiveness. For this reason, there is a need to develop a system that can automatically provide an individual rehabilitation plan that corresponds to the patient's memory ability and emotions.
[0779] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0780] In this invention, the server includes: means for a user to collect past photos and related memories; means for transmitting the collected data from the terminal to the server; means for the server to analyze the received data using a natural language processing module and extract basic information about the episode; means for generating a video using a generative AI model based on the extracted information; means for storing the generated video in cloud storage; means for a user to request video playback; means for the server to deliver video data to the terminal in response to the request; means for the terminal to play the video for the user; means for analyzing the user's voice input and converting it into text data; means for transmitting the converted voice data to the server and generating an appropriate response; means for transmitting the response data generated by the server to the terminal and playing it; means for analyzing the user's voice and facial expressions using an emotion engine to determine the user's emotional state; and means for adjusting the content to be provided next based on the emotional state. This makes it possible to automatically provide rehabilitation tailored to the individual memories and emotions of dementia patients.
[0781] "User" refers to people who use the system to input past photos and memories, or dementia patients and their families.
[0782] "Terminal" refers to the device through which a user accesses the system, uploads photos, and performs voice input. Specifically, this includes smartphones, tablets, and PCs.
[0783] "Server" refers to a computer system that receives, analyzes, stores, and manages data sent by users.
[0784] "Natural language processing module" refers to a software module for analyzing received text data and extracting basic information about an episode.
[0785] "Generative AI model" refers to an artificial intelligence algorithm that automatically generates videos based on extracted information.
[0786] "Cloud storage" refers to an online storage service for storing generated video data.
[0787] An "emotion engine" refers to a computer program that analyzes a user's voice and facial expressions to determine their emotional state and generate appropriate responses and content.
[0788] A "request" refers to an operation by a user to request playback of a video through a terminal.
[0789] "Video" refers to video content generated from past photos and related memories.
[0790] "Text data" refers to data that has been analyzed and converted into text information from a user's voice.
[0791] "Response" refers to an appropriate reply generated by a server in response to a user's input.
[0792] "Content" refers to videos, responses, and other information provided to users.
[0793] "Emotional state" refers to the emotional state such as joy, sadness, excitement, etc., analyzed from the user's voice and facial expressions.
[0794] "Rehabilitation" refers to treatment aimed at stimulating the memories and emotions of dementia patients and maintaining or improving their functions.
[0795] This invention is a system aimed at the rehabilitation of dementia patients, which generates videos based on past photos and reminiscences and provides them to patients. It also uses an emotion engine that recognizes the user's emotions to more effectively provide individual rehabilitation plans.
[0796] Overall flow
[0797] 1. Data Collection Phase
[0798] Users log in to the system using a dedicated device or application, then scan or upload old photos and enter related memories as text or voice data. The device temporarily saves this data in its local storage and then sends it to the server.
[0799] 2. Data Processing Phase
[0800] The server receives the photos and text data sent from the device. The received data is input into a natural language processing (NLP) module to extract basic information about the episode. Based on the extracted information, a generative AI model generates videos containing emotional elements. The generated videos are then stored in cloud storage.
[0801] 3. Content provision phase
[0802] The user requests video playback from a dedicated device or application. The server receives this request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[0803] 4. Dialogue Phase
[0804] The patient can start a dialogue based on the content of the video. For example, if the patient says, "I had a really fun day," the device analyzes the voice input and converts it into text data. The converted data is sent to the server, which then generates an appropriate response and plays it back through the device.
[0805] 5. Emotion Recognition Phase
[0806] The user's voice and facial expression data are analyzed by the emotion engine, which determines the user's emotional state and generates emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[0807] Hardware and software used
[0808] Devices: Smartphones, tablets, computers, etc.
[0809] Server: A high-performance computer that receives, analyzes, stores, and runs generative AI models.
[0810] Natural Language Processing Module: Software that performs text analysis (e.g., SpaCy, NLTK, etc.).
[0811] Generative AI models: Artificial intelligence algorithms that generate videos (e.g., GPT-3, DALL-E).
[0812] Emotion engine: A software module that performs emotion analysis (e.g., IBM Watson, Microsoft Azure Emotion API).
[0813] Cloud storage: An internet storage service for saving video data (e.g., Amazon S3, Google Cloud Storage).
[0814] Specific examples
[0815] For example, if a family member of a dementia patient enters "photos from a cherry blossom viewing party" and "an episode from that time" into the system, the process would be as follows:
[0816] 1. The user (family member) logs in to the application, enters the text "This photo was taken when my family went cherry blossom viewing," and scans the photo.
[0817] 2. The device temporarily stores the photo and text data and sends them to the server.
[0818] 3. The server receives the data and uses a natural language processing module to extract basic information about the episodes on the theme of "cherry blossom viewing." A generative AI model then generates a video based on that information and saves it in cloud storage.
[0819] 4. The user (patient) uses the application to request video playback and watches the video on their device.
[0820] 5. When the user (patient) says, "The cherry blossoms were beautiful at this time," the device captures the voice, converts it into text, and sends it to the server. The server generates an appropriate response (e.g., "They were really beautiful. Your grandmother was with you at the time") and plays it back on the device.
[0821] 6. The emotion engine analyzes the user's emotional state, and if they are happy, it will provide them with the next relevant episode.
[0822] Prompt Sentence Examples
[0823] An example of a prompt for a generative AI model is:
[0824] text
[0825] "Upload a photo of you and your family going cherry blossom viewing, enter the text 'We had a great time enjoying cherry blossom viewing in the park on this day,' and generate a two-minute video with the theme 'Cherry Blossom Park.'"
[0826] In this way, combining emotion recognition capabilities can effectively advance individual rehabilitation for dementia patients. Also, by providing responses that correspond to the user's emotional state, interactions become more natural and effective.
[0827] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0828] Step 1:
[0829] A user logs in to a dedicated terminal or application.
[0830] Specific operation: The user launches the application and logs in by entering their account information (username and password).
[0831] Input: Username, Password.
[0832] Output: The user is logged in.
[0833] Step 2:
[0834] Users scan or upload old photos and enter their memories.
[0835] What happens: The user uses the device's camera to scan a photo or upload an existing photo file, then uses the keyboard to enter text or the microphone to enter voice input.
[0836] Input: Photo files, text data, or audio data.
[0837] Output: Collected photos and text data.
[0838] Step 3:
[0839] The device temporarily stores the collected data in local storage and then transmits it to the server.
[0840] Specific operation: The device temporarily saves the photo and text data as files in local storage, and then sends the data to the server.
[0841] Input: Collected photo files and text data.
[0842] Output: The data sent to the server.
[0843] Step 4:
[0844] The server receives the data sent from the terminal.
[0845] Specific operation: The server receives the photos and text data sent from the device via the network and temporarily stores them in the server's storage.
[0846] Input: Photo files and text data sent from the device.
[0847] Output: Data stored on the server.
[0848] Step 5:
[0849] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode.
[0850] Specific operation: The server inputs the received text data into a natural language processing (NLP) module to extract basic information such as keywords, date and time, location, and names of people.
[0851] Input: Received text data.
[0852] Output: Basic information about the extracted episode.
[0853] Step 6:
[0854] The server generates a video using a generative AI model based on the extracted information.
[0855] How it works: The server inputs the extracted basic information into the generative AI model, builds a video scenario based on it, and then generates a video by combining related images and music.
[0856] Input: Basic information.
[0857] Output: The generated video data.
[0858] Step 7:
[0859] The server stores the generated video in cloud storage.
[0860] Specific operation: The server saves the generated video to a cloud storage service and generates a URL for accessing it.
[0861] Input: The generated video data.
[0862] Output: Videos stored in cloud storage, access URL.
[0863] Step 8:
[0864] The user requests video playback from a dedicated device or application.
[0865] Specific operation: The user logs in to the application again and selects the video they want to play from the generated video list page.
[0866] Input: Login information, video request.
[0867] Output: Video playback request.
[0868] Step 9:
[0869] The server distributes the video data to the terminal.
[0870] Specific operation: The server searches for video data in response to a user request and delivers it to the device in streaming or download format.
[0871] Input: A video playback request.
[0872] Output: Video data delivery to the device.
[0873] Step 10:
[0874] The terminal plays the video and allows the user to watch it.
[0875] Specific operation: The device decodes the received video data and plays it on the display. The user watches the video on the device.
[0876] Input: Received video data.
[0877] Output: The video shown on the display.
[0878] Step 11:
[0879] Users begin to interact based on the content of the video.
[0880] Specific operation: The user talks while watching the video (e.g., "The cherry blossoms were beautiful at this time."). The device captures the user's voice with a microphone.
[0881] Input: User's voice input.
[0882] Output: The captured audio data.
[0883] Step 12:
[0884] The device analyzes the voice, converts it into text data, and sends it to the server.
[0885] Specific operation: The device uses a speech recognition engine to convert the user's voice into text data and send it to the server.
[0886] Input: The captured audio data.
[0887] Output: The text data sent to the server.
[0888] Step 13:
[0889] The server generates an appropriate response based on the data it receives.
[0890] Specific operation: The server understands the context from the received text data and generates an appropriate response using a generative AI model.
[0891] Input: Received text data.
[0892] Output: The generated response data.
[0893] Step 14:
[0894] The terminal plays back the generated response.
[0895] Specific operation: The device converts the response data received from the server into speech using a speech synthesis engine and plays it back to the user.
[0896] Input: The response data received.
[0897] Output: The audio response played.
[0898] Step 15:
[0899] The emotion engine analyzes the user's voice and facial expression data to determine their emotional state.
[0900] How it works: The device captures the user's facial expressions with a camera and records their voice. The emotion engine analyzes this data to determine the user's emotional state.
[0901] Input: User's facial expression data, voice data.
[0902] Output: Emotion tag (e.g., happy, sad, excited).
[0903] Step 16:
[0904] Tailor your next piece of content based on your emotional state.
[0905] Specific operation: Based on the emotion tag, the server adjusts the next video and dialogue response according to the user's emotional state.
[0906] Input: emotion tag.
[0907] Output: A tailored content delivery plan.
[0908] (Application example 2)
[0909] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0910] Conventional rehabilitation systems provide the same content to all workers, including those with dementia and those unfamiliar with work processes, making it difficult to respond appropriately to individual needs and emotional states. Furthermore, there were insufficient means to properly reproduce and visually present past work procedures and episodes, making it difficult for workers to efficiently understand work procedures.
[0911] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting past photos and related memories, means for generating a video based on the collected data, means for providing the generated video to the target person, means for analyzing the user's voice input, means for generating a response based on the analyzed voice data, means for providing the generated response to the user, means for generating a video that reproduces a work procedure based on the collected data, and means for recognizing the user's emotional state and adjusting the content of the video and the response based on the emotional state. This enables appropriate responses according to individual needs and emotional states, and by providing a visual reproduction of past work procedures and episodes, the target person can efficiently understand the information.
[0912] "Past photographs and related reminiscences" are visual information related to the subject's past experiences or episodes and accompanying text or audio data.
[0913] "Collected data" is a collection of information provided by users, including photographs, text data, audio data, etc.
[0914] "Video generation means" refers to technology or devices that use collected data to create videos that convey information visually and audibly.
[0915] The "means for providing a video" refers to a method or device for allowing a target person to view the generated video.
[0916] "Means for analyzing user voice input" refers to technology that converts the user's voice into text data and understands and analyzes its content.
[0917] A "means for generating a response" is a technique or device that determines an appropriate response or next action based on the analyzed voice data.
[0918] "Means for generating videos that reproduce work procedures based on collected data" refers to technology or devices that visually reproduce past work procedures or episodes and provide them as videos in an easy-to-understand format.
[0919] "Means for recognizing the user's emotional state" refers to technology or devices that analyze the user's voice, facial expressions, behavior, etc. to determine and identify their emotional state.
[0920] The "means for adjusting the response content" refers to a technique or device for appropriately changing the video or response content provided depending on the user's emotional state or situation.
[0921] This invention is a rehabilitation system for dementia patients and workers, which aims to generate videos based on past photos and reminiscences and provide them to users. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides an individual rehabilitation plan.
[0922] Program processing explanation
[0923] Data Collection Phase
[0924] Terminal
[0925] Users log in to their account using a dedicated device (smartphone or head-mounted display) and upload past photos and related memories. This data is entered as text or voice data, temporarily saved in local storage, and then sent to the server.
[0926] Data Processing Phase
[0927] server
[0928] The server receives the photos and text data sent from the device. The received data is analyzed using a natural language processing (NLP) module (e.g., SpaCy or Google NLP API) to extract basic information about the episode. Based on this information, an AI-based video generation module (e.g., OpenAI or DeepMind) runs and automatically generates videos based on work procedures and reminiscences. The generated video data is stored in cloud storage (e.g., Google Cloud Storage or AWS S3).
[0929] Content provision phase
[0930] Terminal
[0931] The user requests playback of the generated video from a dedicated terminal. The server receives the request and delivers the video data to the terminal via streaming or download. The terminal then plays the received video and allows the user to watch it.
[0932] Dialogue Phase
[0933] Server and Device
[0934] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[0935] Emotion Recognition Phase
[0936] server
[0937] The user's voice input and facial expression data are analyzed by an emotion engine (e.g., Microsoft Azure Cognitive Services or Amazon Rekognition). The emotion engine determines the user's emotional state and generates emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[0938] Specific examples
[0939] For example, the following flow would be used for a rehabilitation system that supports operations within a factory.
[0940] 1. A worker uses a terminal to upload photos and procedures related to past "factory line maintenance work."
[0941] 2. The server receives the data and uses a video generation AI model to generate a video that reproduces the "factory line maintenance work procedures."
[0942] 3. The worker requests "Maintenance procedure playback" and watches the video.
[0943] 4. The worker asks, "What should I pay attention to in this part?" The device analyzes the voice and sends it to the server.
[0944] 5. The server generates a response saying, "It is important to check the safety devices in this section. In particular, make sure this switch is working," and plays it back through the terminal.
[0945] Prompt Sentence Examples
[0946] Generate a video that recreates a past maintenance procedure on a factory line based on the following data:
[0947] Photo: Maintenance work on a factory line
[0948] Procedure: 1. Turn off the power. 2. Check the safety devices. 3. Clean the conveyor belt. 4. Test the operation.
[0949] Your video should include the following elements:
[0950] 1. Detailed explanation of each step
[0951] 2. Points to pay particular attention to
[0952] 3. Clear visual and audio guidance for easy understanding by workers
[0953] Also, generate appropriate responses according to the worker's emotions.
[0954] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0955] Step 1: Data collection
[0956] The user logs into their account using a dedicated device (smartphone or head-mounted display) and uploads past photos and related memories. The input is photo data and text or voice data, which are temporarily stored in local storage. The output is data that is saved in local storage and then sent to the server. The specific actions that take place in this step are scanning photos, inputting text and voice data, saving the data in local storage, and sending it to the server.
[0957] Step 2: Data reception and analysis
[0958] The server receives the photo and text data sent from the device. The input is the data sent from the device, and the output is the data stored on the server. This data is input into a natural language processing (NLP) module, which extracts basic information about the episode. The specific operations performed here are receiving the data, storing it, and analyzing it using the NLP module.
[0959] Step 3: Video Generation
[0960] The AI video generation module in the server generates videos based on the analyzed episode information. The input is the data analyzed by the NLP module, and the output is the generated video data. The specific operations are as follows: input data into the AI model, generate video, and save it to cloud storage.
[0961] Step 4: Submit your video
[0962] A user requests playback of a generated video from a dedicated device. The input is a video playback request, and the output is video data that is streamed or downloaded to the user. Specific operations include accepting the request, obtaining the video data, and delivering it to the device.
[0963] Step 5: Dialogue analysis
[0964] The device receives and analyzes the audio emitted by the user while watching a video. The input is audio data, and the output is the analysis results converted into text data. The specific operations are to collect audio, analyze the audio, convert it into text data, and send it to the server.
[0965] Step 6: Response Generation
[0966] The server generates an appropriate response based on the received voice data. The input is the analyzed voice data, and the output is the generated response data. The specific operations are analyzing the voice data, determining the response content, and generating the response.
[0967] Step 7: Provide a response
[0968] The generated response is provided to the user through the terminal. The input is the generated response data, and the output is the response voice heard by the user. The specific operations are receiving the response data, generating the voice, and delivering it to the user.
[0969] Step 8: Emotion Recognition
[0970] The emotion engine analyzes the user's voice input and facial expression data. The input is the user's voice and facial expression data, and the output is an emotion tag. Specific operations include collecting voice and facial expression data, analyzing emotions, and generating emotion tags.
[0971] Step 9: Content Adjustment
[0972] The next video or response content is adjusted based on the emotion tag. The input is the emotion tag and the next content data to be provided, and the output is the adjusted video or response content. The specific operations are emotion tag analysis, content adjustment, and next content generation.
[0973] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0974] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0975] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0976] [Third embodiment]
[0977] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0978] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0979] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0980] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0981] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0982] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0983] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0984] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0985] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0986] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0987] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0988] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0989] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and reminiscences and provides them to patients. The system uses artificial intelligence technology to create individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[0990] The program for this system mainly consists of the following phases:
[0991] Data Collection Phase
[0992] Users (dementia patients and their families) input past photos and episodes through a dedicated device or application. Photos are collected by scanning or uploading, and episodes are input as text or voice data. The device temporarily stores the collected data and sends it to a server.
[0993] Data Processing Phase
[0994] The server receives the photos and text data sent from the device. The received data is analyzed by an artificial intelligence model to extract basic information about the episode. Based on this information, an AI video generation module operates and automatically generates videos that include emotional elements and specific locations. The generated video data is stored in cloud storage.
[0995] Content provision phase
[0996] The user requests playback of the generated video from a dedicated device or application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[0997] Dialogue Phase
[0998] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[0999] Specific examples
[1000] For example, if a family member of a dementia patient inputs a photo of the family cherry blossom viewing and an episode from that time into the system, the process would be as follows:
[1001] 1. The user (family member) logs in to the application, scans a photo, and enters the text "This photo was taken when my family went cherry blossom viewing."
[1002] 2. The device sends the photo and text data to the server.
[1003] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[1004] 4. The user (patient) requests video playback and watches the video on their device.
[1005] 5. The user (patient) says, "The cherry blossoms were beautiful at this time." The device analyzes the voice and sends it to the server.
[1006] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[1007] In this way, the system can stimulate past memories and promote brain activity in dementia patients. Furthermore, by utilizing artificial intelligence, it can efficiently provide individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[1008] The processing flow will be explained below.
[1009] Step 1:
[1010] Users (dementia patients and their families) log in to their accounts through a dedicated device or application.
[1011] Step 2:
[1012] Users scan or upload old photos and enter related memories as text or audio data.
[1013] Step 3:
[1014] The device temporarily stores the entered photo and text data in local storage.
[1015] Step 4:
[1016] The device transmits the photos and text data stored in the local storage to the server.
[1017] Step 5:
[1018] The server receives the photo and text data sent from the terminal.
[1019] Step 6:
[1020] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode from the text data.
[1021] Step 7:
[1022] The server uses an AI video generation module to generate videos based on the extracted information, including emotional elements and specific locations.
[1023] Step 8:
[1024] The server stores the generated video data in cloud storage.
[1025] Step 9:
[1026] The user requests playback of the generated video from a dedicated terminal or application.
[1027] Step 10:
[1028] The server receives the user's request and delivers the video data to the terminal in streaming or download format.
[1029] Step 11:
[1030] The terminal plays the distributed video and allows the user to view it.
[1031] Step 12:
[1032] The user (a dementia patient) can start a conversation based on the content of the video, for example, saying, "This day was really fun."
[1033] Step 13:
[1034] The device analyzes the user's voice input and converts it into text data using voice recognition technology.
[1035] Step 14:
[1036] The terminal transmits the analyzed voice data (text) to the server.
[1037] Step 15:
[1038] The server generates a response using a natural language generation (NLG) module based on the received voice data.
[1039] Step 16:
[1040] The server transmits the generated response data to the terminal.
[1041] Step 17:
[1042] The terminal uses a voice synthesis function to reproduce the received response data and convey it to the user.
[1043] Example 1
[1044] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1045] In the rehabilitation of dementia patients, it is important to stimulate individual memories and elicit emotional responses. However, rehabilitation using past photographs and anecdotes can be laborious for some caregivers and medical professionals. Furthermore, these tasks tend to rely on subjective judgment and are difficult to perform consistently. This can result in the insufficient effectiveness of dementia rehabilitation. Therefore, there is a need for an efficient and individually tailored rehabilitation system.
[1046] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1047] In this invention, the server includes: a means for a user to input past photos and episodes; a means for transmitting the input photos and episodes to the server; a means for the server to analyze the received data using an artificial intelligence model; a means for automatically generating videos based on the analysis results; a means for storing the generated video data in cloud storage; a means for a user to request video playback; a means for the server to receive the request and deliver the video data to a terminal; a means for the terminal to play the video; a means for analyzing the user's voice input and converting it into text data; a means for the server to generate a response based on the voice input; and a means for transmitting the generated response to the terminal and playing it as audio. This allows a user to input past photos and episodes and efficiently stimulate the memory of a dementia patient using the generated video. Furthermore, the use of artificial intelligence can provide a consistent rehabilitation plan, reducing the burden on caregivers and medical professionals.
[1048] "User" refers to the person who operates the system and inputs data, such as a dementia patient, their family member, or caregiver.
[1049] A "terminal" is an electronic device that allows a user to input data, receive data from a server, and execute processing.
[1050] A "server" is a central processing system that receives data sent from users and performs processes such as data analysis, video generation, and response generation.
[1051] "Artificial intelligence model" refers to machine learning algorithms and techniques used for complex information processing such as data analysis, video generation, and response generation.
[1052] "Cloud storage" refers to internet-based data storage servers that remotely store generated video data and other data and make it accessible as needed.
[1053] An "episode" refers to a specific memory or event related to a photo entered by the user.
[1054] A "request" refers to an action in which a user requests a specific operation or process from the system.
[1055] "Voice input" refers to voice data input by a user speaking to the system using a microphone or the like.
[1056] "Text data" refers to data in the form of a string of characters obtained by analyzing voice input.
[1057] "Video generation" refers to the process of creating visual and audio content based on photographs and anecdotes.
[1058] "Response generation" refers to the process of generating an appropriate reaction or reply to a user's voice input.
[1059] "Analysis" refers to the processing of collected data to extract useful information and to perform subsequent processing based on that information.
[1060] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and episodes and provides them to patients. The system uses artificial intelligence technology to analyze data entered by users (dementia patients, their families, or caregivers) and create individual rehabilitation plans.
[1061] First, a user uses a dedicated device or application to input past photos and related episodes. For example, the user scans a photo of a family cherry blossom viewing and enters a text description such as "This photo was taken when my family went cherry blossom viewing." This data is temporarily stored by the device and then sent to the server.
[1062] The server receives the photos and text data sent from the device. The received data is analyzed using artificial intelligence models such as TensorFlow and OpenAI's GPT. The analysis extracts basic information about the episode and automatically generates a video that includes emotional elements and a specific location (in this case, cherry blossom viewing). The generated video data is stored in cloud storage such as Amazon S3.
[1063] The user requests playback of the generated video from a dedicated device or application. The server receives the request, retrieves the video data stored in cloud storage, and delivers it to the device via streaming or download. The device plays the received video, and the user (dementia patient) watches it.
[1064] A dementia patient can initiate a dialogue based on the content of the video. For example, they can say, "The cherry blossoms were beautiful at this time." The device analyzes this voice input and converts it into text data. The speech is converted to text using speech recognition technology such as the Google Speech-to-Text API. The converted text data is sent to a server, which uses an artificial intelligence model such as OpenAI's GPT-3 to generate an appropriate response. This response is sent to the device and played back using a text-to-speech engine that converts text to speech.
[1065] A specific example is shown below. For example, if a user inputs "a photo of a family cherry blossom viewing party" and an episode from that event, the system operates as follows.
[1066] 1. The user uses the application to scan a photo of a family cherry blossom viewing party and enters the message, "This photo was taken when my family went cherry blossom viewing."
[1067] 2. The device temporarily stores the photo and text data and sends them to the server.
[1068] 3. The server receives the data and uses an AI model (e.g., TensorFlow) to generate a "cherry blossom viewing"-themed video, which includes footage of cherry blossom viewing and scenes of families having fun.
[1069] 4. The user (patient) requests video playback and watches the video on their device.
[1070] 5. The user (patient) says, "The cherry blossoms were beautiful at this time."
[1071] 6. The device analyzes the voice, converts it into text data, and sends it to the server.
[1072] 7. The server generates a response saying, "It was really beautiful. Your grandma was with you at the time," and plays it back through the device.
[1073] An example prompt is, "Based on photos and stories from this family's cherry blossom viewing, create a video that includes emotional elements and specific locations."
[1074] In this way, the system can stimulate past memories and promote brain activity in dementia patients, and by utilizing artificial intelligence, it can efficiently provide individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[1075] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1076] Step 1:
[1077] User data entry and saving
[1078] The user inputs past photos and episodes using a dedicated device or application. For example, the user scans a photo of a family cherry blossom viewing party and enters a text description such as "This photo is from when we went cherry blossom viewing with my family." The input data (photo and text data) is temporarily saved by the device. The input data is the photo file and text data, and the output is the temporarily saved data.
[1079] Step 2:
[1080] Sending data
[1081] The device sends the saved photos and text data to the server. Specifically, the device sends the data by making an HTTP POST request to the server. The input is the temporarily saved photos and text data, and the output is the data sent to the server.
[1082] Step 3:
[1083] Data reception and analysis
[1084] The server receives the photos and text data sent from the device. The received data is input into an artificial intelligence model such as Google's TensorFlow or OpenAI's GPT, which extracts basic information about the episode from the text data and matches it with the photos to analyze its relevance. The input is the photos and text data sent to the server, and the output is the analyzed basic information about the episode.
[1085] Step 4:
[1086] Video generation
[1087] The server runs an AI video generation module based on the analysis results. It automatically generates videos that include emotional elements or specific locations (e.g., cherry blossom viewing). Specifically, the AI model selects video material based on the analysis results, and the video is edited using an editor. The input is basic information about the analyzed episode, and the output is the generated video data.
[1088] Step 5:
[1089] Saving videos
[1090] The server saves the generated video data in cloud storage (e.g., Amazon S3). It uploads the video data to the cloud using an API. The input is the generated video data, and the output is the video data saved in the cloud storage.
[1091] Step 6:
[1092] Video Request
[1093] A user requests video playback from a dedicated device or application. Specifically, the user presses the play button on the device to send a playback request to the server. The input is the user's playback request, and the output is a playback request to the server.
[1094] Step 7:
[1095] Video data acquisition and distribution
[1096] The server receives a request from a user and retrieves video data from cloud storage. The retrieved video data is then delivered to the device via streaming or download. The input is the video data retrieved from cloud storage, and the output is the video data delivered to the device.
[1097] Step 8:
[1098] Play video
[1099] The device plays the received video data. The video is displayed using a dedicated playback application or a standard video player. The input is the video data delivered to the device, and the output is the played video.
[1100] Step 9:
[1101] Analyzing voice input
[1102] The patient begins a conversation based on the content of the video. For example, they might say, "The cherry blossoms were beautiful at this time." The device analyzes this voice input and converts it into text data using the Google Speech-to-Text API or similar. The input is the patient's voice, and the output is text data.
[1103] Step 10:
[1104] Generate and send the response
[1105] The device sends the converted text data to the server. The server generates an appropriate response based on the received text data using an artificial intelligence model such as OpenAI's GPT. The generated response is sent to the device in text format. The input is text data, and the output is the generated response.
[1106] Step 11:
[1107] Response playback
[1108] The device plays the response sent from the server as audio. It uses a text-to-speech engine to convert the response into audio and plays it back. The input is the generated response text data, and the output is the played audio response.
[1109] (Application example 1)
[1110] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1111] Rehabilitation is important for dementia patients to stimulate their memories and improve their cognitive function. However, effective rehabilitation requires the creation of individual rehabilitation plans, which places a heavy burden on caregivers and medical professionals. Furthermore, there is a problem with brick-and-mortar rehabilitation support, where there is no effective way to provide videos and dialogue based on past memories.
[1112] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1113] In this invention, the server includes a means for collecting past photos and related episodes, a means for generating videos based on the collected data, a means for providing the generated videos to the dementia patient, a means for analyzing a user's voice input, a means for generating responses based on the analyzed voice data, and a means for providing videos and interactive dialogues based on past memories in a physical store, thereby stimulating the memory of the dementia patient and effectively implementing rehabilitation, while reducing the burden on caregivers and medical professionals.
[1114] "Past photos and related episodes" is information provided by a patient with dementia or their family that indicates images taken in the past and memories or events related to the images.
[1115] "Means of collection" refers to the methods and technologies used to collect photos and stories via devices such as tablets and smartphones.
[1116] "Means for generating video" refers to methods and algorithms for creating video using artificial intelligence technology based on collected photos and episodes.
[1117] "Means of delivery to dementia patients" refers to methods of providing the generated videos and other rehabilitation content to dementia patients via devices such as smartphones and tablets.
[1118] "Means for analyzing user voice input" means techniques for collecting voice data using a microphone or other voice capture device and analyzing the voice data.
[1119] "Means for generating a response based on analyzed voice data" refers to a method or algorithm for generating an appropriate reply based on the analyzed user's voice input.
[1120] "Means for providing videos and interactive dialogues based on past memories in physical locations" means technologies and methods for providing videos based on past photos and episodes and related dialogues in physical locations such as cafes and rehabilitation centers.
[1121] This invention is a rehabilitation system for dementia patients, and aims to provide videos and interactive dialogues based on past photos and episodes in a physical store. This system consistently performs everything from collecting photos and episodes to processing the data, providing content, and generating dialogues.
[1122] The system that realizes this application example is made up of a number of hardware and software components, each of which will be described in detail below.
[1123] 1. Data Collection Phase
[1124] Users (dementia patients and their families) enter past photos and episodes through dedicated terminals located in physical stores or smartphone applications. Photos are collected by scanning or uploading, and episodes are entered as text or voice data. The terminal temporarily stores this data and later sends it to a server.
[1125] 2. Data Processing Phase
[1126] The server receives the photos and text data sent from the device. The received data is analyzed by an artificial intelligence model (using TensorFlow, for example) to extract basic information about the episode. Based on this information, an AI video generation module operates and automatically generates videos that include emotional elements and specific locations. The generated video data is stored in cloud storage.
[1127] 3. Content provision phase
[1128] The user then requests playback of the generated video again from a dedicated device or smartphone application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[1129] 4. Dialogue Phase
[1130] The patient with dementia initiates a dialogue based on the content of the video. For example, they might say, "I really enjoyed this day." The device analyzes the voice input and converts it into text data (for example, using a Transformer-based NLP model). The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[1131] Specific examples
[1132] For example, if a family member of a dementia patient inputs a photo of the family cherry blossom viewing and an episode from that time into the system, the process would be as follows:
[1133] 1. The user (family member) logs in to a terminal at a physical store, scans a photo, and enters a text story such as, "This photo was taken when the family went cherry blossom viewing."
[1134] 2. The device sends the photo and text data to the server.
[1135] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[1136] 4. The user (patient) requests video playback and watches the video on a tablet in the store.
[1137] 5. When the user (patient) says, "The cherry blossoms were beautiful at this time," the device analyzes the voice and sends it to the server.
[1138] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[1139] Prompt Sentence Examples
[1140] "If you look at a photo of a family cherry blossom viewing and say the cherry blossoms were beautiful, how will the AI respond?"
[1141] "What kind of comments would you have the AI make when shown photos of past summer festivals?"
[1142] In this way, the system can stimulate past memories in dementia patients, enabling effective rehabilitation while reducing the burden on caregivers and medical professionals.
[1143] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1144] Step 1: Data collection phase
[1145] Specific operation: Users (dementia patients and their families) input past photos and episodes using a dedicated terminal installed in a physical store or a smartphone application. Users scan or upload photos and enter an episode in text, such as "This photo was taken when my family went cherry blossom viewing."
[1146] Input: scanned or uploaded photos, text or audio episodes
[1147] Output: Photos and episode data temporarily saved on the device
[1148] Step 2: Data transmission phase
[1149] Specific operation: The device sends the temporarily stored photos and episode data to the server. During this process, the data is compressed and encrypted as necessary.
[1150] Input: Photos and episode data stored on the device
[1151] Output: Photos and episode data sent to the server
[1152] Step 3: Data analysis phase
[1153] How it works: The server analyzes the received photo and episode data. An artificial intelligence model (using, for example, TensorFlow) extracts basic information about the episode and identifies related themes (e.g., "cherry blossom viewing") based on the analysis results.
[1154] Input: Photos and episode data sent to the server
[1155] Output: Basic information about the analyzed episodes and identified themes
[1156] Step 4: Video Generation Phase
[1157] How it works: Based on the analysis results, the server uses an AI video generation module to automatically generate videos that incorporate emotional elements and specific locations, including footage of actual cherry blossom viewing and scenes of families having fun.
[1158] Input: Basic information about the analyzed episode and the identified themes
[1159] Output: Generated video data
[1160] Step 5: Video saving phase
[1161] Specific operation: The generated video data is stored in cloud storage. Before storing, the video data is compressed and encrypted as necessary.
[1162] Input: Generated video data
[1163] Output: Video data stored in cloud storage
[1164] Step 6: Playback Request Phase
[1165] Specific operation: A user requests playback of the generated video using a dedicated device or a smartphone application. The request is sent to the server.
[1166] Input: User's playback request
[1167] Output: Playback request sent to the server
[1168] Step 7: Video Distribution Phase
[1169] Specific operation: The server receives a playback request and delivers the video data stored in cloud storage to the device in streaming or download format.
[1170] Input: Playback request sent to the server, video data stored in cloud storage
[1171] Output: Video data delivered to the device
[1172] Step 8: Video viewing phase
[1173] Specific operation: The device plays the distributed video and allows the dementia patient to watch it. As the patient watches the video, past memories are stimulated.
[1174] Input: Video data delivered to the device
[1175] Output: Video currently being played by the dementia patient
[1176] Step 9: Audio Analysis Phase
[1177] Specific operation: The dementia patient begins a dialogue based on the content of the video. When a voice input such as "This day was really fun," the device analyzes the voice and converts it into text data.
[1178] Input: Voice input from dementia patients
[1179] Output: Speech input converted to text data
[1180] Step 10: Response generation phase
[1181] Specific operation: The server receives the converted text data and uses artificial intelligence techniques (e.g., a Transformer-based NLP model) to generate an appropriate response. The generated response is then sent from the server to the device.
[1182] Input: Voice input converted to text data
[1183] Output: The server-generated response
[1184] Step 11: Response delivery phase
[1185] Specific operation: The device plays back the response sent from the server and provides it to the dementia patient. Responses such as "It was really beautiful. Your grandmother was with you at the time" are played back.
[1186] Input: The response sent by the server
[1187] Output: Response voice heard by dementia patient
[1188] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1189] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and reminiscences and provides them to patients. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can provide an individual rehabilitation plan more effectively.
[1190] The program for this system mainly consists of the following phases:
[1191] Data Collection Phase
[1192] Users (people with dementia or their families) log in to their accounts through a dedicated device or application. They then scan or upload past photos and enter related memories as text or voice data. The device temporarily stores the collected photos and text data in its local storage and then sends them to the server.
[1193] Data Processing Phase
[1194] The server receives the photos and text data sent from the device. The received data is input into a natural language processing (NLP) module to extract basic information about the episode. Based on this information, an AI-based video generation module operates to automatically generate videos that include emotional elements and specific locations. The generated video data is then stored in cloud storage.
[1195] Content provision phase
[1196] The user requests playback of the generated video from a dedicated device or application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[1197] Dialogue Phase
[1198] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[1199] Emotion Recognition Phase
[1200] The emotion engine analyzes the user's voice input and facial expression data, determining their emotional state and generating emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[1201] Specific examples
[1202] For example, a family member of a dementia patient can enter "photos from cherry blossom viewing" and "anecdotes from that time" into the system, and the process will go as follows:
[1203] 1. The user (family member) logs in to the application, scans a photo, and enters the text "This photo was taken when my family went cherry blossom viewing."
[1204] 2. The device sends the photo and text data to the server.
[1205] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[1206] 4. The user (patient) requests video playback and watches the video on their device.
[1207] 5. The user (patient) says, "The cherry blossoms were beautiful at this time." The device analyzes the voice and sends it to the server.
[1208] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[1209] 7. The emotion engine analyzes the user's emotions and adjusts the next dialogue content based on the results. For example, if the user is happy, it will provide more stories related to happiness.
[1210] In this way, the combination of emotion recognition capabilities can more effectively advance the personalized rehabilitation plan for dementia patients, while providing responses that correspond to the user's emotional state makes interactions more natural and effective.
[1211] The processing flow will be explained below.
[1212] Step 1:
[1213] Users (dementia patients and their families) log in to their accounts through a dedicated device or application.
[1214] Step 2:
[1215] Users scan or upload old photos and enter related memories as text or audio data.
[1216] Step 3:
[1217] The device temporarily stores the entered photo and text data in local storage.
[1218] Step 4:
[1219] The device transmits the photos and text data stored in the local storage to the server.
[1220] Step 5:
[1221] The server receives the photo and text data sent from the terminal.
[1222] Step 6:
[1223] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode from the text data.
[1224] Step 7:
[1225] The server uses an AI video generation module to generate videos based on the extracted information, including emotional elements and specific locations.
[1226] Step 8:
[1227] The server stores the generated video data in cloud storage.
[1228] Step 9:
[1229] The user requests playback of the generated video from a dedicated terminal or application.
[1230] Step 10:
[1231] The server receives the user's request and delivers the video data to the terminal in streaming or download format.
[1232] Step 11:
[1233] The terminal plays the distributed video and allows the user to view it.
[1234] Step 12:
[1235] The user (a dementia patient) can start a conversation based on the content of the video, for example, saying, "This day was really fun."
[1236] Step 13:
[1237] The device analyzes the user's voice input and converts it into text data using voice recognition technology.
[1238] Step 14:
[1239] The terminal transmits the analyzed voice data (text) to the server.
[1240] Step 15:
[1241] The server generates a response using a natural language generation (NLG) module based on the received voice data.
[1242] Step 16:
[1243] The server transmits the generated response data to the terminal.
[1244] Step 17:
[1245] The terminal uses a voice synthesis function to reproduce the received response data and convey it to the user.
[1246] Step 18:
[1247] The device inputs the user's voice and facial expression data into an emotion engine to analyze the user's emotional state. For example, emotion tags such as "joy," "sadness," and "excitement" are generated.
[1248] Step 19:
[1249] The server automatically adjusts the content of the next video and dialogue responses based on the user's emotions recognized by the emotion engine. For example, if the user is happy, the server generates a video that includes more episodes related to happiness.
[1250] Step 20:
[1251] The user can receive continually adapted rehabilitation through updated content.
[1252] The above is the specific processing flow of the system of the invention combined with the emotion engine. In this way, by providing rehabilitation that corresponds to the user's emotions, it is possible to effectively stimulate the memory of dementia patients and realize more natural and effective dialogue.
[1253] Example 2
[1254] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1255] In rehabilitation for dementia patients, there is a need to provide effective stimulation that corresponds to the individual memories and emotions of each patient. Conventional rehabilitation methods have difficulty customizing to correspond to the memories and emotions of each patient, which often limits their effectiveness. For this reason, there is a need to develop a system that can automatically provide an individual rehabilitation plan that corresponds to the patient's memory ability and emotions.
[1256] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1257] In this invention, the server includes: means for a user to collect past photos and related memories; means for transmitting the collected data from the terminal to the server; means for the server to analyze the received data using a natural language processing module and extract basic information about the episode; means for generating a video using a generative AI model based on the extracted information; means for storing the generated video in cloud storage; means for a user to request video playback; means for the server to deliver video data to the terminal in response to the request; means for the terminal to play the video for the user; means for analyzing the user's voice input and converting it into text data; means for transmitting the converted voice data to the server and generating an appropriate response; means for transmitting the response data generated by the server to the terminal and playing it; means for analyzing the user's voice and facial expressions using an emotion engine to determine the user's emotional state; and means for adjusting the content to be provided next based on the emotional state. This makes it possible to automatically provide rehabilitation tailored to the individual memories and emotions of dementia patients.
[1258] "User" refers to people who use the system to input past photos and memories, or dementia patients and their families.
[1259] "Terminal" refers to the device through which a user accesses the system, uploads photos, and performs voice input. Specifically, this includes smartphones, tablets, and PCs.
[1260] "Server" refers to a computer system that receives, analyzes, stores, and manages data sent by users.
[1261] "Natural language processing module" refers to a software module for analyzing received text data and extracting basic information about an episode.
[1262] "Generative AI model" refers to an artificial intelligence algorithm that automatically generates videos based on extracted information.
[1263] "Cloud storage" refers to an online storage service for storing generated video data.
[1264] An "emotion engine" refers to a computer program that analyzes a user's voice and facial expressions to determine their emotional state and generate appropriate responses and content.
[1265] A "request" refers to an operation by a user to request playback of a video through a terminal.
[1266] "Video" refers to video content generated from past photos and related memories.
[1267] "Text data" refers to data that has been analyzed and converted into text information from a user's voice.
[1268] "Response" refers to an appropriate reply generated by a server in response to a user's input.
[1269] "Content" refers to videos, responses, and other information provided to users.
[1270] "Emotional state" refers to the emotional state such as joy, sadness, excitement, etc., analyzed from the user's voice and facial expressions.
[1271] "Rehabilitation" refers to treatment aimed at stimulating the memories and emotions of dementia patients and maintaining or improving their functions.
[1272] This invention is a system aimed at the rehabilitation of dementia patients, which generates videos based on past photos and reminiscences and provides them to patients. It also uses an emotion engine that recognizes the user's emotions to more effectively provide individual rehabilitation plans.
[1273] Overall flow
[1274] 1. Data Collection Phase
[1275] Users log in to the system using a dedicated device or application, then scan or upload old photos and enter related memories as text or voice data. The device temporarily saves this data in its local storage and then sends it to the server.
[1276] 2. Data Processing Phase
[1277] The server receives the photos and text data sent from the device. The received data is input into a natural language processing (NLP) module to extract basic information about the episode. Based on the extracted information, a generative AI model generates videos containing emotional elements. The generated videos are then stored in cloud storage.
[1278] 3. Content provision phase
[1279] The user requests video playback from a dedicated device or application. The server receives this request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[1280] 4. Dialogue Phase
[1281] The patient can start a dialogue based on the content of the video. For example, if the patient says, "I had a really fun day," the device analyzes the voice input and converts it into text data. The converted data is sent to the server, which then generates an appropriate response and plays it back through the device.
[1282] 5. Emotion Recognition Phase
[1283] The user's voice and facial expression data are analyzed by the emotion engine, which determines the user's emotional state and generates emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[1284] Hardware and software used
[1285] Devices: Smartphones, tablets, computers, etc.
[1286] Server: A high-performance computer that receives, analyzes, stores, and runs generative AI models.
[1287] Natural Language Processing Module: Software that performs text analysis (e.g., SpaCy, NLTK, etc.).
[1288] Generative AI models: Artificial intelligence algorithms that generate videos (e.g., GPT-3, DALL-E).
[1289] Emotion engine: A software module that performs emotion analysis (e.g., IBM Watson, Microsoft Azure Emotion API).
[1290] Cloud storage: An internet storage service for saving video data (e.g., Amazon S3, Google Cloud Storage).
[1291] Specific examples
[1292] For example, if a family member of a dementia patient enters "photos from a cherry blossom viewing party" and "an episode from that time" into the system, the process would be as follows:
[1293] 1. The user (family member) logs in to the application, enters the text "This photo was taken when my family went cherry blossom viewing," and scans the photo.
[1294] 2. The device temporarily stores the photo and text data and sends them to the server.
[1295] 3. The server receives the data and uses a natural language processing module to extract basic information about the episodes on the theme of "cherry blossom viewing." A generative AI model then generates a video based on that information and saves it in cloud storage.
[1296] 4. The user (patient) uses the application to request video playback and watches the video on their device.
[1297] 5. When the user (patient) says, "The cherry blossoms were beautiful at this time," the device captures the voice, converts it into text, and sends it to the server. The server generates an appropriate response (e.g., "They were really beautiful. Your grandmother was with you at the time") and plays it back on the device.
[1298] 6. The emotion engine analyzes the user's emotional state, and if they are happy, it will provide them with the next relevant episode.
[1299] Prompt Sentence Examples
[1300] An example of a prompt for a generative AI model is:
[1301] text
[1302] "Upload a photo of you and your family going cherry blossom viewing, enter the text 'We had a great time enjoying cherry blossom viewing in the park on this day,' and generate a two-minute video with the theme 'Cherry Blossom Park.'"
[1303] In this way, combining emotion recognition capabilities can effectively advance individual rehabilitation for dementia patients. Also, by providing responses that correspond to the user's emotional state, interactions become more natural and effective.
[1304] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1305] Step 1:
[1306] A user logs in to a dedicated terminal or application.
[1307] Specific operation: The user launches the application and logs in by entering their account information (username and password).
[1308] Input: Username, Password.
[1309] Output: The user is logged in.
[1310] Step 2:
[1311] Users scan or upload old photos and enter their memories.
[1312] What happens: The user uses the device's camera to scan a photo or upload an existing photo file, then uses the keyboard to enter text or the microphone to enter voice input.
[1313] Input: Photo files, text data, or audio data.
[1314] Output: Collected photos and text data.
[1315] Step 3:
[1316] The device temporarily stores the collected data in local storage and then transmits it to the server.
[1317] Specific operation: The device temporarily saves the photo and text data as files in local storage, and then sends the data to the server.
[1318] Input: Collected photo files and text data.
[1319] Output: The data sent to the server.
[1320] Step 4:
[1321] The server receives the data sent from the terminal.
[1322] Specific operation: The server receives the photos and text data sent from the device via the network and temporarily stores them in the server's storage.
[1323] Input: Photo files and text data sent from the device.
[1324] Output: Data stored on the server.
[1325] Step 5:
[1326] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode.
[1327] Specific operation: The server inputs the received text data into a natural language processing (NLP) module to extract basic information such as keywords, date and time, location, and names of people.
[1328] Input: Received text data.
[1329] Output: Basic information about the extracted episode.
[1330] Step 6:
[1331] The server generates a video using a generative AI model based on the extracted information.
[1332] How it works: The server inputs the extracted basic information into the generative AI model, builds a video scenario based on it, and then generates a video by combining related images and music.
[1333] Input: Basic information.
[1334] Output: The generated video data.
[1335] Step 7:
[1336] The server stores the generated video in cloud storage.
[1337] Specific operation: The server saves the generated video to a cloud storage service and generates a URL for accessing it.
[1338] Input: The generated video data.
[1339] Output: Videos stored in cloud storage, access URL.
[1340] Step 8:
[1341] The user requests video playback from a dedicated device or application.
[1342] Specific operation: The user logs in to the application again and selects the video they want to play from the generated video list page.
[1343] Input: Login information, video request.
[1344] Output: Video playback request.
[1345] Step 9:
[1346] The server distributes the video data to the terminal.
[1347] Specific operation: The server searches for video data in response to a user request and delivers it to the device in streaming or download format.
[1348] Input: A video playback request.
[1349] Output: Video data delivery to the device.
[1350] Step 10:
[1351] The terminal plays the video and allows the user to watch it.
[1352] Specific operation: The device decodes the received video data and plays it on the display. The user watches the video on the device.
[1353] Input: Received video data.
[1354] Output: The video shown on the display.
[1355] Step 11:
[1356] Users begin to interact based on the content of the video.
[1357] Specific operation: The user talks while watching the video (e.g., "The cherry blossoms were beautiful at this time."). The device captures the user's voice with a microphone.
[1358] Input: User's voice input.
[1359] Output: The captured audio data.
[1360] Step 12:
[1361] The device analyzes the voice, converts it into text data, and sends it to the server.
[1362] Specific operation: The device uses a speech recognition engine to convert the user's voice into text data and send it to the server.
[1363] Input: The captured audio data.
[1364] Output: The text data sent to the server.
[1365] Step 13:
[1366] The server generates an appropriate response based on the data it receives.
[1367] Specific operation: The server understands the context from the received text data and generates an appropriate response using a generative AI model.
[1368] Input: Received text data.
[1369] Output: The generated response data.
[1370] Step 14:
[1371] The terminal plays back the generated response.
[1372] Specific operation: The device converts the response data received from the server into speech using a speech synthesis engine and plays it back to the user.
[1373] Input: The response data received.
[1374] Output: The audio response played.
[1375] Step 15:
[1376] The emotion engine analyzes the user's voice and facial expression data to determine their emotional state.
[1377] How it works: The device captures the user's facial expressions with a camera and records their voice. The emotion engine analyzes this data to determine the user's emotional state.
[1378] Input: User's facial expression data, voice data.
[1379] Output: Emotion tag (e.g., happy, sad, excited).
[1380] Step 16:
[1381] Tailor your next piece of content based on your emotional state.
[1382] Specific operation: Based on the emotion tag, the server adjusts the next video and dialogue response according to the user's emotional state.
[1383] Input: emotion tag.
[1384] Output: A tailored content delivery plan.
[1385] (Application example 2)
[1386] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1387] Conventional rehabilitation systems provide the same content to all workers, including those with dementia and those unfamiliar with work processes, making it difficult to respond appropriately to individual needs and emotional states. Furthermore, there were insufficient means to properly reproduce and visually present past work procedures and episodes, making it difficult for workers to efficiently understand work procedures.
[1388] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting past photos and related memories, means for generating a video based on the collected data, means for providing the generated video to the target person, means for analyzing the user's voice input, means for generating a response based on the analyzed voice data, means for providing the generated response to the user, means for generating a video that reproduces a work procedure based on the collected data, and means for recognizing the user's emotional state and adjusting the content of the video and the response based on the emotional state. This enables appropriate responses according to individual needs and emotional states, and by providing a visual reproduction of past work procedures and episodes, the target person can efficiently understand the information.
[1389] "Past photographs and related reminiscences" are visual information related to the subject's past experiences or episodes and accompanying text or audio data.
[1390] "Collected data" is a collection of information provided by users, including photographs, text data, audio data, etc.
[1391] "Video generation means" refers to technology or devices that use collected data to create videos that convey information visually and audibly.
[1392] The "means for providing a video" refers to a method or device for allowing a target person to view the generated video.
[1393] "Means for analyzing user voice input" refers to technology that converts the user's voice into text data and understands and analyzes its content.
[1394] A "means for generating a response" is a technique or device that determines an appropriate response or next action based on the analyzed voice data.
[1395] "Means for generating videos that reproduce work procedures based on collected data" refers to technology or devices that visually reproduce past work procedures or episodes and provide them as videos in an easy-to-understand format.
[1396] "Means for recognizing the user's emotional state" refers to technology or devices that analyze the user's voice, facial expressions, behavior, etc. to determine and identify their emotional state.
[1397] The "means for adjusting the response content" refers to a technique or device for appropriately changing the video or response content provided depending on the user's emotional state or situation.
[1398] This invention is a rehabilitation system for dementia patients and workers, which aims to generate videos based on past photos and reminiscences and provide them to users. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides an individual rehabilitation plan.
[1399] Program processing explanation
[1400] Data Collection Phase
[1401] Terminal
[1402] Users log in to their account using a dedicated device (smartphone or head-mounted display) and upload past photos and related memories. This data is entered as text or voice data, temporarily saved in local storage, and then sent to the server.
[1403] Data Processing Phase
[1404] server
[1405] The server receives the photos and text data sent from the device. The received data is analyzed using a natural language processing (NLP) module (e.g., SpaCy or Google NLP API) to extract basic information about the episode. Based on this information, an AI-based video generation module (e.g., OpenAI or DeepMind) runs and automatically generates videos based on work procedures and reminiscences. The generated video data is stored in cloud storage (e.g., Google Cloud Storage or AWS S3).
[1406] Content provision phase
[1407] Terminal
[1408] The user requests playback of the generated video from a dedicated terminal. The server receives the request and delivers the video data to the terminal via streaming or download. The terminal then plays the received video and allows the user to watch it.
[1409] Dialogue Phase
[1410] Server and Device
[1411] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[1412] Emotion Recognition Phase
[1413] server
[1414] The user's voice input and facial expression data are analyzed by an emotion engine (e.g., Microsoft Azure Cognitive Services or Amazon Rekognition). The emotion engine determines the user's emotional state and generates emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[1415] Specific examples
[1416] For example, the following flow would be used for a rehabilitation system that supports operations within a factory.
[1417] 1. A worker uses a terminal to upload photos and procedures related to past "factory line maintenance work."
[1418] 2. The server receives the data and uses a video generation AI model to generate a video that reproduces the "factory line maintenance work procedures."
[1419] 3. The worker requests "Maintenance procedure playback" and watches the video.
[1420] 4. The worker asks, "What should I pay attention to in this part?" The device analyzes the voice and sends it to the server.
[1421] 5. The server generates a response saying, "It is important to check the safety devices in this section. In particular, make sure this switch is working," and plays it back through the terminal.
[1422] Prompt Sentence Examples
[1423] Generate a video that recreates a past maintenance procedure on a factory line based on the following data:
[1424] Photo: Maintenance work on a factory line
[1425] Procedure: 1. Turn off the power. 2. Check the safety devices. 3. Clean the conveyor belt. 4. Test the operation.
[1426] Your video should include the following elements:
[1427] 1. Detailed explanation of each step
[1428] 2. Points to pay particular attention to
[1429] 3. Clear visual and audio guidance for easy understanding by workers
[1430] Also, generate appropriate responses according to the worker's emotions.
[1431] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1432] Step 1: Data collection
[1433] The user logs into their account using a dedicated device (smartphone or head-mounted display) and uploads past photos and related memories. The input is photo data and text or voice data, which are temporarily stored in local storage. The output is data that is saved in local storage and then sent to the server. The specific actions that take place in this step are scanning photos, inputting text and voice data, saving the data in local storage, and sending it to the server.
[1434] Step 2: Data reception and analysis
[1435] The server receives the photo and text data sent from the device. The input is the data sent from the device, and the output is the data stored on the server. This data is input into a natural language processing (NLP) module, which extracts basic information about the episode. The specific operations performed here are receiving the data, storing it, and analyzing it using the NLP module.
[1436] Step 3: Video Generation
[1437] The AI video generation module in the server generates videos based on the analyzed episode information. The input is the data analyzed by the NLP module, and the output is the generated video data. The specific operations are as follows: input data into the AI model, generate video, and save it to cloud storage.
[1438] Step 4: Submit your video
[1439] A user requests playback of a generated video from a dedicated device. The input is a video playback request, and the output is video data that is streamed or downloaded to the user. Specific operations include accepting the request, obtaining the video data, and delivering it to the device.
[1440] Step 5: Dialogue analysis
[1441] The device receives and analyzes the audio emitted by the user while watching a video. The input is audio data, and the output is the analysis results converted into text data. The specific operations are to collect audio, analyze the audio, convert it into text data, and send it to the server.
[1442] Step 6: Response Generation
[1443] The server generates an appropriate response based on the received voice data. The input is the analyzed voice data, and the output is the generated response data. The specific operations are analyzing the voice data, determining the response content, and generating the response.
[1444] Step 7: Provide a response
[1445] The generated response is provided to the user through the terminal. The input is the generated response data, and the output is the response voice heard by the user. The specific operations are receiving the response data, generating the voice, and delivering it to the user.
[1446] Step 8: Emotion Recognition
[1447] The emotion engine analyzes the user's voice input and facial expression data. The input is the user's voice and facial expression data, and the output is an emotion tag. Specific operations include collecting voice and facial expression data, analyzing emotions, and generating emotion tags.
[1448] Step 9: Content Adjustment
[1449] The next video or response content is adjusted based on the emotion tag. The input is the emotion tag and the next content data to be provided, and the output is the adjusted video or response content. The specific operations are emotion tag analysis, content adjustment, and next content generation.
[1450] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1451] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1452] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1453] [Fourth embodiment]
[1454] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1455] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1456] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1457] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1458] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1459] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1460] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1461] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1462] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1463] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1464] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1465] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1466] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1467] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and reminiscences and provides them to patients. The system uses artificial intelligence technology to create individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[1468] The program for this system mainly consists of the following phases:
[1469] Data Collection Phase
[1470] Users (dementia patients and their families) input past photos and episodes through a dedicated device or application. Photos are collected by scanning or uploading, and episodes are input as text or voice data. The device temporarily stores the collected data and sends it to a server.
[1471] Data Processing Phase
[1472] The server receives the photos and text data sent from the device. The received data is analyzed by an artificial intelligence model to extract basic information about the episode. Based on this information, an AI video generation module operates and automatically generates videos that include emotional elements and specific locations. The generated video data is stored in cloud storage.
[1473] Content provision phase
[1474] The user requests playback of the generated video from a dedicated device or application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[1475] Dialogue Phase
[1476] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[1477] Specific examples
[1478] For example, if a family member of a dementia patient inputs a photo of the family cherry blossom viewing and an episode from that time into the system, the process would be as follows:
[1479] 1. The user (family member) logs in to the application, scans a photo, and enters the text "This photo was taken when my family went cherry blossom viewing."
[1480] 2. The device sends the photo and text data to the server.
[1481] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[1482] 4. The user (patient) requests video playback and watches the video on their device.
[1483] 5. The user (patient) says, "The cherry blossoms were beautiful at this time." The device analyzes the voice and sends it to the server.
[1484] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[1485] In this way, the system can stimulate past memories and promote brain activity in dementia patients. Furthermore, by utilizing artificial intelligence, it can efficiently provide individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[1486] The processing flow will be explained below.
[1487] Step 1:
[1488] Users (dementia patients and their families) log in to their accounts through a dedicated device or application.
[1489] Step 2:
[1490] Users scan or upload old photos and enter related memories as text or audio data.
[1491] Step 3:
[1492] The device temporarily stores the entered photo and text data in local storage.
[1493] Step 4:
[1494] The device transmits the photos and text data stored in the local storage to the server.
[1495] Step 5:
[1496] The server receives the photo and text data sent from the terminal.
[1497] Step 6:
[1498] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode from the text data.
[1499] Step 7:
[1500] The server uses an AI video generation module to generate videos based on the extracted information, including emotional elements and specific locations.
[1501] Step 8:
[1502] The server stores the generated video data in cloud storage.
[1503] Step 9:
[1504] The user requests playback of the generated video from a dedicated terminal or application.
[1505] Step 10:
[1506] The server receives the user's request and delivers the video data to the terminal in streaming or download format.
[1507] Step 11:
[1508] The terminal plays the distributed video and allows the user to view it.
[1509] Step 12:
[1510] The user (a dementia patient) can start a conversation based on the content of the video, for example, saying, "This day was really fun."
[1511] Step 13:
[1512] The device analyzes the user's voice input and converts it into text data using voice recognition technology.
[1513] Step 14:
[1514] The terminal transmits the analyzed voice data (text) to the server.
[1515] Step 15:
[1516] The server generates a response using a natural language generation (NLG) module based on the received voice data.
[1517] Step 16:
[1518] The server transmits the generated response data to the terminal.
[1519] Step 17:
[1520] The terminal uses a voice synthesis function to reproduce the received response data and convey it to the user.
[1521] Example 1
[1522] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1523] In the rehabilitation of dementia patients, it is important to stimulate individual memories and elicit emotional responses. However, rehabilitation using past photographs and anecdotes can be laborious for some caregivers and medical professionals. Furthermore, these tasks tend to rely on subjective judgment and are difficult to perform consistently. This can result in the insufficient effectiveness of dementia rehabilitation. Therefore, there is a need for an efficient and individually tailored rehabilitation system.
[1524] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1525] In this invention, the server includes: a means for a user to input past photos and episodes; a means for transmitting the input photos and episodes to the server; a means for the server to analyze the received data using an artificial intelligence model; a means for automatically generating videos based on the analysis results; a means for storing the generated video data in cloud storage; a means for a user to request video playback; a means for the server to receive the request and deliver the video data to a terminal; a means for the terminal to play the video; a means for analyzing the user's voice input and converting it into text data; a means for the server to generate a response based on the voice input; and a means for transmitting the generated response to the terminal and playing it as audio. This allows a user to input past photos and episodes and efficiently stimulate the memory of a dementia patient using the generated video. Furthermore, the use of artificial intelligence can provide a consistent rehabilitation plan, reducing the burden on caregivers and medical professionals.
[1526] "User" refers to the person who operates the system and inputs data, such as a dementia patient, their family member, or caregiver.
[1527] A "terminal" is an electronic device that allows a user to input data, receive data from a server, and execute processing.
[1528] A "server" is a central processing system that receives data sent from users and performs processes such as data analysis, video generation, and response generation.
[1529] "Artificial intelligence model" refers to machine learning algorithms and techniques used for complex information processing such as data analysis, video generation, and response generation.
[1530] "Cloud storage" refers to internet-based data storage servers that remotely store generated video data and other data and make it accessible as needed.
[1531] An "episode" refers to a specific memory or event related to a photo entered by the user.
[1532] A "request" refers to an action in which a user requests a specific operation or process from the system.
[1533] "Voice input" refers to voice data input by a user speaking to the system using a microphone or the like.
[1534] "Text data" refers to data in the form of a string of characters obtained by analyzing voice input.
[1535] "Video generation" refers to the process of creating visual and audio content based on photographs and anecdotes.
[1536] "Response generation" refers to the process of generating an appropriate reaction or reply to a user's voice input.
[1537] "Analysis" refers to the processing of collected data to extract useful information and to perform subsequent processing based on that information.
[1538] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and episodes and provides them to patients. The system uses artificial intelligence technology to analyze data entered by users (dementia patients, their families, or caregivers) and create individual rehabilitation plans.
[1539] First, a user uses a dedicated device or application to input past photos and related episodes. For example, the user scans a photo of a family cherry blossom viewing and enters a text description such as "This photo was taken when my family went cherry blossom viewing." This data is temporarily stored by the device and then sent to the server.
[1540] The server receives the photos and text data sent from the device. The received data is analyzed using artificial intelligence models such as TensorFlow and OpenAI's GPT. The analysis extracts basic information about the episode and automatically generates a video that includes emotional elements and a specific location (in this case, cherry blossom viewing). The generated video data is stored in cloud storage such as Amazon S3.
[1541] The user requests playback of the generated video from a dedicated device or application. The server receives the request, retrieves the video data stored in cloud storage, and delivers it to the device via streaming or download. The device plays the received video, and the user (dementia patient) watches it.
[1542] A dementia patient can initiate a dialogue based on the content of the video. For example, they can say, "The cherry blossoms were beautiful at this time." The device analyzes this voice input and converts it into text data. The speech is converted to text using speech recognition technology such as the Google Speech-to-Text API. The converted text data is sent to a server, which uses an artificial intelligence model such as OpenAI's GPT-3 to generate an appropriate response. This response is sent to the device and played back using a text-to-speech engine that converts text to speech.
[1543] A specific example is shown below. For example, if a user inputs "a photo of a family cherry blossom viewing party" and an episode from that event, the system operates as follows.
[1544] 1. The user uses the application to scan a photo of a family cherry blossom viewing party and enters the message, "This photo was taken when my family went cherry blossom viewing."
[1545] 2. The device temporarily stores the photo and text data and sends them to the server.
[1546] 3. The server receives the data and uses an AI model (e.g., TensorFlow) to generate a "cherry blossom viewing"-themed video, which includes footage of cherry blossom viewing and scenes of families having fun.
[1547] 4. The user (patient) requests video playback and watches the video on their device.
[1548] 5. The user (patient) says, "The cherry blossoms were beautiful at this time."
[1549] 6. The device analyzes the voice, converts it into text data, and sends it to the server.
[1550] 7. The server generates a response saying, "It was really beautiful. Your grandma was with you at the time," and plays it back through the device.
[1551] An example prompt is, "Based on photos and stories from this family's cherry blossom viewing, create a video that includes emotional elements and specific locations."
[1552] In this way, the system can stimulate past memories and promote brain activity in dementia patients, and by utilizing artificial intelligence, it can efficiently provide individual rehabilitation plans, reducing the burden on caregivers and medical professionals.
[1553] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1554] Step 1:
[1555] User data entry and saving
[1556] The user inputs past photos and episodes using a dedicated device or application. For example, the user scans a photo of a family cherry blossom viewing party and enters a text description such as "This photo is from when we went cherry blossom viewing with my family." The input data (photo and text data) is temporarily saved by the device. The input data is the photo file and text data, and the output is the temporarily saved data.
[1557] Step 2:
[1558] Sending data
[1559] The device sends the saved photos and text data to the server. Specifically, the device sends the data by making an HTTP POST request to the server. The input is the temporarily saved photos and text data, and the output is the data sent to the server.
[1560] Step 3:
[1561] Data reception and analysis
[1562] The server receives the photos and text data sent from the device. The received data is input into an artificial intelligence model such as Google's TensorFlow or OpenAI's GPT, which extracts basic information about the episode from the text data and matches it with the photos to analyze its relevance. The input is the photos and text data sent to the server, and the output is the analyzed basic information about the episode.
[1563] Step 4:
[1564] Video generation
[1565] The server runs an AI video generation module based on the analysis results. It automatically generates videos that include emotional elements or specific locations (e.g., cherry blossom viewing). Specifically, the AI model selects video material based on the analysis results, and the video is edited using an editor. The input is basic information about the analyzed episode, and the output is the generated video data.
[1566] Step 5:
[1567] Saving videos
[1568] The server saves the generated video data in cloud storage (e.g., Amazon S3). It uploads the video data to the cloud using an API. The input is the generated video data, and the output is the video data saved in the cloud storage.
[1569] Step 6:
[1570] Video Request
[1571] A user requests video playback from a dedicated device or application. Specifically, the user presses the play button on the device to send a playback request to the server. The input is the user's playback request, and the output is a playback request to the server.
[1572] Step 7:
[1573] Video data acquisition and distribution
[1574] The server receives a request from a user and retrieves video data from cloud storage. The retrieved video data is then delivered to the device via streaming or download. The input is the video data retrieved from cloud storage, and the output is the video data delivered to the device.
[1575] Step 8:
[1576] Play video
[1577] The device plays the received video data. The video is displayed using a dedicated playback application or a standard video player. The input is the video data delivered to the device, and the output is the played video.
[1578] Step 9:
[1579] Analyzing voice input
[1580] The patient begins a conversation based on the content of the video. For example, they might say, "The cherry blossoms were beautiful at this time." The device analyzes this voice input and converts it into text data using the Google Speech-to-Text API or similar. The input is the patient's voice, and the output is text data.
[1581] Step 10:
[1582] Generate and send the response
[1583] The device sends the converted text data to the server. The server generates an appropriate response based on the received text data using an artificial intelligence model such as OpenAI's GPT. The generated response is sent to the device in text format. The input is text data, and the output is the generated response.
[1584] Step 11:
[1585] Response playback
[1586] The device plays the response sent from the server as audio. It uses a text-to-speech engine to convert the response into audio and plays it back. The input is the generated response text data, and the output is the played audio response.
[1587] (Application example 1)
[1588] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1589] Rehabilitation is important for dementia patients to stimulate their memories and improve their cognitive function. However, effective rehabilitation requires the creation of individual rehabilitation plans, which places a heavy burden on caregivers and medical professionals. Furthermore, there is a problem with brick-and-mortar rehabilitation support, where there is no effective way to provide videos and dialogue based on past memories.
[1590] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1591] In this invention, the server includes a means for collecting past photos and related episodes, a means for generating videos based on the collected data, a means for providing the generated videos to the dementia patient, a means for analyzing a user's voice input, a means for generating responses based on the analyzed voice data, and a means for providing videos and interactive dialogues based on past memories in a physical store, thereby stimulating the memory of the dementia patient and effectively implementing rehabilitation, while reducing the burden on caregivers and medical professionals.
[1592] "Past photos and related episodes" is information provided by a patient with dementia or their family that indicates images taken in the past and memories or events related to the images.
[1593] "Means of collection" refers to the methods and technologies used to collect photos and stories via devices such as tablets and smartphones.
[1594] "Means for generating video" refers to methods and algorithms for creating video using artificial intelligence technology based on collected photos and episodes.
[1595] "Means of delivery to dementia patients" refers to methods of providing the generated videos and other rehabilitation content to dementia patients via devices such as smartphones and tablets.
[1596] "Means for analyzing user voice input" means techniques for collecting voice data using a microphone or other voice capture device and analyzing the voice data.
[1597] "Means for generating a response based on analyzed voice data" refers to a method or algorithm for generating an appropriate reply based on the analyzed user's voice input.
[1598] "Means for providing videos and interactive dialogues based on past memories in physical locations" means technologies and methods for providing videos based on past photos and episodes and related dialogues in physical locations such as cafes and rehabilitation centers.
[1599] This invention is a rehabilitation system for dementia patients, and aims to provide videos and interactive dialogues based on past photos and episodes in a physical store. This system consistently performs everything from collecting photos and episodes to processing the data, providing content, and generating dialogues.
[1600] The system that realizes this application example is made up of a number of hardware and software components, each of which will be described in detail below.
[1601] 1. Data Collection Phase
[1602] Users (dementia patients and their families) enter past photos and episodes through dedicated terminals located in physical stores or smartphone applications. Photos are collected by scanning or uploading, and episodes are entered as text or voice data. The terminal temporarily stores this data and later sends it to a server.
[1603] 2. Data Processing Phase
[1604] The server receives the photos and text data sent from the device. The received data is analyzed by an artificial intelligence model (using TensorFlow, for example) to extract basic information about the episode. Based on this information, an AI video generation module operates and automatically generates videos that include emotional elements and specific locations. The generated video data is stored in cloud storage.
[1605] 3. Content provision phase
[1606] The user then requests playback of the generated video again from a dedicated device or smartphone application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[1607] 4. Dialogue Phase
[1608] The patient with dementia initiates a dialogue based on the content of the video. For example, they might say, "I really enjoyed this day." The device analyzes the voice input and converts it into text data (for example, using a Transformer-based NLP model). The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[1609] Specific examples
[1610] For example, if a family member of a dementia patient inputs a photo of the family cherry blossom viewing and an episode from that time into the system, the process would be as follows:
[1611] 1. The user (family member) logs in to a terminal at a physical store, scans a photo, and enters a text story such as, "This photo was taken when the family went cherry blossom viewing."
[1612] 2. The device sends the photo and text data to the server.
[1613] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[1614] 4. The user (patient) requests video playback and watches the video on a tablet in the store.
[1615] 5. When the user (patient) says, "The cherry blossoms were beautiful at this time," the device analyzes the voice and sends it to the server.
[1616] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[1617] Prompt Sentence Examples
[1618] "If you look at a photo of a family cherry blossom viewing and say the cherry blossoms were beautiful, how will the AI respond?"
[1619] "What kind of comments would you have the AI make when shown photos of past summer festivals?"
[1620] In this way, the system can stimulate past memories in dementia patients, enabling effective rehabilitation while reducing the burden on caregivers and medical professionals.
[1621] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1622] Step 1: Data collection phase
[1623] Specific operation: Users (dementia patients and their families) input past photos and episodes using a dedicated terminal installed in a physical store or a smartphone application. Users scan or upload photos and enter an episode in text, such as "This photo was taken when my family went cherry blossom viewing."
[1624] Input: scanned or uploaded photos, text or audio episodes
[1625] Output: Photos and episode data temporarily saved on the device
[1626] Step 2: Data transmission phase
[1627] Specific operation: The device sends the temporarily stored photos and episode data to the server. During this process, the data is compressed and encrypted as necessary.
[1628] Input: Photos and episode data stored on the device
[1629] Output: Photos and episode data sent to the server
[1630] Step 3: Data analysis phase
[1631] How it works: The server analyzes the received photo and episode data. An artificial intelligence model (using, for example, TensorFlow) extracts basic information about the episode and identifies related themes (e.g., "cherry blossom viewing") based on the analysis results.
[1632] Input: Photos and episode data sent to the server
[1633] Output: Basic information about the analyzed episodes and identified themes
[1634] Step 4: Video Generation Phase
[1635] How it works: Based on the analysis results, the server uses an AI video generation module to automatically generate videos that incorporate emotional elements and specific locations, including footage of actual cherry blossom viewing and scenes of families having fun.
[1636] Input: Basic information about the analyzed episode and the identified themes
[1637] Output: Generated video data
[1638] Step 5: Video saving phase
[1639] Specific operation: The generated video data is stored in cloud storage. Before storing, the video data is compressed and encrypted as necessary.
[1640] Input: Generated video data
[1641] Output: Video data stored in cloud storage
[1642] Step 6: Playback Request Phase
[1643] Specific operation: A user requests playback of the generated video using a dedicated device or a smartphone application. The request is sent to the server.
[1644] Input: User's playback request
[1645] Output: Playback request sent to the server
[1646] Step 7: Video Distribution Phase
[1647] Specific operation: The server receives a playback request and delivers the video data stored in cloud storage to the device in streaming or download format.
[1648] Input: Playback request sent to the server, video data stored in cloud storage
[1649] Output: Video data delivered to the device
[1650] Step 8: Video viewing phase
[1651] Specific operation: The device plays the distributed video and allows the dementia patient to watch it. As the patient watches the video, past memories are stimulated.
[1652] Input: Video data delivered to the device
[1653] Output: Video currently being played by the dementia patient
[1654] Step 9: Audio Analysis Phase
[1655] Specific operation: The dementia patient begins a dialogue based on the content of the video. When a voice input such as "This day was really fun," the device analyzes the voice and converts it into text data.
[1656] Input: Voice input from dementia patients
[1657] Output: Speech input converted to text data
[1658] Step 10: Response generation phase
[1659] Specific operation: The server receives the converted text data and uses artificial intelligence techniques (e.g., a Transformer-based NLP model) to generate an appropriate response. The generated response is then sent from the server to the device.
[1660] Input: Voice input converted to text data
[1661] Output: The server-generated response
[1662] Step 11: Response delivery phase
[1663] Specific operation: The device plays back the response sent from the server and provides it to the dementia patient. Responses such as "It was really beautiful. Your grandmother was with you at the time" are played back.
[1664] Input: The response sent by the server
[1665] Output: Response voice heard by dementia patient
[1666] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1667] This invention is a rehabilitation system for dementia patients that generates videos based on past photos and reminiscences and provides them to patients. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it can provide an individual rehabilitation plan more effectively.
[1668] The program for this system mainly consists of the following phases:
[1669] Data Collection Phase
[1670] Users (people with dementia or their families) log in to their accounts through a dedicated device or application. They then scan or upload past photos and enter related memories as text or voice data. The device temporarily stores the collected photos and text data in its local storage and then sends them to the server.
[1671] Data Processing Phase
[1672] The server receives the photos and text data sent from the device. The received data is input into a natural language processing (NLP) module to extract basic information about the episode. Based on this information, an AI-based video generation module operates to automatically generate videos that include emotional elements and specific locations. The generated video data is then stored in cloud storage.
[1673] Content provision phase
[1674] The user requests playback of the generated video from a dedicated device or application. The server receives the request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[1675] Dialogue Phase
[1676] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[1677] Emotion Recognition Phase
[1678] The emotion engine analyzes the user's voice input and facial expression data, determining their emotional state and generating emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[1679] Specific examples
[1680] For example, a family member of a dementia patient can enter "photos from cherry blossom viewing" and "anecdotes from that time" into the system, and the process will go as follows:
[1681] 1. The user (family member) logs in to the application, scans a photo, and enters the text "This photo was taken when my family went cherry blossom viewing."
[1682] 2. The device sends the photo and text data to the server.
[1683] 3. The server receives the data and uses the AI model to generate a "cherry blossom viewing"-themed video, which includes footage of actual cherry blossom viewing and scenes of families enjoying themselves.
[1684] 4. The user (patient) requests video playback and watches the video on their device.
[1685] 5. The user (patient) says, "The cherry blossoms were beautiful at this time." The device analyzes the voice and sends it to the server.
[1686] 6. The server generates a response saying, "It was really beautiful. Your grandmother was with you at the time," and plays it back through the device.
[1687] 7. The emotion engine analyzes the user's emotions and adjusts the next dialogue content based on the results. For example, if the user is happy, it will provide more stories related to happiness.
[1688] In this way, the combination of emotion recognition capabilities can more effectively advance the personalized rehabilitation plan for dementia patients, while providing responses that correspond to the user's emotional state makes interactions more natural and effective.
[1689] The processing flow will be explained below.
[1690] Step 1:
[1691] Users (dementia patients and their families) log in to their accounts through a dedicated device or application.
[1692] Step 2:
[1693] Users scan or upload old photos and enter related memories as text or audio data.
[1694] Step 3:
[1695] The device temporarily stores the entered photo and text data in local storage.
[1696] Step 4:
[1697] The device transmits the photos and text data stored in the local storage to the server.
[1698] Step 5:
[1699] The server receives the photo and text data sent from the terminal.
[1700] Step 6:
[1701] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode from the text data.
[1702] Step 7:
[1703] The server uses an AI video generation module to generate videos based on the extracted information, including emotional elements and specific locations.
[1704] Step 8:
[1705] The server stores the generated video data in cloud storage.
[1706] Step 9:
[1707] The user requests playback of the generated video from a dedicated terminal or application.
[1708] Step 10:
[1709] The server receives the user's request and delivers the video data to the terminal in streaming or download format.
[1710] Step 11:
[1711] The terminal plays the distributed video and allows the user to view it.
[1712] Step 12:
[1713] The user (a dementia patient) can start a conversation based on the content of the video, for example, saying, "This day was really fun."
[1714] Step 13:
[1715] The device analyzes the user's voice input and converts it into text data using voice recognition technology.
[1716] Step 14:
[1717] The terminal transmits the analyzed voice data (text) to the server.
[1718] Step 15:
[1719] The server generates a response using a natural language generation (NLG) module based on the received voice data.
[1720] Step 16:
[1721] The server transmits the generated response data to the terminal.
[1722] Step 17:
[1723] The terminal uses a voice synthesis function to reproduce the received response data and convey it to the user.
[1724] Step 18:
[1725] The device inputs the user's voice and facial expression data into an emotion engine to analyze the user's emotional state. For example, emotion tags such as "joy," "sadness," and "excitement" are generated.
[1726] Step 19:
[1727] The server automatically adjusts the content of the next video and dialogue responses based on the user's emotions recognized by the emotion engine. For example, if the user is happy, the server generates a video that includes more episodes related to happiness.
[1728] Step 20:
[1729] The user can receive continually adapted rehabilitation through updated content.
[1730] The above is the specific processing flow of the system of the invention combined with the emotion engine. In this way, by providing rehabilitation that corresponds to the user's emotions, it is possible to effectively stimulate the memory of dementia patients and realize more natural and effective dialogue.
[1731] Example 2
[1732] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1733] In rehabilitation for dementia patients, there is a need to provide effective stimulation that corresponds to the individual memories and emotions of each patient. Conventional rehabilitation methods have difficulty customizing to correspond to the memories and emotions of each patient, which often limits their effectiveness. For this reason, there is a need to develop a system that can automatically provide an individual rehabilitation plan that corresponds to the patient's memory ability and emotions.
[1734] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1735] In this invention, the server includes: means for a user to collect past photos and related memories; means for transmitting the collected data from the terminal to the server; means for the server to analyze the received data using a natural language processing module and extract basic information about the episode; means for generating a video using a generative AI model based on the extracted information; means for storing the generated video in cloud storage; means for a user to request video playback; means for the server to deliver video data to the terminal in response to the request; means for the terminal to play the video for the user; means for analyzing the user's voice input and converting it into text data; means for transmitting the converted voice data to the server and generating an appropriate response; means for transmitting the response data generated by the server to the terminal and playing it; means for analyzing the user's voice and facial expressions using an emotion engine to determine the user's emotional state; and means for adjusting the content to be provided next based on the emotional state. This makes it possible to automatically provide rehabilitation tailored to the individual memories and emotions of dementia patients.
[1736] "User" refers to people who use the system to input past photos and memories, or dementia patients and their families.
[1737] "Terminal" refers to the device through which a user accesses the system, uploads photos, and performs voice input. Specifically, this includes smartphones, tablets, and PCs.
[1738] "Server" refers to a computer system that receives, analyzes, stores, and manages data sent by users.
[1739] "Natural language processing module" refers to a software module for analyzing received text data and extracting basic information about an episode.
[1740] "Generative AI model" refers to an artificial intelligence algorithm that automatically generates videos based on extracted information.
[1741] "Cloud storage" refers to an online storage service for storing generated video data.
[1742] An "emotion engine" refers to a computer program that analyzes a user's voice and facial expressions to determine their emotional state and generate appropriate responses and content.
[1743] A "request" refers to an operation by a user to request playback of a video through a terminal.
[1744] "Video" refers to video content generated from past photos and related memories.
[1745] "Text data" refers to data that has been analyzed and converted into text information from a user's voice.
[1746] "Response" refers to an appropriate reply generated by a server in response to a user's input.
[1747] "Content" refers to videos, responses, and other information provided to users.
[1748] "Emotional state" refers to the emotional state such as joy, sadness, excitement, etc., analyzed from the user's voice and facial expressions.
[1749] "Rehabilitation" refers to treatment aimed at stimulating the memories and emotions of dementia patients and maintaining or improving their functions.
[1750] This invention is a system aimed at the rehabilitation of dementia patients, which generates videos based on past photos and reminiscences and provides them to patients. It also uses an emotion engine that recognizes the user's emotions to more effectively provide individual rehabilitation plans.
[1751] Overall flow
[1752] 1. Data Collection Phase
[1753] Users log in to the system using a dedicated device or application, then scan or upload old photos and enter related memories as text or voice data. The device temporarily saves this data in its local storage and then sends it to the server.
[1754] 2. Data Processing Phase
[1755] The server receives the photos and text data sent from the device. The received data is input into a natural language processing (NLP) module to extract basic information about the episode. Based on the extracted information, a generative AI model generates videos containing emotional elements. The generated videos are then stored in cloud storage.
[1756] 3. Content provision phase
[1757] The user requests video playback from a dedicated device or application. The server receives this request and delivers the video data to the device via streaming or download. The device then plays the received video and allows the dementia patient to watch it.
[1758] 4. Dialogue Phase
[1759] The patient can start a dialogue based on the content of the video. For example, if the patient says, "I had a really fun day," the device analyzes the voice input and converts it into text data. The converted data is sent to the server, which then generates an appropriate response and plays it back through the device.
[1760] 5. Emotion Recognition Phase
[1761] The user's voice and facial expression data are analyzed by the emotion engine, which determines the user's emotional state and generates emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[1762] Hardware and software used
[1763] Devices: Smartphones, tablets, computers, etc.
[1764] Server: A high-performance computer that receives, analyzes, stores, and runs generative AI models.
[1765] Natural Language Processing Module: Software that performs text analysis (e.g., SpaCy, NLTK, etc.).
[1766] Generative AI models: Artificial intelligence algorithms that generate videos (e.g., GPT-3, DALL-E).
[1767] Emotion engine: A software module that performs emotion analysis (e.g., IBM Watson, Microsoft Azure Emotion API).
[1768] Cloud storage: An internet storage service for saving video data (e.g., Amazon S3, Google Cloud Storage).
[1769] Specific examples
[1770] For example, if a family member of a dementia patient enters "photos from a cherry blossom viewing party" and "an episode from that time" into the system, the process would be as follows:
[1771] 1. The user (family member) logs in to the application, enters the text "This photo was taken when my family went cherry blossom viewing," and scans the photo.
[1772] 2. The device temporarily stores the photo and text data and sends them to the server.
[1773] 3. The server receives the data and uses a natural language processing module to extract basic information about the episodes on the theme of "cherry blossom viewing." A generative AI model then generates a video based on that information and saves it in cloud storage.
[1774] 4. The user (patient) uses the application to request video playback and watches the video on their device.
[1775] 5. When the user (patient) says, "The cherry blossoms were beautiful at this time," the device captures the voice, converts it into text, and sends it to the server. The server generates an appropriate response (e.g., "They were really beautiful. Your grandmother was with you at the time") and plays it back on the device.
[1776] 6. The emotion engine analyzes the user's emotional state, and if they are happy, it will provide them with the next relevant episode.
[1777] Prompt Sentence Examples
[1778] An example of a prompt for a generative AI model is:
[1779] text
[1780] "Upload a photo of you and your family going cherry blossom viewing, enter the text 'We had a great time enjoying cherry blossom viewing in the park on this day,' and generate a two-minute video with the theme 'Cherry Blossom Park.'"
[1781] In this way, combining emotion recognition capabilities can effectively advance individual rehabilitation for dementia patients. Also, by providing responses that correspond to the user's emotional state, interactions become more natural and effective.
[1782] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1783] Step 1:
[1784] A user logs in to a dedicated terminal or application.
[1785] Specific operation: The user launches the application and logs in by entering their account information (username and password).
[1786] Input: Username, Password.
[1787] Output: The user is logged in.
[1788] Step 2:
[1789] Users scan or upload old photos and enter their memories.
[1790] What happens: The user uses the device's camera to scan a photo or upload an existing photo file, then uses the keyboard to enter text or the microphone to enter voice input.
[1791] Input: Photo files, text data, or audio data.
[1792] Output: Collected photos and text data.
[1793] Step 3:
[1794] The device temporarily stores the collected data in local storage and then transmits it to the server.
[1795] Specific operation: The device temporarily saves the photo and text data as files in local storage, and then sends the data to the server.
[1796] Input: Collected photo files and text data.
[1797] Output: The data sent to the server.
[1798] Step 4:
[1799] The server receives the data sent from the terminal.
[1800] Specific operation: The server receives the photos and text data sent from the device via the network and temporarily stores them in the server's storage.
[1801] Input: Photo files and text data sent from the device.
[1802] Output: Data stored on the server.
[1803] Step 5:
[1804] The server inputs the received data into a natural language processing (NLP) module to extract basic information about the episode.
[1805] Specific operation: The server inputs the received text data into a natural language processing (NLP) module to extract basic information such as keywords, date and time, location, and names of people.
[1806] Input: Received text data.
[1807] Output: Basic information about the extracted episode.
[1808] Step 6:
[1809] The server generates a video using a generative AI model based on the extracted information.
[1810] How it works: The server inputs the extracted basic information into the generative AI model, builds a video scenario based on it, and then generates a video by combining related images and music.
[1811] Input: Basic information.
[1812] Output: The generated video data.
[1813] Step 7:
[1814] The server stores the generated video in cloud storage.
[1815] Specific operation: The server saves the generated video to a cloud storage service and generates a URL for accessing it.
[1816] Input: The generated video data.
[1817] Output: Videos stored in cloud storage, access URL.
[1818] Step 8:
[1819] The user requests video playback from a dedicated device or application.
[1820] Specific operation: The user logs in to the application again and selects the video they want to play from the generated video list page.
[1821] Input: Login information, video request.
[1822] Output: Video playback request.
[1823] Step 9:
[1824] The server distributes the video data to the terminal.
[1825] Specific operation: The server searches for video data in response to a user request and delivers it to the device in streaming or download format.
[1826] Input: A video playback request.
[1827] Output: Video data delivery to the device.
[1828] Step 10:
[1829] The terminal plays the video and allows the user to watch it.
[1830] Specific operation: The device decodes the received video data and plays it on the display. The user watches the video on the device.
[1831] Input: Received video data.
[1832] Output: The video shown on the display.
[1833] Step 11:
[1834] Users begin to interact based on the content of the video.
[1835] Specific operation: The user talks while watching the video (e.g., "The cherry blossoms were beautiful at this time."). The device captures the user's voice with a microphone.
[1836] Input: User's voice input.
[1837] Output: The captured audio data.
[1838] Step 12:
[1839] The device analyzes the voice, converts it into text data, and sends it to the server.
[1840] Specific operation: The device uses a speech recognition engine to convert the user's voice into text data and send it to the server.
[1841] Input: The captured audio data.
[1842] Output: The text data sent to the server.
[1843] Step 13:
[1844] The server generates an appropriate response based on the data it receives.
[1845] Specific operation: The server understands the context from the received text data and generates an appropriate response using a generative AI model.
[1846] Input: Received text data.
[1847] Output: The generated response data.
[1848] Step 14:
[1849] The terminal plays back the generated response.
[1850] Specific operation: The device converts the response data received from the server into speech using a speech synthesis engine and plays it back to the user.
[1851] Input: The response data received.
[1852] Output: The audio response played.
[1853] Step 15:
[1854] The emotion engine analyzes the user's voice and facial expression data to determine their emotional state.
[1855] How it works: The device captures the user's facial expressions with a camera and records their voice. The emotion engine analyzes this data to determine the user's emotional state.
[1856] Input: User's facial expression data, voice data.
[1857] Output: Emotion tag (e.g., happy, sad, excited).
[1858] Step 16:
[1859] Tailor your next piece of content based on your emotional state.
[1860] Specific operation: Based on the emotion tag, the server adjusts the next video and dialogue response according to the user's emotional state.
[1861] Input: emotion tag.
[1862] Output: A tailored content delivery plan.
[1863] (Application example 2)
[1864] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1865] Conventional rehabilitation systems provide the same content to all workers, including those with dementia and those unfamiliar with work processes, making it difficult to respond appropriately to individual needs and emotional states. Furthermore, there were insufficient means to properly reproduce and visually present past work procedures and episodes, making it difficult for workers to efficiently understand work procedures.
[1866] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting past photos and related memories, means for generating a video based on the collected data, means for providing the generated video to the target person, means for analyzing the user's voice input, means for generating a response based on the analyzed voice data, means for providing the generated response to the user, means for generating a video that reproduces a work procedure based on the collected data, and means for recognizing the user's emotional state and adjusting the content of the video and the response based on the emotional state. This enables appropriate responses according to individual needs and emotional states, and by providing a visual reproduction of past work procedures and episodes, the target person can efficiently understand the information.
[1867] "Past photographs and related reminiscences" are visual information related to the subject's past experiences or episodes and accompanying text or audio data.
[1868] "Collected data" is a collection of information provided by users, including photographs, text data, audio data, etc.
[1869] "Video generation means" refers to technology or devices that use collected data to create videos that convey information visually and audibly.
[1870] The "means for providing a video" refers to a method or device for allowing a target person to view the generated video.
[1871] "Means for analyzing user voice input" refers to technology that converts the user's voice into text data and understands and analyzes its content.
[1872] A "means for generating a response" is a technique or device that determines an appropriate response or next action based on the analyzed voice data.
[1873] "Means for generating videos that reproduce work procedures based on collected data" refers to technology or devices that visually reproduce past work procedures or episodes and provide them as videos in an easy-to-understand format.
[1874] "Means for recognizing the user's emotional state" refers to technology or devices that analyze the user's voice, facial expressions, behavior, etc. to determine and identify their emotional state.
[1875] The "means for adjusting the response content" refers to a technique or device for appropriately changing the video or response content provided depending on the user's emotional state or situation.
[1876] This invention is a rehabilitation system for dementia patients and workers, which aims to generate videos based on past photos and reminiscences and provide them to users. Furthermore, by combining it with an emotion engine that recognizes the user's emotions, it provides an individual rehabilitation plan.
[1877] Program processing explanation
[1878] Data Collection Phase
[1879] Terminal
[1880] Users log in to their account using a dedicated device (smartphone or head-mounted display) and upload past photos and related memories. This data is entered as text or voice data, temporarily saved in local storage, and then sent to the server.
[1881] Data Processing Phase
[1882] server
[1883] The server receives the photos and text data sent from the device. The received data is analyzed using a natural language processing (NLP) module (e.g., SpaCy or Google NLP API) to extract basic information about the episode. Based on this information, an AI-based video generation module (e.g., OpenAI or DeepMind) runs and automatically generates videos based on work procedures and reminiscences. The generated video data is stored in cloud storage (e.g., Google Cloud Storage or AWS S3).
[1884] Content provision phase
[1885] Terminal
[1886] The user requests playback of the generated video from a dedicated terminal. The server receives the request and delivers the video data to the terminal via streaming or download. The terminal then plays the received video and allows the user to watch it.
[1887] Dialogue Phase
[1888] Server and Device
[1889] The patient with dementia begins a dialogue based on the content of the video. For example, the patient might say, "I had a really fun day." The device analyzes the voice input and converts it into text data. The converted data is sent to the server, which generates an appropriate response based on the received data. This response is played back through the device, and the dialogue with the patient progresses.
[1890] Emotion Recognition Phase
[1891] server
[1892] The user's voice input and facial expression data are analyzed by an emotion engine (e.g., Microsoft Azure Cognitive Services or Amazon Rekognition). The emotion engine determines the user's emotional state and generates emotion tags such as "happiness," "sadness," and "excitement." Based on this, the content of the next video and the dialogue responses are automatically adjusted.
[1893] Specific examples
[1894] For example, the following flow would be used for a rehabilitation system that supports operations within a factory.
[1895] 1. A worker uses a terminal to upload photos and procedures related to past "factory line maintenance work."
[1896] 2. The server receives the data and uses a video generation AI model to generate a video that reproduces the "factory line maintenance work procedures."
[1897] 3. The worker requests "Maintenance procedure playback" and watches the video.
[1898] 4. The worker asks, "What should I pay attention to in this part?" The device analyzes the voice and sends it to the server.
[1899] 5. The server generates a response saying, "It is important to check the safety devices in this section. In particular, make sure this switch is working," and plays it back through the terminal.
[1900] Prompt Sentence Examples
[1901] Generate a video that recreates a past maintenance procedure on a factory line based on the following data:
[1902] Photo: Maintenance work on a factory line
[1903] Procedure: 1. Turn off the power. 2. Check the safety devices. 3. Clean the conveyor belt. 4. Test the operation.
[1904] Your video should include the following elements:
[1905] 1. Detailed explanation of each step
[1906] 2. Points to pay particular attention to
[1907] 3. Clear visual and audio guidance for easy understanding by workers
[1908] Also, generate appropriate responses according to the worker's emotions.
[1909] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1910] Step 1: Data collection
[1911] The user logs into their account using a dedicated device (smartphone or head-mounted display) and uploads past photos and related memories. The input is photo data and text or voice data, which are temporarily stored in local storage. The output is data that is saved in local storage and then sent to the server. The specific actions that take place in this step are scanning photos, inputting text and voice data, saving the data in local storage, and sending it to the server.
[1912] Step 2: Data reception and analysis
[1913] The server receives the photo and text data sent from the device. The input is the data sent from the device, and the output is the data stored on the server. This data is input into a natural language processing (NLP) module, which extracts basic information about the episode. The specific operations performed here are receiving the data, storing it, and analyzing it using the NLP module.
[1914] Step 3: Video Generation
[1915] The AI video generation module in the server generates videos based on the analyzed episode information. The input is the data analyzed by the NLP module, and the output is the generated video data. The specific operations are as follows: input data into the AI model, generate video, and save it to cloud storage.
[1916] Step 4: Submit your video
[1917] A user requests playback of a generated video from a dedicated device. The input is a video playback request, and the output is video data that is streamed or downloaded to the user. Specific operations include accepting the request, obtaining the video data, and delivering it to the device.
[1918] Step 5: Dialogue analysis
[1919] The device receives and analyzes the audio emitted by the user while watching a video. The input is audio data, and the output is the analysis results converted into text data. The specific operations are to collect audio, analyze the audio, convert it into text data, and send it to the server.
[1920] Step 6: Response Generation
[1921] The server generates an appropriate response based on the received voice data. The input is the analyzed voice data, and the output is the generated response data. The specific operations are analyzing the voice data, determining the response content, and generating the response.
[1922] Step 7: Provide a response
[1923] The generated response is provided to the user through the terminal. The input is the generated response data, and the output is the response voice heard by the user. The specific operations are receiving the response data, generating the voice, and delivering it to the user.
[1924] Step 8: Emotion Recognition
[1925] The emotion engine analyzes the user's voice input and facial expression data. The input is the user's voice and facial expression data, and the output is an emotion tag. Specific operations include collecting voice and facial expression data, analyzing emotions, and generating emotion tags.
[1926] Step 9: Content Adjustment
[1927] The next video or response content is adjusted based on the emotion tag. The input is the emotion tag and the next content data to be provided, and the output is the adjusted video or response content. The specific operations are emotion tag analysis, content adjustment, and next content generation.
[1928] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1929] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1930] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1931] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1932] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1933] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1934] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1935] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1936] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1937] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1938] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1939] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1940] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1941] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1942] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1943] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1944] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1945] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1946] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1947] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1948] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1949] The following is further disclosed regarding the above embodiment.
[1950] (Claim 1)
[1951] A means of collecting historical photographs and related memoirs;
[1952] a means for generating a video based on the collected data;
[1953] A means for providing the generated video to dementia patients;
[1954] means for analyzing a user's voice input;
[1955] means for generating a response based on the analyzed voice data;
[1956] means for providing the generated response to a user;
[1957] A system including:
[1958] (Claim 2)
[1959] The system according to claim 1, which provides videos to stimulate the memory of dementia patients.
[1960] (Claim 3)
[1961] The system of claim 1, which uses artificial intelligence techniques to generate the video.
[1962] "Example 1"
[1963] (Claim 1)
[1964] A means for users to input past photos and episodes;
[1965] means for transmitting the input photos and episodes to a server;
[1966] means for analyzing the received data by the server using an artificial intelligence model;
[1967] A means for automatically generating videos based on the analysis results;
[1968] A means for storing the generated video data in cloud storage;
[1969] a means by which a user may request playback of a video;
[1970] A means for the server to receive the request and deliver the video data to the terminal;
[1971] a means by which the device plays the video;
[1972] means for analyzing and converting a user's voice input into text data;
[1973] means for the server to generate a response based on the speech input;
[1974] means for transmitting the generated response to a terminal and playing it as audio;
[1975] A system including:
[1976] (Claim 2)
[1977] The system according to claim 1, which provides videos to stimulate the memory of dementia patients.
[1978] (Claim 3)
[1979] 10. The system of claim 1, wherein an artificial intelligence model is used to generate the video.
[1980] "Application Example 1"
[1981] (Claim 1)
[1982] A means of collecting historical photographs and related anecdotes;
[1983] a means for generating a video based on the collected data;
[1984] A means for providing the generated video to dementia patients;
[1985] means for analyzing a user's voice input;
[1986] means for generating a response based on the analyzed voice data;
[1987] means for providing the generated response to a user;
[1988] A means for providing memory-based video and interactive interactions within a physical store;
[1989] A system including:
[1990] (Claim 2)
[1991] 10. The system of claim 1, which provides videos and interactive dialogues to stimulate memory and rehabilitate dementia patients.
[1992] (Claim 3)
[1993] The system of claim 1, which uses artificial intelligence techniques for generating video and analyzing audio data.
[1994] "Example 2: Combining Emotion Engines"
[1995] (Claim 1)
[1996] a means for a user to collect historical photographs and associated memories;
[1997] means for transmitting the collected data from the terminal to a server;
[1998] A means for analyzing the data received by the server using a natural language processing module to extract basic information about the episode;
[1999] A means for generating a video using a generative AI model based on the extracted information;
[2000] A means for storing the generated video in cloud storage;
[2001] a means by which a user may request playback of a video;
[2002] A means for the server to deliver video data to the terminal in response to a request;
[2003] a means for the terminal to play the video to the user;
[2004] means for analyzing and converting a user's voice input into text data;
[2005] means for transmitting the converted voice data to a server and generating an appropriate response;
[2006] means for transmitting response data generated by the server to the terminal and playing it back;
[2007] A means for analyzing the user's voice and facial expressions using an emotion engine to determine their emotional state;
[2008] a means for adjusting the content of subsequent content provided based on the emotional state;
[2009] A system including:
[2010] (Claim 2)
[2011] 10. The system of claim 1, which automatically adjusts a personalized rehabilitation plan with emotion recognition capabilities to stimulate memory and rehabilitate dementia patients.
[2012] (Claim 3)
[2013] The system of claim 1, which uses a generative AI model to generate the video.
[2014] "Application example 2 when combining emotion engines"
[2015] (Claim 1)
[2016] A means of collecting historical photographs and related memoirs;
[2017] a means for generating a video based on the collected data;
[2018] A means for providing the generated video to dementia patients;
[2019] means for analyzing a user's voice input;
[2020] means for generating a response based on the analyzed voice data;
[2021] means for providing the generated response to a user;
[2022] A means for generating a video that reproduces a work procedure based on the collected data;
[2023] means for recognizing a user's emotional state and adjusting video content and response content based on the user's emotional state;
[2024] A system including:
[2025] (Claim 2)
[2026] The system according to claim 1, which provides videos to stimulate the memory of dementia patients.
[2027] (Claim 3)
[2028] The system of claim 1, which uses artificial intelligence techniques to generate the video. [Explanation of symbols]
[2029] 10, 210, 310, 410 Data Processing Systems 1...
Claims
1. A means of collecting historical photographs and related memoirs; a means for generating a video based on the collected data; A means for providing the generated video to dementia patients; means for analyzing a user's voice input; means for generating a response based on the analyzed voice data; means for providing the generated response to a user; A system including:
2. The system according to claim 1, wherein videos are provided to stimulate the memory of dementia patients.
3. The system of claim 1, wherein the video is generated using artificial intelligence techniques.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A