System
The system addresses grief and separation by allowing users to interact with virtual representations of loved ones and pets, using data analysis and learning models to recreate memories and enhance emotional well-being.
Patent Information
- Application Number
- JP2024137276
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Individuals grieving the loss of loved ones or pets experience deep sadness and loneliness, and those separated due to illness or other reasons worry about their loved ones, necessitating methods to restore peace of mind and increase hope for life.
A system that allows users to upload data, analyze it using natural language processing and image recognition, build a learning model, and engage in real-time interactive sessions with a virtual representation of their beloved family or pets, enabling emotional stability and happiness through recreated memories.
Enables users to have emotionally stable and happy interactions with past loved ones and pets in a virtual space, providing a sense of happiness and emotional support.
Smart Images

Figure 2026034155000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] People who lose their beloved family members or pets are often plagued by deep sadness and loneliness. Furthermore, when people leave their families behind due to illness or other reasons, they worry about their loved ones. To address these issues, there is a need for methods to restore peace of mind and a sense of happiness, and increase hope for life. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for a user to upload data, a means for a server to receive and store the uploaded data, a means for the server to analyze the stored data, a means for the server to build a learning model from the analyzed data, a means for a user to start an interactive session, a means for the server to generate a response based on input from the user through the interactive session, and a means for a terminal to display the generated response, allowing a user to have a real-time two-way conversation with a beloved family member or pet from the past. This allows the user to recreate memories with their beloved in a virtual space, resulting in emotional stability and a sense of happiness.
[0006] "User" means an individual or end-user who uses the Service to upload data and initiate interactive sessions.
[0007] "Data" means information that a user can upload, including in the form of letters, photos, videos, audio, etc.
[0008] "Upload" is the act of a user sending data to a server and having it stored.
[0009] A "server" is a computer in a system that receives, stores, analyzes, builds learning models, and generates responses.
[0010] "Storage" is the process of permanently retaining uploaded data on a storage medium.
[0011] "Analysis" is the process of processing stored data and extracting meaning and patterns.
[0012] "Data analysis" is the process of extracting meaningful information from uploaded data using natural language processing and image recognition techniques.
[0013] A "learning model" is an algorithm or program that uses machine learning techniques to learn patterns and features extracted from analytical data.
[0014] "Construction" is the process of designing, training, and finally bringing a learning model into a usable state.
[0015] "Interaction Session" refers to a real-time interaction in which a User engages in a virtual interaction with a past family member or pet through the Service.
[0016] A "response" is a reply or reaction to a user's input that a server generates during an interactive session.
[0017] "Display" is the act of the terminal providing the generated response to the user visually or audibly.
[0018] "Terminal" means a device used by a user to conduct an interactive session, including a PC, smartphone, AR / VR device, etc. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The system according to the present invention is a system that allows a user to recreate memories of past loved ones and pets in a virtual space and engage in two-way conversations in real time. Specific embodiments of the system are described below.
[0041] System Overview
[0042] The system mainly consists of the following components:
[0043] A means for users to upload data
[0044] The means by which the server receives, stores, and analyzes uploaded data
[0045] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[0046] A means for the terminal to display the generated response
[0047] This allows users to recreate memories with loved ones in a virtual space, providing emotional stability and a sense of happiness.
[0048] Program processing flow
[0049] Uploading and saving data
[0050] A user uses their device (e.g., PC, smartphone) to access a specific web page or application and upload data such as letters, photos, videos, etc. The uploaded data is received by the server and saved in temporary storage. The server analyzes the data's metadata (e.g., file type, size, content) and saves it appropriately.
[0051] Data analysis and learning
[0052] The server then begins the process of analyzing the stored data. For text data (e.g., letters, conversation logs), natural language processing (NLP) is used to extract meaning and patterns, and for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. Based on this analyzed data, the server builds a learning model.
[0053] Starting an interactive session
[0054] When a user logs in to the service and starts an interactive session, the server loads the previously constructed learning model. For each message or question entered by the user, the server generates an appropriate response and sends this response to the user's device.
[0055] Viewing the response
[0056] The user's device (e.g., PC, smartphone, AR / VR device) displays the response received from the server in real time. AR / VR devices, in particular, generate a virtual environment and present visual and auditory responses to allow the user to interact deeply with the experience.
[0057] Specific examples
[0058] 1. When a user uploads information about their dog, Pochi:
[0059] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[0060] The server receives this data, temporarily stores it, and analyzes and stores the metadata for each file.
[0061] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[0062] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality (energetic and curious) and behavioral patterns (responds to specific commands).
[0063] When the user says, "Pochi, let's play a lot today," the server generates a response saying, "Yes! Let's go to the park!" and sends it to the user's device.
[0064] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[0065] This system allows users to experience emotional stability and happiness through virtual interactions with their beloved family and pets.
[0066] ---
[0067] The processing flow will be explained below.
[0068] Step 1:
[0069] A user uses his / her own terminal (e.g., PC, smartphone) to access a specific web page or application, which provides an interface for uploading data such as letters, photos, and videos.
[0070] Step 2:
[0071] Using the upload interface, users select data related to the deceased person or pet (photos, videos, text, etc.) and initiate the upload, which sends the data to the server.
[0072] Step 3:
[0073] The server receives the uploaded data and stores it in temporary storage. For each file received, it analyzes its metadata (file type, size, content, etc.) and stores it in a database.
[0074] Step 4:
[0075] The server then begins the process of analyzing the stored data. Text data (e.g., letters, conversation logs) is analyzed using natural language processing (NLP), while images and videos are analyzed using image recognition technology. This allows for the extraction of characteristics and behavioral patterns of the pet or deceased person.
[0076] Step 5:
[0077] The server builds a machine learning model based on the analyzed data. Specifically, conversation patterns extracted from text data and behavioral patterns extracted from image and video data are input into the model and it learns.
[0078] Step 6:
[0079] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[0080] Step 7:
[0081] The user enters text into the dialogue window and sends it. For example, they send a message like "Pochi, let's play lots today."
[0082] Step 8:
[0083] The server receives input from the user and uses the learning model to generate an appropriate response, such as "Yes! Let's go to the park!", based on the personality and behavioral patterns of past pets and deceased pets.
[0084] Step 9:
[0085] The server sends the generated response to the user's device, which receives and displays the response. In particular, if the user is using an AR / VR device, a virtual environment is generated and the response is presented visually and audibly.
[0086] Step 10:
[0087] The user receives a response and continues the two-way dialogue by entering questions or messages again. This interactive process allows the user to recreate memories with their beloved family and pets in a virtual space, resulting in emotional stability and a sense of happiness.
[0088] By the above steps, the user can smoothly have a virtual conversation with a loved one, and the object of the present invention can be achieved.
[0089] Example 1
[0090] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0091] In a system where users can recreate memories of their precious family and pets in a virtual space and engage in real-time interactive dialogue, the challenge is to provide a system that allows users to deepen their emotional interactions and feel a sense of happiness. It is also necessary to provide a means for users to intuitively experience responses visually and audibly.
[0092] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0093] In this invention, the server includes means for a user to upload data, means for the server to receive and store the uploaded data, means for the server to analyze the stored data, means for the server to build a learning model from the analyzed data, means for a user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and means for the terminal to generate a virtual environment and present the response visually and audibly. This allows the user to recreate memories with loved ones in a virtual space and engage in two-way dialogue in real time, thereby achieving emotional stability and a sense of happiness.
[0094] "User" refers to an individual who uses the system to recreate memories of their beloved family and pets in a virtual space and engage in real-time interactive dialogue.
[0095] "Server" refers to the central computing resource for receiving, storing, analyzing, and building learning models of data, as well as generating interactive responses.
[0096] "Data" refers to information such as letters, photos, videos, etc. uploaded by users.
[0097] "Upload" refers to the act of a user sending data from their own device to a server.
[0098] "Means for storage" refers to the system's ability to store uploaded data in storage in an appropriate format.
[0099] "Means of analysis" refers to the function of understanding the content of stored data using natural language processing, image recognition technology, etc., and extracting the necessary information.
[0100] "Learning model" refers to an algorithmic model that is generated based on analyzed data and that generates dialogue responses.
[0101] An "interaction session" refers to a series of interactions in which a user has a two-way dialogue with a system in real time.
[0102] "Response" refers to a reply message generated by a server based on user input.
[0103] "Terminal" refers to digital devices used by users, such as PCs, smartphones, and AR / VR devices.
[0104] "Virtual environment" refers to a visually and aurally recreated space generated on a device for users to interact with deeply and emotionally.
[0105] The system of the present invention allows users to recreate memories of their precious family and pets in a virtual space and engage in real-time interactive conversations. To realize this system, the following elements are required:
[0106] User upload of data
[0107] Users use their own devices (e.g., PCs or smartphones) to access a web page or application and upload data such as letters, photos, and videos. Specifically, they use a browser (e.g., GOOGLE CHROME (registered trademark), Safari) or a dedicated application (e.g., iOS, ANDROID (registered trademark) app). Users send this data to the upload page by dragging and dropping it.
[0108] Receiving and storing data by the server
[0109] The server receives and stores data uploaded by users. This is done using a database management system (e.g., MySQL (registered trademark), PostgreSQL) or a file storage service (e.g., Amazon S3). The server analyzes the metadata of the uploaded data (e.g., file type, size, content), and categorizes and stores it appropriately.
[0110] Data analysis and learning model construction
[0111] The server analyzes the stored data and uses natural language processing (NLP) to extract meaning and patterns from text data (such as letters and conversation logs). This analysis uses NLP libraries (e.g., spaCy, NLTK, BERT model). On the other hand, for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. This analysis uses deep learning frameworks (e.g., TENSORFLOW (registered trademark), PyTorch). Based on the analyzed data, the server builds a learning model and prepares to generate dialogue responses.
[0112] Initiating an interactive session and generating responses
[0113] A user logs in to the service and starts an interactive session. The server loads the saved learning model and generates an appropriate response for each message or question from the user using a generative AI model (e.g., GPT-3 (registered trademark), BERT). The generated response is sent to the user's device in real time.
[0114] Displaying the response and generating the virtual environment
[0115] The device displays the response received from the server in real time. When the user checks the response using a PC or smartphone, if an AR / VR device (e.g., Oculus Rift or HTC Vive) is used, a game engine such as Unity or Unreal Engine is used to generate a visually and aurally virtual environment.
[0116] Specific examples
[0117] 1. When a user uploads information about their dog, Pochi:
[0118] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[0119] The server receives this data, analyzes the metadata, and stores it in temporary storage.
[0120] The server analyzes the text of the letter using natural language processing and the photos and videos using image recognition technology.
[0121] The server learns Pochi's characteristics and behavioral patterns from the analysis data and builds a learning model.
[0122] During an interactive session, the user might say, "Pochi, let's play a lot today," and the server would generate and send a response saying, "Yes! Let's go to the park!"
[0123] The user's device displays this response in a virtual environment, allowing the user to relive their memories with Pochi in real time.
[0124] Prompt Sentence Examples
[0125] "Create a scenario for going for a walk based on the characteristics and behavior patterns of my beloved dog Pochi."
[0126] "Enter your Mother's Day letter and generate a response from your mother."
[0127] This system allows users to experience emotional stability and happiness through virtual interactions with their beloved family and pets.
[0128] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0129] Step 1: User uploads data
[0130] Input: A user uses a PC or smartphone to access a specific web page or application.
[0131] How it works: Users drag and drop data files such as photos, videos, letters, etc. onto the upload page.
[0132] Output: The user's device temporarily stores the data and prepares it to be sent to the server.
[0133] Step 2: Server receives and stores data
[0134] Input: Data sent from the user's device.
[0135] How it works: The server receives the data and stores it in a database management system (e.g., MySQL, PostgreSQL) or file storage service (e.g., Amazon S3). The server analyzes the data's metadata (e.g., file type, size, content) and categorizes and stores it appropriately.
[0136] Output: Data is stored on the server and metadata is registered in the database.
[0137] Step 3: Analyzing the text data
[0138] Input: Text data such as letters and conversation logs stored on the server.
[0139] How it works: The server uses natural language processing (NLP) libraries (e.g., spaCy, NLTK, BERT models) to extract meaning and patterns from text data, including tokenization, syntactic analysis, and semantic analysis.
[0140] Output: Extracted semantic and pattern data.
[0141] Step 4: Image and video data analysis
[0142] Input: Image and video data such as photos and videos stored on the server.
[0143] How it works: The server uses deep learning frameworks (e.g., TensorFlow, PyTorch) to analyze image and video data. Specific processes include object detection, feature extraction, and behavioral pattern recognition.
[0144] Output: Analyzed feature and behavioral pattern data.
[0145] Step 5: Building a learning model
[0146] Input: Analysis results of text data and image / video data.
[0147] How it works: The server builds a learning model using a generative AI model (e.g., GPT-3, BERT) based on the analyzed data.
[0148] Output: The constructed learning model.
[0149] Step 6: Prepare for an interactive session
[0150] Input: User login information.
[0151] How it works: A user logs in to the service and begins an interactive session. The server loads a saved training model.
[0152] Output: An interactive session is ready.
[0153] Step 7: Running an interactive session
[0154] Input: User's message or question.
[0155] How it works: The server uses a generative AI model to generate an appropriate response based on the user's input.
[0156] Output: Sends the server-generated response to the user's terminal.
[0157] Step 8: View the response
[0158] Input: The response sent by the server.
[0159] How it works: The user's device displays the response in real time.
[0160] Output: User sees response on terminal.
[0161] Step 9: Generate a Virtual Environment (for AR / VR devices)
[0162] Input: The response sent by the server.
[0163] How it works: The device uses a game engine such as Unity or Unreal Engine to generate a visually and aurally virtual environment.
[0164] Output: The user experiences the virtual interaction through the AR / VR device.
[0165] (Application example 1)
[0166] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0167] Conventional virtual shopping experiences have been little more than a means for users to simply purchase goods, making it difficult for them to experience emotional satisfaction or happiness. Furthermore, it is difficult for users to relive and share memories of their precious family and pets in the virtual space, resulting in insufficient emotional support. To address these issues, a system is needed that allows users to have an emotionally rich experience in the virtual space.
[0168] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0169] In this invention, the server includes means for a user to upload data, means for the server to receive and store the uploaded data, means for the server to analyze the stored data, means for the server to build a learning model from the analyzed data, means for the user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and means for the virtual assistant to act as a guide when the user has a product purchasing experience in a virtual space. This allows the user to share a shopping experience with their beloved family or pet from the past in the virtual space, and enables them to receive emotional support and a sense of happiness.
[0170] "Means for users to upload data" refers to an interface that allows users to send data such as photos, videos, and text to a server via the Internet.
[0171] "Means by which the server receives and stores uploaded data" refers to a system for receiving and securely storing data sent by users on the server.
[0172] "Means by which the server analyzes the data stored on it" refers to the algorithms and programs that analyze the data stored on the server and extract and understand specific information.
[0173] "Means for the server to build a learning model from the analyzed data" refers to the function of training an artificial intelligence or machine learning model based on the analyzed data, enabling future predictions and response generation.
[0174] "Means by which a user initiates an interactive session" refers to the interface by which a user logs into the system and begins an interactive session.
[0175] "Means for the server to generate responses based on input from the user throughout an interactive session" refers to functionality for generating and providing appropriate responses to questions and messages from the user during an interactive session.
[0176] "Means for the terminal to display the response generated" refers to the function for displaying the response sent from the server on the terminal used by the user (such as a smartphone or VR device).
[0177] "Means for a virtual assistant to act as a guide when a user is making a purchase in a virtual space" refers to a function in which a virtual assistant generated based on analytical data provides guidance and advice when a user is making a purchase in a virtual space.
[0178] The system according to the present invention allows users to recreate memories of their precious family and pets in a virtual space, and also allows users to enjoy a shopping experience. Specific embodiments of the system will be described below.
[0179] System Overview
[0180] The system mainly consists of the following components:
[0181] How users upload data
[0182] The means by which the server receives, stores, and analyzes uploaded data
[0183] A means for the server to build a learning model from the analytics data and generate responses based on user input through interactive sessions
[0184] A means for the terminal to display the generated response
[0185] A method for virtual assistants to guide users through a virtual shopping experience
[0186] Program processing overview
[0187] Uploading data
[0188] Users use their smartphones or PCs to upload their photos, videos, letters, and other memorable data to the application, which requires an internet connection and an interface that supports data transmission.
[0189] Data analysis
[0190] The server receives the uploaded data and temporarily stores it. The stored data is then analyzed. Natural language processing (NLP) techniques are used to extract meaning and patterns from text data, while image and video processing is performed using image recognition techniques. The main software tools used include Hugging Face Transformers and OpenCV.
[0191] Building a learning model
[0192] The server then builds a learning model based on the analyzed data. For example, it learns the characteristics of pets and family members based on past conversation logs and photos, and models their emotions and behavioral patterns. This process uses deep learning libraries such as TensorFlow and PyTorch.
[0193] Starting an interactive session
[0194] When a user logs in to the application and starts an interaction session, the server loads the previously built learning model. Based on the user's questions and comments, the virtual assistant generates an appropriate response and sends it to the user's device.
[0195] Viewing the response
[0196] The user's device, such as a smartphone or VR device, displays the response received from the server in real time. In particular, VR devices provide visual and auditory responses so that the user can enjoy shopping with a virtual assistant.
[0197] Specific examples
[0198] For example, if a user wants to relive a memory of their beloved dog:
[0199] 1. Users upload photos, videos and letters of their pet dogs to the application.
[0200] 2. The server receives and analyzes this data to learn your dog's characteristics and behavioral patterns.
[0201] 3. A pet dog is generated as a virtual assistant, and when the user asks, "Pochi, what should we buy today?", the virtual assistant Pochi responds, "I want a new toy!"
[0202] 4. This response is displayed on smartphones and VR devices, allowing users to enjoy shopping with Pochi in a virtual space.
[0203] Prompt Sentence Examples
[0204] "Take a photo and upload it to the app. Also, enter a message about your pet's memory."
[0205] "Visit a virtual store and ask the assistant questions, such as, 'What products do you recommend?'"
[0206] In this way, the system allows users to enjoy an emotionally rich shopping experience in a virtual space.
[0207] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0208] Step 1:
[0209] Uploading data
[0210] User action: The user uploads data such as photos, videos, and text from their smartphone or PC to the application.
[0211] Input: Files uploaded by users (photos, videos, text data)
[0212] Server action: Receives uploaded data and stores it in temporary storage.
[0213] Output: Data stored in the server storage
[0214] Step 2:
[0215] Data analysis
[0216] Server operation: The server analyzes the stored photos, videos, and text data.
[0217] Input: Data stored in the server storage
[0218] Natural Language Processing (NLP): Processing text data to extract meaning and patterns, using techniques such as Hugging Face Transformers.
[0219] Image Recognition: Process images and videos to extract important features and behavioral patterns. Uses OpenCV, etc.
[0220] Output: Analyzed data (text semantics, image and video features)
[0221] Step 3:
[0222] Building a learning model
[0223] Server operation: Builds a learning model based on the analyzed data, which trains the virtual assistant's actions and responses.
[0224] Input: Analyzed data (text semantics, image and video features)
[0225] Data Computation: Use TensorFlow and PyTorch to build learning models from analytical data.
[0226] Output: A trained learning model
[0227] Step 4:
[0228] Starting an interactive session
[0229] User Action: A user logs into an application and begins an interactive session.
[0230] Input: User login information and a request to start a session
[0231] Server operation: The server receives the user's session initiation request and loads the trained learning model.
[0232] Output: Initial response of the interactive session
[0233] Step 5:
[0234] Generating a response
[0235] User Action: The user enters a question or message.
[0236] Input: User question or message
[0237] Server behavior: Based on the user's input, the server uses a learning model to generate an appropriate response.
[0238] Data computation: A generative AI model (e.g., GPT-3) processes the prompt and generates a response.
[0239] Output: The generated response
[0240] Step 6:
[0241] Viewing the response
[0242] Terminal behavior: The user's terminal displays the response received from the server in real time.
[0243] Input: The response sent by the server
[0244] What it does: The smartphone displays text, audio, and images, while the VR device displays a virtual assistant in virtual reality.
[0245] Output: The response displayed on the user's terminal
[0246] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0247] The system of the present invention allows users to recreate memories of their beloved family and pets in a virtual space and engage in real-time interactive dialogue. Furthermore, by incorporating an emotion engine, the system recognizes the user's emotions and generates more appropriate responses, improving the user experience.
[0248] System Overview
[0249] The system mainly consists of the following components:
[0250] A means for users to upload data
[0251] The means by which the server receives, stores, and analyzes uploaded data
[0252] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[0253] A means for the terminal to display the generated response
[0254] Emotion engine that recognizes user emotions
[0255] This allows users to recreate memories with loved ones in a virtual space, providing emotional stability and a sense of happiness.
[0256] Program processing flow
[0257] Uploading and saving data
[0258] A user uses their device (e.g., PC, smartphone) to access a specific web page or application and upload data such as letters, photos, videos, etc. The uploaded data is received by the server and saved in temporary storage. The server analyzes the data's metadata (e.g., file type, size, content) and saves it appropriately.
[0259] Data analysis and learning
[0260] The server then begins the process of analyzing the stored data. For text data (e.g., letters, conversation logs), natural language processing (NLP) is used to extract meaning and patterns, and for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. Based on this analyzed data, the server builds a learning model.
[0261] Emotion Engine Operation
[0262] The server uses an emotion engine to recognize the user's emotions. The emotion engine uses natural language processing and speech recognition technologies to extract emotions from the user's text and voice inputs and determine their emotional state.
[0263] Starting an interactive session
[0264] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[0265] Generate and display the response
[0266] The user enters a message in the dialogue window using text or voice and sends it. The server receives the input from the user and recognizes the user's emotions through an emotion engine. Based on the recognized emotions, the server uses a learning model to generate a more appropriate response. The generated response is sent to the user's device and displayed. In particular, if an AR / VR device is used, a virtual environment is generated and the response is presented visually and audibly.
[0267] Specific examples
[0268] 1. When a user uploads information about their dog, Pochi:
[0269] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[0270] The server receives this data, temporarily stores it, and analyzes and stores the metadata for each file.
[0271] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[0272] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality (energetic and curious) and behavioral patterns (responds to specific commands).
[0273] When the user says, "Pochi, let's play a lot today," the server generates a response saying, "Yes! Let's go to the park!" and sends it to the user's device.
[0274] The server also recognizes the user's emotion of happiness through an emotion engine and generates a more positive response based on that.
[0275] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[0276] Through the above process, users can achieve emotional stability and a sense of happiness through virtual interactions with loved ones. The system of the present invention, in particular, can provide a more sophisticated and emotionally in line with the interactive experience by adding an emotion engine.
[0277] ---
[0278] The processing flow will be explained below.
[0279] Step 1:
[0280] A user uses his / her own terminal (e.g., PC, smartphone) to access a specific web page or application, which provides an interface for uploading data such as letters, photos, and videos.
[0281] Step 2:
[0282] The user uses the upload interface to select data related to the deceased person or pet (photos, videos, text, etc.) and initiates the upload, which is then sent to the server.
[0283] Step 3:
[0284] The server receives the uploaded data and stores it in temporary storage. At the same time, the server analyzes the metadata of each file (file type, size, content) and stores it appropriately in a database.
[0285] Step 4:
[0286] The server then begins the process of analyzing the stored data. First, for text data (such as letters or conversation logs), natural language processing (NLP) is used to analyze the content and extract various patterns and meanings. Next, for image and video data, image recognition technology is used to extract features and behavioral patterns.
[0287] Step 5:
[0288] The server trains a machine learning model based on the analyzed data, inputting conversation patterns extracted from text data and behavioral patterns extracted from image and video data into the model and applying the learning algorithm.
[0289] Step 6:
[0290] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[0291] Step 7:
[0292] The user enters text into the dialogue window and sends it. For example, they send a message like "Pochi, let's play lots today."
[0293] Step 8:
[0294] The server receives input from the user and analyzes the user's emotions through an emotion engine, which uses natural language processing techniques to extract emotions (such as joy, sadness, or anger) from the user's text.
[0295] Step 9:
[0296] The server generates an appropriate response based on the emotion analysis results using a learning model. For example, if the user's message is recognized as joyful, it generates a positive response.
[0297] Step 10:
[0298] The server sends the generated response to the user's device, which receives the response and displays it in real time. In particular, in the case of AR / VR devices, a virtual environment is generated that matches the emotion, and responses are presented visually and audibly.
[0299] Step 11:
[0300] The user can continue the conversation by entering more messages. The server receives the user's input again, analyzes the emotions through the emotion engine, and generates and displays an appropriate response. By repeating this process, the user can continue an emotionally rich virtual conversation with their beloved family or pet.
[0301] ---
[0302] Example 2
[0303] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0304] Until now, there has been no system that recreates memories of past loved ones and pets in a virtual space and allows for real-time interactive dialogue. Furthermore, there was limited technology available to recognize user emotions and generate more appropriate responses, making it difficult to improve the user experience.
[0305] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0306] In this invention, the server includes means for a user to upload information, means for the server to receive and store the uploaded information, means for the server to analyze the stored information, means for the server to build a learning model from the analyzed data, means for a user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and emotion engine means for the server to recognize the user's emotions. This allows the user to recreate memories with loved ones in a virtual space, thereby achieving emotional stability and a sense of happiness.
[0307] "User" means an individual or entity that uses the System to upload information and conduct interactive sessions.
[0308] "Information" refers to any data that users upload to the system, such as letters, photos, videos, etc.
[0309] A "means" refers to a device, process, or method employed to accomplish a particular end.
[0310] The "server" is a central system that receives, stores, analyzes, and builds learning models from information uploaded by users.
[0311] A "learning model" is an artificial intelligence model that the server builds based on the analysis data and is used to generate responses to user input.
[0312] An "interactive session" refers to a process in which a user and a system exchange information in real time.
[0313] A "response" is a reply message generated by a server in response to a user's input.
[0314] "Terminal" refers to the device (e.g., PC, smartphone, AR / VR device) that a user uses to interact with the system.
[0315] An "emotion engine" is an artificial intelligence technique used by the server to recognize the user's emotions and generate appropriate responses.
[0316] "Natural language processing" is a technique used in information analysis to extract meaning and patterns from text data.
[0317] "Image recognition technology" is a technology used in information analysis to extract features and behavioral patterns from image and video data.
[0318] The system of the present invention allows users to recreate memories of their precious family and pets in a virtual space and engage in real-time interactive dialogue. Furthermore, by combining it with an emotion engine, it is possible to recognize the user's emotions and generate more appropriate responses, thereby improving the user experience.
[0319] System Overview
[0320] The system mainly consists of the following components:
[0321] How users upload information
[0322] The means by which the server receives, stores, and analyzes the uploaded information
[0323] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[0324] A means for the terminal to display the generated response
[0325] Emotion engine that recognizes user emotions
[0326] Hardware and Software Configuration
[0327] Server: Use a high-performance server machine (e.g., AWS (registered trademark) EC2) and use an S3 bucket to store data.
[0328] Device: The user uses a PC, smartphone, or AR / VR device.
[0329] Natural Language Processing (NLP) library: Spacy is used to tokenize text data and extract meaning and patterns.
[0330] Image recognition technology: Uses OpenCV and TensorFlow to analyze features and behavioral patterns in images and videos.
[0331] Machine learning libraries: Use TensorFlow and PyTorch to generate learning models from analytical data.
[0332] Emotion Detection Library: Extracts emotions from user text and speech using the NRC Emotion Lexicon and Google® Cloud Speech-to-Text.
[0333] Program processing
[0334] Users use their devices to access web pages or apps and upload data such as letters, photos, and videos. The server receives this data and temporarily stores it in storage. The stored data is analyzed using NLP and image recognition technologies. Based on the analyzed data, the server uses machine learning libraries to generate a learning model.
[0335] When a user logs in to the system and starts an interaction session, the server loads the saved learning model and waits for user input. When the user enters a message in the interaction window using text or voice, the server receives the input and recognizes the emotion using the emotion engine. Based on the recognized emotion, the server generates an appropriate response using the generative AI model and sends it to the device. The device displays the generated response, and the response is presented visually and audibly in the virtual environment, especially when using an AR / VR device.
[0336] Specific examples
[0337] 1. When a user uploads information about their dog, Pochi:
[0338] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[0339] The server receives this data, temporarily stores it, and analyzes and stores the meta information for each.
[0340] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[0341] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality and behavioral patterns.
[0342] When the user says, "Pochi, let's play a lot today," the server generates a response, "Yes! Let's go to the park!" and sends it to the user's device.
[0343] The server recognizes the user's joy through an emotion engine and generates a more positive response.
[0344] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[0345] Prompt Sentence Examples
[0346] "Please tell me how to recognize user emotions and generate appropriate responses."
[0347] "Please explain in detail the process of analyzing the uploaded data."
[0348] "Please tell us specifically what experience you would like to have through an interaction session with your dog."
[0349] Through the above process, users can achieve emotional stability and a sense of happiness through virtual conversations with loved ones. By using an emotion engine, this system provides a more emotionally in line with the user's needs.
[0350] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0351] Step 1:
[0352] Users access a specified web page or app from their device (PC or smartphone) and upload information such as letters, photos, and videos.
[0353] Specific behavior: The user clicks the upload button, selects a file from the file selection dialog, and presses the send button.
[0354] Input: Data uploaded by users, such as letters, photos, videos, etc.
[0355] Output: Data uploaded to the server.
[0356] Step 2:
[0357] The server receives the uploaded information and temporarily stores it in storage.
[0358] Specific behavior: The server receives the HTTP request, saves the file in a temporary directory on the server side, and returns an upload completion message.
[0359] Input: Uploaded information.
[0360] Output: Information stored on the server.
[0361] Step 3:
[0362] The server analyzes the information's metadata (file type, size, content, etc.) and stores it in a database in an appropriate format.
[0363] Specific operation: The server checks the file format, extracts meta information for each file (e.g., 'photo.jpg', size: 2MB, creation date: 2023-10-15), and records it in the database.
[0364] Input: Temporarily stored information.
[0365] Output: Metadata and information stored in a database.
[0366] Step 4:
[0367] The server analyzes the stored information, specifically using natural language processing (NLP) for text data and image recognition technology for images and videos.
[0368] How it works: The server uses NLP libraries (e.g., Spacy) to tokenize text and extract important meaning and patterns. For images and videos, it uses OpenCV, TensorFlow, etc. to analyze facial recognition and behavioral patterns.
[0369] Input: Information stored in a database.
[0370] Output: Parsed data.
[0371] Step 5:
[0372] The server constructs a learning model based on the analysis data.
[0373] Specific operation: The server uses machine learning libraries (e.g., TensorFlow, PyTorch) to learn individual features from the analysis data and generate a unique learning model.
[0374] Input: Parsed data.
[0375] Output: The trained model.
[0376] Step 6:
[0377] The server uses an emotion engine to recognize the user's emotions. The emotion engine uses natural language processing and speech recognition technologies to extract and distinguish emotions from the user's text and voice input.
[0378] Specific behavior: The server calls an emotion detection library (e.g., NRC Emotion Lexicon, Google Cloud Speech-to-Text) to extract emotion labels (e.g., joy, sadness, anger) from text or audio.
[0379] Input: User text and voice input.
[0380] Output: User emotion label.
[0381] Step 7:
[0382] A user logs into a system and begins an interactive session.
[0383] Specific operation: The user enters the username and password on the login page and presses the login button. The server performs authentication, and if successful, displays the start screen of the interactive session.
[0384] Enter your username and password.
[0385] Output: The start screen of an interactive session after successful authentication.
[0386] Step 8:
[0387] The server loads the saved learning model and waits for user input.
[0388] Specific operation: The server loads the appropriate learning model into memory based on the user profile and waits for data input from the user.
[0389] Input: The training model.
[0390] Output: Server in standby state.
[0391] Step 9:
[0392] The user enters a message in the dialogue window and sends it.
[0393] Specific behavior: When the user enters text into the dialogue window and clicks the submit button, the data is sent to the server.
[0394] Input: A text message.
[0395] Output: The text message sent to the server.
[0396] Step 10:
[0397] The server receives the user's input and recognizes the user's emotions through an emotion engine.
[0398] Specific operation: The server analyzes the input text and audio data, extracts emotion labels, and passes them to the next step.
[0399] Input: Text or audio data.
[0400] Output: Extracted emotion labels.
[0401] Step 11:
[0402] The server uses a learning model to generate an appropriate response based on the recognized emotion.
[0403] Specific operation: The server generates a response sentence using a generative AI model (e.g., GPT-3) based on the recognized emotion and stored features.
[0404] Input: emotion labels and a trained model.
[0405] Output: The generated response sentence.
[0406] Step 12:
[0407] The device displays the generated response and, especially when using an AR / VR device, presents the response visually and audibly along with the virtual environment.
[0408] Specific behavior: The server sends the generated response to the user's device, which displays it. If an AR / VR device is used, the device displays the response visually in the virtual environment and also plays an audio response.
[0409] Input: The generated response sentence.
[0410] Output: The response displayed on the user's terminal.
[0411] (Application example 2)
[0412] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0413] Currently, in physical stores, customers have few easy ways to obtain information about products or specific areas within the store, and there is a need to improve the customer experience.In addition, there is a lack of systems that can respond individually to customers' emotions, making it difficult to improve customer satisfaction.
[0414] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0415] In this invention, the server includes a means for a user to upload data, a means for the server to receive and store the uploaded data, and a means for the server to analyze the stored data. This enables information about a specified object in a virtual space to be displayed in a physical store. The server also includes a means for constructing a learning model from the analyzed data, a means for a user to start an interactive session, a means for the server to generate a response based on input from the user through the interactive session, a means for the terminal to display the generated response, and an emotion recognition means for generating a response based on the user's emotion. This enables individual responses based on the user's emotions, which is expected to improve customer satisfaction.
[0416] The "means for users to upload data" refers to an interface that allows users to use their own terminals to send data such as photos, videos, and text to the system.
[0417] "Means for the server to receive and store uploaded data" refers to the function by which the server receives data sent by the user and stores it temporarily or permanently.
[0418] "Means for the server to analyze the stored data" refers to the technology that the server uses to process the stored data and decrypt and analyze its contents.
[0419] "Means by which the server constructs a machine learning model from the analyzed data" means the process by which the server creates a machine learning model based on the analyzed data and uses the model to generate appropriate responses to future data inputs.
[0420] The "means by which a user initiates an interactive session" refers to the interface or functionality by which a user accesses the system and initiates an interaction.
[0421] "Means for a server to generate a response based on input from a user throughout an interactive session" refers to a technique in which a server receives text or voice input from a user and generates a reply based thereon.
[0422] "Means for displaying the response generated by the terminal" refers to a function that allows the user's device (smartphone, smart glasses, etc.) to visually or audibly present the response generated by the server to the user.
[0423] "Means for displaying information about a specified object in a virtual space" refers to a function that allows a user to specify a specific item or area in a physical store and visually provide the user with related information using AR technology.
[0424] "Emotion recognition means for generating responses according to the user's emotions" is a technology that analyzes the user's input (text or voice), identifies the user's emotional state, and generates an appropriate response accordingly.
[0425] The system according to the present invention is a system that allows users to recreate memories of their precious family members and pets in a virtual space and engage in two-way conversations in real time. A specific embodiment of this system is described below.
[0426] Hardware and software used
[0427] This system uses the following hardware and software:
[0428] Hardware: Smartphones, smart glasses
[0429] Software: Google Cloud Vision API, Google Cloud Natural Language API, Python, OpenCV, TextBlob, pyttsx3
[0430] Uploading and saving data
[0431] Users use their smartphones or PCs to access specific web pages or applications and upload data such as photos, videos, and text. The server receives this data and stores it temporarily or permanently. The metadata (file type, size, and content) of the temporarily stored data is also analyzed, and it is then categorized and stored appropriately.
[0432] Data analysis and learning
[0433] The server analyzes the stored data. For photos and videos, it uses the Google Cloud Vision API to analyze the characteristics of specific people and pets. For text data (such as letters and conversation logs), it uses the Google Cloud Natural Language API to extract emotions and meanings, and performs sentiment analysis using TextBlob. This allows the server to build a learning model and prepare to generate responses based on the user's specific emotions and memories.
[0434] Emotion Engine Operation
[0435] The server uses an emotion engine to recognize emotions from the user's text and voice input. It uses the Google Cloud Natural Language API and TextBlob to extract emotions from the user's text and voice input and determine their emotional state.
[0436] Initiating an interactive session and generating responses
[0437] A user logs in to the system and begins an interactive session. The server loads the saved learning model and waits for input from the user. The user inputs a message via text or voice, which the server receives and recognizes the user's emotions through an emotion engine. It then generates a response based on the user's emotions and outputs it to the terminal, either displayed or via voice synthesis.
[0438] Specific examples
[0439] If a user wants to recreate memories with their beloved dog in a virtual space, they can take the following steps.
[0440] 1. Users send data such as photos, videos, and letters of their beloved dog through a dedicated upload page.
[0441] 2. The server receives the data, analyzes the metadata, and saves it. It then uses the Google Cloud Vision API to recognize images and videos, and the Google Cloud Natural Language API to analyze the sentiment of the text data.
[0442] 3. When the user says, "Pochi, what do you want to do today?", the server generates a response, "I want to play in the park!", and displays it on the device. If the server recognizes the user's emotion of joy, it generates an even more positive response.
[0443] Prompt Sentence Examples
[0444] I want to know what memories you have of this product.
[0445] "I want you to recreate the story of when you purchased this product."
[0446] This allows users to experience emotional stability and happiness through virtual interactions with loved ones. The system realized by the invention, in particular by adding an emotion engine, can provide a more sophisticated and emotionally in line with the interactive experience.
[0447] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0448] Step 1: Upload your data
[0449] A user uses his / her own device (e.g., smartphone, PC) to access a specific web page or application and upload photos, videos, text data, etc. The input includes the files and data selected by the user. The server receives this data and temporarily stores it.
[0450] Output: Received and stored data.
[0451] Step 2: Data storage and metadata analysis
[0452] The server receives the uploaded data and stores it in temporary storage. In parallel, the server parses the metadata (file type, size, content) of each uploaded data file. As input, it contains the uploaded data file. The server parses the metadata and stores it in the appropriate format.
[0453] Output: Parsed metadata and organized data files.
[0454] Step 3: Analyze the data
[0455] The server starts the process of analyzing the stored data. For text data, it uses the Google Cloud Natural Language API to extract sentiment and meaning. For images and videos, it uses the Google Cloud Vision API to analyze features of specific people or pets. The input includes the stored data files. The results of the data analysis are feature-extracted data and sentiment analysis results.
[0456] Output: Feature extracted data and sentiment analysis results.
[0457] Step 4: Building a learning model
[0458] The server creates a machine learning model based on the analyzed data. It learns past interactions and behavioral patterns with specific people and pets, and uses the model to generate appropriate responses to future user input. The analyzed data is included as input. The server builds and stores the machine learning model.
[0459] Output: The built machine learning model.
[0460] Step 5: Starting an interactive session
[0461] A user logs into the system and accesses an interface to start an interactive session. When the user logs in, the server loads the saved learning model and waits for input from the user. The input includes the user's login operation. The server loads the learning model and enters an interactive waiting state.
[0462] Output: Training model loaded and waiting.
[0463] Step 6: Receiving User Input and Emotion Recognition
[0464] The user inputs a message using text or voice and sends it. When the server receives it, it uses Google Cloud Natural Language API or TextBlob to recognize the user's emotions. The input contains the user's text or voice data. The server performs emotion recognition and returns the result.
[0465] Output: User's emotional state and analysis results.
[0466] Step 7: Generate a response
[0467] The server generates an appropriate response based on the user's emotional state using the learning model. The input includes the user's emotional state and the learning model. The server sends the generated response to the user's device.
[0468] Output: The generated response.
[0469] Step 8: Display or speak the response
[0470] The terminal receives the response sent by the server and displays it to the user, possibly providing an audible response using a speech synthesis engine (pyttsx3). The input includes the generated response. The terminal presents the response to the user visually or audibly.
[0471] Output: The response presented to the user.
[0472] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0473] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0474] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0475] [Second embodiment]
[0476] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0477] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0478] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0479] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0480] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0481] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0482] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0483] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0484] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0485] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0486] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0487] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0488] The system according to the present invention is a system that allows a user to recreate memories of past loved ones and pets in a virtual space and engage in two-way conversations in real time. Specific embodiments of the system are described below.
[0489] System Overview
[0490] The system mainly consists of the following components:
[0491] A means for users to upload data
[0492] The means by which the server receives, stores, and analyzes uploaded data
[0493] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[0494] A means for the terminal to display the generated response
[0495] This allows users to recreate memories with loved ones in a virtual space, providing emotional stability and a sense of happiness.
[0496] Program processing flow
[0497] Uploading and saving data
[0498] A user uses their device (e.g., PC, smartphone) to access a specific web page or application and upload data such as letters, photos, videos, etc. The uploaded data is received by the server and saved in temporary storage. The server analyzes the data's metadata (e.g., file type, size, content) and saves it appropriately.
[0499] Data analysis and learning
[0500] The server then begins the process of analyzing the stored data. For text data (e.g., letters, conversation logs), natural language processing (NLP) is used to extract meaning and patterns, and for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. Based on this analyzed data, the server builds a learning model.
[0501] Starting an interactive session
[0502] When a user logs in to the service and starts an interactive session, the server loads the previously constructed learning model. For each message or question entered by the user, the server generates an appropriate response and sends this response to the user's device.
[0503] Viewing the response
[0504] The user's device (e.g., PC, smartphone, AR / VR device) displays the response received from the server in real time. AR / VR devices, in particular, generate a virtual environment and present visual and auditory responses to allow the user to interact deeply with the experience.
[0505] Specific examples
[0506] 1. When a user uploads information about their dog, Pochi:
[0507] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[0508] The server receives this data, temporarily stores it, and analyzes and stores the metadata for each file.
[0509] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[0510] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality (energetic and curious) and behavioral patterns (responds to specific commands).
[0511] When the user says, "Pochi, let's play a lot today," the server generates a response saying, "Yes! Let's go to the park!" and sends it to the user's device.
[0512] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[0513] This system allows users to experience emotional stability and happiness through virtual interactions with their beloved family and pets.
[0514] ---
[0515] The processing flow will be explained below.
[0516] Step 1:
[0517] A user uses his / her own terminal (e.g., PC, smartphone) to access a specific web page or application, which provides an interface for uploading data such as letters, photos, and videos.
[0518] Step 2:
[0519] Using the upload interface, users select data related to the deceased person or pet (photos, videos, text, etc.) and initiate the upload, which sends the data to the server.
[0520] Step 3:
[0521] The server receives the uploaded data and stores it in temporary storage. For each file received, it analyzes its metadata (file type, size, content, etc.) and stores it in a database.
[0522] Step 4:
[0523] The server then begins the process of analyzing the stored data. Text data (e.g., letters, conversation logs) is analyzed using natural language processing (NLP), while images and videos are analyzed using image recognition technology. This allows for the extraction of characteristics and behavioral patterns of the pet or deceased person.
[0524] Step 5:
[0525] The server builds a machine learning model based on the analyzed data. Specifically, conversation patterns extracted from text data and behavioral patterns extracted from image and video data are input into the model and it learns.
[0526] Step 6:
[0527] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[0528] Step 7:
[0529] The user enters text into the dialogue window and sends it. For example, they send a message like "Pochi, let's play lots today."
[0530] Step 8:
[0531] The server receives input from the user and uses the learning model to generate an appropriate response, such as "Yes! Let's go to the park!", based on the personality and behavioral patterns of past pets and deceased pets.
[0532] Step 9:
[0533] The server sends the generated response to the user's device, which receives and displays the response. In particular, if the user is using an AR / VR device, a virtual environment is generated and the response is presented visually and audibly.
[0534] Step 10:
[0535] The user receives a response and continues the two-way dialogue by entering questions or messages again. This interactive process allows the user to recreate memories with their beloved family and pets in a virtual space, resulting in emotional stability and a sense of happiness.
[0536] By the above steps, the user can smoothly have a virtual conversation with a loved one, and the object of the present invention can be achieved.
[0537] Example 1
[0538] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0539] In a system where users can recreate memories of their precious family and pets in a virtual space and engage in real-time interactive dialogue, the challenge is to provide a system that allows users to deepen their emotional interactions and feel a sense of happiness. It is also necessary to provide a means for users to intuitively experience responses visually and audibly.
[0540] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0541] In this invention, the server includes means for a user to upload data, means for the server to receive and store the uploaded data, means for the server to analyze the stored data, means for the server to build a learning model from the analyzed data, means for a user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and means for the terminal to generate a virtual environment and present the response visually and audibly. This allows the user to recreate memories with loved ones in a virtual space and engage in two-way dialogue in real time, thereby achieving emotional stability and a sense of happiness.
[0542] "User" refers to an individual who uses the system to recreate memories of their beloved family and pets in a virtual space and engage in real-time interactive dialogue.
[0543] "Server" refers to the central computing resource for receiving, storing, analyzing, and building learning models of data, as well as generating interactive responses.
[0544] "Data" refers to information such as letters, photos, videos, etc. uploaded by users.
[0545] "Upload" refers to the act of a user sending data from their own device to a server.
[0546] "Means for storage" refers to the system's ability to store uploaded data in storage in an appropriate format.
[0547] "Means of analysis" refers to the function of understanding the content of stored data using natural language processing, image recognition technology, etc., and extracting the necessary information.
[0548] "Learning model" refers to an algorithmic model that is generated based on analyzed data and that generates dialogue responses.
[0549] An "interaction session" refers to a series of interactions in which a user has a two-way dialogue with a system in real time.
[0550] "Response" refers to a reply message generated by a server based on user input.
[0551] "Terminal" refers to digital devices used by users, such as PCs, smartphones, and AR / VR devices.
[0552] "Virtual environment" refers to a visually and aurally recreated space generated on a device for users to interact with deeply and emotionally.
[0553] The system of the present invention allows users to recreate memories of their precious family and pets in a virtual space and engage in real-time interactive conversations. To realize this system, the following elements are required:
[0554] User upload of data
[0555] Users use their own devices (e.g., PCs or smartphones) to access a web page or application and upload data such as letters, photos, and videos. Specifically, they use a browser (e.g., Google Chrome or Safari) or a dedicated application (e.g., iOS or Android app). Users drag and drop the data onto the upload page.
[0556] Receiving and storing data by the server
[0557] The server receives and stores data uploaded by users. This is done using a database management system (e.g., MySQL, PostgreSQL) or a file storage service (e.g., Amazon S3). The server analyzes the metadata of the uploaded data (e.g., file type, size, and content) and categorizes and stores it appropriately.
[0558] Data analysis and learning model construction
[0559] The server analyzes the stored data and uses natural language processing (NLP) to extract meaning and patterns from text data (such as letters and conversation logs). This analysis uses NLP libraries (e.g., spaCy, NLTK, BERT model). On the other hand, for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. This analysis uses deep learning frameworks (e.g., TensorFlow, PyTorch). Based on the analyzed data, the server builds a learning model and prepares to generate dialogue responses.
[0560] Initiating an interactive session and generating responses
[0561] A user logs in to the service and starts a conversation session. The server loads the saved learning model and generates appropriate responses for each message or question from the user using a generative AI model (e.g., GPT-3, BERT). The generated responses are sent to the user's device in real time.
[0562] Displaying the response and generating the virtual environment
[0563] The device displays the response received from the server in real time. When the user checks the response using a PC or smartphone, if an AR / VR device (e.g., Oculus Rift or HTC Vive) is used, a game engine such as Unity or Unreal Engine is used to generate a visually and aurally virtual environment.
[0564] Specific examples
[0565] 1. When a user uploads information about their dog, Pochi:
[0566] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[0567] The server receives this data, analyzes the metadata, and stores it in temporary storage.
[0568] The server analyzes the text of the letter using natural language processing and the photos and videos using image recognition technology.
[0569] The server learns Pochi's characteristics and behavioral patterns from the analysis data and builds a learning model.
[0570] During an interactive session, the user might say, "Pochi, let's play a lot today," and the server would generate and send a response saying, "Yes! Let's go to the park!"
[0571] The user's device displays this response in a virtual environment, allowing the user to relive their memories with Pochi in real time.
[0572] Prompt Sentence Examples
[0573] "Create a scenario for going for a walk based on the characteristics and behavior patterns of my beloved dog Pochi."
[0574] "Enter your Mother's Day letter and generate a response from your mother."
[0575] This system allows users to experience emotional stability and happiness through virtual interactions with their beloved family and pets.
[0576] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0577] Step 1: User uploads data
[0578] Input: A user uses a PC or smartphone to access a specific web page or application.
[0579] How it works: Users drag and drop data files such as photos, videos, letters, etc. onto the upload page.
[0580] Output: The user's device temporarily stores the data and prepares it to be sent to the server.
[0581] Step 2: Server receives and stores data
[0582] Input: Data sent from the user's device.
[0583] How it works: The server receives the data and stores it in a database management system (e.g., MySQL, PostgreSQL) or file storage service (e.g., Amazon S3). The server analyzes the data's metadata (e.g., file type, size, content) and categorizes and stores it appropriately.
[0584] Output: Data is stored on the server and metadata is registered in the database.
[0585] Step 3: Analyzing the text data
[0586] Input: Text data such as letters and conversation logs stored on the server.
[0587] How it works: The server uses natural language processing (NLP) libraries (e.g., spaCy, NLTK, BERT models) to extract meaning and patterns from text data, including tokenization, syntactic analysis, and semantic analysis.
[0588] Output: Extracted semantic and pattern data.
[0589] Step 4: Image and video data analysis
[0590] Input: Image and video data such as photos and videos stored on the server.
[0591] How it works: The server uses deep learning frameworks (e.g., TensorFlow, PyTorch) to analyze image and video data. Specific processes include object detection, feature extraction, and behavioral pattern recognition.
[0592] Output: Analyzed feature and behavioral pattern data.
[0593] Step 5: Building a learning model
[0594] Input: Analysis results of text data and image / video data.
[0595] How it works: The server builds a learning model using a generative AI model (e.g., GPT-3, BERT) based on the analyzed data.
[0596] Output: The constructed learning model.
[0597] Step 6: Prepare for an interactive session
[0598] Input: User login information.
[0599] How it works: A user logs in to the service and begins an interactive session. The server loads a saved training model.
[0600] Output: An interactive session is ready.
[0601] Step 7: Running an interactive session
[0602] Input: User's message or question.
[0603] How it works: The server uses a generative AI model to generate an appropriate response based on the user's input.
[0604] Output: Sends the server-generated response to the user's terminal.
[0605] Step 8: View the response
[0606] Input: The response sent by the server.
[0607] How it works: The user's device displays the response in real time.
[0608] Output: User sees response on terminal.
[0609] Step 9: Generate a Virtual Environment (for AR / VR devices)
[0610] Input: The response sent by the server.
[0611] How it works: The device uses a game engine such as Unity or Unreal Engine to generate a visually and aurally virtual environment.
[0612] Output: The user experiences the virtual interaction through the AR / VR device.
[0613] (Application example 1)
[0614] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0615] Conventional virtual shopping experiences have been little more than a means for users to simply purchase goods, making it difficult for them to experience emotional satisfaction or happiness. Furthermore, it is difficult for users to relive and share memories of their precious family and pets in the virtual space, resulting in insufficient emotional support. To address these issues, a system is needed that allows users to have an emotionally rich experience in the virtual space.
[0616] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0617] In this invention, the server includes means for a user to upload data, means for the server to receive and store the uploaded data, means for the server to analyze the stored data, means for the server to build a learning model from the analyzed data, means for the user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and means for the virtual assistant to act as a guide when the user has a product purchasing experience in a virtual space. This allows the user to share a shopping experience with their beloved family or pet from the past in the virtual space, and enables them to receive emotional support and a sense of happiness.
[0618] "Means for users to upload data" refers to an interface that allows users to send data such as photos, videos, and text to a server via the Internet.
[0619] "Means by which the server receives and stores uploaded data" refers to a system for receiving and securely storing data sent by users on the server.
[0620] "Means by which the server analyzes the data stored on it" refers to the algorithms and programs that analyze the data stored on the server and extract and understand specific information.
[0621] "Means for the server to build a learning model from the analyzed data" refers to the function of training an artificial intelligence or machine learning model based on the analyzed data, enabling future predictions and response generation.
[0622] "Means by which a user initiates an interactive session" refers to the interface by which a user logs into the system and begins an interactive session.
[0623] "Means for the server to generate responses based on input from the user throughout an interactive session" refers to functionality for generating and providing appropriate responses to questions and messages from the user during an interactive session.
[0624] "Means for the terminal to display the response generated" refers to the function for displaying the response sent from the server on the terminal used by the user (such as a smartphone or VR device).
[0625] "Means for a virtual assistant to act as a guide when a user is making a purchase in a virtual space" refers to a function in which a virtual assistant generated based on analytical data provides guidance and advice when a user is making a purchase in a virtual space.
[0626] The system according to the present invention allows users to recreate memories of their precious family and pets in a virtual space, and also allows users to enjoy a shopping experience. Specific embodiments of the system will be described below.
[0627] System Overview
[0628] The system mainly consists of the following components:
[0629] How users upload data
[0630] The means by which the server receives, stores, and analyzes uploaded data
[0631] A means for the server to build a learning model from the analytics data and generate responses based on user input through interactive sessions
[0632] A means for the terminal to display the generated response
[0633] A method for virtual assistants to guide users through a virtual shopping experience
[0634] Program processing overview
[0635] Uploading data
[0636] Users use their smartphones or PCs to upload their photos, videos, letters, and other memorable data to the application, which requires an internet connection and an interface that supports data transmission.
[0637] Data analysis
[0638] The server receives the uploaded data and temporarily stores it. The stored data is then analyzed. Natural language processing (NLP) techniques are used to extract meaning and patterns from text data, while image and video processing is performed using image recognition techniques. The main software tools used include Hugging Face Transformers and OpenCV.
[0639] Building a learning model
[0640] The server then builds a learning model based on the analyzed data. For example, it learns the characteristics of pets and family members based on past conversation logs and photos, and models their emotions and behavioral patterns. This process uses deep learning libraries such as TensorFlow and PyTorch.
[0641] Starting an interactive session
[0642] When a user logs in to the application and starts an interaction session, the server loads the previously built learning model. Based on the user's questions and comments, the virtual assistant generates an appropriate response and sends it to the user's device.
[0643] Viewing the response
[0644] The user's device, such as a smartphone or VR device, displays the response received from the server in real time. In particular, VR devices provide visual and auditory responses so that the user can enjoy shopping with a virtual assistant.
[0645] Specific examples
[0646] For example, if a user wants to relive a memory of their beloved dog:
[0647] 1. Users upload photos, videos and letters of their pet dogs to the application.
[0648] 2. The server receives and analyzes this data to learn your dog's characteristics and behavioral patterns.
[0649] 3. A pet dog is generated as a virtual assistant, and when the user asks, "Pochi, what should we buy today?", the virtual assistant Pochi responds, "I want a new toy!"
[0650] 4. This response is displayed on smartphones and VR devices, allowing users to enjoy shopping with Pochi in a virtual space.
[0651] Prompt Sentence Examples
[0652] "Take a photo and upload it to the app. Also, enter a message about your pet's memory."
[0653] "Visit a virtual store and ask the assistant questions, such as, 'What products do you recommend?'"
[0654] In this way, the system allows users to enjoy an emotionally rich shopping experience in a virtual space.
[0655] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0656] Step 1:
[0657] Uploading data
[0658] User action: The user uploads data such as photos, videos, and text from their smartphone or PC to the application.
[0659] Input: Files uploaded by users (photos, videos, text data)
[0660] Server action: Receives uploaded data and stores it in temporary storage.
[0661] Output: Data stored in the server storage
[0662] Step 2:
[0663] Data analysis
[0664] Server operation: The server analyzes the stored photos, videos, and text data.
[0665] Input: Data stored in the server storage
[0666] Natural Language Processing (NLP): Processing text data to extract meaning and patterns, using techniques such as Hugging Face Transformers.
[0667] Image Recognition: Process images and videos to extract important features and behavioral patterns. Uses OpenCV, etc.
[0668] Output: Analyzed data (text semantics, image and video features)
[0669] Step 3:
[0670] Building a learning model
[0671] Server operation: Builds a learning model based on the analyzed data, which trains the virtual assistant's actions and responses.
[0672] Input: Analyzed data (text semantics, image and video features)
[0673] Data Computation: Use TensorFlow and PyTorch to build learning models from analytical data.
[0674] Output: A trained learning model
[0675] Step 4:
[0676] Starting an interactive session
[0677] User Action: A user logs into an application and begins an interactive session.
[0678] Input: User login information and a request to start a session
[0679] Server operation: The server receives the user's session initiation request and loads the trained learning model.
[0680] Output: Initial response of the interactive session
[0681] Step 5:
[0682] Generating a response
[0683] User Action: The user enters a question or message.
[0684] Input: User question or message
[0685] Server behavior: Based on the user's input, the server uses a learning model to generate an appropriate response.
[0686] Data computation: A generative AI model (e.g., GPT-3) processes the prompt and generates a response.
[0687] Output: The generated response
[0688] Step 6:
[0689] Viewing the response
[0690] Terminal behavior: The user's terminal displays the response received from the server in real time.
[0691] Input: The response sent by the server
[0692] What it does: The smartphone displays text, audio, and images, while the VR device displays a virtual assistant in virtual reality.
[0693] Output: The response displayed on the user's terminal
[0694] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0695] The system of the present invention allows users to recreate memories of their beloved family and pets in a virtual space and engage in real-time interactive dialogue. Furthermore, by incorporating an emotion engine, the system recognizes the user's emotions and generates more appropriate responses, improving the user experience.
[0696] System Overview
[0697] The system mainly consists of the following components:
[0698] A means for users to upload data
[0699] The means by which the server receives, stores, and analyzes uploaded data
[0700] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[0701] A means for the terminal to display the generated response
[0702] Emotion engine that recognizes user emotions
[0703] This allows users to recreate memories with loved ones in a virtual space, providing emotional stability and a sense of happiness.
[0704] Program processing flow
[0705] Uploading and saving data
[0706] A user uses their device (e.g., PC, smartphone) to access a specific web page or application and upload data such as letters, photos, videos, etc. The uploaded data is received by the server and saved in temporary storage. The server analyzes the data's metadata (e.g., file type, size, content) and saves it appropriately.
[0707] Data analysis and learning
[0708] The server then begins the process of analyzing the stored data. For text data (e.g., letters, conversation logs), natural language processing (NLP) is used to extract meaning and patterns, and for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. Based on this analyzed data, the server builds a learning model.
[0709] Emotion Engine Operation
[0710] The server uses an emotion engine to recognize the user's emotions. The emotion engine uses natural language processing and speech recognition technologies to extract emotions from the user's text and voice inputs and determine their emotional state.
[0711] Starting an interactive session
[0712] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[0713] Generate and display the response
[0714] The user enters a message in the dialogue window using text or voice and sends it. The server receives the input from the user and recognizes the user's emotions through an emotion engine. Based on the recognized emotions, the server uses a learning model to generate a more appropriate response. The generated response is sent to the user's device and displayed. In particular, if an AR / VR device is used, a virtual environment is generated and the response is presented visually and audibly.
[0715] Specific examples
[0716] 1. When a user uploads information about their dog, Pochi:
[0717] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[0718] The server receives this data, temporarily stores it, and analyzes and stores the metadata for each file.
[0719] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[0720] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality (energetic and curious) and behavioral patterns (responds to specific commands).
[0721] When the user says, "Pochi, let's play a lot today," the server generates a response saying, "Yes! Let's go to the park!" and sends it to the user's device.
[0722] The server also recognizes the user's emotion of happiness through an emotion engine and generates a more positive response based on that.
[0723] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[0724] Through the above process, users can achieve emotional stability and a sense of happiness through virtual interactions with loved ones. The system of the present invention, in particular, can provide a more sophisticated and emotionally in line with the interactive experience by adding an emotion engine.
[0725] ---
[0726] The processing flow will be explained below.
[0727] Step 1:
[0728] A user uses his / her own terminal (e.g., PC, smartphone) to access a specific web page or application, which provides an interface for uploading data such as letters, photos, and videos.
[0729] Step 2:
[0730] The user uses the upload interface to select data related to the deceased person or pet (photos, videos, text, etc.) and initiates the upload, which is then sent to the server.
[0731] Step 3:
[0732] The server receives the uploaded data and stores it in temporary storage. At the same time, the server analyzes the metadata of each file (file type, size, content) and stores it appropriately in a database.
[0733] Step 4:
[0734] The server then begins the process of analyzing the stored data. First, for text data (such as letters or conversation logs), natural language processing (NLP) is used to analyze the content and extract various patterns and meanings. Next, for image and video data, image recognition technology is used to extract features and behavioral patterns.
[0735] Step 5:
[0736] The server trains a machine learning model based on the analyzed data, inputting conversation patterns extracted from text data and behavioral patterns extracted from image and video data into the model and applying the learning algorithm.
[0737] Step 6:
[0738] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[0739] Step 7:
[0740] The user enters text into the dialogue window and sends it. For example, they send a message like "Pochi, let's play lots today."
[0741] Step 8:
[0742] The server receives input from the user and analyzes the user's emotions through an emotion engine, which uses natural language processing techniques to extract emotions (such as joy, sadness, or anger) from the user's text.
[0743] Step 9:
[0744] The server generates an appropriate response based on the emotion analysis results using a learning model. For example, if the user's message is recognized as joyful, it generates a positive response.
[0745] Step 10:
[0746] The server sends the generated response to the user's device, which receives the response and displays it in real time. In particular, in the case of AR / VR devices, a virtual environment is generated that matches the emotion, and responses are presented visually and audibly.
[0747] Step 11:
[0748] The user can continue the conversation by entering more messages. The server receives the user's input again, analyzes the emotions through the emotion engine, and generates and displays an appropriate response. By repeating this process, the user can continue an emotionally rich virtual conversation with their beloved family or pet.
[0749] ---
[0750] Example 2
[0751] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0752] Until now, there has been no system that recreates memories of past loved ones and pets in a virtual space and allows for real-time interactive dialogue. Furthermore, there was limited technology available to recognize user emotions and generate more appropriate responses, making it difficult to improve the user experience.
[0753] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0754] In this invention, the server includes means for a user to upload information, means for the server to receive and store the uploaded information, means for the server to analyze the stored information, means for the server to build a learning model from the analyzed data, means for a user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and emotion engine means for the server to recognize the user's emotions. This allows the user to recreate memories with loved ones in a virtual space, thereby achieving emotional stability and a sense of happiness.
[0755] "User" means an individual or entity that uses the System to upload information and conduct interactive sessions.
[0756] "Information" refers to any data that users upload to the system, such as letters, photos, videos, etc.
[0757] A "means" refers to a device, process, or method employed to accomplish a particular end.
[0758] The "server" is a central system that receives, stores, analyzes, and builds learning models from information uploaded by users.
[0759] A "learning model" is an artificial intelligence model that the server builds based on the analysis data and is used to generate responses to user input.
[0760] An "interactive session" refers to a process in which a user and a system exchange information in real time.
[0761] A "response" is a reply message generated by a server in response to a user's input.
[0762] "Terminal" refers to the device (e.g., PC, smartphone, AR / VR device) that a user uses to interact with the system.
[0763] An "emotion engine" is an artificial intelligence technique used by the server to recognize the user's emotions and generate appropriate responses.
[0764] "Natural language processing" is a technique used in information analysis to extract meaning and patterns from text data.
[0765] "Image recognition technology" is a technology used in information analysis to extract features and behavioral patterns from image and video data.
[0766] The system of the present invention allows users to recreate memories of their precious family and pets in a virtual space and engage in real-time interactive dialogue. Furthermore, by combining it with an emotion engine, it is possible to recognize the user's emotions and generate more appropriate responses, thereby improving the user experience.
[0767] System Overview
[0768] The system mainly consists of the following components:
[0769] How users upload information
[0770] The means by which the server receives, stores, and analyzes the uploaded information
[0771] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[0772] A means for the terminal to display the generated response
[0773] Emotion engine that recognizes user emotions
[0774] Hardware and Software Configuration
[0775] Server: Use a high-performance server machine (e.g., AWS EC2) and use an S3 bucket to store data.
[0776] Device: The user uses a PC, smartphone, or AR / VR device.
[0777] Natural Language Processing (NLP) library: Spacy is used to tokenize text data and extract meaning and patterns.
[0778] Image recognition technology: Uses OpenCV and TensorFlow to analyze features and behavioral patterns in images and videos.
[0779] Machine learning libraries: Use TensorFlow and PyTorch to generate learning models from analytical data.
[0780] Emotion Detection Library: Extracts emotions from user text and speech using the NRC Emotion Lexicon and Google Cloud Speech-to-Text.
[0781] Program processing
[0782] Users use their devices to access web pages or apps and upload data such as letters, photos, and videos. The server receives this data and temporarily stores it in storage. The stored data is analyzed using NLP and image recognition technologies. Based on the analyzed data, the server uses machine learning libraries to generate a learning model.
[0783] When a user logs in to the system and starts an interaction session, the server loads the saved learning model and waits for user input. When the user enters a message in the interaction window using text or voice, the server receives the input and recognizes the emotion using the emotion engine. Based on the recognized emotion, the server generates an appropriate response using the generative AI model and sends it to the device. The device displays the generated response, and the response is presented visually and audibly in the virtual environment, especially when using an AR / VR device.
[0784] Specific examples
[0785] 1. When a user uploads information about their dog, Pochi:
[0786] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[0787] The server receives this data, temporarily stores it, and analyzes and stores the meta information for each.
[0788] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[0789] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality and behavioral patterns.
[0790] When the user says, "Pochi, let's play a lot today," the server generates a response, "Yes! Let's go to the park!" and sends it to the user's device.
[0791] The server recognizes the user's joy through an emotion engine and generates a more positive response.
[0792] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[0793] Prompt Sentence Examples
[0794] "Please tell me how to recognize user emotions and generate appropriate responses."
[0795] "Please explain in detail the process of analyzing the uploaded data."
[0796] "Please tell us specifically what experience you would like to have through an interaction session with your dog."
[0797] Through the above process, users can achieve emotional stability and a sense of happiness through virtual conversations with loved ones. By using an emotion engine, this system provides a more emotionally in line with the user's needs.
[0798] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0799] Step 1:
[0800] Users access a specified web page or app from their device (PC or smartphone) and upload information such as letters, photos, and videos.
[0801] Specific behavior: The user clicks the upload button, selects a file from the file selection dialog, and presses the send button.
[0802] Input: Data uploaded by users, such as letters, photos, videos, etc.
[0803] Output: Data uploaded to the server.
[0804] Step 2:
[0805] The server receives the uploaded information and temporarily stores it in storage.
[0806] Specific behavior: The server receives the HTTP request, saves the file in a temporary directory on the server side, and returns an upload completion message.
[0807] Input: Uploaded information.
[0808] Output: Information stored on the server.
[0809] Step 3:
[0810] The server analyzes the information's metadata (file type, size, content, etc.) and stores it in a database in an appropriate format.
[0811] Specific operation: The server checks the file format, extracts meta information for each file (e.g., 'photo.jpg', size: 2MB, creation date: 2023-10-15), and records it in the database.
[0812] Input: Temporarily stored information.
[0813] Output: Metadata and information stored in a database.
[0814] Step 4:
[0815] The server analyzes the stored information, specifically using natural language processing (NLP) for text data and image recognition technology for images and videos.
[0816] How it works: The server uses NLP libraries (e.g., Spacy) to tokenize text and extract important meaning and patterns. For images and videos, it uses OpenCV, TensorFlow, etc. to analyze facial recognition and behavioral patterns.
[0817] Input: Information stored in a database.
[0818] Output: Parsed data.
[0819] Step 5:
[0820] The server constructs a learning model based on the analysis data.
[0821] Specific operation: The server uses machine learning libraries (e.g., TensorFlow, PyTorch) to learn individual features from the analysis data and generate a unique learning model.
[0822] Input: Parsed data.
[0823] Output: The trained model.
[0824] Step 6:
[0825] The server uses an emotion engine to recognize the user's emotions. The emotion engine uses natural language processing and speech recognition technologies to extract and distinguish emotions from the user's text and voice input.
[0826] Specific behavior: The server calls an emotion detection library (e.g., NRC Emotion Lexicon, Google Cloud Speech-to-Text) to extract emotion labels (e.g., joy, sadness, anger) from text or audio.
[0827] Input: User text and voice input.
[0828] Output: User emotion label.
[0829] Step 7:
[0830] A user logs into a system and begins an interactive session.
[0831] Specific operation: The user enters the username and password on the login page and presses the login button. The server performs authentication, and if successful, displays the start screen of the interactive session.
[0832] Enter your username and password.
[0833] Output: The start screen of an interactive session after successful authentication.
[0834] Step 8:
[0835] The server loads the saved learning model and waits for user input.
[0836] Specific operation: The server loads the appropriate learning model into memory based on the user profile and waits for data input from the user.
[0837] Input: The training model.
[0838] Output: Server in standby state.
[0839] Step 9:
[0840] The user enters a message in the dialogue window and sends it.
[0841] Specific behavior: When the user enters text into the dialogue window and clicks the submit button, the data is sent to the server.
[0842] Input: A text message.
[0843] Output: The text message sent to the server.
[0844] Step 10:
[0845] The server receives the user's input and recognizes the user's emotions through an emotion engine.
[0846] Specific operation: The server analyzes the input text and audio data, extracts emotion labels, and passes them to the next step.
[0847] Input: Text or audio data.
[0848] Output: Extracted emotion labels.
[0849] Step 11:
[0850] The server uses a learning model to generate an appropriate response based on the recognized emotion.
[0851] Specific operation: The server generates a response sentence using a generative AI model (e.g., GPT-3) based on the recognized emotion and stored features.
[0852] Input: emotion labels and a trained model.
[0853] Output: The generated response sentence.
[0854] Step 12:
[0855] The device displays the generated response and, especially when using an AR / VR device, presents the response visually and audibly along with the virtual environment.
[0856] Specific behavior: The server sends the generated response to the user's device, which displays it. If an AR / VR device is used, the device displays the response visually in the virtual environment and also plays an audio response.
[0857] Input: The generated response sentence.
[0858] Output: The response displayed on the user's terminal.
[0859] (Application example 2)
[0860] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0861] Currently, in physical stores, customers have few easy ways to obtain information about products or specific areas within the store, and there is a need to improve the customer experience.In addition, there is a lack of systems that can respond individually to customers' emotions, making it difficult to improve customer satisfaction.
[0862] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0863] In this invention, the server includes a means for a user to upload data, a means for the server to receive and store the uploaded data, and a means for the server to analyze the stored data. This enables information about a specified object in a virtual space to be displayed in a physical store. The server also includes a means for constructing a learning model from the analyzed data, a means for a user to start an interactive session, a means for the server to generate a response based on input from the user through the interactive session, a means for the terminal to display the generated response, and an emotion recognition means for generating a response based on the user's emotion. This enables individual responses based on the user's emotions, which is expected to improve customer satisfaction.
[0864] The "means for users to upload data" refers to an interface that allows users to use their own terminals to send data such as photos, videos, and text to the system.
[0865] "Means for the server to receive and store uploaded data" refers to the function by which the server receives data sent by the user and stores it temporarily or permanently.
[0866] "Means for the server to analyze the stored data" refers to the technology that the server uses to process the stored data and decrypt and analyze its contents.
[0867] "Means by which the server constructs a machine learning model from the analyzed data" means the process by which the server creates a machine learning model based on the analyzed data and uses the model to generate appropriate responses to future data inputs.
[0868] The "means by which a user initiates an interactive session" refers to the interface or functionality by which a user accesses the system and initiates an interaction.
[0869] "Means for a server to generate a response based on input from a user throughout an interactive session" refers to a technique in which a server receives text or voice input from a user and generates a reply based thereon.
[0870] "Means for displaying the response generated by the terminal" refers to a function that allows the user's device (smartphone, smart glasses, etc.) to visually or audibly present the response generated by the server to the user.
[0871] "Means for displaying information about a specified object in a virtual space" refers to a function that allows a user to specify a specific item or area in a physical store and visually provide the user with related information using AR technology.
[0872] "Emotion recognition means for generating responses according to the user's emotions" is a technology that analyzes the user's input (text or voice), identifies the user's emotional state, and generates an appropriate response accordingly.
[0873] The system according to the present invention is a system that allows users to recreate memories of their precious family members and pets in a virtual space and engage in two-way conversations in real time. A specific embodiment of this system is described below.
[0874] Hardware and software used
[0875] This system uses the following hardware and software:
[0876] Hardware: Smartphones, smart glasses
[0877] Software: Google Cloud Vision API, Google Cloud Natural Language API, Python, OpenCV, TextBlob, pyttsx3
[0878] Uploading and saving data
[0879] Users use their smartphones or PCs to access specific web pages or applications and upload data such as photos, videos, and text. The server receives this data and stores it temporarily or permanently. The metadata (file type, size, and content) of the temporarily stored data is also analyzed, and it is then categorized and stored appropriately.
[0880] Data analysis and learning
[0881] The server analyzes the stored data. For photos and videos, it uses the Google Cloud Vision API to analyze the characteristics of specific people and pets. For text data (such as letters and conversation logs), it uses the Google Cloud Natural Language API to extract emotions and meanings, and performs sentiment analysis using TextBlob. This allows the server to build a learning model and prepare to generate responses based on the user's specific emotions and memories.
[0882] Emotion Engine Operation
[0883] The server uses an emotion engine to recognize emotions from the user's text and voice input. It uses the Google Cloud Natural Language API and TextBlob to extract emotions from the user's text and voice input and determine their emotional state.
[0884] Initiating an interactive session and generating responses
[0885] A user logs in to the system and begins an interactive session. The server loads the saved learning model and waits for input from the user. The user inputs a message via text or voice, which the server receives and recognizes the user's emotions through an emotion engine. It then generates a response based on the user's emotions and outputs it to the terminal, either displayed or via voice synthesis.
[0886] Specific examples
[0887] If a user wants to recreate memories with their beloved dog in a virtual space, they can take the following steps.
[0888] 1. Users send data such as photos, videos, and letters of their beloved dog through a dedicated upload page.
[0889] 2. The server receives the data, analyzes the metadata, and saves it. It then uses the Google Cloud Vision API to recognize images and videos, and the Google Cloud Natural Language API to analyze the sentiment of the text data.
[0890] 3. When the user says, "Pochi, what do you want to do today?", the server generates a response, "I want to play in the park!", and displays it on the device. If the server recognizes the user's emotion of joy, it generates an even more positive response.
[0891] Prompt Sentence Examples
[0892] I want to know what memories you have of this product.
[0893] "I want you to recreate the story of when you purchased this product."
[0894] This allows users to experience emotional stability and happiness through virtual interactions with loved ones. The system realized by the invention, in particular by adding an emotion engine, can provide a more sophisticated and emotionally in line with the interactive experience.
[0895] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0896] Step 1: Upload your data
[0897] A user uses his / her own device (e.g., smartphone, PC) to access a specific web page or application and upload photos, videos, text data, etc. The input includes the files and data selected by the user. The server receives this data and temporarily stores it.
[0898] Output: Received and stored data.
[0899] Step 2: Data storage and metadata analysis
[0900] The server receives the uploaded data and stores it in temporary storage. In parallel, the server parses the metadata (file type, size, content) of each uploaded data file. As input, it contains the uploaded data file. The server parses the metadata and stores it in the appropriate format.
[0901] Output: Parsed metadata and organized data files.
[0902] Step 3: Analyze the data
[0903] The server starts the process of analyzing the stored data. For text data, it uses the Google Cloud Natural Language API to extract sentiment and meaning. For images and videos, it uses the Google Cloud Vision API to analyze features of specific people or pets. The input includes the stored data files. The results of the data analysis are feature-extracted data and sentiment analysis results.
[0904] Output: Feature extracted data and sentiment analysis results.
[0905] Step 4: Building a learning model
[0906] The server creates a machine learning model based on the analyzed data. It learns past interactions and behavioral patterns with specific people and pets, and uses the model to generate appropriate responses to future user input. The analyzed data is included as input. The server builds and stores the machine learning model.
[0907] Output: The built machine learning model.
[0908] Step 5: Starting an interactive session
[0909] A user logs into the system and accesses an interface to start an interactive session. When the user logs in, the server loads the saved learning model and waits for input from the user. The input includes the user's login operation. The server loads the learning model and enters an interactive waiting state.
[0910] Output: Training model loaded and waiting.
[0911] Step 6: Receiving User Input and Emotion Recognition
[0912] The user inputs a message using text or voice and sends it. When the server receives it, it uses Google Cloud Natural Language API or TextBlob to recognize the user's emotions. The input contains the user's text or voice data. The server performs emotion recognition and returns the result.
[0913] Output: User's emotional state and analysis results.
[0914] Step 7: Generate a response
[0915] The server generates an appropriate response based on the user's emotional state using the learning model. The input includes the user's emotional state and the learning model. The server sends the generated response to the user's device.
[0916] Output: The generated response.
[0917] Step 8: Display or speak the response
[0918] The terminal receives the response sent by the server and displays it to the user, possibly providing an audible response using a speech synthesis engine (pyttsx3). The input includes the generated response. The terminal presents the response to the user visually or audibly.
[0919] Output: The response presented to the user.
[0920] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0921] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0922] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0923] [Third embodiment]
[0924] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0925] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0926] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0927] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0928] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0929] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0930] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0931] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0932] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0933] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0934] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0935] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0936] The system according to the present invention is a system that allows a user to recreate memories of past loved ones and pets in a virtual space and engage in two-way conversations in real time. Specific embodiments of the system are described below.
[0937] System Overview
[0938] The system mainly consists of the following components:
[0939] A means for users to upload data
[0940] The means by which the server receives, stores, and analyzes uploaded data
[0941] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[0942] A means for the terminal to display the generated response
[0943] This allows users to recreate memories with loved ones in a virtual space, providing emotional stability and a sense of happiness.
[0944] Program processing flow
[0945] Uploading and saving data
[0946] A user uses their device (e.g., PC, smartphone) to access a specific web page or application and upload data such as letters, photos, videos, etc. The uploaded data is received by the server and saved in temporary storage. The server analyzes the data's metadata (e.g., file type, size, content) and saves it appropriately.
[0947] Data analysis and learning
[0948] The server then begins the process of analyzing the stored data. For text data (e.g., letters, conversation logs), natural language processing (NLP) is used to extract meaning and patterns, and for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. Based on this analyzed data, the server builds a learning model.
[0949] Starting an interactive session
[0950] When a user logs in to the service and starts an interactive session, the server loads the previously constructed learning model. For each message or question entered by the user, the server generates an appropriate response and sends this response to the user's device.
[0951] Viewing the response
[0952] The user's device (e.g., PC, smartphone, AR / VR device) displays the response received from the server in real time. AR / VR devices, in particular, generate a virtual environment and present visual and auditory responses to allow the user to interact deeply with the experience.
[0953] Specific examples
[0954] 1. When a user uploads information about their dog, Pochi:
[0955] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[0956] The server receives this data, temporarily stores it, and analyzes and stores the metadata for each file.
[0957] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[0958] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality (energetic and curious) and behavioral patterns (responds to specific commands).
[0959] When the user says, "Pochi, let's play a lot today," the server generates a response saying, "Yes! Let's go to the park!" and sends it to the user's device.
[0960] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[0961] This system allows users to experience emotional stability and happiness through virtual interactions with their beloved family and pets.
[0962] ---
[0963] The processing flow will be explained below.
[0964] Step 1:
[0965] A user uses his / her own terminal (e.g., PC, smartphone) to access a specific web page or application, which provides an interface for uploading data such as letters, photos, and videos.
[0966] Step 2:
[0967] Using the upload interface, users select data related to the deceased person or pet (photos, videos, text, etc.) and initiate the upload, which sends the data to the server.
[0968] Step 3:
[0969] The server receives the uploaded data and stores it in temporary storage. For each file received, it analyzes its metadata (file type, size, content, etc.) and stores it in a database.
[0970] Step 4:
[0971] The server then begins the process of analyzing the stored data. Text data (e.g., letters, conversation logs) is analyzed using natural language processing (NLP), while images and videos are analyzed using image recognition technology. This allows for the extraction of characteristics and behavioral patterns of the pet or deceased person.
[0972] Step 5:
[0973] The server builds a machine learning model based on the analyzed data. Specifically, conversation patterns extracted from text data and behavioral patterns extracted from image and video data are input into the model and it learns.
[0974] Step 6:
[0975] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[0976] Step 7:
[0977] The user enters text into the dialogue window and sends it. For example, they send a message like "Pochi, let's play lots today."
[0978] Step 8:
[0979] The server receives input from the user and uses the learning model to generate an appropriate response, such as "Yes! Let's go to the park!", based on the personality and behavioral patterns of past pets and deceased pets.
[0980] Step 9:
[0981] The server sends the generated response to the user's device, which receives and displays the response. In particular, if the user is using an AR / VR device, a virtual environment is generated and the response is presented visually and audibly.
[0982] Step 10:
[0983] The user receives a response and continues the two-way dialogue by entering questions or messages again. This interactive process allows the user to recreate memories with their beloved family and pets in a virtual space, resulting in emotional stability and a sense of happiness.
[0984] By the above steps, the user can smoothly have a virtual conversation with a loved one, and the object of the present invention can be achieved.
[0985] Example 1
[0986] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0987] In a system where users can recreate memories of their precious family and pets in a virtual space and engage in real-time interactive dialogue, the challenge is to provide a system that allows users to deepen their emotional interactions and feel a sense of happiness. It is also necessary to provide a means for users to intuitively experience responses visually and audibly.
[0988] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0989] In this invention, the server includes means for a user to upload data, means for the server to receive and store the uploaded data, means for the server to analyze the stored data, means for the server to build a learning model from the analyzed data, means for a user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and means for the terminal to generate a virtual environment and present the response visually and audibly. This allows the user to recreate memories with loved ones in a virtual space and engage in two-way dialogue in real time, thereby achieving emotional stability and a sense of happiness.
[0990] "User" refers to an individual who uses the system to recreate memories of their beloved family and pets in a virtual space and engage in real-time interactive dialogue.
[0991] "Server" refers to the central computing resource for receiving, storing, analyzing, and building learning models of data, as well as generating interactive responses.
[0992] "Data" refers to information such as letters, photos, videos, etc. uploaded by users.
[0993] "Upload" refers to the act of a user sending data from their own device to a server.
[0994] "Means for storage" refers to the system's ability to store uploaded data in storage in an appropriate format.
[0995] "Means of analysis" refers to the function of understanding the content of stored data using natural language processing, image recognition technology, etc., and extracting the necessary information.
[0996] "Learning model" refers to an algorithmic model that is generated based on analyzed data and that generates dialogue responses.
[0997] An "interaction session" refers to a series of interactions in which a user has a two-way dialogue with a system in real time.
[0998] "Response" refers to a reply message generated by a server based on user input.
[0999] "Terminal" refers to digital devices used by users, such as PCs, smartphones, and AR / VR devices.
[1000] "Virtual environment" refers to a visually and aurally recreated space generated on a device for users to interact with deeply and emotionally.
[1001] The system of the present invention allows users to recreate memories of their precious family and pets in a virtual space and engage in real-time interactive conversations. To realize this system, the following elements are required:
[1002] User upload of data
[1003] Users use their own devices (e.g., PCs or smartphones) to access a web page or application and upload data such as letters, photos, and videos. Specifically, they use a browser (e.g., Google Chrome or Safari) or a dedicated application (e.g., iOS or Android app). Users drag and drop the data onto the upload page.
[1004] Receiving and storing data by the server
[1005] The server receives and stores data uploaded by users. This is done using a database management system (e.g., MySQL, PostgreSQL) or a file storage service (e.g., Amazon S3). The server analyzes the metadata of the uploaded data (e.g., file type, size, and content) and categorizes and stores it appropriately.
[1006] Data analysis and learning model construction
[1007] The server analyzes the stored data and uses natural language processing (NLP) to extract meaning and patterns from text data (such as letters and conversation logs). This analysis uses NLP libraries (e.g., spaCy, NLTK, BERT model). On the other hand, for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. This analysis uses deep learning frameworks (e.g., TensorFlow, PyTorch). Based on the analyzed data, the server builds a learning model and prepares to generate dialogue responses.
[1008] Initiating an interactive session and generating responses
[1009] A user logs in to the service and starts a conversation session. The server loads the saved learning model and generates appropriate responses for each message or question from the user using a generative AI model (e.g., GPT-3, BERT). The generated responses are sent to the user's device in real time.
[1010] Displaying the response and generating the virtual environment
[1011] The device displays the response received from the server in real time. When the user checks the response using a PC or smartphone, if an AR / VR device (e.g., Oculus Rift or HTC Vive) is used, a game engine such as Unity or Unreal Engine is used to generate a visually and aurally virtual environment.
[1012] Specific examples
[1013] 1. When a user uploads information about their dog, Pochi:
[1014] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[1015] The server receives this data, analyzes the metadata, and stores it in temporary storage.
[1016] The server analyzes the text of the letter using natural language processing and the photos and videos using image recognition technology.
[1017] The server learns Pochi's characteristics and behavioral patterns from the analysis data and builds a learning model.
[1018] During an interactive session, the user might say, "Pochi, let's play a lot today," and the server would generate and send a response saying, "Yes! Let's go to the park!"
[1019] The user's device displays this response in a virtual environment, allowing the user to relive their memories with Pochi in real time.
[1020] Prompt Sentence Examples
[1021] "Create a scenario for going for a walk based on the characteristics and behavior patterns of my beloved dog Pochi."
[1022] "Enter your Mother's Day letter and generate a response from your mother."
[1023] This system allows users to experience emotional stability and happiness through virtual interactions with their beloved family and pets.
[1024] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1025] Step 1: User uploads data
[1026] Input: A user uses a PC or smartphone to access a specific web page or application.
[1027] How it works: Users drag and drop data files such as photos, videos, letters, etc. onto the upload page.
[1028] Output: The user's device temporarily stores the data and prepares it to be sent to the server.
[1029] Step 2: Server receives and stores data
[1030] Input: Data sent from the user's device.
[1031] How it works: The server receives the data and stores it in a database management system (e.g., MySQL, PostgreSQL) or file storage service (e.g., Amazon S3). The server analyzes the data's metadata (e.g., file type, size, content) and categorizes and stores it appropriately.
[1032] Output: Data is stored on the server and metadata is registered in the database.
[1033] Step 3: Analyzing the text data
[1034] Input: Text data such as letters and conversation logs stored on the server.
[1035] How it works: The server uses natural language processing (NLP) libraries (e.g., spaCy, NLTK, BERT models) to extract meaning and patterns from text data, including tokenization, syntactic analysis, and semantic analysis.
[1036] Output: Extracted semantic and pattern data.
[1037] Step 4: Image and video data analysis
[1038] Input: Image and video data such as photos and videos stored on the server.
[1039] How it works: The server uses deep learning frameworks (e.g., TensorFlow, PyTorch) to analyze image and video data. Specific processes include object detection, feature extraction, and behavioral pattern recognition.
[1040] Output: Analyzed feature and behavioral pattern data.
[1041] Step 5: Building a learning model
[1042] Input: Analysis results of text data and image / video data.
[1043] How it works: The server builds a learning model using a generative AI model (e.g., GPT-3, BERT) based on the analyzed data.
[1044] Output: The constructed learning model.
[1045] Step 6: Prepare for an interactive session
[1046] Input: User login information.
[1047] How it works: A user logs in to the service and begins an interactive session. The server loads a saved training model.
[1048] Output: An interactive session is ready.
[1049] Step 7: Running an interactive session
[1050] Input: User's message or question.
[1051] How it works: The server uses a generative AI model to generate an appropriate response based on the user's input.
[1052] Output: Sends the server-generated response to the user's terminal.
[1053] Step 8: View the response
[1054] Input: The response sent by the server.
[1055] How it works: The user's device displays the response in real time.
[1056] Output: User sees response on terminal.
[1057] Step 9: Generate a Virtual Environment (for AR / VR devices)
[1058] Input: The response sent by the server.
[1059] How it works: The device uses a game engine such as Unity or Unreal Engine to generate a visually and aurally virtual environment.
[1060] Output: The user experiences the virtual interaction through the AR / VR device.
[1061] (Application example 1)
[1062] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1063] Conventional virtual shopping experiences have been little more than a means for users to simply purchase goods, making it difficult for them to experience emotional satisfaction or happiness. Furthermore, it is difficult for users to relive and share memories of their precious family and pets in the virtual space, resulting in insufficient emotional support. To address these issues, a system is needed that allows users to have an emotionally rich experience in the virtual space.
[1064] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1065] In this invention, the server includes means for a user to upload data, means for the server to receive and store the uploaded data, means for the server to analyze the stored data, means for the server to build a learning model from the analyzed data, means for the user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and means for the virtual assistant to act as a guide when the user has a product purchasing experience in a virtual space. This allows the user to share a shopping experience with their beloved family or pet from the past in the virtual space, and enables them to receive emotional support and a sense of happiness.
[1066] "Means for users to upload data" refers to an interface that allows users to send data such as photos, videos, and text to a server via the Internet.
[1067] "Means by which the server receives and stores uploaded data" refers to a system for receiving and securely storing data sent by users on the server.
[1068] "Means by which the server analyzes the data stored on it" refers to the algorithms and programs that analyze the data stored on the server and extract and understand specific information.
[1069] "Means for the server to build a learning model from the analyzed data" refers to the function of training an artificial intelligence or machine learning model based on the analyzed data, enabling future predictions and response generation.
[1070] "Means by which a user initiates an interactive session" refers to the interface by which a user logs into the system and begins an interactive session.
[1071] "Means for the server to generate responses based on input from the user throughout an interactive session" refers to functionality for generating and providing appropriate responses to questions and messages from the user during an interactive session.
[1072] "Means for the terminal to display the response generated" refers to the function for displaying the response sent from the server on the terminal used by the user (such as a smartphone or VR device).
[1073] "Means for a virtual assistant to act as a guide when a user is making a purchase in a virtual space" refers to a function in which a virtual assistant generated based on analytical data provides guidance and advice when a user is making a purchase in a virtual space.
[1074] The system according to the present invention allows users to recreate memories of their precious family and pets in a virtual space, and also allows users to enjoy a shopping experience. Specific embodiments of the system will be described below.
[1075] System Overview
[1076] The system mainly consists of the following components:
[1077] How users upload data
[1078] The means by which the server receives, stores, and analyzes uploaded data
[1079] A means for the server to build a learning model from the analytics data and generate responses based on user input through interactive sessions
[1080] A means for the terminal to display the generated response
[1081] A method for virtual assistants to guide users through a virtual shopping experience
[1082] Program processing overview
[1083] Uploading data
[1084] Users use their smartphones or PCs to upload their photos, videos, letters, and other memorable data to the application, which requires an internet connection and an interface that supports data transmission.
[1085] Data analysis
[1086] The server receives the uploaded data and temporarily stores it. The stored data is then analyzed. Natural language processing (NLP) techniques are used to extract meaning and patterns from text data, while image and video processing is performed using image recognition techniques. The main software tools used include Hugging Face Transformers and OpenCV.
[1087] Building a learning model
[1088] The server then builds a learning model based on the analyzed data. For example, it learns the characteristics of pets and family members based on past conversation logs and photos, and models their emotions and behavioral patterns. This process uses deep learning libraries such as TensorFlow and PyTorch.
[1089] Starting an interactive session
[1090] When a user logs in to the application and starts an interaction session, the server loads the previously built learning model. Based on the user's questions and comments, the virtual assistant generates an appropriate response and sends it to the user's device.
[1091] Viewing the response
[1092] The user's device, such as a smartphone or VR device, displays the response received from the server in real time. In particular, VR devices provide visual and auditory responses so that the user can enjoy shopping with a virtual assistant.
[1093] Specific examples
[1094] For example, if a user wants to relive a memory of their beloved dog:
[1095] 1. Users upload photos, videos and letters of their pet dogs to the application.
[1096] 2. The server receives and analyzes this data to learn your dog's characteristics and behavioral patterns.
[1097] 3. A pet dog is generated as a virtual assistant, and when the user asks, "Pochi, what should we buy today?", the virtual assistant Pochi responds, "I want a new toy!"
[1098] 4. This response is displayed on smartphones and VR devices, allowing users to enjoy shopping with Pochi in a virtual space.
[1099] Prompt Sentence Examples
[1100] "Take a photo and upload it to the app. Also, enter a message about your pet's memory."
[1101] "Visit a virtual store and ask the assistant questions, such as, 'What products do you recommend?'"
[1102] In this way, the system allows users to enjoy an emotionally rich shopping experience in a virtual space.
[1103] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1104] Step 1:
[1105] Uploading data
[1106] User action: The user uploads data such as photos, videos, and text from their smartphone or PC to the application.
[1107] Input: Files uploaded by users (photos, videos, text data)
[1108] Server action: Receives uploaded data and stores it in temporary storage.
[1109] Output: Data stored in the server storage
[1110] Step 2:
[1111] Data analysis
[1112] Server operation: The server analyzes the stored photos, videos, and text data.
[1113] Input: Data stored in the server storage
[1114] Natural Language Processing (NLP): Processing text data to extract meaning and patterns, using techniques such as Hugging Face Transformers.
[1115] Image Recognition: Process images and videos to extract important features and behavioral patterns. Uses OpenCV, etc.
[1116] Output: Analyzed data (text semantics, image and video features)
[1117] Step 3:
[1118] Building a learning model
[1119] Server operation: Builds a learning model based on the analyzed data, which trains the virtual assistant's actions and responses.
[1120] Input: Analyzed data (text semantics, image and video features)
[1121] Data Computation: Use TensorFlow and PyTorch to build learning models from analytical data.
[1122] Output: A trained learning model
[1123] Step 4:
[1124] Starting an interactive session
[1125] User Action: A user logs into an application and begins an interactive session.
[1126] Input: User login information and a request to start a session
[1127] Server operation: The server receives the user's session initiation request and loads the trained learning model.
[1128] Output: Initial response of the interactive session
[1129] Step 5:
[1130] Generating a response
[1131] User Action: The user enters a question or message.
[1132] Input: User question or message
[1133] Server behavior: Based on the user's input, the server uses a learning model to generate an appropriate response.
[1134] Data computation: A generative AI model (e.g., GPT-3) processes the prompt and generates a response.
[1135] Output: The generated response
[1136] Step 6:
[1137] Viewing the response
[1138] Terminal behavior: The user's terminal displays the response received from the server in real time.
[1139] Input: The response sent by the server
[1140] What it does: The smartphone displays text, audio, and images, while the VR device displays a virtual assistant in virtual reality.
[1141] Output: The response displayed on the user's terminal
[1142] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1143] The system of the present invention allows users to recreate memories of their beloved family and pets in a virtual space and engage in real-time interactive dialogue. Furthermore, by incorporating an emotion engine, the system recognizes the user's emotions and generates more appropriate responses, improving the user experience.
[1144] System Overview
[1145] The system mainly consists of the following components:
[1146] A means for users to upload data
[1147] The means by which the server receives, stores, and analyzes uploaded data
[1148] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[1149] A means for the terminal to display the generated response
[1150] Emotion engine that recognizes user emotions
[1151] This allows users to recreate memories with loved ones in a virtual space, providing emotional stability and a sense of happiness.
[1152] Program processing flow
[1153] Uploading and saving data
[1154] A user uses their device (e.g., PC, smartphone) to access a specific web page or application and upload data such as letters, photos, videos, etc. The uploaded data is received by the server and saved in temporary storage. The server analyzes the data's metadata (e.g., file type, size, content) and saves it appropriately.
[1155] Data analysis and learning
[1156] The server then begins the process of analyzing the stored data. For text data (e.g., letters, conversation logs), natural language processing (NLP) is used to extract meaning and patterns, and for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. Based on this analyzed data, the server builds a learning model.
[1157] Emotion Engine Operation
[1158] The server uses an emotion engine to recognize the user's emotions. The emotion engine uses natural language processing and speech recognition technologies to extract emotions from the user's text and voice inputs and determine their emotional state.
[1159] Starting an interactive session
[1160] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[1161] Generate and display the response
[1162] The user enters a message in the dialogue window using text or voice and sends it. The server receives the input from the user and recognizes the user's emotions through an emotion engine. Based on the recognized emotions, the server uses a learning model to generate a more appropriate response. The generated response is sent to the user's device and displayed. In particular, if an AR / VR device is used, a virtual environment is generated and the response is presented visually and audibly.
[1163] Specific examples
[1164] 1. When a user uploads information about their dog, Pochi:
[1165] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[1166] The server receives this data, temporarily stores it, and analyzes and stores the metadata for each file.
[1167] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[1168] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality (energetic and curious) and behavioral patterns (responds to specific commands).
[1169] When the user says, "Pochi, let's play a lot today," the server generates a response saying, "Yes! Let's go to the park!" and sends it to the user's device.
[1170] The server also recognizes the user's emotion of happiness through an emotion engine and generates a more positive response based on that.
[1171] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[1172] Through the above process, users can achieve emotional stability and a sense of happiness through virtual interactions with loved ones. The system of the present invention, in particular, can provide a more sophisticated and emotionally in line with the interactive experience by adding an emotion engine.
[1173] ---
[1174] The processing flow will be explained below.
[1175] Step 1:
[1176] A user uses his / her own terminal (e.g., PC, smartphone) to access a specific web page or application, which provides an interface for uploading data such as letters, photos, and videos.
[1177] Step 2:
[1178] The user uses the upload interface to select data related to the deceased person or pet (photos, videos, text, etc.) and initiates the upload, which is then sent to the server.
[1179] Step 3:
[1180] The server receives the uploaded data and stores it in temporary storage. At the same time, the server analyzes the metadata of each file (file type, size, content) and stores it appropriately in a database.
[1181] Step 4:
[1182] The server then begins the process of analyzing the stored data. First, for text data (such as letters or conversation logs), natural language processing (NLP) is used to analyze the content and extract various patterns and meanings. Next, for image and video data, image recognition technology is used to extract features and behavioral patterns.
[1183] Step 5:
[1184] The server trains a machine learning model based on the analyzed data, inputting conversation patterns extracted from text data and behavioral patterns extracted from image and video data into the model and applying the learning algorithm.
[1185] Step 6:
[1186] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[1187] Step 7:
[1188] The user enters text into the dialogue window and sends it. For example, they send a message like "Pochi, let's play lots today."
[1189] Step 8:
[1190] The server receives input from the user and analyzes the user's emotions through an emotion engine, which uses natural language processing techniques to extract emotions (such as joy, sadness, or anger) from the user's text.
[1191] Step 9:
[1192] The server generates an appropriate response based on the emotion analysis results using a learning model. For example, if the user's message is recognized as joyful, it generates a positive response.
[1193] Step 10:
[1194] The server sends the generated response to the user's device, which receives the response and displays it in real time. In particular, in the case of AR / VR devices, a virtual environment is generated that matches the emotion, and responses are presented visually and audibly.
[1195] Step 11:
[1196] The user can continue the conversation by entering more messages. The server receives the user's input again, analyzes the emotions through the emotion engine, and generates and displays an appropriate response. By repeating this process, the user can continue an emotionally rich virtual conversation with their beloved family or pet.
[1197] ---
[1198] Example 2
[1199] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1200] Until now, there has been no system that recreates memories of past loved ones and pets in a virtual space and allows for real-time interactive dialogue. Furthermore, there was limited technology available to recognize user emotions and generate more appropriate responses, making it difficult to improve the user experience.
[1201] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1202] In this invention, the server includes means for a user to upload information, means for the server to receive and store the uploaded information, means for the server to analyze the stored information, means for the server to build a learning model from the analyzed data, means for a user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and emotion engine means for the server to recognize the user's emotions. This allows the user to recreate memories with loved ones in a virtual space, thereby achieving emotional stability and a sense of happiness.
[1203] "User" means an individual or entity that uses the System to upload information and conduct interactive sessions.
[1204] "Information" refers to any data that users upload to the system, such as letters, photos, videos, etc.
[1205] A "means" refers to a device, process, or method employed to accomplish a particular end.
[1206] The "server" is a central system that receives, stores, analyzes, and builds learning models from information uploaded by users.
[1207] A "learning model" is an artificial intelligence model that the server builds based on the analysis data and is used to generate responses to user input.
[1208] An "interactive session" refers to a process in which a user and a system exchange information in real time.
[1209] A "response" is a reply message generated by a server in response to a user's input.
[1210] "Terminal" refers to the device (e.g., PC, smartphone, AR / VR device) that a user uses to interact with the system.
[1211] An "emotion engine" is an artificial intelligence technique used by the server to recognize the user's emotions and generate appropriate responses.
[1212] "Natural language processing" is a technique used in information analysis to extract meaning and patterns from text data.
[1213] "Image recognition technology" is a technology used in information analysis to extract features and behavioral patterns from image and video data.
[1214] The system of the present invention allows users to recreate memories of their precious family and pets in a virtual space and engage in real-time interactive dialogue. Furthermore, by combining it with an emotion engine, it is possible to recognize the user's emotions and generate more appropriate responses, thereby improving the user experience.
[1215] System Overview
[1216] The system mainly consists of the following components:
[1217] How users upload information
[1218] The means by which the server receives, stores, and analyzes the uploaded information
[1219] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[1220] A means for the terminal to display the generated response
[1221] Emotion engine that recognizes user emotions
[1222] Hardware and Software Configuration
[1223] Server: Use a high-performance server machine (e.g., AWS EC2) and use an S3 bucket to store data.
[1224] Device: The user uses a PC, smartphone, or AR / VR device.
[1225] Natural Language Processing (NLP) library: Spacy is used to tokenize text data and extract meaning and patterns.
[1226] Image recognition technology: Uses OpenCV and TensorFlow to analyze features and behavioral patterns in images and videos.
[1227] Machine learning libraries: Use TensorFlow and PyTorch to generate learning models from analytical data.
[1228] Emotion Detection Library: Extracts emotions from user text and speech using the NRC Emotion Lexicon and Google Cloud Speech-to-Text.
[1229] Program processing
[1230] Users use their devices to access web pages or apps and upload data such as letters, photos, and videos. The server receives this data and temporarily stores it in storage. The stored data is analyzed using NLP and image recognition technologies. Based on the analyzed data, the server uses machine learning libraries to generate a learning model.
[1231] When a user logs in to the system and starts an interaction session, the server loads the saved learning model and waits for user input. When the user enters a message in the interaction window using text or voice, the server receives the input and recognizes the emotion using the emotion engine. Based on the recognized emotion, the server generates an appropriate response using the generative AI model and sends it to the device. The device displays the generated response, and the response is presented visually and audibly in the virtual environment, especially when using an AR / VR device.
[1232] Specific examples
[1233] 1. When a user uploads information about their dog, Pochi:
[1234] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[1235] The server receives this data, temporarily stores it, and analyzes and stores the meta information for each.
[1236] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[1237] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality and behavioral patterns.
[1238] When the user says, "Pochi, let's play a lot today," the server generates a response, "Yes! Let's go to the park!" and sends it to the user's device.
[1239] The server recognizes the user's joy through an emotion engine and generates a more positive response.
[1240] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[1241] Prompt Sentence Examples
[1242] "Please tell me how to recognize user emotions and generate appropriate responses."
[1243] "Please explain in detail the process of analyzing the uploaded data."
[1244] "Please tell us specifically what experience you would like to have through an interaction session with your dog."
[1245] Through the above process, users can achieve emotional stability and a sense of happiness through virtual conversations with loved ones. By using an emotion engine, this system provides a more emotionally in line with the user's needs.
[1246] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1247] Step 1:
[1248] Users access a specified web page or app from their device (PC or smartphone) and upload information such as letters, photos, and videos.
[1249] Specific behavior: The user clicks the upload button, selects a file from the file selection dialog, and presses the send button.
[1250] Input: Data uploaded by users, such as letters, photos, videos, etc.
[1251] Output: Data uploaded to the server.
[1252] Step 2:
[1253] The server receives the uploaded information and temporarily stores it in storage.
[1254] Specific behavior: The server receives the HTTP request, saves the file in a temporary directory on the server side, and returns an upload completion message.
[1255] Input: Uploaded information.
[1256] Output: Information stored on the server.
[1257] Step 3:
[1258] The server analyzes the information's metadata (file type, size, content, etc.) and stores it in a database in an appropriate format.
[1259] Specific operation: The server checks the file format, extracts meta information for each file (e.g., 'photo.jpg', size: 2MB, creation date: 2023-10-15), and records it in the database.
[1260] Input: Temporarily stored information.
[1261] Output: Metadata and information stored in a database.
[1262] Step 4:
[1263] The server analyzes the stored information, specifically using natural language processing (NLP) for text data and image recognition technology for images and videos.
[1264] How it works: The server uses NLP libraries (e.g., Spacy) to tokenize text and extract important meaning and patterns. For images and videos, it uses OpenCV, TensorFlow, etc. to analyze facial recognition and behavioral patterns.
[1265] Input: Information stored in a database.
[1266] Output: Parsed data.
[1267] Step 5:
[1268] The server constructs a learning model based on the analysis data.
[1269] Specific operation: The server uses machine learning libraries (e.g., TensorFlow, PyTorch) to learn individual features from the analysis data and generate a unique learning model.
[1270] Input: Parsed data.
[1271] Output: The trained model.
[1272] Step 6:
[1273] The server uses an emotion engine to recognize the user's emotions. The emotion engine uses natural language processing and speech recognition technologies to extract and distinguish emotions from the user's text and voice input.
[1274] Specific behavior: The server calls an emotion detection library (e.g., NRC Emotion Lexicon, Google Cloud Speech-to-Text) to extract emotion labels (e.g., joy, sadness, anger) from text or audio.
[1275] Input: User text and voice input.
[1276] Output: User emotion label.
[1277] Step 7:
[1278] A user logs into a system and begins an interactive session.
[1279] Specific operation: The user enters the username and password on the login page and presses the login button. The server performs authentication, and if successful, displays the start screen of the interactive session.
[1280] Enter your username and password.
[1281] Output: The start screen of an interactive session after successful authentication.
[1282] Step 8:
[1283] The server loads the saved learning model and waits for user input.
[1284] Specific operation: The server loads the appropriate learning model into memory based on the user profile and waits for data input from the user.
[1285] Input: The training model.
[1286] Output: Server in standby state.
[1287] Step 9:
[1288] The user enters a message in the dialogue window and sends it.
[1289] Specific behavior: When the user enters text into the dialogue window and clicks the submit button, the data is sent to the server.
[1290] Input: A text message.
[1291] Output: The text message sent to the server.
[1292] Step 10:
[1293] The server receives the user's input and recognizes the user's emotions through an emotion engine.
[1294] Specific operation: The server analyzes the input text and audio data, extracts emotion labels, and passes them to the next step.
[1295] Input: Text or audio data.
[1296] Output: Extracted emotion labels.
[1297] Step 11:
[1298] The server uses a learning model to generate an appropriate response based on the recognized emotion.
[1299] Specific operation: The server generates a response sentence using a generative AI model (e.g., GPT-3) based on the recognized emotion and stored features.
[1300] Input: emotion labels and a trained model.
[1301] Output: The generated response sentence.
[1302] Step 12:
[1303] The device displays the generated response and, especially when using an AR / VR device, presents the response visually and audibly along with the virtual environment.
[1304] Specific behavior: The server sends the generated response to the user's device, which displays it. If an AR / VR device is used, the device displays the response visually in the virtual environment and also plays an audio response.
[1305] Input: The generated response sentence.
[1306] Output: The response displayed on the user's terminal.
[1307] (Application example 2)
[1308] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1309] Currently, in physical stores, customers have few easy ways to obtain information about products or specific areas within the store, and there is a need to improve the customer experience.In addition, there is a lack of systems that can respond individually to customers' emotions, making it difficult to improve customer satisfaction.
[1310] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1311] In this invention, the server includes a means for a user to upload data, a means for the server to receive and store the uploaded data, and a means for the server to analyze the stored data. This enables information about a specified object in a virtual space to be displayed in a physical store. The server also includes a means for constructing a learning model from the analyzed data, a means for a user to start an interactive session, a means for the server to generate a response based on input from the user through the interactive session, a means for the terminal to display the generated response, and an emotion recognition means for generating a response based on the user's emotion. This enables individual responses based on the user's emotions, which is expected to improve customer satisfaction.
[1312] The "means for users to upload data" refers to an interface that allows users to use their own terminals to send data such as photos, videos, and text to the system.
[1313] "Means for the server to receive and store uploaded data" refers to the function by which the server receives data sent by the user and stores it temporarily or permanently.
[1314] "Means for the server to analyze the stored data" refers to the technology that the server uses to process the stored data and decrypt and analyze its contents.
[1315] "Means by which the server constructs a machine learning model from the analyzed data" means the process by which the server creates a machine learning model based on the analyzed data and uses the model to generate appropriate responses to future data inputs.
[1316] The "means by which a user initiates an interactive session" refers to the interface or functionality by which a user accesses the system and initiates an interaction.
[1317] "Means for a server to generate a response based on input from a user throughout an interactive session" refers to a technique in which a server receives text or voice input from a user and generates a reply based thereon.
[1318] "Means for displaying the response generated by the terminal" refers to a function that allows the user's device (smartphone, smart glasses, etc.) to visually or audibly present the response generated by the server to the user.
[1319] "Means for displaying information about a specified object in a virtual space" refers to a function that allows a user to specify a specific item or area in a physical store and visually provide the user with related information using AR technology.
[1320] "Emotion recognition means for generating responses according to the user's emotions" is a technology that analyzes the user's input (text or voice), identifies the user's emotional state, and generates an appropriate response accordingly.
[1321] The system according to the present invention is a system that allows users to recreate memories of their precious family members and pets in a virtual space and engage in two-way conversations in real time. A specific embodiment of this system is described below.
[1322] Hardware and software used
[1323] This system uses the following hardware and software:
[1324] Hardware: Smartphones, smart glasses
[1325] Software: Google Cloud Vision API, Google Cloud Natural Language API, Python, OpenCV, TextBlob, pyttsx3
[1326] Uploading and saving data
[1327] Users use their smartphones or PCs to access specific web pages or applications and upload data such as photos, videos, and text. The server receives this data and stores it temporarily or permanently. The metadata (file type, size, and content) of the temporarily stored data is also analyzed, and it is then categorized and stored appropriately.
[1328] Data analysis and learning
[1329] The server analyzes the stored data. For photos and videos, it uses the Google Cloud Vision API to analyze the characteristics of specific people and pets. For text data (such as letters and conversation logs), it uses the Google Cloud Natural Language API to extract emotions and meanings, and performs sentiment analysis using TextBlob. This allows the server to build a learning model and prepare to generate responses based on the user's specific emotions and memories.
[1330] Emotion Engine Operation
[1331] The server uses an emotion engine to recognize emotions from the user's text and voice input. It uses the Google Cloud Natural Language API and TextBlob to extract emotions from the user's text and voice input and determine their emotional state.
[1332] Initiating an interactive session and generating responses
[1333] A user logs in to the system and begins an interactive session. The server loads the saved learning model and waits for input from the user. The user inputs a message via text or voice, which the server receives and recognizes the user's emotions through an emotion engine. It then generates a response based on the user's emotions and outputs it to the terminal, either displayed or via voice synthesis.
[1334] Specific examples
[1335] If a user wants to recreate memories with their beloved dog in a virtual space, they can take the following steps.
[1336] 1. Users send data such as photos, videos, and letters of their beloved dog through a dedicated upload page.
[1337] 2. The server receives the data, analyzes the metadata, and saves it. It then uses the Google Cloud Vision API to recognize images and videos, and the Google Cloud Natural Language API to analyze the sentiment of the text data.
[1338] 3. When the user says, "Pochi, what do you want to do today?", the server generates a response, "I want to play in the park!", and displays it on the device. If the server recognizes the user's emotion of joy, it generates an even more positive response.
[1339] Prompt Sentence Examples
[1340] I want to know what memories you have of this product.
[1341] "I want you to recreate the story of when you purchased this product."
[1342] This allows users to experience emotional stability and happiness through virtual interactions with loved ones. The system realized by the invention, in particular by adding an emotion engine, can provide a more sophisticated and emotionally in line with the interactive experience.
[1343] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1344] Step 1: Upload your data
[1345] A user uses his / her own device (e.g., smartphone, PC) to access a specific web page or application and upload photos, videos, text data, etc. The input includes the files and data selected by the user. The server receives this data and temporarily stores it.
[1346] Output: Received and stored data.
[1347] Step 2: Data storage and metadata analysis
[1348] The server receives the uploaded data and stores it in temporary storage. In parallel, the server parses the metadata (file type, size, content) of each uploaded data file. As input, it contains the uploaded data file. The server parses the metadata and stores it in the appropriate format.
[1349] Output: Parsed metadata and organized data files.
[1350] Step 3: Analyze the data
[1351] The server starts the process of analyzing the stored data. For text data, it uses the Google Cloud Natural Language API to extract sentiment and meaning. For images and videos, it uses the Google Cloud Vision API to analyze features of specific people or pets. The input includes the stored data files. The results of the data analysis are feature-extracted data and sentiment analysis results.
[1352] Output: Feature extracted data and sentiment analysis results.
[1353] Step 4: Building a learning model
[1354] The server creates a machine learning model based on the analyzed data. It learns past interactions and behavioral patterns with specific people and pets, and uses the model to generate appropriate responses to future user input. The analyzed data is included as input. The server builds and stores the machine learning model.
[1355] Output: The built machine learning model.
[1356] Step 5: Starting an interactive session
[1357] A user logs into the system and accesses an interface to start an interactive session. When the user logs in, the server loads the saved learning model and waits for input from the user. The input includes the user's login operation. The server loads the learning model and enters an interactive waiting state.
[1358] Output: Training model loaded and waiting.
[1359] Step 6: Receiving User Input and Emotion Recognition
[1360] The user inputs a message using text or voice and sends it. When the server receives it, it uses Google Cloud Natural Language API or TextBlob to recognize the user's emotions. The input contains the user's text or voice data. The server performs emotion recognition and returns the result.
[1361] Output: User's emotional state and analysis results.
[1362] Step 7: Generate a response
[1363] The server generates an appropriate response based on the user's emotional state using the learning model. The input includes the user's emotional state and the learning model. The server sends the generated response to the user's device.
[1364] Output: The generated response.
[1365] Step 8: Display or speak the response
[1366] The terminal receives the response sent by the server and displays it to the user, possibly providing an audible response using a speech synthesis engine (pyttsx3). The input includes the generated response. The terminal presents the response to the user visually or audibly.
[1367] Output: The response presented to the user.
[1368] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1369] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1370] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1371] [Fourth embodiment]
[1372] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1373] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1374] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1375] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1376] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1377] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1378] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1379] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1380] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1381] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1382] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1383] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1384] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1385] The system according to the present invention is a system that allows a user to recreate memories of past loved ones and pets in a virtual space and engage in two-way conversations in real time. Specific embodiments of the system are described below.
[1386] System Overview
[1387] The system mainly consists of the following components:
[1388] A means for users to upload data
[1389] The means by which the server receives, stores, and analyzes uploaded data
[1390] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[1391] A means for the terminal to display the generated response
[1392] This allows users to recreate memories with loved ones in a virtual space, providing emotional stability and a sense of happiness.
[1393] Program processing flow
[1394] Uploading and saving data
[1395] A user uses their device (e.g., PC, smartphone) to access a specific web page or application and upload data such as letters, photos, videos, etc. The uploaded data is received by the server and saved in temporary storage. The server analyzes the data's metadata (e.g., file type, size, content) and saves it appropriately.
[1396] Data analysis and learning
[1397] The server then begins the process of analyzing the stored data. For text data (e.g., letters, conversation logs), natural language processing (NLP) is used to extract meaning and patterns, and for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. Based on this analyzed data, the server builds a learning model.
[1398] Starting an interactive session
[1399] When a user logs in to the service and starts an interactive session, the server loads the previously constructed learning model. For each message or question entered by the user, the server generates an appropriate response and sends this response to the user's device.
[1400] Viewing the response
[1401] The user's device (e.g., PC, smartphone, AR / VR device) displays the response received from the server in real time. AR / VR devices, in particular, generate a virtual environment and present visual and auditory responses to allow the user to interact deeply with the experience.
[1402] Specific examples
[1403] 1. When a user uploads information about their dog, Pochi:
[1404] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[1405] The server receives this data, temporarily stores it, and analyzes and stores the metadata for each file.
[1406] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[1407] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality (energetic and curious) and behavioral patterns (responds to specific commands).
[1408] When the user says, "Pochi, let's play a lot today," the server generates a response saying, "Yes! Let's go to the park!" and sends it to the user's device.
[1409] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[1410] This system allows users to experience emotional stability and happiness through virtual interactions with their beloved family and pets.
[1411] ---
[1412] The processing flow will be explained below.
[1413] Step 1:
[1414] A user uses his / her own terminal (e.g., PC, smartphone) to access a specific web page or application, which provides an interface for uploading data such as letters, photos, and videos.
[1415] Step 2:
[1416] Using the upload interface, users select data related to the deceased person or pet (photos, videos, text, etc.) and initiate the upload, which sends the data to the server.
[1417] Step 3:
[1418] The server receives the uploaded data and stores it in temporary storage. For each file received, it analyzes its metadata (file type, size, content, etc.) and stores it in a database.
[1419] Step 4:
[1420] The server then begins the process of analyzing the stored data. Text data (e.g., letters, conversation logs) is analyzed using natural language processing (NLP), while images and videos are analyzed using image recognition technology. This allows for the extraction of characteristics and behavioral patterns of the pet or deceased person.
[1421] Step 5:
[1422] The server builds a machine learning model based on the analyzed data. Specifically, conversation patterns extracted from text data and behavioral patterns extracted from image and video data are input into the model and it learns.
[1423] Step 6:
[1424] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[1425] Step 7:
[1426] The user enters text into the dialogue window and sends it. For example, they send a message like "Pochi, let's play lots today."
[1427] Step 8:
[1428] The server receives input from the user and uses the learning model to generate an appropriate response, such as "Yes! Let's go to the park!", based on the personality and behavioral patterns of past pets and deceased pets.
[1429] Step 9:
[1430] The server sends the generated response to the user's device, which receives and displays the response. In particular, if the user is using an AR / VR device, a virtual environment is generated and the response is presented visually and audibly.
[1431] Step 10:
[1432] The user receives a response and continues the two-way dialogue by entering questions or messages again. This interactive process allows the user to recreate memories with their beloved family and pets in a virtual space, resulting in emotional stability and a sense of happiness.
[1433] By the above steps, the user can smoothly have a virtual conversation with a loved one, and the object of the present invention can be achieved.
[1434] Example 1
[1435] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1436] In a system where users can recreate memories of their precious family and pets in a virtual space and engage in real-time interactive dialogue, the challenge is to provide a system that allows users to deepen their emotional interactions and feel a sense of happiness. It is also necessary to provide a means for users to intuitively experience responses visually and audibly.
[1437] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1438] In this invention, the server includes means for a user to upload data, means for the server to receive and store the uploaded data, means for the server to analyze the stored data, means for the server to build a learning model from the analyzed data, means for a user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and means for the terminal to generate a virtual environment and present the response visually and audibly. This allows the user to recreate memories with loved ones in a virtual space and engage in two-way dialogue in real time, thereby achieving emotional stability and a sense of happiness.
[1439] "User" refers to an individual who uses the system to recreate memories of their beloved family and pets in a virtual space and engage in real-time interactive dialogue.
[1440] "Server" refers to the central computing resource for receiving, storing, analyzing, and building learning models of data, as well as generating interactive responses.
[1441] "Data" refers to information such as letters, photos, videos, etc. uploaded by users.
[1442] "Upload" refers to the act of a user sending data from their own device to a server.
[1443] "Means for storage" refers to the system's ability to store uploaded data in storage in an appropriate format.
[1444] "Means of analysis" refers to the function of understanding the content of stored data using natural language processing, image recognition technology, etc., and extracting the necessary information.
[1445] "Learning model" refers to an algorithmic model that is generated based on analyzed data and that generates dialogue responses.
[1446] An "interaction session" refers to a series of interactions in which a user has a two-way dialogue with a system in real time.
[1447] "Response" refers to a reply message generated by a server based on user input.
[1448] "Terminal" refers to digital devices used by users, such as PCs, smartphones, and AR / VR devices.
[1449] "Virtual environment" refers to a visually and aurally recreated space generated on a device for users to interact with deeply and emotionally.
[1450] The system of the present invention allows users to recreate memories of their precious family and pets in a virtual space and engage in real-time interactive conversations. To realize this system, the following elements are required:
[1451] User upload of data
[1452] Users use their own devices (e.g., PCs or smartphones) to access a web page or application and upload data such as letters, photos, and videos. Specifically, they use a browser (e.g., Google Chrome or Safari) or a dedicated application (e.g., iOS or Android app). Users drag and drop the data onto the upload page.
[1453] Receiving and storing data by the server
[1454] The server receives and stores data uploaded by users. This is done using a database management system (e.g., MySQL, PostgreSQL) or a file storage service (e.g., Amazon S3). The server analyzes the metadata of the uploaded data (e.g., file type, size, and content) and categorizes and stores it appropriately.
[1455] Data analysis and learning model construction
[1456] The server analyzes the stored data and uses natural language processing (NLP) to extract meaning and patterns from text data (such as letters and conversation logs). This analysis uses NLP libraries (e.g., spaCy, NLTK, BERT model). On the other hand, for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. This analysis uses deep learning frameworks (e.g., TensorFlow, PyTorch). Based on the analyzed data, the server builds a learning model and prepares to generate dialogue responses.
[1457] Initiating an interactive session and generating responses
[1458] A user logs in to the service and starts a conversation session. The server loads the saved learning model and generates appropriate responses for each message or question from the user using a generative AI model (e.g., GPT-3, BERT). The generated responses are sent to the user's device in real time.
[1459] Displaying the response and generating the virtual environment
[1460] The device displays the response received from the server in real time. When the user checks the response using a PC or smartphone, if an AR / VR device (e.g., Oculus Rift or HTC Vive) is used, a game engine such as Unity or Unreal Engine is used to generate a visually and aurally virtual environment.
[1461] Specific examples
[1462] 1. When a user uploads information about their dog, Pochi:
[1463] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[1464] The server receives this data, analyzes the metadata, and stores it in temporary storage.
[1465] The server analyzes the text of the letter using natural language processing and the photos and videos using image recognition technology.
[1466] The server learns Pochi's characteristics and behavioral patterns from the analysis data and builds a learning model.
[1467] During an interactive session, the user might say, "Pochi, let's play a lot today," and the server would generate and send a response saying, "Yes! Let's go to the park!"
[1468] The user's device displays this response in a virtual environment, allowing the user to relive their memories with Pochi in real time.
[1469] Prompt Sentence Examples
[1470] "Create a scenario for going for a walk based on the characteristics and behavior patterns of my beloved dog Pochi."
[1471] "Enter your Mother's Day letter and generate a response from your mother."
[1472] This system allows users to experience emotional stability and happiness through virtual interactions with their beloved family and pets.
[1473] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1474] Step 1: User uploads data
[1475] Input: A user uses a PC or smartphone to access a specific web page or application.
[1476] How it works: Users drag and drop data files such as photos, videos, letters, etc. onto the upload page.
[1477] Output: The user's device temporarily stores the data and prepares it to be sent to the server.
[1478] Step 2: Server receives and stores data
[1479] Input: Data sent from the user's device.
[1480] How it works: The server receives the data and stores it in a database management system (e.g., MySQL, PostgreSQL) or file storage service (e.g., Amazon S3). The server analyzes the data's metadata (e.g., file type, size, content) and categorizes and stores it appropriately.
[1481] Output: Data is stored on the server and metadata is registered in the database.
[1482] Step 3: Analyzing the text data
[1483] Input: Text data such as letters and conversation logs stored on the server.
[1484] How it works: The server uses natural language processing (NLP) libraries (e.g., spaCy, NLTK, BERT models) to extract meaning and patterns from text data, including tokenization, syntactic analysis, and semantic analysis.
[1485] Output: Extracted semantic and pattern data.
[1486] Step 4: Image and video data analysis
[1487] Input: Image and video data such as photos and videos stored on the server.
[1488] How it works: The server uses deep learning frameworks (e.g., TensorFlow, PyTorch) to analyze image and video data. Specific processes include object detection, feature extraction, and behavioral pattern recognition.
[1489] Output: Analyzed feature and behavioral pattern data.
[1490] Step 5: Building a learning model
[1491] Input: Analysis results of text data and image / video data.
[1492] How it works: The server builds a learning model using a generative AI model (e.g., GPT-3, BERT) based on the analyzed data.
[1493] Output: The constructed learning model.
[1494] Step 6: Prepare for an interactive session
[1495] Input: User login information.
[1496] How it works: A user logs in to the service and begins an interactive session. The server loads a saved training model.
[1497] Output: An interactive session is ready.
[1498] Step 7: Running an interactive session
[1499] Input: User's message or question.
[1500] How it works: The server uses a generative AI model to generate an appropriate response based on the user's input.
[1501] Output: Sends the server-generated response to the user's terminal.
[1502] Step 8: View the response
[1503] Input: The response sent by the server.
[1504] How it works: The user's device displays the response in real time.
[1505] Output: User sees response on terminal.
[1506] Step 9: Generate a Virtual Environment (for AR / VR devices)
[1507] Input: The response sent by the server.
[1508] How it works: The device uses a game engine such as Unity or Unreal Engine to generate a visually and aurally virtual environment.
[1509] Output: The user experiences the virtual interaction through the AR / VR device.
[1510] (Application example 1)
[1511] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1512] Conventional virtual shopping experiences have been little more than a means for users to simply purchase goods, making it difficult for them to experience emotional satisfaction or happiness. Furthermore, it is difficult for users to relive and share memories of their precious family and pets in the virtual space, resulting in insufficient emotional support. To address these issues, a system is needed that allows users to have an emotionally rich experience in the virtual space.
[1513] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1514] In this invention, the server includes means for a user to upload data, means for the server to receive and store the uploaded data, means for the server to analyze the stored data, means for the server to build a learning model from the analyzed data, means for the user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and means for the virtual assistant to act as a guide when the user has a product purchasing experience in a virtual space. This allows the user to share a shopping experience with their beloved family or pet from the past in the virtual space, and enables them to receive emotional support and a sense of happiness.
[1515] "Means for users to upload data" refers to an interface that allows users to send data such as photos, videos, and text to a server via the Internet.
[1516] "Means by which the server receives and stores uploaded data" refers to a system for receiving and securely storing data sent by users on the server.
[1517] "Means by which the server analyzes the data stored on it" refers to the algorithms and programs that analyze the data stored on the server and extract and understand specific information.
[1518] "Means for the server to build a learning model from the analyzed data" refers to the function of training an artificial intelligence or machine learning model based on the analyzed data, enabling future predictions and response generation.
[1519] "Means by which a user initiates an interactive session" refers to the interface by which a user logs into the system and begins an interactive session.
[1520] "Means for the server to generate responses based on input from the user throughout an interactive session" refers to functionality for generating and providing appropriate responses to questions and messages from the user during an interactive session.
[1521] "Means for the terminal to display the response generated" refers to the function for displaying the response sent from the server on the terminal used by the user (such as a smartphone or VR device).
[1522] "Means for a virtual assistant to act as a guide when a user is making a purchase in a virtual space" refers to a function in which a virtual assistant generated based on analytical data provides guidance and advice when a user is making a purchase in a virtual space.
[1523] The system according to the present invention allows users to recreate memories of their precious family and pets in a virtual space, and also allows users to enjoy a shopping experience. Specific embodiments of the system will be described below.
[1524] System Overview
[1525] The system mainly consists of the following components:
[1526] How users upload data
[1527] The means by which the server receives, stores, and analyzes uploaded data
[1528] A means for the server to build a learning model from the analytics data and generate responses based on user input through interactive sessions
[1529] A means for the terminal to display the generated response
[1530] A method for virtual assistants to guide users through a virtual shopping experience
[1531] Program processing overview
[1532] Uploading data
[1533] Users use their smartphones or PCs to upload their photos, videos, letters, and other memorable data to the application, which requires an internet connection and an interface that supports data transmission.
[1534] Data analysis
[1535] The server receives the uploaded data and temporarily stores it. The stored data is then analyzed. Natural language processing (NLP) techniques are used to extract meaning and patterns from text data, while image and video processing is performed using image recognition techniques. The main software tools used include Hugging Face Transformers and OpenCV.
[1536] Building a learning model
[1537] The server then builds a learning model based on the analyzed data. For example, it learns the characteristics of pets and family members based on past conversation logs and photos, and models their emotions and behavioral patterns. This process uses deep learning libraries such as TensorFlow and PyTorch.
[1538] Starting an interactive session
[1539] When a user logs in to the application and starts an interaction session, the server loads the previously built learning model. Based on the user's questions and comments, the virtual assistant generates an appropriate response and sends it to the user's device.
[1540] Viewing the response
[1541] The user's device, such as a smartphone or VR device, displays the response received from the server in real time. In particular, VR devices provide visual and auditory responses so that the user can enjoy shopping with a virtual assistant.
[1542] Specific examples
[1543] For example, if a user wants to relive a memory of their beloved dog:
[1544] 1. Users upload photos, videos and letters of their pet dogs to the application.
[1545] 2. The server receives and analyzes this data to learn your dog's characteristics and behavioral patterns.
[1546] 3. A pet dog is generated as a virtual assistant, and when the user asks, "Pochi, what should we buy today?", the virtual assistant Pochi responds, "I want a new toy!"
[1547] 4. This response is displayed on smartphones and VR devices, allowing users to enjoy shopping with Pochi in a virtual space.
[1548] Prompt Sentence Examples
[1549] "Take a photo and upload it to the app. Also, enter a message about your pet's memory."
[1550] "Visit a virtual store and ask the assistant questions, such as, 'What products do you recommend?'"
[1551] In this way, the system allows users to enjoy an emotionally rich shopping experience in a virtual space.
[1552] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1553] Step 1:
[1554] Uploading data
[1555] User action: The user uploads data such as photos, videos, and text from their smartphone or PC to the application.
[1556] Input: Files uploaded by users (photos, videos, text data)
[1557] Server action: Receives uploaded data and stores it in temporary storage.
[1558] Output: Data stored in the server storage
[1559] Step 2:
[1560] Data analysis
[1561] Server operation: The server analyzes the stored photos, videos, and text data.
[1562] Input: Data stored in the server storage
[1563] Natural Language Processing (NLP): Processing text data to extract meaning and patterns, using techniques such as Hugging Face Transformers.
[1564] Image Recognition: Process images and videos to extract important features and behavioral patterns. Uses OpenCV, etc.
[1565] Output: Analyzed data (text semantics, image and video features)
[1566] Step 3:
[1567] Building a learning model
[1568] Server operation: Builds a learning model based on the analyzed data, which trains the virtual assistant's actions and responses.
[1569] Input: Analyzed data (text semantics, image and video features)
[1570] Data Computation: Use TensorFlow and PyTorch to build learning models from analytical data.
[1571] Output: A trained learning model
[1572] Step 4:
[1573] Starting an interactive session
[1574] User Action: A user logs into an application and begins an interactive session.
[1575] Input: User login information and a request to start a session
[1576] Server operation: The server receives the user's session initiation request and loads the trained learning model.
[1577] Output: Initial response of the interactive session
[1578] Step 5:
[1579] Generating a response
[1580] User Action: The user enters a question or message.
[1581] Input: User question or message
[1582] Server behavior: Based on the user's input, the server uses a learning model to generate an appropriate response.
[1583] Data computation: A generative AI model (e.g., GPT-3) processes the prompt and generates a response.
[1584] Output: The generated response
[1585] Step 6:
[1586] Viewing the response
[1587] Terminal behavior: The user's terminal displays the response received from the server in real time.
[1588] Input: The response sent by the server
[1589] What it does: The smartphone displays text, audio, and images, while the VR device displays a virtual assistant in virtual reality.
[1590] Output: The response displayed on the user's terminal
[1591] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1592] The system of the present invention allows users to recreate memories of their beloved family and pets in a virtual space and engage in real-time interactive dialogue. Furthermore, by incorporating an emotion engine, the system recognizes the user's emotions and generates more appropriate responses, improving the user experience.
[1593] System Overview
[1594] The system mainly consists of the following components:
[1595] A means for users to upload data
[1596] The means by which the server receives, stores, and analyzes uploaded data
[1597] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[1598] A means for the terminal to display the generated response
[1599] Emotion engine that recognizes user emotions
[1600] This allows users to recreate memories with loved ones in a virtual space, providing emotional stability and a sense of happiness.
[1601] Program processing flow
[1602] Uploading and saving data
[1603] A user uses their device (e.g., PC, smartphone) to access a specific web page or application and upload data such as letters, photos, videos, etc. The uploaded data is received by the server and saved in temporary storage. The server analyzes the data's metadata (e.g., file type, size, content) and saves it appropriately.
[1604] Data analysis and learning
[1605] The server then begins the process of analyzing the stored data. For text data (e.g., letters, conversation logs), natural language processing (NLP) is used to extract meaning and patterns, and for images and videos, image recognition technology is used to analyze the characteristics and behavioral patterns of pets and deceased individuals. Based on this analyzed data, the server builds a learning model.
[1606] Emotion Engine Operation
[1607] The server uses an emotion engine to recognize the user's emotions. The emotion engine uses natural language processing and speech recognition technologies to extract emotions from the user's text and voice inputs and determine their emotional state.
[1608] Starting an interactive session
[1609] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[1610] Generate and display the response
[1611] The user enters a message in the dialogue window using text or voice and sends it. The server receives the input from the user and recognizes the user's emotions through an emotion engine. Based on the recognized emotions, the server uses a learning model to generate a more appropriate response. The generated response is sent to the user's device and displayed. In particular, if an AR / VR device is used, a virtual environment is generated and the response is presented visually and audibly.
[1612] Specific examples
[1613] 1. When a user uploads information about their dog, Pochi:
[1614] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[1615] The server receives this data, temporarily stores it, and analyzes and stores the metadata for each file.
[1616] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[1617] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality (energetic and curious) and behavioral patterns (responds to specific commands).
[1618] When the user says, "Pochi, let's play a lot today," the server generates a response saying, "Yes! Let's go to the park!" and sends it to the user's device.
[1619] The server also recognizes the user's emotion of happiness through an emotion engine and generates a more positive response based on that.
[1620] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[1621] Through the above process, users can achieve emotional stability and a sense of happiness through virtual interactions with loved ones. The system of the present invention, in particular, can provide a more sophisticated and emotionally in line with the interactive experience by adding an emotion engine.
[1622] ---
[1623] The processing flow will be explained below.
[1624] Step 1:
[1625] A user uses his / her own terminal (e.g., PC, smartphone) to access a specific web page or application, which provides an interface for uploading data such as letters, photos, and videos.
[1626] Step 2:
[1627] The user uses the upload interface to select data related to the deceased person or pet (photos, videos, text, etc.) and initiates the upload, which is then sent to the server.
[1628] Step 3:
[1629] The server receives the uploaded data and stores it in temporary storage. At the same time, the server analyzes the metadata of each file (file type, size, content) and stores it appropriately in a database.
[1630] Step 4:
[1631] The server then begins the process of analyzing the stored data. First, for text data (such as letters or conversation logs), natural language processing (NLP) is used to analyze the content and extract various patterns and meanings. Next, for image and video data, image recognition technology is used to extract features and behavioral patterns.
[1632] Step 5:
[1633] The server trains a machine learning model based on the analyzed data, inputting conversation patterns extracted from text data and behavioral patterns extracted from image and video data into the model and applying the learning algorithm.
[1634] Step 6:
[1635] A user logs into the service and begins an interactive session. The server loads the saved learning model and waits for input from the user.
[1636] Step 7:
[1637] The user enters text into the dialogue window and sends it. For example, they send a message like "Pochi, let's play lots today."
[1638] Step 8:
[1639] The server receives input from the user and analyzes the user's emotions through an emotion engine, which uses natural language processing techniques to extract emotions (such as joy, sadness, or anger) from the user's text.
[1640] Step 9:
[1641] The server generates an appropriate response based on the emotion analysis results using a learning model. For example, if the user's message is recognized as joyful, it generates a positive response.
[1642] Step 10:
[1643] The server sends the generated response to the user's device, which receives the response and displays it in real time. In particular, in the case of AR / VR devices, a virtual environment is generated that matches the emotion, and responses are presented visually and audibly.
[1644] Step 11:
[1645] The user can continue the conversation by entering more messages. The server receives the user's input again, analyzes the emotions through the emotion engine, and generates and displays an appropriate response. By repeating this process, the user can continue an emotionally rich virtual conversation with their beloved family or pet.
[1646] ---
[1647] Example 2
[1648] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1649] Until now, there has been no system that recreates memories of past loved ones and pets in a virtual space and allows for real-time interactive dialogue. Furthermore, there was limited technology available to recognize user emotions and generate more appropriate responses, making it difficult to improve the user experience.
[1650] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1651] In this invention, the server includes means for a user to upload information, means for the server to receive and store the uploaded information, means for the server to analyze the stored information, means for the server to build a learning model from the analyzed data, means for a user to start an interactive session, means for the server to generate a response based on input from the user through the interactive session, means for the terminal to display the generated response, and emotion engine means for the server to recognize the user's emotions. This allows the user to recreate memories with loved ones in a virtual space, thereby achieving emotional stability and a sense of happiness.
[1652] "User" means an individual or entity that uses the System to upload information and conduct interactive sessions.
[1653] "Information" refers to any data that users upload to the system, such as letters, photos, videos, etc.
[1654] A "means" refers to a device, process, or method employed to accomplish a particular end.
[1655] The "server" is a central system that receives, stores, analyzes, and builds learning models from information uploaded by users.
[1656] A "learning model" is an artificial intelligence model that the server builds based on the analysis data and is used to generate responses to user input.
[1657] An "interactive session" refers to a process in which a user and a system exchange information in real time.
[1658] A "response" is a reply message generated by a server in response to a user's input.
[1659] "Terminal" refers to the device (e.g., PC, smartphone, AR / VR device) that a user uses to interact with the system.
[1660] An "emotion engine" is an artificial intelligence technique used by the server to recognize the user's emotions and generate appropriate responses.
[1661] "Natural language processing" is a technique used in information analysis to extract meaning and patterns from text data.
[1662] "Image recognition technology" is a technology used in information analysis to extract features and behavioral patterns from image and video data.
[1663] The system of the present invention allows users to recreate memories of their precious family and pets in a virtual space and engage in real-time interactive dialogue. Furthermore, by combining it with an emotion engine, it is possible to recognize the user's emotions and generate more appropriate responses, thereby improving the user experience.
[1664] System Overview
[1665] The system mainly consists of the following components:
[1666] How users upload information
[1667] The means by which the server receives, stores, and analyzes the uploaded information
[1668] A means for the server to build a learning model from the analyzed data and generate responses based on input from the user through an interactive session
[1669] A means for the terminal to display the generated response
[1670] Emotion engine that recognizes user emotions
[1671] Hardware and Software Configuration
[1672] Server: Use a high-performance server machine (e.g., AWS EC2) and use an S3 bucket to store data.
[1673] Device: The user uses a PC, smartphone, or AR / VR device.
[1674] Natural Language Processing (NLP) library: Spacy is used to tokenize text data and extract meaning and patterns.
[1675] Image recognition technology: Uses OpenCV and TensorFlow to analyze features and behavioral patterns in images and videos.
[1676] Machine learning libraries: Use TensorFlow and PyTorch to generate learning models from analytical data.
[1677] Emotion Detection Library: Extracts emotions from user text and speech using the NRC Emotion Lexicon and Google Cloud Speech-to-Text.
[1678] Program processing
[1679] Users use their devices to access web pages or apps and upload data such as letters, photos, and videos. The server receives this data and temporarily stores it in storage. The stored data is analyzed using NLP and image recognition technologies. Based on the analyzed data, the server uses machine learning libraries to generate a learning model.
[1680] When a user logs in to the system and starts an interaction session, the server loads the saved learning model and waits for user input. When the user enters a message in the interaction window using text or voice, the server receives the input and recognizes the emotion using the emotion engine. Based on the recognized emotion, the server generates an appropriate response using the generative AI model and sends it to the device. The device displays the generated response, and the response is presented visually and audibly in the virtual environment, especially when using an AR / VR device.
[1681] Specific examples
[1682] 1. When a user uploads information about their dog, Pochi:
[1683] Users can access a dedicated upload page to send photos, videos, letters, etc. of Pochi.
[1684] The server receives this data, temporarily stores it, and analyzes and stores the meta information for each.
[1685] The server analyzes the stored data, analyzing the text of letters using natural language processing and processing photos and videos using image recognition technology.
[1686] Based on the data obtained from the analysis, the server trains a learning model to learn Pochi's personality and behavioral patterns.
[1687] When the user says, "Pochi, let's play a lot today," the server generates a response, "Yes! Let's go to the park!" and sends it to the user's device.
[1688] The server recognizes the user's joy through an emotion engine and generates a more positive response.
[1689] The user's AR device displays this response in a virtual environment, allowing the user to realistically relive their memories with Pochi.
[1690] Prompt Sentence Examples
[1691] "Please tell me how to recognize user emotions and generate appropriate responses."
[1692] "Please explain in detail the process of analyzing the uploaded data."
[1693] "Please tell us specifically what experience you would like to have through an interaction session with your dog."
[1694] Through the above process, users can achieve emotional stability and a sense of happiness through virtual conversations with loved ones. By using an emotion engine, this system provides a more emotionally in line with the user's needs.
[1695] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1696] Step 1:
[1697] Users access a specified web page or app from their device (PC or smartphone) and upload information such as letters, photos, and videos.
[1698] Specific behavior: The user clicks the upload button, selects a file from the file selection dialog, and presses the send button.
[1699] Input: Data uploaded by users, such as letters, photos, videos, etc.
[1700] Output: Data uploaded to the server.
[1701] Step 2:
[1702] The server receives the uploaded information and temporarily stores it in storage.
[1703] Specific behavior: The server receives the HTTP request, saves the file in a temporary directory on the server side, and returns an upload completion message.
[1704] Input: Uploaded information.
[1705] Output: Information stored on the server.
[1706] Step 3:
[1707] The server analyzes the information's metadata (file type, size, content, etc.) and stores it in a database in an appropriate format.
[1708] Specific operation: The server checks the file format, extracts meta information for each file (e.g., 'photo.jpg', size: 2MB, creation date: 2023-10-15), and records it in the database.
[1709] Input: Temporarily stored information.
[1710] Output: Metadata and information stored in a database.
[1711] Step 4:
[1712] The server analyzes the stored information, specifically using natural language processing (NLP) for text data and image recognition technology for images and videos.
[1713] How it works: The server uses NLP libraries (e.g., Spacy) to tokenize text and extract important meaning and patterns. For images and videos, it uses OpenCV, TensorFlow, etc. to analyze facial recognition and behavioral patterns.
[1714] Input: Information stored in a database.
[1715] Output: Parsed data.
[1716] Step 5:
[1717] The server constructs a learning model based on the analysis data.
[1718] Specific operation: The server uses machine learning libraries (e.g., TensorFlow, PyTorch) to learn individual features from the analysis data and generate a unique learning model.
[1719] Input: Parsed data.
[1720] Output: The trained model.
[1721] Step 6:
[1722] The server uses an emotion engine to recognize the user's emotions. The emotion engine uses natural language processing and speech recognition technologies to extract and distinguish emotions from the user's text and voice input.
[1723] Specific behavior: The server calls an emotion detection library (e.g., NRC Emotion Lexicon, Google Cloud Speech-to-Text) to extract emotion labels (e.g., joy, sadness, anger) from text or audio.
[1724] Input: User text and voice input.
[1725] Output: User emotion label.
[1726] Step 7:
[1727] A user logs into a system and begins an interactive session.
[1728] Specific operation: The user enters the username and password on the login page and presses the login button. The server performs authentication, and if successful, displays the start screen of the interactive session.
[1729] Enter your username and password.
[1730] Output: The start screen of an interactive session after successful authentication.
[1731] Step 8:
[1732] The server loads the saved learning model and waits for user input.
[1733] Specific operation: The server loads the appropriate learning model into memory based on the user profile and waits for data input from the user.
[1734] Input: The training model.
[1735] Output: Server in standby state.
[1736] Step 9:
[1737] The user enters a message in the dialogue window and sends it.
[1738] Specific behavior: When the user enters text into the dialogue window and clicks the submit button, the data is sent to the server.
[1739] Input: A text message.
[1740] Output: The text message sent to the server.
[1741] Step 10:
[1742] The server receives the user's input and recognizes the user's emotions through an emotion engine.
[1743] Specific operation: The server analyzes the input text and audio data, extracts emotion labels, and passes them to the next step.
[1744] Input: Text or audio data.
[1745] Output: Extracted emotion labels.
[1746] Step 11:
[1747] The server uses a learning model to generate an appropriate response based on the recognized emotion.
[1748] Specific operation: The server generates a response sentence using a generative AI model (e.g., GPT-3) based on the recognized emotion and stored features.
[1749] Input: emotion labels and a trained model.
[1750] Output: The generated response sentence.
[1751] Step 12:
[1752] The device displays the generated response and, especially when using an AR / VR device, presents the response visually and audibly along with the virtual environment.
[1753] Specific behavior: The server sends the generated response to the user's device, which displays it. If an AR / VR device is used, the device displays the response visually in the virtual environment and also plays an audio response.
[1754] Input: The generated response sentence.
[1755] Output: The response displayed on the user's terminal.
[1756] (Application example 2)
[1757] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1758] Currently, in physical stores, customers have few easy ways to obtain information about products or specific areas within the store, and there is a need to improve the customer experience.In addition, there is a lack of systems that can respond individually to customers' emotions, making it difficult to improve customer satisfaction.
[1759] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1760] In this invention, the server includes a means for a user to upload data, a means for the server to receive and store the uploaded data, and a means for the server to analyze the stored data. This enables information about a specified object in a virtual space to be displayed in a physical store. The server also includes a means for constructing a learning model from the analyzed data, a means for a user to start an interactive session, a means for the server to generate a response based on input from the user through the interactive session, a means for the terminal to display the generated response, and an emotion recognition means for generating a response based on the user's emotion. This enables individual responses based on the user's emotions, which is expected to improve customer satisfaction.
[1761] The "means for users to upload data" refers to an interface that allows users to use their own terminals to send data such as photos, videos, and text to the system.
[1762] "Means for the server to receive and store uploaded data" refers to the function by which the server receives data sent by the user and stores it temporarily or permanently.
[1763] "Means for the server to analyze the stored data" refers to the technology that the server uses to process the stored data and decrypt and analyze its contents.
[1764] "Means by which the server constructs a machine learning model from the analyzed data" means the process by which the server creates a machine learning model based on the analyzed data and uses the model to generate appropriate responses to future data inputs.
[1765] The "means by which a user initiates an interactive session" refers to the interface or functionality by which a user accesses the system and initiates an interaction.
[1766] "Means for a server to generate a response based on input from a user throughout an interactive session" refers to a technique in which a server receives text or voice input from a user and generates a reply based thereon.
[1767] "Means for displaying the response generated by the terminal" refers to a function that allows the user's device (smartphone, smart glasses, etc.) to visually or audibly present the response generated by the server to the user.
[1768] "Means for displaying information about a specified object in a virtual space" refers to a function that allows a user to specify a specific item or area in a physical store and visually provide the user with related information using AR technology.
[1769] "Emotion recognition means for generating responses according to the user's emotions" is a technology that analyzes the user's input (text or voice), identifies the user's emotional state, and generates an appropriate response accordingly.
[1770] The system according to the present invention is a system that allows users to recreate memories of their precious family members and pets in a virtual space and engage in two-way conversations in real time. A specific embodiment of this system is described below.
[1771] Hardware and software used
[1772] This system uses the following hardware and software:
[1773] Hardware: Smartphones, smart glasses
[1774] Software: Google Cloud Vision API, Google Cloud Natural Language API, Python, OpenCV, TextBlob, pyttsx3
[1775] Uploading and saving data
[1776] Users use their smartphones or PCs to access specific web pages or applications and upload data such as photos, videos, and text. The server receives this data and stores it temporarily or permanently. The metadata (file type, size, and content) of the temporarily stored data is also analyzed, and it is then categorized and stored appropriately.
[1777] Data analysis and learning
[1778] The server analyzes the stored data. For photos and videos, it uses the Google Cloud Vision API to analyze the characteristics of specific people and pets. For text data (such as letters and conversation logs), it uses the Google Cloud Natural Language API to extract emotions and meanings, and performs sentiment analysis using TextBlob. This allows the server to build a learning model and prepare to generate responses based on the user's specific emotions and memories.
[1779] Emotion Engine Operation
[1780] The server uses an emotion engine to recognize emotions from the user's text and voice input. It uses the Google Cloud Natural Language API and TextBlob to extract emotions from the user's text and voice input and determine their emotional state.
[1781] Initiating an interactive session and generating responses
[1782] A user logs in to the system and begins an interactive session. The server loads the saved learning model and waits for input from the user. The user inputs a message via text or voice, which the server receives and recognizes the user's emotions through an emotion engine. It then generates a response based on the user's emotions and outputs it to the terminal, either displayed or via voice synthesis.
[1783] Specific examples
[1784] If a user wants to recreate memories with their beloved dog in a virtual space, they can take the following steps.
[1785] 1. Users send data such as photos, videos, and letters of their beloved dog through a dedicated upload page.
[1786] 2. The server receives the data, analyzes the metadata, and saves it. It then uses the Google Cloud Vision API to recognize images and videos, and the Google Cloud Natural Language API to analyze the sentiment of the text data.
[1787] 3. When the user says, "Pochi, what do you want to do today?", the server generates a response, "I want to play in the park!", and displays it on the device. If the server recognizes the user's emotion of joy, it generates an even more positive response.
[1788] Prompt Sentence Examples
[1789] I want to know what memories you have of this product.
[1790] "I want you to recreate the story of when you purchased this product."
[1791] This allows users to experience emotional stability and happiness through virtual interactions with loved ones. The system realized by the invention, in particular by adding an emotion engine, can provide a more sophisticated and emotionally in line with the interactive experience.
[1792] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1793] Step 1: Upload your data
[1794] A user uses his / her own device (e.g., smartphone, PC) to access a specific web page or application and upload photos, videos, text data, etc. The input includes the files and data selected by the user. The server receives this data and temporarily stores it.
[1795] Output: Received and stored data.
[1796] Step 2: Data storage and metadata analysis
[1797] The server receives the uploaded data and stores it in temporary storage. In parallel, the server parses the metadata (file type, size, content) of each uploaded data file. As input, it contains the uploaded data file. The server parses the metadata and stores it in the appropriate format.
[1798] Output: Parsed metadata and organized data files.
[1799] Step 3: Analyze the data
[1800] The server starts the process of analyzing the stored data. For text data, it uses the Google Cloud Natural Language API to extract sentiment and meaning. For images and videos, it uses the Google Cloud Vision API to analyze features of specific people or pets. The input includes the stored data files. The results of the data analysis are feature-extracted data and sentiment analysis results.
[1801] Output: Feature extracted data and sentiment analysis results.
[1802] Step 4: Building a learning model
[1803] The server creates a machine learning model based on the analyzed data. It learns past interactions and behavioral patterns with specific people and pets, and uses the model to generate appropriate responses to future user input. The analyzed data is included as input. The server builds and stores the machine learning model.
[1804] Output: The built machine learning model.
[1805] Step 5: Starting an interactive session
[1806] A user logs into the system and accesses an interface to start an interactive session. When the user logs in, the server loads the saved learning model and waits for input from the user. The input includes the user's login operation. The server loads the learning model and enters an interactive waiting state.
[1807] Output: Training model loaded and waiting.
[1808] Step 6: Receiving User Input and Emotion Recognition
[1809] The user inputs a message using text or voice and sends it. When the server receives it, it uses Google Cloud Natural Language API or TextBlob to recognize the user's emotions. The input contains the user's text or voice data. The server performs emotion recognition and returns the result.
[1810] Output: User's emotional state and analysis results.
[1811] Step 7: Generate a response
[1812] The server generates an appropriate response based on the user's emotional state using the learning model. The input includes the user's emotional state and the learning model. The server sends the generated response to the user's device.
[1813] Output: The generated response.
[1814] Step 8: Display or speak the response
[1815] The terminal receives the response sent by the server and displays it to the user, possibly providing an audible response using a speech synthesis engine (pyttsx3). The input includes the generated response. The terminal presents the response to the user visually or audibly.
[1816] Output: The response presented to the user.
[1817] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1818] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1819] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1820] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1821] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1822] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1823] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1824] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1825] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1826] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1827] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1828] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1829] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1830] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1831] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1832] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1833] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1834] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1835] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1836] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1837] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1838] The following is further disclosed regarding the above embodiment.
[1839] (Claim 1)
[1840] a means for users to upload data;
[1841] a means by which the server receives and stores the uploaded data;
[1842] a means for the server to analyze the stored data;
[1843] A means for the server to construct a learning model from the analysis data;
[1844] a means for a user to initiate an interactive session;
[1845] means for the server to generate a response based on input from the user through an interactive session;
[1846] means for the terminal to display the generated response;
[1847] A system including:
[1848] (Claim 2)
[1849] 2. The system of claim 1, wherein natural language processing is used to analyze the data.
[1850] (Claim 3)
[1851] The system according to claim 1, wherein image recognition technology is used for data analysis.
[1852] "Example 1"
[1853] (Claim 1)
[1854] a means for users to upload data;
[1855] a means by which the server receives and stores the uploaded data;
[1856] a means for the server to analyze the stored data;
[1857] A means for the server to construct a learning model from the analysis data;
[1858] a means for a user to initiate an interactive session;
[1859] means for the server to generate a response based on input from the user through an interactive session;
[1860] means for the terminal to display the generated response;
[1861] a means for the terminal to generate a virtual environment and present visual and auditory responses;
[1862] A system including:
[1863] (Claim 2)
[1864] 2. The system of claim 1, wherein natural language processing is used to analyze the data.
[1865] (Claim 3)
[1866] The system according to claim 1, wherein image recognition technology is used for data analysis.
[1867] "Application Example 1"
[1868] (Claim 1)
[1869] a means for users to upload data;
[1870] a means by which the server receives and stores the uploaded data;
[1871] a means for the server to analyze the stored data;
[1872] A means for the server to construct a learning model from the analysis data;
[1873] a means for a user to initiate an interactive session;
[1874] means for the server to generate a response based on input from the user through an interactive session;
[1875] means for the terminal to display the generated response;
[1876] A means for a virtual assistant to act as a guide when a user is experiencing purchasing an item in a virtual space;
[1877] A system including:
[1878] (Claim 2)
[1879] 2. The system of claim 1, wherein natural language processing is used to analyze the data.
[1880] (Claim 3)
[1881] The system according to claim 1, wherein image recognition technology is used for data analysis.
[1882] "Example 2: Combining Emotion Engines"
[1883] (Claim 1)
[1884] a means for users to upload information;
[1885] a means by which the server receives and stores the uploaded information;
[1886] a means by which the server analyzes the stored information; and
[1887] A means for the server to construct a learning model from the analysis data;
[1888] a means for a user to initiate an interactive session;
[1889] means for the server to generate a response based on input from the user through an interactive session;
[1890] means for the terminal to display the generated response;
[1891] An emotion engine means for the server to recognize the emotion of the user;
[1892] A system including:
[1893] (Claim 2)
[1894] 2. The system of claim 1, wherein natural language processing is used to analyze the information.
[1895] (Claim 3)
[1896] 2. The system according to claim 1, wherein image recognition technology is used for information analysis.
[1897] "Application example 2 when combining emotion engines"
[1898] (Claim 1)
[1899] a means for users to upload data;
[1900] a means by which the server receives and stores the uploaded data;
[1901] a means for the server to analyze the stored data;
[1902] A means for the server to construct a learning model from the analysis data;
[1903] a means for a user to initiate an interactive session;
[1904] means for the server to generate a response based on input from the user through an interactive session;
[1905] means for the terminal to display the generated response;
[1906] means for displaying information about a specified object in the virtual space;
[1907] emotion recognition means for generating a response according to the emotion of the user;
[1908] A system including:
[1909] (Claim 2)
[1910] 2. The system of claim 1, wherein natural language processing is used to analyze the data.
[1911] (Claim 3)
[1912] The system according to claim 1, wherein image recognition technology is used for data analysis. [Explanation of symbols]
[1913] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for users to upload data; a means by which the server receives and stores the uploaded data; a means for the server to analyze the stored data; A means for the server to construct a learning model from the analysis data; a means for a user to initiate an interactive session; means for the server to generate a response based on input from the user through an interactive session; means for the terminal to display the generated response; A system including:
2. 10. The system of claim 1, wherein natural language processing is used to analyze the data.
3. The system according to claim 1, wherein image recognition technology is used for data analysis.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A