system

JP2026085761APending Publication Date: 2026-05-25SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-11-13
Publication Date
2026-05-25

AI Technical Summary

Technical Problem

In modern society, the increasing number of people living alone or in isolated lives leads to loneliness, especially during meals, which can increase psychological stress, and conventional shared meal experiences face challenges such as time adjustment and hindered communication due to differing hobbies and preferences.

Method used

A system that analyzes images of meals using a generation AI model to create a virtual avatar eating a similar meal, adjusts the avatar's movements to match the user's pace, and generates personalized conversations using natural language processing technology to provide a shared dining experience.

Benefits of technology

The system reduces feelings of loneliness and enhances mealtime satisfaction by offering a virtual dining partner that adapts to the user's eating pace and engages in meaningful conversations, creating a comfortable and fulfilling dining environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026085761000001_ABST
    Figure 2026085761000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of receiving images taken by the user, A means for analyzing the received image and identifying elements within the image, A means for generating similar meal images based on identified elements, A means of displaying avatars using generated images and providing a shared dining experience, A means for analyzing user behavior and adjusting avatar behavior, A means of generating and providing casual conversation using natural language processing technology, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0004] , , , ,

[0005] , , , , ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, with the increase in the number of people living alone or in isolated lives, loneliness has become a serious social problem. In particular, being alone during meals can increase psychological stress. Conventional shared meal experiences have problems such as time adjustment and pacing with others, and furthermore, smooth communication may be hindered by differences in the hobbies and preferences of the people at the same table. There is a need for means to solve such problems and provide more fulfilling mental health.

Means for Solving the Problems

[0005] This invention provides a system that offers a shared dining experience by analyzing images of meals taken by the user using a generation AI model and creating an avatar that eats a similar meal. This system analyzes the user's eating speed in real time and adapts the avatar's movements accordingly. Furthermore, it generates and provides personalized small talk using natural language processing technology based on the user's voice input and profile information, thereby realizing a friendly and interesting conversation for the user. This system makes it possible to reduce feelings of loneliness and create a comfortable dining environment.

[0006] "User-captured images" refer to digital data, including visual information of food, that a user acquires using their own device and sends to the system.

[0007] "Analysis" is the process of using received digital data to identify its content and characteristics and extract information from it.

[0008] "Generating similar meal images" is the process of creating new digital images that are similar in appearance and composition to the original meal, based on the information obtained through analysis.

[0009] "Displaying avatars and providing a shared dining experience" is a process in which a virtual character appears on the digital screen using a generated image of the meal, acting as the user's visual dining partner and creating a shared experience during meals.

[0010] "Analyzing user behavior and adjusting avatar behavior" refers to the process of monitoring and analyzing the user's behavior while eating in real time, and then appropriately changing the avatar's movements and reactions based on that data.

[0011] "Natural language processing technology" is a technology that enables computers to analyze, generate, and understand natural human language, and uses it to realize meaningful dialogue between humans and computers.

[0012] "Generating and providing casual conversation" refers to the process by which a system creates appropriate conversation content based on user information and presents it to the user as audio or text. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system for virtually realizing the experience of sharing a meal with a user, with the aim of reducing feelings of loneliness. This system provides users with a sense of psychological security by allowing them to share a meal through an avatar. The embodiments for carrying out the invention will be described in detail below.

[0035] The user first takes a picture of their meal using their device. The device sends this image to the server. The server analyzes the received image and identifies the type and characteristics of the food it contains. Based on the analysis results, the server has the function to generate a similar meal image. This image resembles what the user is eating and is used for the avatar to participate in the shared meal.

[0036] Next, the server sends the generated meal image to the terminal, which uses it to display the avatar on the user interface. The avatar operates in real time to provide the user with a visual and interactive shared-meal experience. Furthermore, the terminal captures the user's eating movements through its camera and sends this data to the server. The server analyzes this data to determine the user's eating speed and pace. This allows the avatar's movements to be adjusted to harmonize with the user's movements.

[0037] Furthermore, the server uses natural language processing technology to generate conversations with users. Specifically, based on the user's voice input and pre-registered profile information, it selects topics of interest to the user and generates appropriate small talk accordingly. In this way, the avatar facilitates natural conversations with users and makes the shared dining experience more enjoyable.

[0038] For example, if a user is eating pasta, the server generates a similar pasta image for the avatar based on the image of the pasta. The avatar then uses this image to perform the action of eating pasta and initiates a conversation with the user, such as "What are you thinking about while you eat today?" This system functionality allows users to enjoy a meal with a virtual partner even when they are eating alone.

[0039] Thus, this invention not only alleviates the user's feelings of loneliness, but also enriches mealtimes and enhances their mental satisfaction.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] The user takes a picture of their meal using their device and uploads the image to the server via the application. This image file contains details about the meal.

[0043] Step 2:

[0044] The server processes the received images and extracts features such as the type, color, and shape of the food through image analysis algorithms. This analysis generates data to identify the food.

[0045] Step 3:

[0046] The server generates similar meal images for the avatar based on the data obtained through analysis. Using a generation AI model, it digitally constructs meal images that visually resemble the user's actual meal.

[0047] Step 4:

[0048] The server sends the generated similar meal image to the user's terminal. The terminal receives this image, displays an avatar on the user interface, and starts the shared meal simulation.

[0049] Step 5:

[0050] The device uses its camera to record the user's eating habits in real time. This data is sent to a server and serves as foundational data for analyzing the user's eating speed and rhythm.

[0051] Step 6:

[0052] The server analyzes the collected motion data and adjusts the avatar's movements to match the user's eating pace. This changes the avatar's behavior so that it eats at the same speed as the user.

[0053] Step 7:

[0054] The server uses natural language processing technology to generate personalized conversation content based on the user's voice input and profile information. This information is sent to the terminal, and the avatar begins a conversation with the user.

[0055] Step 8:

[0056] The device uses the generated conversation content to control the avatar and facilitate interaction with the user. The avatar brings up topics that are familiar to the user and enriches the shared dining experience through conversation.

[0057] Step 9:

[0058] The device detects that the user has finished eating and reports this to the server. The server ends the shared meal session and presents the user with an option to provide feedback.

[0059] (Example 1)

[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0061] In society, eating alone can be a factor that increases feelings of loneliness. Furthermore, eating alone lacks opportunities for social interaction, which can lower individual life satisfaction. This invention aims to alleviate such feelings of loneliness and make solitary mealtimes more fulfilling.

[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0063] In this invention, the server includes means for receiving images captured by an information-providing terminal, means for analyzing the received images and identifying objects within the images, and means for generating similar meal images based on the identified objects. This makes it possible to share meals with end users through virtual representations and provide a sense of psychological reassurance.

[0064] A "terminal that provides information" is a device for users to input or manipulate information, and in this case, it is a device that has the function of taking and transmitting images of a meal.

[0065] An "object" is a specific element present in the received image that is identified through analysis, and in this system, it mainly refers to food.

[0066] "Virtual representation" refers to avatars and simulations that are visually displayed on a screen using computer technology, providing an interactive shared dining experience with the user.

[0067] "Free conversation" refers to dialogue-style conversation content generated using natural language processing technology, providing discourse based on the user's interests and concerns.

[0068] An "end user" refers to the person who uses this system to share their dining experience, taking pictures and interacting with the system.

[0069] The embodiments for carrying out this invention are shown below.

[0070] First, the user takes a picture of their meal using a device that provides information, such as a smartphone or tablet. This image is then sent to a server via Wi-Fi or a mobile network. The device is equipped with a camera module and a communication module.

[0071] The server uses hardware and image analysis software (e.g., TENSORFLOW® or OpenCV) to process the received image data, identify objects within the image, and determine the type of meal. Based on these identified objects, the server uses a generative AI model (e.g., Stable Diffusion) to generate similar meal images. During this process, the prompt "Generate an image of a dish similar to XX (XX is the name of the analyzed meal)" is input.

[0072] Next, the server sends the generated virtual representation image data to the terminal. The terminal uses 3D modeling software (for example, Unity) to visually process this image and display it as a virtual representation in the user interface. Through the virtual dining scene displayed on the terminal's screen, the user can enjoy an interactive shared dining experience. This makes it possible to enjoy a meal with a virtual partner, even when dining alone.

[0073] Furthermore, the server utilizes natural language processing technology (e.g., GPT-4®) to enable free-flowing conversation between the end user and the virtual representation. It analyzes the user's voice input and pre-entered profile information to provide personalized conversation themes and content. This allows users to converse flexibly on topics of interest and make mealtimes more enjoyable.

[0074] For example, if a user is eating pasta, the server analyzes the image they take and generates a virtual representation of similar pasta. Based on this, the avatar on the device can perform the action of eating pasta and ask the user, "Have you tried any pasta recipes today?" In this way, even though the user is actually alone, they can gain a sense of psychological comfort through virtual interaction.

[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0076] Step 1:

[0077] The user launches a camera application on their device and takes a picture of their meal. This operation provides an input means for capturing image data. The device saves the captured image and retains the image data in a file format (e.g., JPEG or PNG) in preparation for the next processing step.

[0078] Step 2:

[0079] The device sends the saved image data to the server via the network (Wi-Fi or mobile data). In this step, the image data is transferred to the server in the form of an HTTP POST request. The input from the device is the image data, and the output is a response indicating that the transmission is complete.

[0080] Step 3:

[0081] The server analyzes the received image data. Specifically, it uses image analysis software (for example, OpenCV or TensorFlow) to identify the type of meal. The input is image data received from the terminal, and the type of meal is output as the result of the analysis. Features within the image are detected, and objects are recognized using a predetermined algorithm.

[0082] Step 4:

[0083] The server utilizes a generative AI model (e.g., Stable Diffusion) based on the analyzed meal type to generate images of similar meals. In this step, the AI ​​model is input using a prompt in the form of "Generate images of dishes similar to XX (where XX is the name of the analyzed meal)." The model then outputs the generated meal images.

[0084] Step 5:

[0085] The server sends the generated similar meal image to the terminal. In this process, the generated image is sent to the terminal as an HTTP response. The output is that the terminal receives the image data.

[0086] Step 6:

[0087] The terminal processes the received food images using 3D modeling software (e.g., Unity) and displays them as a virtual representation in the user interface. The input is the received image data, and the output is the displayed virtual representation. A visual interface is generated in which the avatar performs the action of eating.

[0088] Step 7:

[0089] The device uses its camera to video record the user's actions while eating and immediately transmits the data to the server. The input is video data, and the output is the transmission of that data to the server. The recorded action data is analyzed in real time.

[0090] Step 8:

[0091] The server analyzes the received motion data to determine the user's eating pace. The input is motion data, and the output is pace information based on that motion. The motion analysis algorithm evaluates the user's motion speed and synchronizes it with the motion speed of a virtual representation.

[0092] Step 9:

[0093] The server uses natural language processing techniques (e.g., GPT-4) to generate free-flowing conversations tailored to the user. Input consists of the user's past voice inputs and profile information, while output is the conversation content and themes presented to the user. This facilitates engaging dialogue with the user.

[0094] (Application Example 1)

[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] In modern society, the number of people who eat alone is increasing, which is leading to increased feelings of loneliness and psychological dissatisfaction. Furthermore, because opportunities to enjoy meals with others in the real world are limited, users cannot easily enjoy the experience of eating together. Additionally, even in virtual dining experiences, users may struggle to engage in natural and enjoyable conversations.

[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0098] In this invention, the server includes means for receiving video footage captured by the user, means for analyzing the received video footage and identifying objects within the video footage, and means for displaying a virtual character and providing a shared dining experience using the generated video footage. This allows the user to alleviate feelings of loneliness through shared dining with a virtual character and easily obtain diverse shared dining experiences.

[0099] A "user" refers to someone who uses this system to virtually share a dining experience.

[0100] "Video" refers to image data captured by the user using their device while eating.

[0101] "Object" refers to specific meals or related items included in the received video.

[0102] "Food images" refer to image data generated that resembles the food that the user is actually eating.

[0103] A "virtual character" refers to a digital avatar displayed using generated food images to create a shared dining experience with the user.

[0104] "Actions" refer to a series of actions performed by the user or virtual character while eating.

[0105] "Natural language processing technology" refers to the technology that enables computers to understand human language and generate meaningful conversations.

[0106] "Informal conversation" refers to communication using everyday, familiar topics.

[0107] "Conversation content" refers to the content of the conversation generated during the interaction with the user.

[0108] The system that realizes this application example allows users to experience shared dining in a virtual environment by linking their terminal with a server. The system mainly consists of the following elements:

[0109] 1. User's device: The user's device is equipped with a camera that captures images of the meal and sends those images to the server. The device can be a smartphone, smart glasses, or a head-mounted display.

[0110] 2. Server Function: The server executes a program developed in Python to analyze received video and identify objects related to food. Specifically, it uses PIL (Python Imaging Library) for video analysis and data processing. The server then generates food images similar to the identified meal. The generated images may utilize basic image generation models, or in some cases, generative AI models.

[0111] 3. Virtual Character Movement and Conversation: Based on the generated food images, the server displays a virtual character on the user's device to provide a shared dining experience. The device's camera captures the user's movements in real time and sends this data to the server. The server analyzes this data and synchronizes the virtual character's movements with the user's movements.

[0112] 4. Conversation generation using natural language processing: The server uses natural language processing techniques to generate conversations with the user. These conversations are based on the user's voice input and profile information, making the shared dining experience richer and more natural.

[0113] As a concrete example, if a user is eating pasta alone at home and uses this system, a virtual character will appear on the screen and begin a conversation about pasta. The virtual character will adjust its behavior to match the user's eating pace and continue the conversation with questions such as, "What are you thinking about while you eat today?"

[0114] Examples of prompt messages include: "Analyze the user's video of their meal, identify the type of food, and represent a similar meal with a virtual character. Next, analyze the user's conversation from the voice input and generate relevant conversation based on that information."

[0115] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0116] Step 1:

[0117] The user uses their device to film their meal and sends the video from the device to the server. The input data is the video captured by the user's camera, and the output is the data transfer to the server. The device requires a stable internet connection and sends the video to the server immediately after filming.

[0118] Step 2:

[0119] The server analyzes the received video of the meal and identifies objects within the video. The input here is the video data sent in step 1, and the output is information about the objects related to the identified meal. The server uses image analysis libraries such as PIL to process the data and identify the type of meal.

[0120] Step 3:

[0121] The server generates images of similar foods based on the identified object information. The input in this step is the object information obtained from step 2, and the output is the generated images of similar foods. The server uses a simple generative AI model to create images that resemble what the user is eating.

[0122] Step 4:

[0123] The generated food images are sent from the server to the user's terminal, which then uses them to display a virtual character. The input is the generated food image data, and the output is the display of the virtual character on the user interface. The terminal receives this image data and displays it immediately.

[0124] Step 5:

[0125] The user's actions while eating are captured by the device's camera, and the device sends this action data to the server. The input here is the user's video activity, and the output is the transfer of action data to the server. The device continuously monitors the user's actions and records the data.

[0126] Step 6:

[0127] The server analyzes the received motion data and synchronizes the virtual character's movements with the user's movements. The input in this step is the motion data from step 5, and the output is the adjusted virtual character's movements. The server determines the user's eating pace and calculates the character's movements accordingly.

[0128] Step 7:

[0129] The server uses natural language processing technology to generate conversations with the user. Input is the user's voice input or existing profile information, and output is the generated conversation content. A generative AI model is used for this process, forming informal conversations based on prompts.

[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0131] This invention aims to provide a more personalized dining experience by combining an emotion engine with a system designed to create a virtual shared dining experience with the user. This system recognizes the user's emotions during meals and adjusts the avatar's responses and conversation content accordingly, thereby creating a comfortable dining environment for the user.

[0132] The user takes a picture of their meal using their device and sends the image to the server. The server analyzes the image and identifies the characteristics of the meal. Furthermore, the server generates similar meal images based on the collected image data and sends them to the device for use in the avatar's shared meal experience.

[0133] The device displays an avatar on the user interface, providing the user with a shared dining experience. Furthermore, the device uses a camera and microphone to monitor the user's facial expressions and voice tone in real time. This data is sent to a server, where an emotion engine analyzes the user's emotions. Based on the analyzed emotion information, the server adjusts the avatar's movements and conversation content.

[0134] For example, if a user is smiling, the server analyzes that positive emotion and sets the avatar to offer more cheerful topics. Conversely, if a user appears stressed, the server takes that emotion into consideration and has the avatar offer encouraging or relaxing topics.

[0135] The avatar uses natural language processing technology to interact with users and respond in a way that is sensitive to their emotions. Furthermore, the emotion engine learns from past user emotional data and predicts future emotional states, allowing the avatar's responses to be pre-adjusted. For example, if past data predicts that a user is prone to stress on certain days of the week, the avatar can prepare relaxing topics for those days.

[0136] Thus, the present invention makes it possible to reduce feelings of loneliness and mental stress, and to make dining a richer experience, by providing an interactive dining experience based on the user's emotions.

[0137] The following describes the processing flow.

[0138] Step 1:

[0139] The user takes a picture of their meal with their device and uploads it to the server through the application. This image contains information including details about the meal.

[0140] Step 2:

[0141] The server processes the received meal images using an image analysis algorithm to identify the type and characteristics of the meal and extract data. This analysis completes the analysis of the meal's contents.

[0142] Step 3:

[0143] The server generates similar meal images based on the analysis data and sends these images to the terminal for use in displaying the user's avatar.

[0144] Step 4:

[0145] The device displays images of similar meals generated alongside the avatar on the user interface, initiating a shared meal experience with the avatar. During this process, the user virtually dines with the avatar.

[0146] Step 5:

[0147] The device uses its camera and microphone to record the user's facial expressions and voice tone in real time. This data is sent to a server for emotion recognition.

[0148] Step 6:

[0149] The server uses an emotion engine to analyze the user's facial expressions and voice data to determine the user's emotional state. Based on this information, it prepares to adjust the avatar's reactions and conversation content.

[0150] Step 7:

[0151] Based on the sentiment analysis results, the server selects a suitable conversation topic for the user, generates conversation content using natural language processing technology, and sends it to the terminal.

[0152] Step 8:

[0153] The terminal applies the conversation content received from the server to the avatar and begins interacting with the user. At this time, the avatar's tone and topics will reflect the user's emotions.

[0154] Step 9:

[0155] The server learns from past emotional data and pre-adjusts the avatar's behavior and conversation based on predicted emotional states in future sessions. This continuous learning makes it possible to further personalize the user experience.

[0156] (Example 2)

[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0158] In modern society, eating is not merely about nutrition; it is an important act that involves psychological and social elements. However, especially when eating alone, feelings of loneliness and mental stress can occur, which can reduce satisfaction with meals. This invention aims to solve these problems and provide users with a richer and more fulfilling dining experience.

[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0160] In this invention, the server includes means for receiving video acquired by the user, means for analyzing the received video and identifying characteristics within the video, and means for generating similar visual data based on the identified characteristics. This allows the user to add visual and emotional interactive elements to existing meals, thereby reducing feelings of loneliness and alleviating stress.

[0161] "Users" refers to individuals or groups who operate the system and enjoy the experience.

[0162] "Video" refers to visual data that users capture and transmit to the system.

[0163] A "server" refers to a computing device that forms the core of a system and is responsible for receiving, analyzing, and providing information.

[0164] "Receiving" refers to the act of a server acquiring data sent by a user.

[0165] "Analysis" refers to the process of extracting and interpreting information from received data.

[0166] "Characteristics" refer to identifiable elements or features contained within images or data.

[0167] "Visual data" refers to digital image information generated based on analyzed characteristics.

[0168] A "virtual character" refers to a virtual personality created within a system for interacting with users.

[0169] "Empathic experience" refers to the emotional interaction that users have with virtual characters within a virtual environment, through shared perceptions.

[0170] "Actions" refer to the reactions and responses that the virtual character performs towards the user.

[0171] "Natural language processing technology" refers to the technology that enables computers to understand human language and generate appropriate dialogue.

[0172] "Emotional data" refers to the results of collecting and analyzing information that indicates the emotional state of users.

[0173] "Prediction" refers to the act of estimating future results or states based on past data.

[0174] This invention is a system that provides users with an interactive and emotionally engaging dining experience. The system consists of a terminal, a server, and multiple software components.

[0175] The user first uses the device's camera to capture video of their meal. The device also has a function to monitor the user's facial expressions and voice tone in real time. This makes it possible to capture the emotions the user is feeling while eating.

[0176] The acquired video is sent from the terminal to the server. The server uses image recognition software (e.g., TensorFlow or OpenCV) to analyze the received video. This analysis identifies the characteristics and features of the food. Based on the identified data, the server uses generative AI models (e.g., GANs) to generate similar visual data. This visual data is used to enrich the user's visual experience.

[0177] The server generates dialogue with a virtual character using natural language processing techniques (e.g., GPT-4) based on the analyzed data. The generated conversation is then delivered to the user via the terminal. For example, the prompt "How are you feeling today?" can be used to initiate the conversation.

[0178] Furthermore, monitoring data from the device is sent to the server's emotion analysis engine. This engine analyzes the user's emotions in real time and adjusts the virtual character's behavior and conversation content based on the analysis results. For example, if the user is smiling, it provides positive conversation content, and if the user is feeling stressed, it selects relaxing topics.

[0179] Furthermore, the server uses past emotional data as a learning platform to predict future emotional states. Based on these predictions, it can anticipate what emotions a user might feel on specific days of the week or in certain situations, and adjust the virtual character's reactions in advance. This ensures that users always have a comfortable and pleasant dining experience.

[0180] In this way, this system can deliver value to users beyond simply eating by providing a personalized dining experience based on their emotions and behavior.

[0181] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0182] Step 1:

[0183] The user takes a picture of their meal using the device's camera. The captured image is input as digital data from the camera device into the device's application. The device then sends this image to the server. As output, the video data is transferred to the server.

[0184] Step 2:

[0185] The server analyzes the received video data. Image recognition software (e.g., TensorFlow, OpenCV) is used for this analysis. The server identifies the characteristics of the food from the video data received as input and extracts characteristic elements (e.g., types of ingredients, types of dishes) based on that. The output is a dataset containing the characteristics of the identified dishes.

[0186] Step 3:

[0187] The server uses a generative AI model (e.g., GAN) to generate similar visual data based on the analyzed data. The characteristic data from step 2 is used as input. Based on this data, the server generates visual data for similar meals. The output is the generated similar visual data.

[0188] Step 4:

[0189] The server sends the generated visual data to the terminal. On the terminal, the virtual character is displayed as a customized avatar, providing a visual experience of shared dining. The generated visual data is interactively displayed to the user through the user interface. The output is a user interface that includes a visually rich avatar.

[0190] Step 5:

[0191] The device uses its built-in camera and microphone to monitor the user's facial expressions and voice tone in real time. Changes in facial expressions and voice signals are acquired as input data. This data is sent to a server, where an emotion analysis engine performs the analysis. The output is a recognition of the user's current emotional state.

[0192] Step 6:

[0193] The server uses generative AI models and natural language processing techniques to adjust the dialogue of the virtual character based on the analyzed emotional state. The input is the emotional data from step 5, and the server generates corresponding prompt sentences. This provides a conversation that matches the user's feelings. The output is the conversation text adapted for the user.

[0194] Step 7:

[0195] The server uses a database to learn from past user sentiment data and predict future emotional states. In this process, past sentiment data and patterns are used as input. Based on the prediction, preparations are made to pre-adjust the responses of the virtual character. The output is a response scenario corresponding to the predicted user's emotional tendencies.

[0196] (Application Example 2)

[0197] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0198] The challenge lies in reducing the feelings of loneliness and mental stress users experience during meals, and providing a richer dining experience. Furthermore, it is desirable to transform mealtime into an enjoyable experience by enabling personalized content selection based on user emotions.

[0199] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0200] This invention includes a server that analyzes the user's visual and auditory information in real time and dynamically adjusts content selection based on the results; a server that provides an interactive dining experience using virtual characters; and a server that generates and provides dialogue that is sensitive to the user's emotions using natural language processing technology. This allows the user to receive personalized responses and conversations during meals, alleviating feelings of loneliness and allowing them to have an enjoyable time.

[0201] A "user" is an individual who uses this system while consuming food.

[0202] "Image" refers to digital visual information captured by the user, including the contents of the meal and related objects.

[0203] "Visual information" refers to information such as the user's facial expressions and gestures, which are acquired through a camera device.

[0204] "Auditory information" refers to information about the user's pronunciation and surrounding sounds acquired through a microphone.

[0205] "Real-time analysis" refers to a process that rapidly processes information as soon as it is acquired and provides immediate feedback on the results.

[0206] "Content" refers to media information such as music, videos, and dialogues provided to users.

[0207] "Dynamic adjustment" refers to the process of changing the system's response and behavior in response to the user's emotions and circumstances.

[0208] A "virtual character" refers to an artificial character with a personality that interacts with the user on a digital interface.

[0209] "Natural language processing technology" is a technology that enables computers to understand and respond to human language.

[0210] An "interactive shared dining experience" refers to a simulated experience in which the user and a virtual character interact while sharing a meal.

[0211] The system that realizes this application is built on a foundation of mobile devices such as smartphones and tablets, and a cloud server. During a meal, the user takes a picture of their facial expression using the device's camera, and this visual information is sent to the server. The server uses an emotion analysis model built with machine learning libraries such as TensorFlow to analyze the received visual and auditory information in real time. The analyzed emotion data is input into a generative AI model, which generates content and dialogue tailored to the user.

[0212] Furthermore, the server uses natural language processing technology to enable smooth interaction with the user via a platform like Dialogflow. This allows the virtual character to adjust its actions and conversations in response to the user's real-time emotions, providing a personalized dining experience. For example, if the user wants to relax during a meal, a prompt such as "Play some relaxing music" is input into the generating AI model, and calming music is played on the device.

[0213] This application configuration allows users to enjoy personalized services that cater to their emotions and preferences, rather than simply receiving mechanically generated content. For example, when a user uses this system to take a break from their busy daily life, they can be refreshed with soothing music and encouraging messages from a virtual character.

[0214] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0215] Step 1:

[0216] The user uses the device's camera to take a picture of their facial expression while eating. At this time, the device receives the captured image data as input and prepares to send the visual information to the server.

[0217] Step 2:

[0218] The server initializes an emotion analysis model to analyze the received visual information. Here, image data is used as input, and facial expression analysis is performed using a TensorFlow-based model. The analysis result outputs the user's emotional state.

[0219] Step 3:

[0220] The server uses a generative AI model based on the analyzed sentiment data to select appropriate content and generate dialogue. It receives sentiment data as input, generates prompts via a natural language processing platform such as Dialogflow, and determines the content and dialogue to be played back to the user as output.

[0221] Step 4:

[0222] The terminal receives prompts and content instructions sent from the server and displays a virtual character on the interface based on them. The virtual character provides the user with personalized conversations and content based on the analysis results.

[0223] Step 5:

[0224] Users enjoy interacting with virtual characters while viewing content provided through their devices. As long as the interaction with the user continues, the device continuously uses its camera and microphone to collect further emotional data and transmit it to the server in real time. This allows the entire system to remain dynamically responsive.

[0225] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0226] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0227] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0228] [Second Embodiment]

[0229] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0230] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0231] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0232] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0233] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0234] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0235] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0236] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0237] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0238] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0239] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0240] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0241] This invention is a system for virtually realizing the experience of sharing a meal with a user, with the aim of reducing feelings of loneliness. This system provides users with a sense of psychological security by allowing them to share a meal through an avatar. The embodiments for carrying out the invention will be described in detail below.

[0242] The user first takes a picture of their meal using their device. The device sends this image to the server. The server analyzes the received image and identifies the type and characteristics of the food it contains. Based on the analysis results, the server has the function to generate a similar meal image. This image resembles what the user is eating and is used for the avatar to participate in the shared meal.

[0243] Next, the server sends the generated meal image to the terminal, which uses it to display the avatar on the user interface. The avatar operates in real time to provide the user with a visual and interactive shared-meal experience. Furthermore, the terminal captures the user's eating movements through its camera and sends this data to the server. The server analyzes this data to determine the user's eating speed and pace. This allows the avatar's movements to be adjusted to harmonize with the user's movements.

[0244] Furthermore, the server uses natural language processing technology to generate conversations with users. Specifically, based on the user's voice input and pre-registered profile information, it selects topics of interest to the user and generates appropriate small talk accordingly. In this way, the avatar facilitates natural conversations with users and makes the shared dining experience more enjoyable.

[0245] For example, if a user is eating pasta, the server generates a similar pasta image for the avatar based on the image of the pasta. The avatar then uses this image to perform the action of eating pasta and initiates a conversation with the user, such as "What are you thinking about while you eat today?" This system functionality allows users to enjoy a meal with a virtual partner even when they are eating alone.

[0246] Thus, this invention not only alleviates the user's feelings of loneliness, but also enriches mealtimes and enhances their mental satisfaction.

[0247] The following describes the processing flow.

[0248] Step 1:

[0249] The user takes a picture of their meal using their device and uploads the image to the server via the application. This image file contains details about the meal.

[0250] Step 2:

[0251] The server processes the received images and extracts features such as the type, color, and shape of the food through image analysis algorithms. This analysis generates data to identify the food.

[0252] Step 3:

[0253] The server generates similar meal images for the avatar based on the data obtained through analysis. Using a generation AI model, it digitally constructs meal images that visually resemble the user's actual meal.

[0254] Step 4:

[0255] The server sends the generated similar meal image to the user's terminal. The terminal receives this image, displays an avatar on the user interface, and starts the shared meal simulation.

[0256] Step 5:

[0257] The device uses its camera to record the user's eating habits in real time. This data is sent to a server and serves as foundational data for analyzing the user's eating speed and rhythm.

[0258] Step 6:

[0259] The server analyzes the collected motion data and adjusts the avatar's movements to match the user's eating pace. This changes the avatar's behavior so that it eats at the same speed as the user.

[0260] Step 7:

[0261] The server uses natural language processing technology to generate personalized conversation content based on the user's voice input and profile information. This information is sent to the terminal, and the avatar begins a conversation with the user.

[0262] Step 8:

[0263] The device uses the generated conversation content to control the avatar and facilitate interaction with the user. The avatar brings up topics that are familiar to the user and enriches the shared dining experience through conversation.

[0264] Step 9:

[0265] The device detects that the user has finished eating and reports this to the server. The server ends the shared meal session and presents the user with an option to provide feedback.

[0266] (Example 1)

[0267] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0268] In society, eating alone can be a factor that increases feelings of loneliness. Furthermore, eating alone lacks opportunities for social interaction, which can lower individual life satisfaction. This invention aims to alleviate such feelings of loneliness and make solitary mealtimes more fulfilling.

[0269] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0270] In this invention, the server includes means for receiving images captured by an information-providing terminal, means for analyzing the received images and identifying objects within the images, and means for generating similar meal images based on the identified objects. This makes it possible to share meals with end users through virtual representations and provide a sense of psychological reassurance.

[0271] A "terminal that provides information" is a device for users to input or manipulate information, and in this case, it is a device that has the function of taking and transmitting images of a meal.

[0272] An "object" is a specific element present in the received image that is identified through analysis, and in this system, it mainly refers to food.

[0273] "Virtual representation" refers to avatars and simulations that are visually displayed on a screen using computer technology, providing an interactive shared dining experience with the user.

[0274] "Free conversation" refers to dialogue-style conversation content generated using natural language processing technology, providing discourse based on the user's interests and concerns.

[0275] An "end user" refers to the person who uses this system to share their dining experience, taking pictures and interacting with the system.

[0276] The embodiments for carrying out this invention are shown below.

[0277] First, the user takes a picture of their meal using a device that provides information, such as a smartphone or tablet. This image is then sent to a server via Wi-Fi or a mobile network. The device is equipped with a camera module and a communication module.

[0278] The server uses hardware and image analysis software (such as TensorFlow or OpenCV) to process the received image data, identify objects within the image, and determine the type of meal. Based on these identified objects, the server uses a generative AI model (such as Stable Diffusion) to generate similar meal images. During this process, the prompt "Generate an image of a dish similar to XX (where XX is the name of the analyzed meal)" is input.

[0279] Subsequently, the server transmits the generated image data of the virtual representation to the terminal. The terminal visually processes this image using 3D modeling software (e.g., Unity) and displays it as a virtual representation on the user interface. The user can obtain an interactive meal-sharing experience through the virtual meal scene displayed on the terminal's display. This enables the user to enjoy the meal with a virtual partner even when dining alone.

[0280] Furthermore, the server utilizes natural language processing technology (e.g., GPT-4) to enable free conversation between the end user and the virtual representation. By analyzing the user's voice input and pre-entered profile information, it provides personalized conversation topics and content. This allows the user to converse flexibly about topics of interest and enjoy meal time more.

[0281] For example, when the user is eating pasta, the captured image is analyzed by the server, and an image of similar pasta is generated as a virtual representation. Based on this, the avatar on the terminal can perform the action of eating pasta and ask the user, "Have you tried the pasta recipe today?" In this way, even though the user is actually alone, they can gain a mental sense of security through virtual communication.

[0282] The flow of the specific process in Example 1 will be described using FIG. 11.

[0283] Step 1:

[0284] The user launches the camera application on their terminal and takes a picture of the meal. This operation provides an input means for capturing image data. The terminal saves the captured image and holds the image data in a file format (e.g., JPEG or PNG) to prepare for the next process.

[0285] Step 2:

[0286] The terminal sends the saved image data to the server via a network (Wi-Fi or mobile data). In this step, a specific operation is performed where the image data is transferred to the server in the form of an HTTP POST request. The input from the terminal is the image data, and the output is a response indicating that the transmission has been completed.

[0287] Step 3:

[0288] The server analyzes the received image data. Specifically, it uses image analysis software (such as OpenCV or TensorFlow) to identify the type of food. The input is the image data received from the terminal, and the output as a result of the analysis is the type of food. A process is carried out to detect features within the image and recognize objects using a predetermined algorithm.

[0289] Step 4: [[ID=—13]]

[0290] The server utilizes a generative AI model (such as Stable Diffusion) based on the analyzed type of food to generate images of similar foods. In this step, a prompt sentence in the form of "Please generate an image of a dish similar to XX (XX is the name of the analyzed food)" is used as input to the AI model. The model outputs the generated food images.

[0291] Step 5:

[0292] The server sends the generated similar food images to the terminal. In this process, the generated image is sent to the terminal as an HTTP response. The output is that the terminal receives the image data.

[0293] Step 6: [[ID=—29]]

[0294] The terminal processes the received food images using 3D modeling software (e.g., Unity) and displays them as a virtual representation in the user interface. The input is the received image data, and the output is the displayed virtual representation. A visual interface is generated in which the avatar performs the action of eating.

[0295] Step 7:

[0296] The device uses its camera to video record the user's actions while eating and immediately transmits the data to the server. The input is video data, and the output is the transmission of that data to the server. The recorded action data is analyzed in real time.

[0297] Step 8:

[0298] The server analyzes the received motion data to determine the user's eating pace. The input is motion data, and the output is pace information based on that motion. The motion analysis algorithm evaluates the user's motion speed and synchronizes it with the motion speed of a virtual representation.

[0299] Step 9:

[0300] The server uses natural language processing techniques (e.g., GPT-4) to generate free-flowing conversations tailored to the user. Input consists of the user's past voice inputs and profile information, while output is the conversation content and themes presented to the user. This facilitates engaging dialogue with the user.

[0301] (Application Example 1)

[0302] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0303] In modern society, the number of people who eat alone is increasing. Along with this, there is a problem that loneliness increases and psychological dissatisfaction occurs. In addition, since the time and place for enjoying meals with others in the real world are limited, there is an issue that users cannot easily enjoy the experience of eating together. Furthermore, there is a problem that users cannot experience natural and enjoyable conversations even in virtual shared eating experiences.

[0304] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0305] In this invention, the server includes means for receiving the video captured by the user, means for analyzing the received video and identifying the objects in the video, and means for using the generated video to display a virtual character and provide a shared eating experience. As a result, the user can reduce loneliness through sharing a meal with the virtual character and can easily obtain various shared eating experiences.

[0306] The "user" refers to a person who virtually shares a meal experience using this system.

[0307] The "video" refers to the image data captured by the user using a terminal during a meal.

[0308] The "object" refers to specific foods and related items included in the received video.

[0309] The "food video" refers to what is generated by creating image data similar to the food the user is actually eating.

[0310] The "virtual character" refers to a digital avatar displayed using the generated food video to realize a shared eating experience with the user.

[0311] The "action" refers to a series of behaviors performed by the user or the virtual character during a meal.

[0312] "Natural language processing technology" refers to the technology that enables computers to understand human language and generate meaningful conversations.

[0313] "Informal conversation" refers to communication using everyday, familiar topics.

[0314] "Conversation content" refers to the content of the conversation generated during the interaction with the user.

[0315] The system that realizes this application example allows users to experience shared dining in a virtual environment by linking their terminal with a server. The system mainly consists of the following elements:

[0316] 1. User's device: The user's device is equipped with a camera that captures images of the meal and sends those images to the server. The device can be a smartphone, smart glasses, or a head-mounted display.

[0317] 2. Server Function: The server executes a program developed in Python to analyze received video and identify objects related to food. Specifically, it uses PIL (Python Imaging Library) for video analysis and data processing. The server then generates food images similar to the identified meal. The generated images may utilize basic image generation models, or in some cases, generative AI models.

[0318] 3. Virtual Character Movement and Conversation: Based on the generated food images, the server displays a virtual character on the user's device to provide a shared dining experience. The device's camera captures the user's movements in real time and sends this data to the server. The server analyzes this data and synchronizes the virtual character's movements with the user's movements.

[0319] 4. Conversation generation using natural language processing: The server uses natural language processing techniques to generate conversations with the user. These conversations are based on the user's voice input and profile information, making the shared dining experience richer and more natural.

[0320] As a concrete example, if a user is eating pasta alone at home and uses this system, a virtual character will appear on the screen and begin a conversation about pasta. The virtual character will adjust its behavior to match the user's eating pace and continue the conversation with questions such as, "What are you thinking about while you eat today?"

[0321] Examples of prompt messages include: "Analyze the user's video of their meal, identify the type of food, and represent a similar meal with a virtual character. Next, analyze the user's conversation from the voice input and generate relevant conversation based on that information."

[0322] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0323] Step 1:

[0324] The user uses their device to film their meal and sends the video from the device to the server. The input data is the video captured by the user's camera, and the output is the data transfer to the server. The device requires a stable internet connection and sends the video to the server immediately after filming.

[0325] Step 2:

[0326] The server analyzes the received video of the meal and identifies objects within the video. The input here is the video data sent in step 1, and the output is information about the objects related to the identified meal. The server uses image analysis libraries such as PIL to process the data and identify the type of meal.

[0327] Step 3:

[0328] The server generates images of similar foods based on the identified object information. The input in this step is the object information obtained from step 2, and the output is the generated images of similar foods. The server uses a simple generative AI model to create images that resemble what the user is eating.

[0329] Step 4:

[0330] The generated food images are sent from the server to the user's terminal, which then uses them to display a virtual character. The input is the generated food image data, and the output is the display of the virtual character on the user interface. The terminal receives this image data and displays it immediately.

[0331] Step 5:

[0332] The user's actions while eating are captured by the device's camera, and the device sends this action data to the server. The input here is the user's video activity, and the output is the transfer of action data to the server. The device continuously monitors the user's actions and records the data.

[0333] Step 6:

[0334] The server analyzes the received motion data and synchronizes the virtual character's movements with the user's movements. The input in this step is the motion data from step 5, and the output is the adjusted virtual character's movements. The server determines the user's eating pace and calculates the character's movements accordingly.

[0335] Step 7:

[0336] The server uses natural language processing technology to generate conversations with the user. Input is the user's voice input or existing profile information, and output is the generated conversation content. A generative AI model is used for this process, forming informal conversations based on prompts.

[0337] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0338] This invention aims to provide a more personalized dining experience by combining an emotion engine with a system designed to create a virtual shared dining experience with the user. This system recognizes the user's emotions during meals and adjusts the avatar's responses and conversation content accordingly, thereby creating a comfortable dining environment for the user.

[0339] The user takes a picture of their meal using their device and sends the image to the server. The server analyzes the image and identifies the characteristics of the meal. Furthermore, the server generates similar meal images based on the collected image data and sends them to the device for use in the avatar's shared meal experience.

[0340] The device displays an avatar on the user interface, providing the user with a shared dining experience. Furthermore, the device uses a camera and microphone to monitor the user's facial expressions and voice tone in real time. This data is sent to a server, where an emotion engine analyzes the user's emotions. Based on the analyzed emotion information, the server adjusts the avatar's movements and conversation content.

[0341] For example, if a user is smiling, the server analyzes that positive emotion and sets the avatar to offer more cheerful topics. Conversely, if a user appears stressed, the server takes that emotion into consideration and has the avatar offer encouraging or relaxing topics.

[0342] The avatar uses natural language processing technology to interact with users and respond in a way that is sensitive to their emotions. Furthermore, the emotion engine learns from past user emotional data and predicts future emotional states, allowing the avatar's responses to be pre-adjusted. For example, if past data predicts that a user is prone to stress on certain days of the week, the avatar can prepare relaxing topics for those days.

[0343] Thus, the present invention makes it possible to reduce feelings of loneliness and mental stress, and to make dining a richer experience, by providing an interactive dining experience based on the user's emotions.

[0344] The following describes the processing flow.

[0345] Step 1:

[0346] The user takes a picture of their meal with their device and uploads it to the server through the application. This image contains information including details about the meal.

[0347] Step 2:

[0348] The server processes the received meal images using an image analysis algorithm to identify the type and characteristics of the meal and extract data. This analysis completes the analysis of the meal's contents.

[0349] Step 3:

[0350] The server generates similar meal images based on the analysis data and sends these images to the terminal for use in displaying the user's avatar.

[0351] Step 4:

[0352] The device displays images of similar meals generated alongside the avatar on the user interface, initiating a shared meal experience with the avatar. During this process, the user virtually dines with the avatar.

[0353] Step 5:

[0354] The device uses its camera and microphone to record the user's facial expressions and voice tone in real time. This data is sent to a server for emotion recognition.

[0355] Step 6:

[0356] The server uses an emotion engine to analyze the user's facial expressions and voice data to determine the user's emotional state. Based on this information, it prepares to adjust the avatar's reactions and conversation content.

[0357] Step 7:

[0358] Based on the sentiment analysis results, the server selects a suitable conversation topic for the user, generates conversation content using natural language processing technology, and sends it to the terminal.

[0359] Step 8:

[0360] The terminal applies the conversation content received from the server to the avatar and begins interacting with the user. At this time, the avatar's tone and topics will reflect the user's emotions.

[0361] Step 9:

[0362] The server learns from past emotional data and pre-adjusts the avatar's behavior and conversation based on predicted emotional states in future sessions. This continuous learning makes it possible to further personalize the user experience.

[0363] (Example 2)

[0364] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0365] In modern society, eating is not merely about nutrition; it is an important act that involves psychological and social elements. However, especially when eating alone, feelings of loneliness and mental stress can occur, which can reduce satisfaction with meals. This invention aims to solve these problems and provide users with a richer and more fulfilling dining experience.

[0366] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0367] In this invention, the server includes means for receiving video acquired by the user, means for analyzing the received video and identifying characteristics within the video, and means for generating similar visual data based on the identified characteristics. This allows the user to add visual and emotional interactive elements to existing meals, thereby reducing feelings of loneliness and alleviating stress.

[0368] "Users" refers to individuals or groups who operate the system and enjoy the experience.

[0369] "Video" refers to visual data that users capture and transmit to the system.

[0370] A "server" refers to a computing device that forms the core of a system and is responsible for receiving, analyzing, and providing information.

[0371] "Receiving" refers to the act of a server acquiring data sent by a user.

[0372] "Analysis" refers to the process of extracting and interpreting information from received data.

[0373] "Characteristics" refer to identifiable elements or features contained within images or data.

[0374] "Visual data" refers to digital image information generated based on analyzed characteristics.

[0375] A "virtual character" refers to a virtual personality created within a system for interacting with users.

[0376] "Empathic experience" refers to the emotional interaction that users have with virtual characters within a virtual environment, through shared perceptions.

[0377] "Actions" refer to the reactions and responses that the virtual character performs towards the user.

[0378] "Natural language processing technology" refers to the technology that enables computers to understand human language and generate appropriate dialogue.

[0379] "Emotional data" refers to the results of collecting and analyzing information that indicates the emotional state of users.

[0380] "Prediction" refers to the act of estimating future results or states based on past data.

[0381] This invention is a system that provides users with an interactive and emotionally engaging dining experience. The system consists of a terminal, a server, and multiple software components.

[0382] The user first uses the device's camera to capture video of their meal. The device also has a function to monitor the user's facial expressions and voice tone in real time. This makes it possible to capture the emotions the user is feeling while eating.

[0383] The acquired video is sent from the terminal to the server. The server uses image recognition software (e.g., TensorFlow or OpenCV) to analyze the received video. This analysis identifies the characteristics and features of the food. Based on the identified data, the server uses generative AI models (e.g., GANs) to generate similar visual data. This visual data is used to enrich the user's visual experience.

[0384] The server generates dialogue with a virtual character using natural language processing techniques (e.g., GPT-4) based on the analyzed data. The generated conversation is then delivered to the user via the terminal. For example, the prompt "How are you feeling today?" can be used to initiate the conversation.

[0385] Furthermore, monitoring data from the device is sent to the server's emotion analysis engine. This engine analyzes the user's emotions in real time and adjusts the virtual character's behavior and conversation content based on the analysis results. For example, if the user is smiling, it provides positive conversation content, and if the user is feeling stressed, it selects relaxing topics.

[0386] Furthermore, the server uses past emotional data as a learning platform to predict future emotional states. Based on these predictions, it can anticipate what emotions a user might feel on specific days of the week or in certain situations, and adjust the virtual character's reactions in advance. This ensures that users always have a comfortable and pleasant dining experience.

[0387] In this way, this system can deliver value to users beyond simply eating by providing a personalized dining experience based on their emotions and behavior.

[0388] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0389] Step 1:

[0390] The user takes a picture of their meal using the device's camera. The captured image is input as digital data from the camera device into the device's application. The device then sends this image to the server. As output, the video data is transferred to the server.

[0391] Step 2:

[0392] The server analyzes the received video data. Image recognition software (e.g., TensorFlow, OpenCV) is used for this analysis. The server identifies the characteristics of the food from the video data received as input and extracts characteristic elements (e.g., types of ingredients, types of dishes) based on that. The output is a dataset containing the characteristics of the identified dishes.

[0393] Step 3:

[0394] The server uses a generative AI model (e.g., GAN) to generate similar visual data based on the analyzed data. The characteristic data from step 2 is used as input. Based on this data, the server generates visual data for similar meals. The output is the generated similar visual data.

[0395] Step 4:

[0396] The server sends the generated visual data to the terminal. On the terminal, the virtual character is displayed as a customized avatar, providing a visual experience of shared dining. The generated visual data is interactively displayed to the user through the user interface. The output is a user interface that includes a visually rich avatar.

[0397] Step 5:

[0398] The device uses its built-in camera and microphone to monitor the user's facial expressions and voice tone in real time. Changes in facial expressions and voice signals are acquired as input data. This data is sent to a server, where an emotion analysis engine performs the analysis. The output is a recognition of the user's current emotional state.

[0399] Step 6:

[0400] The server uses generative AI models and natural language processing techniques to adjust the dialogue of the virtual character based on the analyzed emotional state. The input is the emotional data from step 5, and the server generates corresponding prompt sentences. This provides a conversation that matches the user's feelings. The output is the conversation text adapted for the user.

[0401] Step 7:

[0402] The server uses a database to learn from past user sentiment data and predict future emotional states. In this process, past sentiment data and patterns are used as input. Based on the prediction, preparations are made to pre-adjust the responses of the virtual character. The output is a response scenario corresponding to the predicted user's emotional tendencies.

[0403] (Application Example 2)

[0404] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0405] The challenge lies in reducing the feelings of loneliness and mental stress users experience during meals, and providing a richer dining experience. Furthermore, it is desirable to transform mealtime into an enjoyable experience by enabling personalized content selection based on user emotions.

[0406] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0407] This invention includes a server that analyzes the user's visual and auditory information in real time and dynamically adjusts content selection based on the results; a server that provides an interactive dining experience using virtual characters; and a server that generates and provides dialogue that is sensitive to the user's emotions using natural language processing technology. This allows the user to receive personalized responses and conversations during meals, alleviating feelings of loneliness and allowing them to have an enjoyable time.

[0408] A "user" is an individual who uses this system while consuming food.

[0409] "Image" refers to digital visual information captured by the user, including the contents of the meal and related objects.

[0410] "Visual information" refers to information such as the user's facial expressions and gestures, which are acquired through a camera device.

[0411] "Auditory information" refers to information about the user's pronunciation and surrounding sounds acquired through a microphone.

[0412] "Real-time analysis" refers to a process that rapidly processes information as soon as it is acquired and provides immediate feedback on the results.

[0413] "Content" refers to media information such as music, videos, and dialogues provided to users.

[0414] "Dynamic adjustment" refers to the process of changing the system's response and behavior in response to the user's emotions and circumstances.

[0415] A "virtual character" refers to an artificial character with a personality that interacts with the user on a digital interface.

[0416] "Natural language processing technology" is a technology that enables computers to understand and respond to human language.

[0417] An "interactive shared dining experience" refers to a simulated experience in which the user and a virtual character interact while sharing a meal.

[0418] The system that realizes this application is built on a foundation of mobile devices such as smartphones and tablets, and a cloud server. During a meal, the user takes a picture of their facial expression using the device's camera, and this visual information is sent to the server. The server uses an emotion analysis model built with machine learning libraries such as TensorFlow to analyze the received visual and auditory information in real time. The analyzed emotion data is input into a generative AI model, which generates content and dialogue tailored to the user.

[0419] Furthermore, the server uses natural language processing technology to enable smooth interaction with the user via a platform like Dialogflow. This allows the virtual character to adjust its actions and conversations in response to the user's real-time emotions, providing a personalized dining experience. For example, if the user wants to relax during a meal, a prompt such as "Play some relaxing music" is input into the generating AI model, and calming music is played on the device.

[0420] This application configuration allows users to enjoy personalized services that cater to their emotions and preferences, rather than simply receiving mechanically generated content. For example, when a user uses this system to take a break from their busy daily life, they can be refreshed with soothing music and encouraging messages from a virtual character.

[0421] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0422] Step 1:

[0423] The user uses the device's camera to take a picture of their facial expression while eating. At this time, the device receives the captured image data as input and prepares to send the visual information to the server.

[0424] Step 2:

[0425] The server initializes an emotion analysis model to analyze the received visual information. Here, image data is used as input, and facial expression analysis is performed using a TensorFlow-based model. The analysis result outputs the user's emotional state.

[0426] Step 3:

[0427] The server uses a generative AI model based on the analyzed sentiment data to select appropriate content and generate dialogue. It receives sentiment data as input, generates prompts via a natural language processing platform such as Dialogflow, and determines the content and dialogue to be played back to the user as output.

[0428] Step 4:

[0429] The terminal receives prompts and content instructions sent from the server and displays a virtual character on the interface based on them. The virtual character provides the user with personalized conversations and content based on the analysis results.

[0430] Step 5:

[0431] Users enjoy interacting with virtual characters while viewing content provided through their devices. As long as the interaction with the user continues, the device continuously uses its camera and microphone to collect further emotional data and transmit it to the server in real time. This allows the entire system to remain dynamically responsive.

[0432] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0433] The data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0434] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0435] [Third Embodiment]

[0436] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0437] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0438] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0439] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0440] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0441] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0442] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0443] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0444] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0445] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0446] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0447] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0448] This invention is a system for virtually realizing the experience of sharing a meal with a user, with the aim of reducing feelings of loneliness. This system provides users with a sense of psychological security by allowing them to share a meal through an avatar. The embodiments for carrying out the invention will be described in detail below.

[0449] The user first takes a picture of their meal using their device. The device sends this image to the server. The server analyzes the received image and identifies the type and characteristics of the food it contains. Based on the analysis results, the server has the function to generate a similar meal image. This image resembles what the user is eating and is used for the avatar to participate in the shared meal.

[0450] Next, the server sends the generated meal image to the terminal, which uses it to display the avatar on the user interface. The avatar operates in real time to provide the user with a visual and interactive shared-meal experience. Furthermore, the terminal captures the user's eating movements through its camera and sends this data to the server. The server analyzes this data to determine the user's eating speed and pace. This allows the avatar's movements to be adjusted to harmonize with the user's movements.

[0451] Furthermore, the server uses natural language processing technology to generate conversations with users. Specifically, based on the user's voice input and pre-registered profile information, it selects topics of interest to the user and generates appropriate small talk accordingly. In this way, the avatar facilitates natural conversations with users and makes the shared dining experience more enjoyable.

[0452] For example, if a user is eating pasta, the server generates a similar pasta image for the avatar based on the image of the pasta. The avatar then uses this image to perform the action of eating pasta and initiates a conversation with the user, such as "What are you thinking about while you eat today?" This system functionality allows users to enjoy a meal with a virtual partner even when they are eating alone.

[0453] Thus, this invention not only alleviates the user's feelings of loneliness, but also enriches mealtimes and enhances their mental satisfaction.

[0454] The following describes the processing flow.

[0455] Step 1:

[0456] The user takes a picture of their meal using their device and uploads the image to the server via the application. This image file contains details about the meal.

[0457] Step 2:

[0458] The server processes the received images and extracts features such as the type, color, and shape of the food through image analysis algorithms. This analysis generates data to identify the food.

[0459] Step 3:

[0460] The server generates similar meal images for the avatar based on the data obtained through analysis. Using a generation AI model, it digitally constructs meal images that visually resemble the user's actual meal.

[0461] Step 4:

[0462] The server sends the generated similar meal image to the user's terminal. The terminal receives this image, displays an avatar on the user interface, and starts the shared meal simulation.

[0463] Step 5:

[0464] The device uses its camera to record the user's eating habits in real time. This data is sent to a server and serves as foundational data for analyzing the user's eating speed and rhythm.

[0465] Step 6:

[0466] The server analyzes the collected motion data and adjusts the avatar's movements to match the user's eating pace. This changes the avatar's behavior so that it eats at the same speed as the user.

[0467] Step 7:

[0468] The server uses natural language processing technology to generate personalized conversation content based on the user's voice input and profile information. This information is sent to the terminal, and the avatar begins a conversation with the user.

[0469] Step 8:

[0470] The device uses the generated conversation content to control the avatar and facilitate interaction with the user. The avatar brings up topics that are familiar to the user and enriches the shared dining experience through conversation.

[0471] Step 9:

[0472] The device detects that the user has finished eating and reports this to the server. The server ends the shared meal session and presents the user with an option to provide feedback.

[0473] (Example 1)

[0474] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0475] In society, eating alone can be a factor that increases feelings of loneliness. Furthermore, eating alone lacks opportunities for social interaction, which can lower individual life satisfaction. This invention aims to alleviate such feelings of loneliness and make solitary mealtimes more fulfilling.

[0476] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0477] In this invention, the server includes means for receiving images captured by an information-providing terminal, means for analyzing the received images and identifying objects within the images, and means for generating similar meal images based on the identified objects. This makes it possible to share meals with end users through virtual representations and provide a sense of psychological reassurance.

[0478] A "terminal that provides information" is a device for users to input or manipulate information, and in this case, it is a device that has the function of taking and transmitting images of a meal.

[0479] An "object" is a specific element present in the received image that is identified through analysis, and in this system, it mainly refers to food.

[0480] "Virtual representation" refers to avatars and simulations that are visually displayed on a screen using computer technology, providing an interactive shared dining experience with the user.

[0481] "Free conversation" refers to dialogue-style conversation content generated using natural language processing technology, providing discourse based on the user's interests and concerns.

[0482] An "end user" refers to the person who uses this system to share their dining experience, taking pictures and interacting with the system.

[0483] The embodiments for carrying out this invention are shown below.

[0484] First, the user takes a picture of their meal using a device that provides information, such as a smartphone or tablet. This image is then sent to a server via Wi-Fi or a mobile network. The device is equipped with a camera module and a communication module.

[0485] The server uses hardware and image analysis software (such as TensorFlow or OpenCV) to process the received image data, identify objects within the image, and determine the type of meal. Based on these identified objects, the server uses a generative AI model (such as Stable Diffusion) to generate similar meal images. During this process, the prompt "Generate an image of a dish similar to XX (where XX is the name of the analyzed meal)" is input.

[0486] Next, the server sends the generated virtual representation image data to the terminal. The terminal uses 3D modeling software (for example, Unity) to visually process this image and display it as a virtual representation in the user interface. Through the virtual dining scene displayed on the terminal's screen, the user can enjoy an interactive shared dining experience. This makes it possible to enjoy a meal with a virtual partner, even when dining alone.

[0487] Furthermore, the server utilizes natural language processing technology (e.g., GPT-4) to enable free-flowing conversation between the end user and the virtual representation. It analyzes the user's voice input and pre-entered profile information to provide personalized conversation themes and content. This allows users to converse flexibly on topics of interest, making mealtimes more enjoyable.

[0488] For example, if a user is eating pasta, the server analyzes the image they take and generates a virtual representation of similar pasta. Based on this, the avatar on the device can perform the action of eating pasta and ask the user, "Have you tried any pasta recipes today?" In this way, even though the user is actually alone, they can gain a sense of psychological comfort through virtual interaction.

[0489] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0490] Step 1:

[0491] The user launches a camera application on their device and takes a picture of their meal. This operation provides an input means for capturing image data. The device saves the captured image and retains the image data in a file format (e.g., JPEG or PNG) in preparation for the next processing step.

[0492] Step 2:

[0493] The device sends the saved image data to the server via the network (Wi-Fi or mobile data). In this step, the image data is transferred to the server in the form of an HTTP POST request. The input from the device is the image data, and the output is a response indicating that the transmission is complete.

[0494] Step 3:

[0495] The server analyzes the received image data. Specifically, it uses image analysis software (for example, OpenCV or TensorFlow) to identify the type of meal. The input is image data received from the terminal, and the type of meal is output as the result of the analysis. Features within the image are detected, and objects are recognized using a predetermined algorithm.

[0496] Step 4:

[0497] The server utilizes a generative AI model (e.g., Stable Diffusion) based on the analyzed meal type to generate images of similar meals. In this step, the AI ​​model is input using a prompt in the form of "Generate images of dishes similar to XX (where XX is the name of the analyzed meal)." The model then outputs the generated meal images.

[0498] Step 5:

[0499] The server sends the generated similar meal image to the terminal. In this process, the generated image is sent to the terminal as an HTTP response. The output is that the terminal receives the image data.

[0500] Step 6:

[0501] The terminal processes the received food images using 3D modeling software (e.g., Unity) and displays them as a virtual representation in the user interface. The input is the received image data, and the output is the displayed virtual representation. A visual interface is generated in which the avatar performs the action of eating.

[0502] Step 7:

[0503] The device uses its camera to video record the user's actions while eating and immediately transmits the data to the server. The input is video data, and the output is the transmission of that data to the server. The recorded action data is analyzed in real time.

[0504] Step 8:

[0505] The server analyzes the received motion data to determine the user's eating pace. The input is motion data, and the output is pace information based on that motion. The motion analysis algorithm evaluates the user's motion speed and synchronizes it with the motion speed of a virtual representation.

[0506] Step 9:

[0507] The server uses natural language processing techniques (e.g., GPT-4) to generate free-flowing conversations tailored to the user. Input consists of the user's past voice inputs and profile information, while output is the conversation content and themes presented to the user. This facilitates engaging dialogue with the user.

[0508] (Application Example 1)

[0509] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0510] In modern society, the number of people who eat alone is increasing, which is leading to increased feelings of loneliness and psychological dissatisfaction. Furthermore, because opportunities to enjoy meals with others in the real world are limited, users cannot easily enjoy the experience of eating together. Additionally, even in virtual dining experiences, users may struggle to engage in natural and enjoyable conversations.

[0511] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0512] In this invention, the server includes means for receiving video footage captured by the user, means for analyzing the received video footage and identifying objects within the video footage, and means for displaying a virtual character and providing a shared dining experience using the generated video footage. This allows the user to alleviate feelings of loneliness through shared dining with a virtual character and easily obtain diverse shared dining experiences.

[0513] A "user" refers to someone who uses this system to virtually share a dining experience.

[0514] "Video" refers to image data captured by the user using their device while eating.

[0515] "Object" refers to specific meals or related items included in the received video.

[0516] "Food images" refer to image data generated that resembles the food that the user is actually eating.

[0517] A "virtual character" refers to a digital avatar displayed using generated food images to create a shared dining experience with the user.

[0518] "Actions" refer to a series of actions performed by the user or virtual character while eating.

[0519] "Natural language processing technology" refers to the technology that enables computers to understand human language and generate meaningful conversations.

[0520] "Informal conversation" refers to communication using everyday, familiar topics.

[0521] "Conversation content" refers to the content of the conversation generated during the interaction with the user.

[0522] The system that realizes this application example allows users to experience shared dining in a virtual environment by linking their terminal with a server. The system mainly consists of the following elements:

[0523] 1. User's device: The user's device is equipped with a camera that captures images of the meal and sends those images to the server. The device can be a smartphone, smart glasses, or a head-mounted display.

[0524] 2. Server Function: The server executes a program developed in Python to analyze received video and identify objects related to food. Specifically, it uses PIL (Python Imaging Library) for video analysis and data processing. The server then generates food images similar to the identified meal. The generated images may utilize basic image generation models, or in some cases, generative AI models.

[0525] 3. Virtual Character Movement and Conversation: Based on the generated food images, the server displays a virtual character on the user's device to provide a shared dining experience. The device's camera captures the user's movements in real time and sends this data to the server. The server analyzes this data and synchronizes the virtual character's movements with the user's movements.

[0526] 4. Conversation generation using natural language processing: The server uses natural language processing techniques to generate conversations with the user. These conversations are based on the user's voice input and profile information, making the shared dining experience richer and more natural.

[0527] As a concrete example, if a user is eating pasta alone at home and uses this system, a virtual character will appear on the screen and begin a conversation about pasta. The virtual character will adjust its behavior to match the user's eating pace and continue the conversation with questions such as, "What are you thinking about while you eat today?"

[0528] Examples of prompt messages include: "Analyze the user's video of their meal, identify the type of food, and represent a similar meal with a virtual character. Next, analyze the user's conversation from the voice input and generate relevant conversation based on that information."

[0529] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0530] Step 1:

[0531] The user uses their device to film their meal and sends the video from the device to the server. The input data is the video captured by the user's camera, and the output is the data transfer to the server. The device requires a stable internet connection and sends the video to the server immediately after filming.

[0532] Step 2:

[0533] The server analyzes the received video of the meal and identifies objects within the video. The input here is the video data sent in step 1, and the output is information about the objects related to the identified meal. The server uses image analysis libraries such as PIL to process the data and identify the type of meal.

[0534] Step 3:

[0535] The server generates images of similar foods based on the identified object information. The input in this step is the object information obtained from step 2, and the output is the generated images of similar foods. The server uses a simple generative AI model to create images that resemble what the user is eating.

[0536] Step 4:

[0537] The generated food images are sent from the server to the user's terminal, which then uses them to display a virtual character. The input is the generated food image data, and the output is the display of the virtual character on the user interface. The terminal receives this image data and displays it immediately.

[0538] Step 5:

[0539] The user's actions while eating are captured by the device's camera, and the device sends this action data to the server. The input here is the user's video activity, and the output is the transfer of action data to the server. The device continuously monitors the user's actions and records the data.

[0540] Step 6:

[0541] The server analyzes the received motion data and synchronizes the virtual character's movements with the user's movements. The input in this step is the motion data from step 5, and the output is the adjusted virtual character's movements. The server determines the user's eating pace and calculates the character's movements accordingly.

[0542] Step 7:

[0543] The server uses natural language processing technology to generate conversations with the user. Input is the user's voice input or existing profile information, and output is the generated conversation content. A generative AI model is used for this process, forming informal conversations based on prompts.

[0544] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0545] This invention aims to provide a more personalized dining experience by combining an emotion engine with a system designed to create a virtual shared dining experience with the user. This system recognizes the user's emotions during meals and adjusts the avatar's responses and conversation content accordingly, thereby creating a comfortable dining environment for the user.

[0546] The user takes a picture of their meal using their device and sends the image to the server. The server analyzes the image and identifies the characteristics of the meal. Furthermore, the server generates similar meal images based on the collected image data and sends them to the device for use in the avatar's shared meal experience.

[0547] The device displays an avatar on the user interface, providing the user with a shared dining experience. Furthermore, the device uses a camera and microphone to monitor the user's facial expressions and voice tone in real time. This data is sent to a server, where an emotion engine analyzes the user's emotions. Based on the analyzed emotion information, the server adjusts the avatar's movements and conversation content.

[0548] For example, if a user is smiling, the server analyzes that positive emotion and sets the avatar to offer more cheerful topics. Conversely, if a user appears stressed, the server takes that emotion into consideration and has the avatar offer encouraging or relaxing topics.

[0549] The avatar uses natural language processing technology to interact with users and respond in a way that is sensitive to their emotions. Furthermore, the emotion engine learns from past user emotional data and predicts future emotional states, allowing the avatar's responses to be pre-adjusted. For example, if past data predicts that a user is prone to stress on certain days of the week, the avatar can prepare relaxing topics for those days.

[0550] Thus, the present invention makes it possible to reduce feelings of loneliness and mental stress, and to make dining a richer experience, by providing an interactive dining experience based on the user's emotions.

[0551] The following describes the processing flow.

[0552] Step 1:

[0553] The user takes a picture of their meal with their device and uploads it to the server through the application. This image contains information including details about the meal.

[0554] Step 2:

[0555] The server processes the received meal images using an image analysis algorithm to identify the type and characteristics of the meal and extract data. This analysis completes the analysis of the meal's contents.

[0556] Step 3:

[0557] The server generates similar meal images based on the analysis data and sends these images to the terminal for use in displaying the user's avatar.

[0558] Step 4:

[0559] The device displays images of similar meals generated alongside the avatar on the user interface, initiating a shared meal experience with the avatar. During this process, the user virtually dines with the avatar.

[0560] Step 5:

[0561] The device uses its camera and microphone to record the user's facial expressions and voice tone in real time. This data is sent to a server for emotion recognition.

[0562] Step 6:

[0563] The server uses an emotion engine to analyze the user's facial expressions and voice data to determine the user's emotional state. Based on this information, it prepares to adjust the avatar's reactions and conversation content.

[0564] Step 7:

[0565] Based on the sentiment analysis results, the server selects a suitable conversation topic for the user, generates conversation content using natural language processing technology, and sends it to the terminal.

[0566] Step 8:

[0567] The terminal applies the conversation content received from the server to the avatar and begins interacting with the user. At this time, the avatar's tone and topics will reflect the user's emotions.

[0568] Step 9:

[0569] The server learns from past emotional data and pre-adjusts the avatar's behavior and conversation based on predicted emotional states in future sessions. This continuous learning makes it possible to further personalize the user experience.

[0570] (Example 2)

[0571] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0572] In modern society, eating is not merely about nutrition; it is an important act that involves psychological and social elements. However, especially when eating alone, feelings of loneliness and mental stress can occur, which can reduce satisfaction with meals. This invention aims to solve these problems and provide users with a richer and more fulfilling dining experience.

[0573] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0574] In this invention, the server includes means for receiving video acquired by the user, means for analyzing the received video and identifying characteristics within the video, and means for generating similar visual data based on the identified characteristics. This allows the user to add visual and emotional interactive elements to existing meals, thereby reducing feelings of loneliness and alleviating stress.

[0575] "Users" refers to individuals or groups who operate the system and enjoy the experience.

[0576] "Video" refers to visual data that users capture and transmit to the system.

[0577] A "server" refers to a computing device that forms the core of a system and is responsible for receiving, analyzing, and providing information.

[0578] "Receiving" refers to the act of a server acquiring data sent by a user.

[0579] "Analysis" refers to the process of extracting and interpreting information from received data.

[0580] "Characteristics" refer to identifiable elements or features contained within images or data.

[0581] "Visual data" refers to digital image information generated based on analyzed characteristics.

[0582] A "virtual character" refers to a virtual personality created within a system for interacting with users.

[0583] "Empathic experience" refers to the emotional interaction that users have with virtual characters within a virtual environment, through shared perceptions.

[0584] "Actions" refer to the reactions and responses that the virtual character performs towards the user.

[0585] "Natural language processing technology" refers to the technology that enables computers to understand human language and generate appropriate dialogue.

[0586] "Emotional data" refers to the results of collecting and analyzing information that indicates the emotional state of users.

[0587] "Prediction" refers to the act of estimating future results or states based on past data.

[0588] This invention is a system that provides users with an interactive and emotionally engaging dining experience. The system consists of a terminal, a server, and multiple software components.

[0589] The user first uses the device's camera to capture video of their meal. The device also has a function to monitor the user's facial expressions and voice tone in real time. This makes it possible to capture the emotions the user is feeling while eating.

[0590] The acquired video is sent from the terminal to the server. The server uses image recognition software (e.g., TensorFlow or OpenCV) to analyze the received video. This analysis identifies the characteristics and features of the food. Based on the identified data, the server uses generative AI models (e.g., GANs) to generate similar visual data. This visual data is used to enrich the user's visual experience.

[0591] The server generates dialogue with a virtual character using natural language processing techniques (e.g., GPT-4) based on the analyzed data. The generated conversation is then delivered to the user via the terminal. For example, the prompt "How are you feeling today?" can be used to initiate the conversation.

[0592] Furthermore, monitoring data from the device is sent to the server's emotion analysis engine. This engine analyzes the user's emotions in real time and adjusts the virtual character's behavior and conversation content based on the analysis results. For example, if the user is smiling, it provides positive conversation content, and if the user is feeling stressed, it selects relaxing topics.

[0593] Furthermore, the server uses past emotional data as a learning platform to predict future emotional states. Based on these predictions, it can anticipate what emotions a user might feel on specific days of the week or in certain situations, and adjust the virtual character's reactions in advance. This ensures that users always have a comfortable and pleasant dining experience.

[0594] In this way, this system can deliver value to users beyond simply eating by providing a personalized dining experience based on their emotions and behavior.

[0595] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0596] Step 1:

[0597] The user takes a picture of their meal using the device's camera. The captured image is input as digital data from the camera device into the device's application. The device then sends this image to the server. As output, the video data is transferred to the server.

[0598] Step 2:

[0599] The server analyzes the received video data. Image recognition software (e.g., TensorFlow, OpenCV) is used for this analysis. The server identifies the characteristics of the food from the video data received as input and extracts characteristic elements (e.g., types of ingredients, types of dishes) based on that. The output is a dataset containing the characteristics of the identified dishes.

[0600] Step 3:

[0601] The server uses a generative AI model (e.g., GAN) to generate similar visual data based on the analyzed data. The characteristic data from step 2 is used as input. Based on this data, the server generates visual data for similar meals. The output is the generated similar visual data.

[0602] Step 4:

[0603] The server sends the generated visual data to the terminal. On the terminal, the virtual character is displayed as a customized avatar, providing a visual experience of shared dining. The generated visual data is interactively displayed to the user through the user interface. The output is a user interface that includes a visually rich avatar.

[0604] Step 5:

[0605] The device uses its built-in camera and microphone to monitor the user's facial expressions and voice tone in real time. Changes in facial expressions and voice signals are acquired as input data. This data is sent to a server, where an emotion analysis engine performs the analysis. The output is a recognition of the user's current emotional state.

[0606] Step 6:

[0607] The server uses generative AI models and natural language processing techniques to adjust the dialogue of the virtual character based on the analyzed emotional state. The input is the emotional data from step 5, and the server generates corresponding prompt sentences. This provides a conversation that matches the user's feelings. The output is the conversation text adapted for the user.

[0608] Step 7:

[0609] The server uses a database to learn from past user sentiment data and predict future emotional states. In this process, past sentiment data and patterns are used as input. Based on the prediction, preparations are made to pre-adjust the responses of the virtual character. The output is a response scenario corresponding to the predicted user's emotional tendencies.

[0610] (Application Example 2)

[0611] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0612] The challenge lies in reducing the feelings of loneliness and mental stress users experience during meals, and providing a richer dining experience. Furthermore, it is desirable to transform mealtime into an enjoyable experience by enabling personalized content selection based on user emotions.

[0613] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0614] This invention includes a server that analyzes the user's visual and auditory information in real time and dynamically adjusts content selection based on the results; a server that provides an interactive dining experience using virtual characters; and a server that generates and provides dialogue that is sensitive to the user's emotions using natural language processing technology. This allows the user to receive personalized responses and conversations during meals, alleviating feelings of loneliness and allowing them to have an enjoyable time.

[0615] A "user" is an individual who uses this system while consuming food.

[0616] "Image" refers to digital visual information captured by the user, including the contents of the meal and related objects.

[0617] "Visual information" refers to information such as the user's facial expressions and gestures, which are acquired through a camera device.

[0618] "Auditory information" refers to information about the user's pronunciation and surrounding sounds acquired through a microphone.

[0619] "Real-time analysis" refers to a process that rapidly processes information as soon as it is acquired and provides immediate feedback on the results.

[0620] "Content" refers to media information such as music, videos, and dialogues provided to users.

[0621] "Dynamic adjustment" refers to the process of changing the system's response and behavior in response to the user's emotions and circumstances.

[0622] A "virtual character" refers to an artificial character with a personality that interacts with the user on a digital interface.

[0623] "Natural language processing technology" is a technology that enables computers to understand and respond to human language.

[0624] An "interactive shared dining experience" refers to a simulated experience in which the user and a virtual character interact while sharing a meal.

[0625] The system that realizes this application is built on a foundation of mobile devices such as smartphones and tablets, and a cloud server. During a meal, the user takes a picture of their facial expression using the device's camera, and this visual information is sent to the server. The server uses an emotion analysis model built with machine learning libraries such as TensorFlow to analyze the received visual and auditory information in real time. The analyzed emotion data is input into a generative AI model, which generates content and dialogue tailored to the user.

[0626] Furthermore, the server uses natural language processing technology to enable smooth interaction with the user via a platform like Dialogflow. This allows the virtual character to adjust its actions and conversations in response to the user's real-time emotions, providing a personalized dining experience. For example, if the user wants to relax during a meal, a prompt such as "Play some relaxing music" is input into the generating AI model, and calming music is played on the device.

[0627] This application configuration allows users to enjoy personalized services that cater to their emotions and preferences, rather than simply receiving mechanically generated content. For example, when a user uses this system to take a break from their busy daily life, they can be refreshed with soothing music and encouraging messages from a virtual character.

[0628] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0629] Step 1:

[0630] The user uses the device's camera to take a picture of their facial expression while eating. At this time, the device receives the captured image data as input and prepares to send the visual information to the server.

[0631] Step 2:

[0632] The server initializes an emotion analysis model to analyze the received visual information. Here, image data is used as input, and facial expression analysis is performed using a TensorFlow-based model. The analysis result outputs the user's emotional state.

[0633] Step 3:

[0634] The server uses a generative AI model based on the analyzed sentiment data to select appropriate content and generate dialogue. It receives sentiment data as input, generates prompts via a natural language processing platform such as Dialogflow, and determines the content and dialogue to be played back to the user as output.

[0635] Step 4:

[0636] The terminal receives prompts and content instructions sent from the server and displays a virtual character on the interface based on them. The virtual character provides the user with personalized conversations and content based on the analysis results.

[0637] Step 5:

[0638] Users enjoy interacting with virtual characters while viewing content provided through their devices. As long as the interaction with the user continues, the device continuously uses its camera and microphone to collect further emotional data and transmit it to the server in real time. This allows the entire system to remain dynamically responsive.

[0639] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0640] The data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0641] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0642] [Fourth Embodiment]

[0643] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0644] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0645] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0646] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0647] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0648] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0649] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0650] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0651] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0652] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0653] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0654] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0655] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0656] This invention is a system for virtually realizing the experience of sharing a meal with a user, with the aim of reducing feelings of loneliness. This system provides users with a sense of psychological security by allowing them to share a meal through an avatar. The embodiments for carrying out the invention will be described in detail below.

[0657] The user first takes a picture of their meal using their device. The device sends this image to the server. The server analyzes the received image and identifies the type and characteristics of the food it contains. Based on the analysis results, the server has the function to generate a similar meal image. This image resembles what the user is eating and is used for the avatar to participate in the shared meal.

[0658] Next, the server sends the generated meal image to the terminal, which uses it to display the avatar on the user interface. The avatar operates in real time to provide the user with a visual and interactive shared-meal experience. Furthermore, the terminal captures the user's eating movements through its camera and sends this data to the server. The server analyzes this data to determine the user's eating speed and pace. This allows the avatar's movements to be adjusted to harmonize with the user's movements.

[0659] Furthermore, the server uses natural language processing technology to generate conversations with users. Specifically, based on the user's voice input and pre-registered profile information, it selects topics of interest to the user and generates appropriate small talk accordingly. In this way, the avatar facilitates natural conversations with users and makes the shared dining experience more enjoyable.

[0660] For example, if a user is eating pasta, the server generates a similar pasta image for the avatar based on the image of the pasta. The avatar then uses this image to perform the action of eating pasta and initiates a conversation with the user, such as "What are you thinking about while you eat today?" This system functionality allows users to enjoy a meal with a virtual partner even when they are eating alone.

[0661] Thus, this invention not only alleviates the user's feelings of loneliness, but also enriches mealtimes and enhances their mental satisfaction.

[0662] The following describes the processing flow.

[0663] Step 1:

[0664] The user takes a picture of their meal using their device and uploads the image to the server via the application. This image file contains details about the meal.

[0665] Step 2:

[0666] The server processes the received images and extracts features such as the type, color, and shape of the food through image analysis algorithms. This analysis generates data to identify the food.

[0667] Step 3:

[0668] The server generates similar meal images for the avatar based on the data obtained through analysis. Using a generation AI model, it digitally constructs meal images that visually resemble the user's actual meal.

[0669] Step 4:

[0670] The server sends the generated similar meal image to the user's terminal. The terminal receives this image, displays an avatar on the user interface, and starts the shared meal simulation.

[0671] Step 5:

[0672] The device uses its camera to record the user's eating habits in real time. This data is sent to a server and serves as foundational data for analyzing the user's eating speed and rhythm.

[0673] Step 6:

[0674] The server analyzes the collected motion data and adjusts the avatar's movements to match the user's eating pace. This changes the avatar's behavior so that it eats at the same speed as the user.

[0675] Step 7:

[0676] The server uses natural language processing technology to generate personalized conversation content based on the user's voice input and profile information. This information is sent to the terminal, and the avatar begins a conversation with the user.

[0677] Step 8:

[0678] The device uses the generated conversation content to control the avatar and facilitate interaction with the user. The avatar brings up topics that are familiar to the user and enriches the shared dining experience through conversation.

[0679] Step 9:

[0680] The device detects that the user has finished eating and reports this to the server. The server ends the shared meal session and presents the user with an option to provide feedback.

[0681] (Example 1)

[0682] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0683] In society, eating alone can be a factor that increases feelings of loneliness. Furthermore, eating alone lacks opportunities for social interaction, which can lower individual life satisfaction. This invention aims to alleviate such feelings of loneliness and make solitary mealtimes more fulfilling.

[0684] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0685] In this invention, the server includes means for receiving images captured by an information-providing terminal, means for analyzing the received images and identifying objects within the images, and means for generating similar meal images based on the identified objects. This makes it possible to share meals with end users through virtual representations and provide a sense of psychological reassurance.

[0686] A "terminal that provides information" is a device for users to input or manipulate information, and in this case, it is a device that has the function of taking and transmitting images of a meal.

[0687] An "object" is a specific element present in the received image that is identified through analysis, and in this system, it mainly refers to food.

[0688] "Virtual representation" refers to avatars and simulations that are visually displayed on a screen using computer technology, providing an interactive shared dining experience with the user.

[0689] "Free conversation" refers to dialogue-style conversation content generated using natural language processing technology, providing discourse based on the user's interests and concerns.

[0690] An "end user" refers to the person who uses this system to share their dining experience, taking pictures and interacting with the system.

[0691] The embodiments for carrying out this invention are shown below.

[0692] First, the user takes a picture of their meal using a device that provides information, such as a smartphone or tablet. This image is then sent to a server via Wi-Fi or a mobile network. The device is equipped with a camera module and a communication module.

[0693] The server uses hardware and image analysis software (such as TensorFlow or OpenCV) to process the received image data, identify objects within the image, and determine the type of meal. Based on these identified objects, the server uses a generative AI model (such as Stable Diffusion) to generate similar meal images. During this process, the prompt "Generate an image of a dish similar to XX (where XX is the name of the analyzed meal)" is input.

[0694] Next, the server sends the generated virtual representation image data to the terminal. The terminal uses 3D modeling software (for example, Unity) to visually process this image and display it as a virtual representation in the user interface. Through the virtual dining scene displayed on the terminal's screen, the user can enjoy an interactive shared dining experience. This makes it possible to enjoy a meal with a virtual partner, even when dining alone.

[0695] Furthermore, the server utilizes natural language processing technology (e.g., GPT-4) to enable free-flowing conversation between the end user and the virtual representation. It analyzes the user's voice input and pre-entered profile information to provide personalized conversation themes and content. This allows users to converse flexibly on topics of interest, making mealtimes more enjoyable.

[0696] For example, if a user is eating pasta, the server analyzes the image they take and generates a virtual representation of similar pasta. Based on this, the avatar on the device can perform the action of eating pasta and ask the user, "Have you tried any pasta recipes today?" In this way, even though the user is actually alone, they can gain a sense of psychological comfort through virtual interaction.

[0697] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0698] Step 1:

[0699] The user launches a camera application on their device and takes a picture of their meal. This operation provides an input means for capturing image data. The device saves the captured image and retains the image data in a file format (e.g., JPEG or PNG) in preparation for the next processing step.

[0700] Step 2:

[0701] The device sends the saved image data to the server via the network (Wi-Fi or mobile data). In this step, the image data is transferred to the server in the form of an HTTP POST request. The input from the device is the image data, and the output is a response indicating that the transmission is complete.

[0702] Step 3:

[0703] The server analyzes the received image data. Specifically, it uses image analysis software (for example, OpenCV or TensorFlow) to identify the type of meal. The input is image data received from the terminal, and the type of meal is output as the result of the analysis. Features within the image are detected, and objects are recognized using a predetermined algorithm.

[0704] Step 4:

[0705] The server utilizes a generative AI model (e.g., Stable Diffusion) based on the analyzed meal type to generate images of similar meals. In this step, the AI ​​model is input using a prompt in the form of "Generate images of dishes similar to XX (where XX is the name of the analyzed meal)." The model then outputs the generated meal images.

[0706] Step 5:

[0707] The server sends the generated similar meal image to the terminal. In this process, the generated image is sent to the terminal as an HTTP response. The output is that the terminal receives the image data.

[0708] Step 6:

[0709] The terminal processes the received food images using 3D modeling software (e.g., Unity) and displays them as a virtual representation in the user interface. The input is the received image data, and the output is the displayed virtual representation. A visual interface is generated in which the avatar performs the action of eating.

[0710] Step 7:

[0711] The device uses its camera to video record the user's actions while eating and immediately transmits the data to the server. The input is video data, and the output is the transmission of that data to the server. The recorded action data is analyzed in real time.

[0712] Step 8:

[0713] The server analyzes the received motion data to determine the user's eating pace. The input is motion data, and the output is pace information based on that motion. The motion analysis algorithm evaluates the user's motion speed and synchronizes it with the motion speed of a virtual representation.

[0714] Step 9:

[0715] The server uses natural language processing techniques (e.g., GPT-4) to generate free-flowing conversations tailored to the user. Input consists of the user's past voice inputs and profile information, while output is the conversation content and themes presented to the user. This facilitates engaging dialogue with the user.

[0716] (Application Example 1)

[0717] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0718] In modern society, the number of people who eat alone is increasing, which is leading to increased feelings of loneliness and psychological dissatisfaction. Furthermore, because opportunities to enjoy meals with others in the real world are limited, users cannot easily enjoy the experience of eating together. Additionally, even in virtual dining experiences, users may struggle to engage in natural and enjoyable conversations.

[0719] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0720] In this invention, the server includes means for receiving video footage captured by the user, means for analyzing the received video footage and identifying objects within the video footage, and means for displaying a virtual character and providing a shared dining experience using the generated video footage. This allows the user to alleviate feelings of loneliness through shared dining with a virtual character and easily obtain diverse shared dining experiences.

[0721] A "user" refers to someone who uses this system to virtually share a dining experience.

[0722] "Video" refers to image data captured by the user using their device while eating.

[0723] "Object" refers to specific meals or related items included in the received video.

[0724] "Food images" refer to image data generated that resembles the food that the user is actually eating.

[0725] A "virtual character" refers to a digital avatar displayed using generated food images to create a shared dining experience with the user.

[0726] "Actions" refer to a series of actions performed by the user or virtual character while eating.

[0727] "Natural language processing technology" refers to the technology that enables computers to understand human language and generate meaningful conversations.

[0728] "Informal conversation" refers to communication using everyday, familiar topics.

[0729] "Conversation content" refers to the content of the conversation generated during the interaction with the user.

[0730] The system that realizes this application example allows users to experience shared dining in a virtual environment by linking their terminal with a server. The system mainly consists of the following elements:

[0731] 1. User's device: The user's device is equipped with a camera that captures images of the meal and sends those images to the server. The device can be a smartphone, smart glasses, or a head-mounted display.

[0732] 2. Server Function: The server executes a program developed in Python to analyze received video and identify objects related to food. Specifically, it uses PIL (Python Imaging Library) for video analysis and data processing. The server then generates food images similar to the identified meal. The generated images may utilize basic image generation models, or in some cases, generative AI models.

[0733] 3. Virtual Character Movement and Conversation: Based on the generated food images, the server displays a virtual character on the user's device to provide a shared dining experience. The device's camera captures the user's movements in real time and sends this data to the server. The server analyzes this data and synchronizes the virtual character's movements with the user's movements.

[0734] 4. Conversation generation using natural language processing: The server uses natural language processing techniques to generate conversations with the user. These conversations are based on the user's voice input and profile information, making the shared dining experience richer and more natural.

[0735] As a concrete example, if a user is eating pasta alone at home and uses this system, a virtual character will appear on the screen and begin a conversation about pasta. The virtual character will adjust its behavior to match the user's eating pace and continue the conversation with questions such as, "What are you thinking about while you eat today?"

[0736] Examples of prompt messages include: "Analyze the user's video of their meal, identify the type of food, and represent a similar meal with a virtual character. Next, analyze the user's conversation from the voice input and generate relevant conversation based on that information."

[0737] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0738] Step 1:

[0739] The user uses their device to film their meal and sends the video from the device to the server. The input data is the video captured by the user's camera, and the output is the data transfer to the server. The device requires a stable internet connection and sends the video to the server immediately after filming.

[0740] Step 2:

[0741] The server analyzes the received video of the meal and identifies objects within the video. The input here is the video data sent in step 1, and the output is information about the objects related to the identified meal. The server uses image analysis libraries such as PIL to process the data and identify the type of meal.

[0742] Step 3:

[0743] The server generates images of similar foods based on the identified object information. The input in this step is the object information obtained from step 2, and the output is the generated images of similar foods. The server uses a simple generative AI model to create images that resemble what the user is eating.

[0744] Step 4:

[0745] The generated food images are sent from the server to the user's terminal, which then uses them to display a virtual character. The input is the generated food image data, and the output is the display of the virtual character on the user interface. The terminal receives this image data and displays it immediately.

[0746] Step 5:

[0747] The user's actions while eating are captured by the device's camera, and the device sends this action data to the server. The input here is the user's video activity, and the output is the transfer of action data to the server. The device continuously monitors the user's actions and records the data.

[0748] Step 6:

[0749] The server analyzes the received motion data and synchronizes the virtual character's movements with the user's movements. The input in this step is the motion data from step 5, and the output is the adjusted virtual character's movements. The server determines the user's eating pace and calculates the character's movements accordingly.

[0750] Step 7:

[0751] The server uses natural language processing technology to generate conversations with the user. Input is the user's voice input or existing profile information, and output is the generated conversation content. A generative AI model is used for this process, forming informal conversations based on prompts.

[0752] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0753] This invention aims to provide a more personalized dining experience by combining an emotion engine with a system designed to create a virtual shared dining experience with the user. This system recognizes the user's emotions during meals and adjusts the avatar's responses and conversation content accordingly, thereby creating a comfortable dining environment for the user.

[0754] The user takes a picture of their meal using their device and sends the image to the server. The server analyzes the image and identifies the characteristics of the meal. Furthermore, the server generates similar meal images based on the collected image data and sends them to the device for use in the avatar's shared meal experience.

[0755] The device displays an avatar on the user interface, providing the user with a shared dining experience. Furthermore, the device uses a camera and microphone to monitor the user's facial expressions and voice tone in real time. This data is sent to a server, where an emotion engine analyzes the user's emotions. Based on the analyzed emotion information, the server adjusts the avatar's movements and conversation content.

[0756] For example, if a user is smiling, the server analyzes that positive emotion and sets the avatar to offer more cheerful topics. Conversely, if a user appears stressed, the server takes that emotion into consideration and has the avatar offer encouraging or relaxing topics.

[0757] The avatar uses natural language processing technology to interact with users and respond in a way that is sensitive to their emotions. Furthermore, the emotion engine learns from past user emotional data and predicts future emotional states, allowing the avatar's responses to be pre-adjusted. For example, if past data predicts that a user is prone to stress on certain days of the week, the avatar can prepare relaxing topics for those days.

[0758] Thus, the present invention makes it possible to reduce feelings of loneliness and mental stress, and to make dining a richer experience, by providing an interactive dining experience based on the user's emotions.

[0759] The following describes the processing flow.

[0760] Step 1:

[0761] The user takes a picture of their meal with their device and uploads it to the server through the application. This image contains information including details about the meal.

[0762] Step 2:

[0763] The server processes the received meal images using an image analysis algorithm to identify the type and characteristics of the meal and extract data. This analysis completes the analysis of the meal's contents.

[0764] Step 3:

[0765] The server generates similar meal images based on the analysis data and sends these images to the terminal for use in displaying the user's avatar.

[0766] Step 4:

[0767] The device displays images of similar meals generated alongside the avatar on the user interface, initiating a shared meal experience with the avatar. During this process, the user virtually dines with the avatar.

[0768] Step 5:

[0769] The device uses its camera and microphone to record the user's facial expressions and voice tone in real time. This data is sent to a server for emotion recognition.

[0770] Step 6:

[0771] The server uses an emotion engine to analyze the user's facial expressions and voice data to determine the user's emotional state. Based on this information, it prepares to adjust the avatar's reactions and conversation content.

[0772] Step 7:

[0773] Based on the sentiment analysis results, the server selects a suitable conversation topic for the user, generates conversation content using natural language processing technology, and sends it to the terminal.

[0774] Step 8:

[0775] The terminal applies the conversation content received from the server to the avatar and begins interacting with the user. At this time, the avatar's tone and topics will reflect the user's emotions.

[0776] Step 9:

[0777] The server learns from past emotional data and pre-adjusts the avatar's behavior and conversation based on predicted emotional states in future sessions. This continuous learning makes it possible to further personalize the user experience.

[0778] (Example 2)

[0779] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0780] In modern society, eating is not merely about nutrition; it is an important act that involves psychological and social elements. However, especially when eating alone, feelings of loneliness and mental stress can occur, which can reduce satisfaction with meals. This invention aims to solve these problems and provide users with a richer and more fulfilling dining experience.

[0781] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0782] In this invention, the server includes means for receiving video acquired by the user, means for analyzing the received video and identifying characteristics within the video, and means for generating similar visual data based on the identified characteristics. This allows the user to add visual and emotional interactive elements to existing meals, thereby reducing feelings of loneliness and alleviating stress.

[0783] "Users" refers to individuals or groups who operate the system and enjoy the experience.

[0784] "Video" refers to visual data that users capture and transmit to the system.

[0785] A "server" refers to a computing device that forms the core of a system and is responsible for receiving, analyzing, and providing information.

[0786] "Receiving" refers to the act of a server acquiring data sent by a user.

[0787] "Analysis" refers to the process of extracting and interpreting information from received data.

[0788] "Characteristics" refer to identifiable elements or features contained within images or data.

[0789] "Visual data" refers to digital image information generated based on analyzed characteristics.

[0790] A "virtual character" refers to a virtual personality created within a system for interacting with users.

[0791] "Empathic experience" refers to the emotional interaction that users have with virtual characters within a virtual environment, through shared perceptions.

[0792] "Actions" refer to the reactions and responses that the virtual character performs towards the user.

[0793] "Natural language processing technology" refers to the technology that enables computers to understand human language and generate appropriate dialogue.

[0794] "Emotional data" refers to the results of collecting and analyzing information that indicates the emotional state of users.

[0795] "Prediction" refers to the act of estimating future results or states based on past data.

[0796] This invention is a system that provides users with an interactive and emotionally engaging dining experience. The system consists of a terminal, a server, and multiple software components.

[0797] The user first uses the device's camera to capture video of their meal. The device also has a function to monitor the user's facial expressions and voice tone in real time. This makes it possible to capture the emotions the user is feeling while eating.

[0798] The acquired video is sent from the terminal to the server. The server uses image recognition software (e.g., TensorFlow or OpenCV) to analyze the received video. This analysis identifies the characteristics and features of the food. Based on the identified data, the server uses generative AI models (e.g., GANs) to generate similar visual data. This visual data is used to enrich the user's visual experience.

[0799] The server generates dialogue with a virtual character using natural language processing techniques (e.g., GPT-4) based on the analyzed data. The generated conversation is then delivered to the user via the terminal. For example, the prompt "How are you feeling today?" can be used to initiate the conversation.

[0800] Furthermore, monitoring data from the device is sent to the server's emotion analysis engine. This engine analyzes the user's emotions in real time and adjusts the virtual character's behavior and conversation content based on the analysis results. For example, if the user is smiling, it provides positive conversation content, and if the user is feeling stressed, it selects relaxing topics.

[0801] Furthermore, the server uses past emotional data as a learning platform to predict future emotional states. Based on these predictions, it can anticipate what emotions a user might feel on specific days of the week or in certain situations, and adjust the virtual character's reactions in advance. This ensures that users always have a comfortable and pleasant dining experience.

[0802] In this way, this system can deliver value to users beyond simply eating by providing a personalized dining experience based on their emotions and behavior.

[0803] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0804] Step 1:

[0805] The user takes a picture of their meal using the device's camera. The captured image is input as digital data from the camera device into the device's application. The device then sends this image to the server. As output, the video data is transferred to the server.

[0806] Step 2:

[0807] The server analyzes the received video data. Image recognition software (e.g., TensorFlow, OpenCV) is used for this analysis. The server identifies the characteristics of the food from the video data received as input and extracts characteristic elements (e.g., types of ingredients, types of dishes) based on that. The output is a dataset containing the characteristics of the identified dishes.

[0808] Step 3:

[0809] The server uses a generative AI model (e.g., GAN) to generate similar visual data based on the analyzed data. The characteristic data from step 2 is used as input. Based on this data, the server generates visual data for similar meals. The output is the generated similar visual data.

[0810] Step 4:

[0811] The server sends the generated visual data to the terminal. On the terminal, the virtual character is displayed as a customized avatar, providing a visual experience of shared dining. The generated visual data is interactively displayed to the user through the user interface. The output is a user interface that includes a visually rich avatar.

[0812] Step 5:

[0813] The device uses its built-in camera and microphone to monitor the user's facial expressions and voice tone in real time. Changes in facial expressions and voice signals are acquired as input data. This data is sent to a server, where an emotion analysis engine performs the analysis. The output is a recognition of the user's current emotional state.

[0814] Step 6:

[0815] The server uses generative AI models and natural language processing techniques to adjust the dialogue of the virtual character based on the analyzed emotional state. The input is the emotional data from step 5, and the server generates corresponding prompt sentences. This provides a conversation that matches the user's feelings. The output is the conversation text adapted for the user.

[0816] Step 7:

[0817] The server uses a database to learn from past user sentiment data and predict future emotional states. In this process, past sentiment data and patterns are used as input. Based on the prediction, preparations are made to pre-adjust the responses of the virtual character. The output is a response scenario corresponding to the predicted user's emotional tendencies.

[0818] (Application Example 2)

[0819] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0820] The challenge lies in reducing the feelings of loneliness and mental stress users experience during meals, and providing a richer dining experience. Furthermore, it is desirable to transform mealtime into an enjoyable experience by enabling personalized content selection based on user emotions.

[0821] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0822] This invention includes a server that analyzes the user's visual and auditory information in real time and dynamically adjusts content selection based on the results; a server that provides an interactive dining experience using virtual characters; and a server that generates and provides dialogue that is sensitive to the user's emotions using natural language processing technology. This allows the user to receive personalized responses and conversations during meals, alleviating feelings of loneliness and allowing them to have an enjoyable time.

[0823] A "user" is an individual who uses this system while consuming food.

[0824] "Image" refers to digital visual information captured by the user, including the contents of the meal and related objects.

[0825] "Visual information" refers to information such as the user's facial expressions and gestures, which are acquired through a camera device.

[0826] "Auditory information" refers to information about the user's pronunciation and surrounding sounds acquired through a microphone.

[0827] "Real-time analysis" refers to a process that rapidly processes information as soon as it is acquired and provides immediate feedback on the results.

[0828] "Content" refers to media information such as music, videos, and dialogues provided to users.

[0829] "Dynamic adjustment" refers to the process of changing the system's response and behavior in response to the user's emotions and circumstances.

[0830] A "virtual character" refers to an artificial character with a personality that interacts with the user on a digital interface.

[0831] "Natural language processing technology" is a technology that enables computers to understand and respond to human language.

[0832] An "interactive shared dining experience" refers to a simulated experience in which the user and a virtual character interact while sharing a meal.

[0833] The system that realizes this application is built on a foundation of mobile devices such as smartphones and tablets, and a cloud server. During a meal, the user takes a picture of their facial expression using the device's camera, and this visual information is sent to the server. The server uses an emotion analysis model built with machine learning libraries such as TensorFlow to analyze the received visual and auditory information in real time. The analyzed emotion data is input into a generative AI model, which generates content and dialogue tailored to the user.

[0834] Furthermore, the server uses natural language processing technology to enable smooth interaction with the user via a platform like Dialogflow. This allows the virtual character to adjust its actions and conversations in response to the user's real-time emotions, providing a personalized dining experience. For example, if the user wants to relax during a meal, a prompt such as "Play some relaxing music" is input into the generating AI model, and calming music is played on the device.

[0835] This application configuration allows users to enjoy personalized services that cater to their emotions and preferences, rather than simply receiving mechanically generated content. For example, when a user uses this system to take a break from their busy daily life, they can be refreshed with soothing music and encouraging messages from a virtual character.

[0836] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0837] Step 1:

[0838] The user uses the device's camera to take a picture of their facial expression while eating. At this time, the device receives the captured image data as input and prepares to send the visual information to the server.

[0839] Step 2:

[0840] The server initializes an emotion analysis model to analyze the received visual information. Here, image data is used as input, and facial expression analysis is performed using a TensorFlow-based model. The analysis result outputs the user's emotional state.

[0841] Step 3:

[0842] The server uses a generative AI model based on the analyzed sentiment data to select appropriate content and generate dialogue. It receives sentiment data as input, generates prompts via a natural language processing platform such as Dialogflow, and determines the content and dialogue to be played back to the user as output.

[0843] Step 4:

[0844] The terminal receives prompts and content instructions sent from the server and displays a virtual character on the interface based on them. The virtual character provides the user with personalized conversations and content based on the analysis results.

[0845] Step 5:

[0846] Users enjoy interacting with virtual characters while viewing content provided through their devices. As long as the interaction with the user continues, the device continuously uses its camera and microphone to collect further emotional data and transmit it to the server in real time. This allows the entire system to remain dynamically responsive.

[0847] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0848] The data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include those described above. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions shown by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0849] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0850] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0851] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0852] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0853] The inside of the Emotion Map 400 represents what's in your mind, while the outside represents what you're doing. Therefore, the further you go out the 400-coordinate scale, the more visible your emotions become (the more they manifest in your actions).

[0854] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0855] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0856] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0857] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0858] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0859] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0860] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0861] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0862] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0863] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0864] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0865] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0866] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0867] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0868] The following is further disclosed regarding the embodiments described above.

[0869] (Claim 1)

[0870] A means of receiving images taken by the user,

[0871] A means for analyzing the received image and identifying elements within the image,

[0872] A means for generating similar meal images based on identified elements,

[0873] A means of displaying avatars using generated images and providing a shared dining experience,

[0874] A means for analyzing user behavior and adjusting avatar behavior,

[0875] A means of generating and providing casual conversation using natural language processing technology,

[0876] A system that includes this.

[0877] (Claim 2)

[0878] The system according to claim 1, which analyzes the user's eating speed, acquires data in real time, and dynamically adjusts the avatar's operating speed based on that data.

[0879] (Claim 3)

[0880] The system according to claim 1, which analyzes the user's voice input and profile information to personalize conversation topics and provide conversation content based on specific interests and concerns.

[0881] "Example 1"

[0882] (Claim 1)

[0883] A means for receiving images taken by a terminal that provides information,

[0884] A means for analyzing the received image and identifying objects within the image,

[0885] A means for generating similar meal images based on identified objects,

[0886] A means of displaying a virtual representation using generated images and providing a meal-sharing experience,

[0887] A means for analyzing end-user behavior and adjusting the behavior of the virtual representation,

[0888] A means of generating and providing free conversation using processing technology,

[0889] A system that includes this.

[0890] (Claim 2)

[0891] The system according to claim 1, which instantly acquires information when analyzing the end user's eating speed and dynamically adjusts the operating speed of a virtual representation based on that information.

[0892] (Claim 3)

[0893] The system according to claim 1, which analyzes the end user's voice input and personal information to personalize the themes of free conversation and provide conversation content based on specific interests and concerns.

[0894] "Application Example 1"

[0895] (Claim 1)

[0896] A means of receiving video footage taken by the user,

[0897] A means for analyzing received video and identifying objects within the video,

[0898] A means for generating similar food images based on identified objects,

[0899] A means of displaying virtual characters using generated video and providing a shared dining experience,

[0900] A means for analyzing user behavior and adjusting the behavior of a virtual character,

[0901] A means of generating and providing informal conversations using natural language processing technology,

[0902] A means for analyzing user conversations and generating related conversations,

[0903] A system that includes this.

[0904] (Claim 2)

[0905] The system according to claim 1, which analyzes the user's feeding speed, acquires information in real time, and dynamically adjusts the movement speed of the virtual character based on that information.

[0906] (Claim 3)

[0907] The system according to claim 1, which analyzes the user's voice input and profile information to personalize the themes of informal conversations and provide conversation content based on specific interests and concerns.

[0908] "Example 2 of combining an emotion engine"

[0909] (Claim 1)

[0910] A means of receiving video acquired by the user,

[0911] A means for analyzing received video and identifying characteristics within the video,

[0912] A means for generating similar visual data based on identified characteristics,

[0913] A means of using the generated data to display virtual characters and provide an empathetic experience,

[0914] A means of analyzing the user's emotional state and adjusting the behavior of the virtual character,

[0915] A means of generating and providing dialogue using natural language processing technology,

[0916] A means of learning user emotional data, predicting future emotional states, and adjusting the reactions of a virtual character,

[0917] A system that includes this.

[0918] (Claim 2)

[0919] The system according to claim 1, which analyzes the user's progress in eating, acquires information in real time, and dynamically adjusts the action speed of the virtual character based on that information.

[0920] (Claim 3)

[0921] The system according to claim 1, which analyzes the user's voice information and profile information to personalize the topic of conversation and provide conversation content based on specific interests and concerns.

[0922] "Application example 2 when combining with an emotional engine"

[0923] (Claim 1)

[0924] A means of receiving images taken by the user,

[0925] A means for analyzing the received image and identifying elements within the image,

[0926] A means for generating similar meal images based on identified elements,

[0927] A means of displaying virtual characters using generated images and providing a shared dining experience,

[0928] A means for analyzing user behavior and adjusting the behavior of a virtual character,

[0929] A means of generating and providing dialogue using natural language processing technology,

[0930] A means of analyzing the user's visual and auditory information in real time and dynamically adjusting content selection based on the results,

[0931] A system that includes this.

[0932] (Claim 2)

[0933] The system according to claim 1, which analyzes the user's eating habits and dynamically adjusts the movement speed of the virtual character and the display of content.

[0934] (Claim 3)

[0935] The system according to claim 1, which analyzes the user's acoustic input and personal information to personalize the topic of conversation and provide conversation content based on specific interests and concerns. [Explanation of symbols]

[0936] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of receiving images taken by the user, A means for analyzing the received image and identifying elements within the image, A means for generating similar meal images based on identified elements, A means of displaying avatars using generated images and providing a shared dining experience, A means for analyzing user behavior and adjusting avatar behavior, A means of generating and providing casual conversation using natural language processing technology, A system that includes this.

2. The system according to claim 1, which analyzes the user's eating speed, acquires data in real time, and dynamically adjusts the avatar's movement speed based on that data.

3. The system according to claim 1, which analyzes the user's voice input and profile information to personalize conversation topics and provide conversation content based on specific interests and concerns.