system
The system addresses social isolation during meals by using image recognition and natural language processing to create a shared dining experience with pace-matched avatars, enhancing user interaction and emotional support.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Increasing social isolation and loneliness due to solitary meals, exacerbated by differing meal times and paces, necessitate a system that provides a natural and interactive shared meal experience.
A system that uses image recognition to identify meals, generates an avatar eating similarly, adjusts its movements to match the user's pace, and engages in natural dialogue through natural language processing to alleviate loneliness.
Enables a shared dining experience, reducing feelings of loneliness and providing emotional support through interactive conversations with avatars.
Smart Images

Figure 2026068366000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern society where the number of people living alone is increasing, social loneliness deaths have become a problem due to feelings of loneliness and lack of communication during meals. Also, it is difficult to provide a satisfactory shared meal experience considering differences in time and meal pace when it is difficult to have a meal with real people. In such a situation, there is a need to provide users with a natural and interactive shared meal experience to relieve feelings of loneliness.
Means for Solving the Problems
[0005] This invention provides a system that receives images of meals taken by users through their devices, identifies the type of meal using image recognition technology, and generates an avatar that eats a similar meal based on the identified meal. The generated avatar can adjust its movements to match the user's eating speed, enabling a natural shared dining experience. Furthermore, by using natural language processing to enable natural dialogue with the user and converting voice data into text to generate responses, feelings of loneliness can be reduced. In this way, it is possible to alleviate users' feelings of loneliness and provide a solution to the problem of social isolation leading to death.
[0006] A "user" refers to an individual who uses the system to experience sharing a meal with others.
[0007] A "device" refers to a computing device, such as a smartphone or PC, that a user uses to access this service, upload images of their meals, or engage in voice conversations.
[0008] "Image recognition means" refers to the technology and algorithms used to identify the type of food from a received image.
[0009] "Avatar generation means" refers to the technology, process, or algorithm for creating an avatar, which is a virtual character that eats a meal similar to the user's meal.
[0010] "Motion adjustment means" refers to a technology or process that allows an avatar to adjust its movements to match the user's eating speed, providing a natural shared dining experience.
[0011] "Natural language processing means" refers to technologies and algorithms that convert speech data into text data and generate appropriate responses in order to enable natural and effective dialogue between users and avatars. [Brief explanation of the drawing]
[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0014] First, the terms used in the following description will be explained.
[0015] In the following embodiments, a tagged processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0016] In the following embodiments, a tagged RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0017] In the following embodiments, a tagged storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0018] In the following embodiments, a tagged communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0020] [First Embodiment]
[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0033] This invention provides an online communal dining service designed to alleviate feelings of loneliness when users eat alone and address the problem of social isolation. This system is primarily implemented using terminals, servers, and communication technologies between them.
[0034] Users launch a dedicated application on their smartphones or PCs. They take pictures of the food they are eating with their device's camera and send them to the server via the application. At this time, users can also set conversation topics and the personality of their avatar character.
[0035] The server uses image recognition technology to analyze the received food images. This identifies the type and characteristics of the food and generates an avatar eating a meal similar to the one the user is eating. Image analysis is performed based on information such as color, shape, and arrangement.
[0036] Once an avatar is generated, the server controls its animation to match the user's eating pace. Furthermore, the server uses natural language processing to create conversation scenarios to support interaction with the user and determines what the avatar should say at the appropriate time.
[0037] To enable natural dialogue, the terminal receives the user's voice input, converts it into text data, and sends it to the server. The server analyzes the conversation based on the text data, generates the avatar's next response, and sends it to the terminal. The terminal uses speech synthesis technology to have the avatar speak that response aloud, which is then played for the user to hear.
[0038] As a concrete example, consider a scenario where a user is eating a salad for dinner. The user uploads an image of the salad, and the server uses image recognition technology to identify the salad and generates an avatar eating it. The avatar then speaks to the user, saying things like, "That salad looks delicious, what dressing are you using?" and responds to the user's answer with something like, "That sounds delicious, I'd like to try it too."
[0039] In this way, users can enjoy meals with their avatars without feeling lonely, even when dining alone. This system aims not only to alleviate feelings of loneliness but also to provide users with an enjoyable conversational experience.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] Users launch the shared meal app on their device, take photos of their meals, and upload them to the app.
[0043] Step 2:
[0044] The device sends images taken by the user directly to the server. It also simultaneously sends information about the conversation topic and avatar's personality that the user has set.
[0045] Step 3:
[0046] The server executes an image recognition algorithm to analyze the received food images. This process identifies the type and characteristics of the food (e.g., color and shape) from the image.
[0047] Step 4:
[0048] Based on the identified meal information, the server determines an avatar meal similar to the one the user is eating, and then generates the avatar based on that information.
[0049] Step 5:
[0050] The server constructs conversation scenarios based on the user's settings and prepares avatar responses using a natural language processing system. The avatar's movements are configured to adapt to the user's eating speed.
[0051] Step 6:
[0052] The server sends the generated avatar data and the corresponding conversation scenario to the terminal.
[0053] Step 7:
[0054] The device displays an avatar and begins to act in accordance with the user's eating speed. The avatar also speaks to the user based on a conversation scenario received from the server.
[0055] Step 8:
[0056] The user initiates a conversation by responding to a question from the avatar. The device receives the voice response.
[0057] Step 9:
[0058] The terminal converts the user's voice input into text using speech recognition and sends it to the server.
[0059] Step 10:
[0060] The server processes the received text and generates an appropriate response to the user's reply. The response is then sent back to the terminal.
[0061] Step 11:
[0062] The device uses speech synthesis technology to convert responses from the server into speech, which is then presented to the user through an avatar.
[0063] Step 12:
[0064] This conversation continues until the user finishes, at which point the user either closes the app or ends the conversation.
[0065] Step 13:
[0066] The server saves session data, activates the feedback function, and prepares to collect additional input and comments from the user.
[0067] (Example 1)
[0068] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0069] In modern society, changes in individual lifestyles and work styles are leading to an increase in people experiencing feelings of loneliness and social isolation. These feelings are particularly pronounced in situations where people frequently eat alone. There is a need for systems that alleviate the associated psychological burden and provide opportunities for people to enjoy richer social interactions.
[0070] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0071] In this invention, the server includes an image analysis means for receiving images of food taken by the user through an information processing device and identifying the type of food from the images; a virtual character generation means for generating a virtual character that consumes similar food based on the identified type of food; and an action adjustment means for the generated virtual character to move in accordance with the user's eating speed. This makes it possible for the user to enjoy a meal with a virtual character even when eating alone, reducing feelings of loneliness and providing an enjoyable conversational experience.
[0072] An "information processing device" is a device that processes image data and input information captured by a user and has the function of communicating with a server.
[0073] "Image analysis means" refers to technical methods and processes for recognizing objects based on specific features from received image data and analyzing that data.
[0074] A "virtual character generation method" refers to a technology or system that creates an animated character that interacts with the user digitally and mimics their actions, based on analyzed data.
[0075] "Motion adjustment means" refers to technology that adapts the movements and behavior of the generated virtual character in real time to the user's actions and pace.
[0076] "Natural language processing means" refers to computer programs and algorithms that understand the context of user utterances and dialogues and generate appropriate responses.
[0077] A "generative AI model" is a machine learning model used to generate new information or content based on data.
[0078] "Conversation control means" refers to technologies that manage the flow and content of conversations so that virtual characters provide consistent responses.
[0079] This invention is an online communal dining system aimed at making meals more enjoyable and reducing feelings of loneliness for users. Users launch a dedicated application using a smartphone or PC, which is an information processing device. This application has a camera function, allowing users to take pictures of their meals. The captured images are temporarily stored on the device and sent to the server along with the conversation topic selected by the user and the personality settings of their avatar.
[0080] The server uses image analysis techniques to process the received image data. Specifically, it utilizes machine learning models to analyze the types and characteristics of food from the images. For example, a convolutional neural network (CNN) can be used. Based on the analysis results, the server generates a virtual character that consumes food similar to the user's meal. This character is adjusted to operate in real time in accordance with the user's eating pace and movements.
[0081] Interaction between the virtual character and the user is achieved using natural language processing technology. The server creates conversation scenarios using a generated AI model and controls the virtual character to ensure consistent speech. Additionally, voice input from the user is converted into text data by the terminal and sent to the server. Based on this text data, the server generates appropriate responses for the user and uses speech synthesis technology to vocalize what the virtual character would say.
[0082] As a concrete example, consider a scenario where a user is eating a salad for dinner. The user takes a picture of the salad with their smartphone and selects "healthy eating habits" as the conversation topic. The server identifies the salad from the image and generates a virtual character eating the salad. Based on the user's selection, the character starts a conversation such as, "That's a healthy meal, are you trying any special recipes?" and provides dialogue that responds to the user's answers.
[0083] An example of a prompt message is, "Generate a dialogue scenario for a virtual character, assuming the user is eating a healthy salad." In this way, the system alleviates feelings of loneliness and provides the user with an enjoyable dining experience.
[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0085] Step 1:
[0086] The user launches a dedicated application on the information processing device and takes a picture of their meal. The captured image is temporarily stored on the terminal. The input is image data of the meal acquired by the camera, and the output is image data ready to be sent to the server. Specifically, the user operates the camera function using the user interface to capture and save a still image.
[0087] Step 2:
[0088] The device sends captured image data, along with user-defined conversation topics and avatar personality information, to the server. Input consists of user-provided images and settings, while output is a data packet containing this information. Specifically, the data is packetized and transmitted using a secure communication protocol.
[0089] Step 3:
[0090] The server receives the image data and uses image analysis tools to identify the type and characteristics of the food. The input is image data and user configuration information, and the output is information about the type and characteristics of the food. Here, an image analysis algorithm (e.g., a convolutional neural network) is applied to analyze the patterns of color and shape.
[0091] Step 4:
[0092] The server generates a virtual character to consume similar foods based on identified food information. The input is identified food information, and the output is model data for the virtual character. This includes specific actions to create the character's appearance and movement patterns using the generated AI model.
[0093] Step 5:
[0094] The server adjusts the virtual character's movements to match the user's eating pace and generates conversation scenarios using natural language processing. Inputs include user-provided settings and food information, while outputs consist of dynamically controlled animation sequences and conversation scenario text. Specific operations include adjusting the character's movement speed and creating dialogue for appropriate timing.
[0095] Step 6:
[0096] The device receives voice input from the user and converts it into text data. The input is the user's voice data, and the output is text data. Specifically, it uses a microphone to acquire voice and speech recognition software (e.g., a voice-to-text API) to convert it into text.
[0097] Step 7:
[0098] The server generates a virtual character response to the user based on text data and sends it to the terminal. The input is the user's text data, and the output is the virtual character's response text. The specific operation here is the process of generating a context-appropriate response using a generative AI model.
[0099] Step 8:
[0100] The terminal uses speech synthesis technology to convert the received response text into speech, and a virtual character speaks it aloud. The input is response text data, and the output is audio data. Specifically, it uses a text-to-speech engine to convert the text into speech and plays it back through the speaker.
[0101] (Application Example 1)
[0102] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0103] This invention aims to provide a system that reduces feelings of loneliness and makes the dining experience more enjoyable, given the increasing number of people who eat alone in modern society. In particular, it seeks to help users maintain their interest in food and increase social interaction through interaction with virtual characters and virtual store experiences in augmented reality spaces.
[0104] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0105] In this invention, the server includes: an image analysis means for receiving images of meals acquired by the user through an information processing device and identifying the type of meal from the images; a virtual character generation means for generating a virtual character eating a similar meal based on the identified type of meal; an action control means for the generated virtual character to move in accordance with the user's eating pace; and a means for displaying a virtual space on the user's visual device using augmented reality technology and conducting a conversation within that virtual space. This enables the user to enjoy a realistic and interactive dining experience by utilizing the virtual space.
[0106] "Image analysis means" refers to a technology that receives images of meals acquired by a user through an information processing device and has the function of identifying the type of meal from those images.
[0107] "Virtual character generation means" refers to a technology for generating virtual characters that eat similar meals based on a specified type of meal.
[0108] "Motion control means" refers to a technology that adjusts the generated virtual character to move in accordance with the user's eating pace.
[0109] "Natural language processing means" refers to technologies that perform language analysis and generation to enable natural conversations between users and virtual characters.
[0110] Augmented reality technology is a technology that displays a virtual space on the user's visual device and allows them to converse with virtual characters within that virtual space.
[0111] An "information processing device" is an electronic device used by a user to acquire images and process the received data.
[0112] A "visual device" is a device that uses augmented reality technology to display a virtual space to the user.
[0113] A "virtual space" is a digital environment constructed within the user's visual environment using augmented reality technology.
[0114] This invention is implemented using an information processing device and related technologies to improve the experience of users dining alone. The system primarily utilizes image analysis means, virtual character generation means, motion control means, natural language processing means, and augmented reality technology.
[0115] The server receives images of meals taken by users using information processing devices such as smartphones and personal computers. These images undergo data analysis using the Google® Cloud Vision API to identify the type of meal. At this stage, features such as color and shape are analyzed, and information on similar meals is retrieved from the database.
[0116] Once the type of meal is identified, the server uses a generative AI model to generate a virtual character common to the user's meals. This virtual character is animated and controlled on platforms such as Unity or Unreal Engine, expressing movements that correspond to the user's eating pace.
[0117] Furthermore, the server generates conversation content through natural language processing technology. In this process, OpenAI's GPT model is utilized to determine the flow of the conversation and the appropriate timing for what to say. Dialogue scenarios with the user are generated based on prompt sentences such as, "This looks delicious! What's the secret ingredient in this dish?"
[0118] Furthermore, augmented reality technology plays a role in facilitating interaction by displaying a virtual space on the user's visual devices (such as smart glasses or head-mounted displays). Virtual characters provide a realistic dining experience, allowing users to enjoy interactive conversations even when alone.
[0119] For example, if a user is trying a new cheese fondue, a virtual character might say, "This cheese has a unique flavor, doesn't it? Do you have any particular preferences when choosing bread?" enriching the dining experience. In this way, an environment is created where users can immerse themselves in their meal without feeling isolated.
[0120] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0121] Step 1:
[0122] Users take pictures of their meals using the camera on their devices, such as smartphones or personal computers. The captured images are sent to a server via a dedicated application on the device. In this input process, the user's identification information is added along with the transmitted image data.
[0123] Step 2:
[0124] The server uses the received meal image data to perform image analysis using the Google Cloud Vision API. As a result of the analysis, data regarding the type and characteristics of the meal is extracted. Specifically, the color, shape, and arrangement of ingredients are analyzed and classified into specific categories. The output is information regarding the identified type of meal.
[0125] Step 3:
[0126] The server generates a virtual character using a generative AI model based on the image analysis results. During the generation process, it utilizes OpenAI's GPT model to generate character information related to the user's meal. This virtual character includes conversation topics and personality settings related to the type of meal. The output consists of the digital model of this virtual character and conversation prompts.
[0127] Step 4:
[0128] The server uses Unity or Unreal Engine to control the generated virtual character's movements. During this process, a sequence of movements is designed based on the user's eating pace, and the virtual character's animations are adjusted accordingly. The input is the user's eating pace information, and the output is the character's movements synchronized with the user's pace.
[0129] Step 5:
[0130] The server generates conversational content for virtual characters using natural language processing. It constructs a concrete conversational flow based on prompts obtained from a generating AI model, taking nuances and timing into consideration. The input is user-generated text data, and the output is the virtual character's response data based on that input.
[0131] Step 6:
[0132] The device displays a virtual space generated using augmented reality technology on the user's visual device. Through smart glasses or a head-mounted display, the user experiences real-time interaction with a virtual character. The output consists of a visual virtual space and synthesized speech from the character.
[0133] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0134] This invention provides an online shared dining service designed to alleviate feelings of loneliness and provide emotional support to users when they are eating alone. This invention achieves more personalized interaction by combining an emotion recognition engine. The system primarily utilizes terminals, servers, and communication technologies between them.
[0135] Users launch a dedicated application on a device such as a smartphone or PC. First, users can take a picture of their meal with their camera and send the image to the server via the app. Furthermore, users can also set conversation topics and the personality of their avatar.
[0136] The server uses advanced image recognition algorithms to analyze food images. This identifies the type and characteristics of the food and generates an avatar that eats with the user. The generated avatar can move in sync with the user's eating speed.
[0137] Furthermore, the system incorporates an emotion recognition engine. The server analyzes the user's voice tone and facial expressions, evaluating the user's emotional state in real time. This information is reflected in the avatar's responses and actions, providing conversations appropriate to the user's emotions. For example, if the user appears depressed, the avatar can continue the conversation with words of encouragement or humor.
[0138] As a concrete example, consider a scenario where a user is eating a salad for dinner. When the user uploads an image of the salad, the server recognizes the salad and generates a suitable avatar. When the user says, "I'm tired today," and the emotion engine detects a downcast tone, the server sends a response such as, "That must have been tough, let's talk about something relaxing together." In this way, users can enjoy a shared dining experience that is empathetic and doesn't make them feel lonely.
[0139] This system aims to create a more fulfilling mealtime experience by alleviating users' feelings of loneliness and providing appropriate responses tailored to their individual emotions.
[0140] The following describes the processing flow.
[0141] Step 1:
[0142] Users launch the shared meal app on their device, take photos of their meal before starting, and upload them to the app.
[0143] Step 2:
[0144] The device sends the captured image of the meal, along with the user's chosen conversation topics and avatar personality, to the server.
[0145] Step 3:
[0146] The server analyzes the transmitted food images using image recognition technology to identify the type and characteristics of the food.
[0147] Step 4:
[0148] The server generates avatars that eat similar meals based on the identified meal type and prepares the avatar data. At this time, it configures the avatars to operate in accordance with the user's eating speed.
[0149] Step 5:
[0150] The emotion recognition engine analyzes the user's voice data and, if necessary, facial expression data from camera footage in real time to determine the user's emotional state.
[0151] Step 6:
[0152] The server generates conversational content suitable for the user based on the emotional state obtained from the emotion recognition engine, and constructs response scenarios for the avatar.
[0153] Step 7:
[0154] The server sends the generated avatar and response scenario to the terminal.
[0155] Step 8:
[0156] The device displays an avatar on the screen, begins performing actions that match the user's eating speed, and the avatar speaks to the user based on a pre-set scenario.
[0157] Step 9:
[0158] The user communicates with the avatar using voice, and the device receives that audio.
[0159] Step 10:
[0160] The terminal converts the user's voice input into text data and sends it to the server.
[0161] Step 11:
[0162] The server analyzes the text data and generates an appropriate response that takes into account the user's reply. In doing so, it also creates a response that reflects the user's emotional state.
[0163] Step 12:
[0164] The server sends a response to the terminal, and the terminal uses speech synthesis to have an avatar speak the content to the user.
[0165] Step 13:
[0166] When the user finishes their meal and ends the session, the app displays a screen requesting feedback.
[0167] Step 14:
[0168] The server collects user feedback and stores it in appropriate data storage to help improve future services.
[0169] (Example 2)
[0170] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0171] Traditionally, eating alone has often been associated with feelings of loneliness, making it difficult to satisfy users' emotional needs. Furthermore, there has been a need for ways to enrich and personalize the general food and beverage consumption experience and provide emotional support.
[0172] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0173] In this invention, the server includes: image recognition means for receiving images of food and beverages acquired by the user through an information processing device and identifying the type of food and beverage from the images; character generation means for generating a virtual character that consumes similar food and beverages based on the identified type of food and beverage; action adjustment means for the generated virtual character to move in accordance with the user's food and beverage consumption rate; emotion analysis means for analyzing the user's voice tone and facial expressions and recognizing the user's emotions; and natural language processing means for the virtual character to engage in natural dialogue with the user in accordance with the analyzed emotions. As a result, even when eating alone, the user can receive emotional support from the virtual character, experience personalized interaction, and reduce feelings of loneliness.
[0174] An "information processing device" is a device that allows users to input digital data and process or communicate that information, and includes smartphones and personal computers.
[0175] "Image recognition means" refers to technologies and algorithms for analyzing received image data and identifying its content and characteristics.
[0176] A "virtual character" is a digital agent created within a digital environment and designed to interact with users.
[0177] "Character generation means" refers to technologies and algorithms for constructing virtual characters based on user information.
[0178] "Motion adjustment means" refers to technology for adjusting the movements of a virtual character to match the user's actions and speed.
[0179] "Emotional analysis means" refers to technology for analyzing and evaluating a user's emotional state from their voice and facial expressions.
[0180] "Natural language processing means" refers to technologies that process information necessary for users and systems to engage in natural language conversations.
[0181] This invention provides a system that allows users to alleviate feelings of loneliness and receive emotional support while eating alone. Specifically, users can utilize this system by using a dedicated application on an information processing device.
[0182] The user takes a picture of their food or drink using their device's camera and sends the image to the server via the application. The server uses image recognition technology to process the received image. Software such as TENSORFLOW® or OpenCV is used to identify the type of food or drink from the image. Based on this identified information, the server generates a virtual character. 3D modeling software such as Blender can be used for this generation.
[0183] The generated virtual character is adjusted to move in accordance with the user's food and drink consumption rate. The server also captures the user's voice and facial expressions through the terminal's microphone and camera, and analyzes the user's emotional state using emotion analysis tools. This process utilizes PyTorch for voice analysis and Dlib for facial expression analysis.
[0184] Based on these analysis results, the server uses natural language processing technology to generate a response appropriate to the user's emotions. The generated response is then delivered to the user in real time through a virtual character.
[0185] As a concrete example of its use, if a user says "I'm tired today" during a meal, the server can analyze this and a virtual character can respond with encouraging words such as "You've worked hard, let's take a break." In this way, users can enjoy an emotionally supportive dining experience even when dining alone.
[0186] Examples of prompts to input into the generation AI model include, "Please generate a virtual character that can provide emotional support when a user is eating alone," and "Please analyze images of food and drinks and suggest a character that suits the user."
[0187] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0188] Step 1:
[0189] The user launches a dedicated application on the information processing device and takes pictures of food and beverages using the device's camera. The food and beverage images, as input, are sent to the server via the app. Specifically, the user taps a button within the app, activates the camera, captures the image, and uploads it.
[0190] Step 2:
[0191] The server processes the received images using image recognition technology. It performs data analysis on input images of food and beverages using TensorFlow and OpenCV to identify the type and characteristics of the food and beverages. The output provides information about the identified food and beverages. Specifically, the server analyzes the image pixel data and performs matching by comparing it with a known database of food and beverages.
[0192] Step 3:
[0193] The server generates virtual characters based on identified food and beverage information. Using the food and beverage information as input, it generates 3D characters using modeling tools such as Blender. The output is a model of a virtual character associated with the food and beverage. In terms of specific operation, the character's appearance is adjusted according to the selected template, creating synergy with the food and beverage.
[0194] Step 4:
[0195] The generated virtual character's movements are adjusted to match the user's eating speed. The user's eating speed, as input, is analyzed from the user's operation history and video data, and the character's animation is adjusted using a movement adjustment mechanism. The output is character movements tailored to the user. Specifically, the user's movement patterns are analyzed, and the character mimics table manners.
[0196] Step 5:
[0197] The device acquires the user's voice and facial expressions through the microphone and camera and transmits them to the server. It collects the user's voice and video as input data and transfers it to the server in real time. Specifically, the microphone's voice recording function and the camera's video recording function work in conjunction.
[0198] Step 6:
[0199] The server uses this data to perform emotion analysis. It processes audio and facial expression data as input using PyTorch or Dlib to identify the user's emotional state. The output is the result of the user's emotion analysis. Specifically, the server performs audio waveform analysis and facial feature point detection to calculate an emotional index.
[0200] Step 7:
[0201] The server considers the user's emotional state and generates appropriate responses using natural language processing. It utilizes a generative AI model based on the emotional analysis results as input to generate conversational content tailored to the user. The output is a text response to the user. In its specific operation, the server uses an AI-powered conversation generation engine to create voice or text responses through a virtual character.
[0202] (Application Example 2)
[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0204] In modern society, the number of people dining alone is increasing, and the loneliness they experience in such situations is becoming a mental burden. There is a need for ways to provide emotional satisfaction and social connection even when customers visit restaurants alone. In particular, for customers who regularly dine alone, the challenge is to provide an interactive dining experience in physical restaurants that they can enjoy without feeling lonely.
[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0206] In this invention, the server includes an image recognition means for receiving images of food and beverages taken by the user through an information processing device and identifying the type of food and beverage from the image; a virtual character generation means for generating a virtual character representing a similar food and beverage based on the identified type of food and beverage; and an emotion recognition means for analyzing voice data and nonverbal signs obtained from the user and evaluating the user's emotional state. This makes it possible to alleviate feelings of loneliness and provide an enjoyable dining experience to customers who visit a restaurant and dine alone through natural conversation using a virtual character.
[0207] An "information processing device" generally refers to a device such as a computer, tablet, or smartphone that inputs, processes, and outputs information.
[0208] "Food and beverages" refers to food and beverages that have been prepared or supplied for human consumption.
[0209] "Image recognition means" refers to a means that uses computer vision technology to analyze images and recognize specific objects or features from them.
[0210] A "virtual character generation method" is a means of generating a character that mimics the form of a human being in a digital environment based on specific conditions, using computer technology.
[0211] "Operation adjustment means" refers to means for adjusting the appropriate timing and operation according to the user's actions and circumstances.
[0212] An "emotion recognition tool" is a tool that analyzes a user's voice, facial expressions, and actions to evaluate their emotional state.
[0213] A "generative AI model" is an artificial intelligence technology that learns from large amounts of data and generates appropriate responses and content in response to various inputs.
[0214] "Natural language processing methods" refer to methods that utilize technologies for computers to understand, process, and generate human language.
[0215] A "dialogue provision method" refers to a technology or system for providing dialogue through interaction with the user.
[0216] To implement this invention, a terminal acting as an information processing device is first required. The user uses this terminal to take pictures of food and beverages and transmits them to the server. The server, acting as an image recognition means, analyzes the received images via computer vision technology to identify the type of food and beverage. Then, a virtual person generation means generates a virtual person to eat and drink with the user based on the identified food and beverage.
[0217] This virtual character moves in accordance with the user's eating and drinking speed using motion adjustment means. Furthermore, it analyzes the user's voice and facial expressions through emotion recognition means to evaluate the user's emotional state in real time. Natural language processing means incorporating a generative AI model generates appropriate responses according to the user's emotional state, providing natural and emotionally resonant dialogue.
[0218] As a concrete example, consider a customer who visits a restaurant alone and orders a dessert. The customer takes a picture of the dessert with their device's camera and sends it to the server. The server recognizes the dessert using image recognition and generates a virtual person. When the customer says in a cheerful voice, "Today was a good day!", emotion recognition captures the positive tone. Based on this information, the generative AI model generates a response such as, "That's wonderful! What happened?" and continues the conversation.
[0219] The specific technologies used include Google Cloud Vision API for image recognition, Google Cloud Speech-to-Text for speech-to-text conversion, Azure® Emotion API for emotion recognition, and AI models such as GPT-3® for natural language generation.
[0220] As an example of a prompt for a generative AI model, in a situation where a customer says, "I want to have a good time today," the following would be used:
[0221] "The customer is speaking in a cheerful tone and says they want to have a good time today. Please bring up some lighthearted topics."
[0222] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0223] Step 1:
[0224] The user uses a device to take pictures of food and drinks. The input is image data from the device's camera. The output is the transmission of the image data to the server. The device prepares the captured images to be sent to the server in the appropriate format.
[0225] Step 2:
[0226] The server receives image data and uses image recognition to identify the type of food or beverage. The input is image data, and the output is the recognized type of food or beverage information. The server uses the Google Cloud Vision API for image recognition, analyzing objects in the image to classify the food or beverage.
[0227] Step 3:
[0228] The server generates a virtual character using a virtual character generation mechanism based on the recognized type of food and beverage. The input is information about the type of food and beverage, and the output is the generated virtual character data. The server designs a virtual character for an appropriate eating and drinking situation and sets visual and behavioral parameters.
[0229] Step 4:
[0230] The user sends voice and facial expressions to the server via their device. Input consists of the user's voice and facial expression data. Output is the transmission of data to the server for emotion analysis. The user's device captures this data in real time and sends it to the server for analysis.
[0231] Step 5:
[0232] The server uses emotion recognition to evaluate the user's emotional state from their voice and facial expressions. The input consists of voice data and facial expression data, and the output is information about the user's emotional state. The server uses the Azure Emotion API to identify emotions from the input data and record the state.
[0233] Step 6:
[0234] The server uses a generative AI model to generate appropriate responses based on the user's emotional state and food / drink consumption. The input is emotional state information and virtual character data, and the output is the generated response message. The server uses a GPT-3 model to construct natural-sounding conversational content based on the prompt sentences.
[0235] Step 7:
[0236] The terminal receives response messages from the server and provides dialogue to the user through a virtual character. The input is the response message from the server, and the output is the dialogue content displayed to the user. The terminal reproduces the obtained dialogue content on the screen as actions of the virtual character and conveys it to the user.
[0237] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0238] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0239] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0240] [Second Embodiment]
[0241] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0242] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0243] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0244] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0245] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0246] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0247] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0248] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0249] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0250] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0251] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0252] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0253] This invention provides an online communal dining service designed to alleviate feelings of loneliness when users eat alone and address the problem of social isolation. This system is primarily implemented using terminals, servers, and communication technologies between them.
[0254] Users launch a dedicated application on their smartphones or PCs. They take pictures of the food they are eating with their device's camera and send them to the server via the application. At this time, users can also set conversation topics and the personality of their avatar character.
[0255] The server uses image recognition technology to analyze the received food images. This identifies the type and characteristics of the food and generates an avatar eating a meal similar to the one the user is eating. Image analysis is performed based on information such as color, shape, and arrangement.
[0256] Once an avatar is generated, the server controls its animation to match the user's eating pace. Furthermore, the server uses natural language processing to create conversation scenarios to support interaction with the user and determines what the avatar should say at the appropriate time.
[0257] To enable natural dialogue, the terminal receives the user's voice input, converts it into text data, and sends it to the server. The server analyzes the conversation based on the text data, generates the avatar's next response, and sends it to the terminal. The terminal uses speech synthesis technology to have the avatar speak that response aloud, which is then played for the user to hear.
[0258] As a concrete example, consider a scenario where a user is eating a salad for dinner. The user uploads an image of the salad, and the server uses image recognition technology to identify the salad and generates an avatar eating it. The avatar then speaks to the user, saying things like, "That salad looks delicious, what dressing are you using?" and responds to the user's answer with something like, "That sounds delicious, I'd like to try it too."
[0259] In this way, users can enjoy meals with their avatars without feeling lonely, even when dining alone. This system aims not only to alleviate feelings of loneliness but also to provide users with an enjoyable conversational experience.
[0260] The following describes the processing flow.
[0261] Step 1:
[0262] Users launch the shared meal app on their device, take photos of their meals, and upload them to the app.
[0263] Step 2:
[0264] The device sends images taken by the user directly to the server. It also simultaneously sends information about the conversation topic and avatar's personality that the user has set.
[0265] Step 3:
[0266] The server executes an image recognition algorithm to analyze the received food images. This process identifies the type and characteristics of the food (e.g., color and shape) from the image.
[0267] Step 4:
[0268] Based on the identified meal information, the server determines an avatar meal similar to the one the user is eating, and then generates the avatar based on that information.
[0269] Step 5:
[0270] The server constructs conversation scenarios based on the user's settings and prepares avatar responses using a natural language processing system. The avatar's movements are configured to adapt to the user's eating speed.
[0271] Step 6:
[0272] The server sends the generated avatar data and the corresponding conversation scenario to the terminal.
[0273] Step 7:
[0274] The device displays an avatar and begins to act in accordance with the user's eating speed. The avatar also speaks to the user based on a conversation scenario received from the server.
[0275] Step 8:
[0276] The user starts a conversation by answering the avatar's question. The terminal receives the response in voice.
[0277] Step 9:
[0278] The terminal converts the user's voice input into text using the voice recognition function and sends it to the server.
[0279] Step 10:
[0280] The server processes the received text and generates an appropriate response to the user's response. The response content is sent back to the terminal again.
[0281] Step 11:
[0282] The terminal converts the response from the server into voice using voice synthesis technology and makes the avatar speak to the user.
[0283] Step 12:
[0284] This conversation exchange continues until the user ends it, and when it ends, the user performs an operation to end the app or close the conversation.
[0285] Step 13:
[0286] The server saves the session data, activates the feedback function, and prepares to collect additional input and feedback from the user.
[0287] (Example 1)
[0288] Next, Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0289] In modern society, changes in individual lifestyles and work styles are leading to an increase in people experiencing feelings of loneliness and social isolation. These feelings are particularly pronounced in situations where people frequently eat alone. There is a need for systems that alleviate the associated psychological burden and provide opportunities for people to enjoy richer social interactions.
[0290] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0291] In this invention, the server includes an image analysis means for receiving images of food taken by the user through an information processing device and identifying the type of food from the images; a virtual character generation means for generating a virtual character that consumes similar food based on the identified type of food; and an action adjustment means for the generated virtual character to move in accordance with the user's eating speed. This makes it possible for the user to enjoy a meal with a virtual character even when eating alone, reducing feelings of loneliness and providing an enjoyable conversational experience.
[0292] An "information processing device" is a device that processes image data and input information captured by a user and has the function of communicating with a server.
[0293] "Image analysis means" refers to technical methods and processes for recognizing objects based on specific features from received image data and analyzing that data.
[0294] A "virtual character generation method" refers to a technology or system that creates an animated character that interacts with the user digitally and mimics their actions, based on analyzed data.
[0295] "Motion adjustment means" refers to technology that adapts the movements and behavior of the generated virtual character in real time to the user's actions and pace.
[0296] "Natural language processing means" refers to computer programs and algorithms that understand the context of user utterances and dialogues and generate appropriate responses.
[0297] A "generative AI model" is a machine learning model used to generate new information or content based on data.
[0298] "Conversation control means" refers to technologies that manage the flow and content of conversations so that virtual characters provide consistent responses.
[0299] This invention is an online communal dining system aimed at making meals more enjoyable and reducing feelings of loneliness for users. Users launch a dedicated application using a smartphone or PC, which is an information processing device. This application has a camera function, allowing users to take pictures of their meals. The captured images are temporarily stored on the device and sent to the server along with the conversation topic selected by the user and the personality settings of their avatar.
[0300] The server uses image analysis techniques to process the received image data. Specifically, it utilizes machine learning models to analyze the types and characteristics of food from the images. For example, a convolutional neural network (CNN) can be used. Based on the analysis results, the server generates a virtual character that consumes food similar to the user's meal. This character is adjusted to operate in real time in accordance with the user's eating pace and movements.
[0301] Interaction between the virtual character and the user is achieved using natural language processing technology. The server creates conversation scenarios using a generated AI model and controls the virtual character to ensure consistent speech. Additionally, voice input from the user is converted into text data by the terminal and sent to the server. Based on this text data, the server generates appropriate responses for the user and uses speech synthesis technology to vocalize what the virtual character would say.
[0302] As a specific example, consider the case where a user is having a salad for dinner. The user takes a picture of the salad with a smartphone and selects "healthy eating lifestyle" as the conversation topic. The server identifies the salad from the image and generates a virtual character eating the salad. Based on the user's selection, the character starts a conversation such as "That's a healthy meal. Are you trying any special recipes?" and provides a dialogue according to the user's answer.
[0303] An example of a prompt sentence is "Please generate a dialogue scenario for a virtual character in the setting where the user is eating a healthy salad." In this way, the system reduces the user's sense of loneliness and provides an enjoyable dining experience.
[0304] The flow of the specific process in Example 1 will be described using FIG. 11.
[0305] Step 1:
[0306] The user starts a dedicated application on the information processing device and takes a picture of the meal. The captured image is temporarily saved on the terminal. The input is the image data of the meal obtained by the camera, and the output is the image data ready to be sent to the server. The specific operation here is to use the user interface to operate the camera function, capture a still image, and save it. <{
[0307] Step 2:
[0308] The terminal sends the captured image data together with information about the conversation topic and avatar personality set by the user to the server. The input is the image and setting information by the user, and the output is a data packet containing this information. The specific operation is to packetize the data and send it using a secure communication protocol.
[0309] Step 3:
[0310] The server receives the image data and uses image analysis tools to identify the type and characteristics of the food. The input is image data and user configuration information, and the output is information about the type and characteristics of the food. Here, an image analysis algorithm (e.g., a convolutional neural network) is applied to analyze the patterns of color and shape.
[0311] Step 4:
[0312] The server generates a virtual character to consume similar foods based on identified food information. The input is identified food information, and the output is model data for the virtual character. This includes specific actions to create the character's appearance and movement patterns using the generated AI model.
[0313] Step 5:
[0314] The server adjusts the virtual character's movements to match the user's eating pace and generates conversation scenarios using natural language processing. Inputs include user-provided settings and food information, while outputs consist of dynamically controlled animation sequences and conversation scenario text. Specific operations include adjusting the character's movement speed and creating dialogue for appropriate timing.
[0315] Step 6:
[0316] The device receives voice input from the user and converts it into text data. The input is the user's voice data, and the output is text data. Specifically, it uses a microphone to acquire voice and speech recognition software (e.g., a voice-to-text API) to convert it into text.
[0317] Step 7:
[0318] The server generates a virtual character response to the user based on text data and sends it to the terminal. The input is the user's text data, and the output is the virtual character's response text. The specific operation here is the process of generating a context-appropriate response using a generative AI model.
[0319] Step 8:
[0320] The terminal uses speech synthesis technology to convert the received response text into speech, and a virtual character speaks it aloud. The input is response text data, and the output is audio data. Specifically, it uses a text-to-speech engine to convert the text into speech and plays it back through the speaker.
[0321] (Application Example 1)
[0322] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0323] This invention aims to provide a system that reduces feelings of loneliness and makes the dining experience more enjoyable, given the increasing number of people who eat alone in modern society. In particular, it seeks to help users maintain their interest in food and increase social interaction through interaction with virtual characters and virtual store experiences in augmented reality spaces.
[0324] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0325] In this invention, the server includes: an image analysis means for receiving images of meals acquired by the user through an information processing device and identifying the type of meal from the images; a virtual character generation means for generating a virtual character eating a similar meal based on the identified type of meal; an action control means for the generated virtual character to move in accordance with the user's eating pace; and a means for displaying a virtual space on the user's visual device using augmented reality technology and conducting a conversation within that virtual space. This enables the user to enjoy a realistic and interactive dining experience by utilizing the virtual space.
[0326] "Image analysis means" refers to a technology that receives images of meals acquired by a user through an information processing device and has the function of identifying the type of meal from those images.
[0327] "Virtual character generation means" refers to a technology for generating virtual characters that eat similar meals based on a specified type of meal.
[0328] "Motion control means" refers to a technology that adjusts the generated virtual character to move in accordance with the user's eating pace.
[0329] "Natural language processing means" refers to technologies that perform language analysis and generation to enable natural conversations between users and virtual characters.
[0330] Augmented reality technology is a technology that displays a virtual space on the user's visual device and allows them to converse with virtual characters within that virtual space.
[0331] An "information processing device" is an electronic device used by a user to acquire images and process the received data.
[0332] A "visual device" is a device that uses augmented reality technology to display a virtual space to the user.
[0333] A "virtual space" is a digital environment constructed within the user's visual environment using augmented reality technology.
[0334] This invention is implemented using an information processing device and related technologies to improve the experience of users dining alone. The system primarily utilizes image analysis means, virtual character generation means, motion control means, natural language processing means, and augmented reality technology.
[0335] The server receives images of meals taken by users using information processing devices such as smartphones and personal computers. These images undergo data analysis using the Google Cloud Vision API to identify the type of meal. At this stage, features such as color and shape are analyzed, and information on similar meals is retrieved from the database.
[0336] Once the type of meal is identified, the server uses a generative AI model to generate a virtual character common to the user's meals. This virtual character is animated and controlled on platforms such as Unity or Unreal Engine, expressing movements that correspond to the user's eating pace.
[0337] Furthermore, the server generates conversation content through natural language processing technology. In this process, OpenAI's GPT model is utilized to determine the flow of the conversation and the appropriate timing for what to say. Dialogue scenarios with the user are generated based on prompt sentences such as, "This looks delicious! What's the secret ingredient in this dish?"
[0338] Furthermore, augmented reality technology plays a role in facilitating interaction by displaying a virtual space on the user's visual devices (such as smart glasses or head-mounted displays). Virtual characters provide a realistic dining experience, allowing users to enjoy interactive conversations even when alone.
[0339] For example, if a user is trying a new cheese fondue, a virtual character might say, "This cheese has a unique flavor, doesn't it? Do you have any particular preferences when choosing bread?" enriching the dining experience. In this way, an environment is created where users can immerse themselves in their meal without feeling isolated.
[0340] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0341] Step 1:
[0342] Users take pictures of their meals using the camera on their devices, such as smartphones or personal computers. The captured images are sent to a server via a dedicated application on the device. In this input process, the user's identification information is added along with the transmitted image data.
[0343] Step 2:
[0344] The server uses the received meal image data to perform image analysis using the Google Cloud Vision API. As a result of the analysis, data regarding the type and characteristics of the meal is extracted. Specifically, the color, shape, and arrangement of ingredients are analyzed and classified into specific categories. The output is information regarding the identified type of meal.
[0345] Step 3:
[0346] The server generates a virtual character using a generative AI model based on the image analysis results. During the generation process, it utilizes OpenAI's GPT model to generate character information related to the user's meal. This virtual character includes conversation topics and personality settings related to the type of meal. The output consists of the digital model of this virtual character and conversation prompts.
[0347] Step 4:
[0348] The server uses Unity or Unreal Engine to control the generated virtual character's movements. During this process, a sequence of movements is designed based on the user's eating pace, and the virtual character's animations are adjusted accordingly. The input is the user's eating pace information, and the output is the character's movements synchronized with the user's pace.
[0349] Step 5:
[0350] The server generates conversational content for virtual characters using natural language processing. It constructs a concrete conversational flow based on prompts obtained from a generating AI model, taking nuances and timing into consideration. The input is user-generated text data, and the output is the virtual character's response data based on that input.
[0351] Step 6:
[0352] The device displays a virtual space generated using augmented reality technology on the user's visual device. Through smart glasses or a head-mounted display, the user experiences real-time interaction with a virtual character. The output consists of a visual virtual space and synthesized speech from the character.
[0353] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0354] This invention provides an online shared dining service designed to alleviate feelings of loneliness and provide emotional support to users when they are eating alone. This invention achieves more personalized interaction by combining an emotion recognition engine. The system primarily utilizes terminals, servers, and communication technologies between them.
[0355] Users launch a dedicated application on a device such as a smartphone or PC. First, users can take a picture of their meal with their camera and send the image to the server via the app. Furthermore, users can also set conversation topics and the personality of their avatar.
[0356] The server uses advanced image recognition algorithms to analyze food images. This identifies the type and characteristics of the food and generates an avatar that eats with the user. The generated avatar can move in sync with the user's eating speed.
[0357] Furthermore, the system incorporates an emotion recognition engine. The server analyzes the user's voice tone and facial expressions, evaluating the user's emotional state in real time. This information is reflected in the avatar's responses and actions, providing conversations appropriate to the user's emotions. For example, if the user appears depressed, the avatar can continue the conversation with words of encouragement or humor.
[0358] As a concrete example, consider a scenario where a user is eating a salad for dinner. When the user uploads an image of the salad, the server recognizes the salad and generates a suitable avatar. When the user says, "I'm tired today," and the emotion engine detects a downcast tone, the server sends a response such as, "That must have been tough, let's talk about something relaxing together." In this way, users can enjoy a shared dining experience that is empathetic and doesn't make them feel lonely.
[0359] This system aims to create a more fulfilling mealtime experience by alleviating users' feelings of loneliness and providing appropriate responses tailored to their individual emotions.
[0360] The following describes the processing flow.
[0361] Step 1:
[0362] Users launch the shared meal app on their device, take photos of their meal before starting, and upload them to the app.
[0363] Step 2:
[0364] The device sends the captured image of the meal, along with the user's chosen conversation topics and avatar personality, to the server.
[0365] Step 3:
[0366] The server analyzes the transmitted food images using image recognition technology to identify the type and characteristics of the food.
[0367] Step 4:
[0368] The server generates avatars that eat similar meals based on the identified meal type and prepares the avatar data. At this time, it configures the avatars to operate in accordance with the user's eating speed.
[0369] Step 5:
[0370] The emotion recognition engine analyzes the user's voice data and, if necessary, facial expression data from camera footage in real time to determine the user's emotional state.
[0371] Step 6:
[0372] The server generates conversational content suitable for the user based on the emotional state obtained from the emotion recognition engine, and constructs response scenarios for the avatar.
[0373] Step 7:
[0374] The server sends the generated avatar and response scenario to the terminal.
[0375] Step 8:
[0376] The device displays an avatar on the screen, begins performing actions that match the user's eating speed, and the avatar speaks to the user based on a pre-set scenario.
[0377] Step 9:
[0378] The user communicates with the avatar using voice, and the device receives that audio.
[0379] Step 10:
[0380] The terminal converts the user's voice input into text data and sends it to the server.
[0381] Step 11:
[0382] The server analyzes the text data and generates an appropriate response that takes into account the user's reply. In doing so, it also creates a response that reflects the user's emotional state.
[0383] Step 12:
[0384] The server sends a response to the terminal, and the terminal uses speech synthesis to have an avatar speak the content to the user.
[0385] Step 13:
[0386] When the user finishes their meal and ends the session, the app displays a screen requesting feedback.
[0387] Step 14:
[0388] The server collects user feedback and stores it in appropriate data storage to help improve future services.
[0389] (Example 2)
[0390] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0391] Traditionally, eating alone has often been associated with feelings of loneliness, making it difficult to satisfy users' emotional needs. Furthermore, there has been a need for ways to enrich and personalize the general food and beverage consumption experience and provide emotional support.
[0392] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0393] In this invention, the server includes: image recognition means for receiving images of food and beverages acquired by the user through an information processing device and identifying the type of food and beverage from the images; character generation means for generating a virtual character that consumes similar food and beverages based on the identified type of food and beverage; action adjustment means for the generated virtual character to move in accordance with the user's food and beverage consumption rate; emotion analysis means for analyzing the user's voice tone and facial expressions and recognizing the user's emotions; and natural language processing means for the virtual character to engage in natural dialogue with the user in accordance with the analyzed emotions. As a result, even when eating alone, the user can receive emotional support from the virtual character, experience personalized interaction, and reduce feelings of loneliness.
[0394] An "information processing device" is a device that allows users to input digital data and process or communicate that information, and includes smartphones and personal computers.
[0395] "Image recognition means" refers to technologies and algorithms for analyzing received image data and identifying its content and characteristics.
[0396] A "virtual character" is a digital agent created within a digital environment and designed to interact with users.
[0397] "Character generation means" refers to technologies and algorithms for constructing virtual characters based on user information.
[0398] "Motion adjustment means" refers to technology for adjusting the movements of a virtual character to match the user's actions and speed.
[0399] "Emotional analysis means" refers to technology for analyzing and evaluating a user's emotional state from their voice and facial expressions.
[0400] "Natural language processing means" refers to technologies that process information necessary for users and systems to engage in natural language conversations.
[0401] This invention provides a system that allows users to alleviate feelings of loneliness and receive emotional support while eating alone. Specifically, users can utilize this system by using a dedicated application on an information processing device.
[0402] The user takes a picture of their food or drink using their device's camera and sends the image to the server via the application. The server uses image recognition technology to process the received image. This involves using software such as TensorFlow or OpenCV to identify the type of food or drink from the image. Based on this identified information, the server generates a virtual character. 3D modeling software such as Blender can be used for this generation.
[0403] The generated virtual character is adjusted to move in accordance with the user's food and drink consumption rate. The server also captures the user's voice and facial expressions through the terminal's microphone and camera, and analyzes the user's emotional state using emotion analysis tools. This process utilizes PyTorch for voice analysis and Dlib for facial expression analysis.
[0404] Based on these analysis results, the server uses natural language processing technology to generate a response appropriate to the user's emotions. The generated response is then delivered to the user in real time through a virtual character.
[0405] As a concrete example of its use, if a user says "I'm tired today" during a meal, the server can analyze this and a virtual character can respond with encouraging words such as "You've worked hard, let's take a break." In this way, users can enjoy an emotionally supportive dining experience even when dining alone.
[0406] Examples of prompts to input into the generation AI model include, "Please generate a virtual character that can provide emotional support when a user is eating alone," and "Please analyze images of food and drinks and suggest a character that suits the user."
[0407] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0408] Step 1:
[0409] The user launches a dedicated application on the information processing device and takes pictures of food and beverages using the device's camera. The food and beverage images, as input, are sent to the server via the app. Specifically, the user taps a button within the app, activates the camera, captures the image, and uploads it.
[0410] Step 2:
[0411] The server processes the received images using image recognition technology. It performs data analysis on input images of food and beverages using TensorFlow and OpenCV to identify the type and characteristics of the food and beverages. The output provides information about the identified food and beverages. Specifically, the server analyzes the image pixel data and performs matching by comparing it with a known database of food and beverages.
[0412] Step 3:
[0413] The server generates virtual characters based on identified food and beverage information. Using the food and beverage information as input, it generates 3D characters using modeling tools such as Blender. The output is a model of a virtual character associated with the food and beverage. In terms of specific operation, the character's appearance is adjusted according to the selected template, creating synergy with the food and beverage.
[0414] Step 4:
[0415] The generated virtual character's movements are adjusted to match the user's eating speed. The user's eating speed, as input, is analyzed from the user's operation history and video data, and the character's animation is adjusted using a movement adjustment mechanism. The output is character movements tailored to the user. Specifically, the user's movement patterns are analyzed, and the character mimics table manners.
[0416] Step 5:
[0417] The device acquires the user's voice and facial expressions through the microphone and camera and transmits them to the server. It collects the user's voice and video as input data and transfers it to the server in real time. Specifically, the microphone's voice recording function and the camera's video recording function work in conjunction.
[0418] Step 6:
[0419] The server uses this data to perform emotion analysis. It processes audio and facial expression data as input using PyTorch or Dlib to identify the user's emotional state. The output is the result of the user's emotion analysis. Specifically, the server performs audio waveform analysis and facial feature point detection to calculate an emotional index.
[0420] Step 7:
[0421] The server considers the user's emotional state and generates appropriate responses using natural language processing. It utilizes a generative AI model based on the emotional analysis results as input to generate conversational content tailored to the user. The output is a text response to the user. In its specific operation, the server uses an AI-powered conversation generation engine to create voice or text responses through a virtual character.
[0422] (Application Example 2)
[0423] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0424] In modern society, the number of people dining alone is increasing, and the loneliness they experience in such situations is becoming a mental burden. There is a need for ways to provide emotional satisfaction and social connection even when customers visit restaurants alone. In particular, for customers who regularly dine alone, the challenge is to provide an interactive dining experience in physical restaurants that they can enjoy without feeling lonely.
[0425] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0426] In this invention, the server includes an image recognition means for receiving images of food and beverages taken by the user through an information processing device and identifying the type of food and beverage from the image; a virtual character generation means for generating a virtual character representing a similar food and beverage based on the identified type of food and beverage; and an emotion recognition means for analyzing voice data and nonverbal signs obtained from the user and evaluating the user's emotional state. This makes it possible to alleviate feelings of loneliness and provide an enjoyable dining experience to customers who visit a restaurant and dine alone through natural conversation using a virtual character.
[0427] An "information processing device" generally refers to a device such as a computer, tablet, or smartphone that inputs, processes, and outputs information.
[0428] "Food and beverages" refers to food and beverages that have been prepared or supplied for human consumption.
[0429] "Image recognition means" refers to a means that uses computer vision technology to analyze images and recognize specific objects or features from them.
[0430] A "virtual character generation method" is a means of generating a character that mimics the form of a human being in a digital environment based on specific conditions, using computer technology.
[0431] "Operation adjustment means" refers to means for adjusting the appropriate timing and operation according to the user's actions and circumstances.
[0432] An "emotion recognition tool" is a tool that analyzes a user's voice, facial expressions, and actions to evaluate their emotional state.
[0433] A "generative AI model" is an artificial intelligence technology that learns from large amounts of data and generates appropriate responses and content in response to various inputs.
[0434] "Natural language processing methods" refer to methods that utilize technologies for computers to understand, process, and generate human language.
[0435] A "dialogue provision method" refers to a technology or system for providing dialogue through interaction with the user.
[0436] To implement this invention, a terminal acting as an information processing device is first required. The user uses this terminal to take pictures of food and beverages and transmits them to the server. The server, acting as an image recognition means, analyzes the received images via computer vision technology to identify the type of food and beverage. Then, a virtual person generation means generates a virtual person to eat and drink with the user based on the identified food and beverage.
[0437] This virtual character moves in accordance with the user's eating and drinking speed using motion adjustment means. Furthermore, it analyzes the user's voice and facial expressions through emotion recognition means to evaluate the user's emotional state in real time. Natural language processing means incorporating a generative AI model generates appropriate responses according to the user's emotional state, providing natural and emotionally resonant dialogue.
[0438] As a concrete example, consider a customer who visits a restaurant alone and orders a dessert. The customer takes a picture of the dessert with their device's camera and sends it to the server. The server recognizes the dessert using image recognition and generates a virtual person. When the customer says in a cheerful voice, "Today was a good day!", emotion recognition captures the positive tone. Based on this information, the generative AI model generates a response such as, "That's wonderful! What happened?" and continues the conversation.
[0439] The specific technologies used include Google Cloud Vision API for image recognition, Google Cloud Speech-to-Text for speech-to-text conversion, Azure Emotion API for emotion recognition, and AI models such as GPT-3 for natural language generation.
[0440] As an example of a prompt for a generative AI model, in a situation where a customer says, "I want to have a good time today," the following would be used:
[0441] "The customer is speaking in a cheerful tone and says they want to have a good time today. Please bring up some lighthearted topics."
[0442] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0443] Step 1:
[0444] The user uses a device to take pictures of food and drinks. The input is image data from the device's camera. The output is the transmission of the image data to the server. The device prepares the captured images to be sent to the server in the appropriate format.
[0445] Step 2:
[0446] The server receives image data and uses image recognition to identify the type of food or beverage. The input is image data, and the output is the recognized type of food or beverage information. The server uses the Google Cloud Vision API for image recognition, analyzing objects in the image to classify the food or beverage.
[0447] Step 3:
[0448] The server generates a virtual character using a virtual character generation mechanism based on the recognized type of food and beverage. The input is information about the type of food and beverage, and the output is the generated virtual character data. The server designs a virtual character for an appropriate eating and drinking situation and sets visual and behavioral parameters.
[0449] Step 4:
[0450] The user sends voice and facial expressions to the server via their device. Input consists of the user's voice and facial expression data. Output is the transmission of data to the server for emotion analysis. The user's device captures this data in real time and sends it to the server for analysis.
[0451] Step 5:
[0452] The server uses emotion recognition to evaluate the user's emotional state from their voice and facial expressions. The input consists of voice data and facial expression data, and the output is information about the user's emotional state. The server uses the Azure Emotion API to identify emotions from the input data and record the state.
[0453] Step 6:
[0454] The server uses a generative AI model to generate appropriate responses based on the user's emotional state and food / drink consumption. The input is emotional state information and virtual character data, and the output is the generated response message. The server uses a GPT-3 model to construct natural-sounding conversational content based on the prompt sentences.
[0455] Step 7:
[0456] The terminal receives response messages from the server and provides dialogue to the user through a virtual character. The input is the response message from the server, and the output is the dialogue content displayed to the user. The terminal reproduces the obtained dialogue content on the screen as actions of the virtual character and conveys it to the user.
[0457] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0458] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0459] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0460] [Third Embodiment]
[0461] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0462] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0463] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0464] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0465] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0466] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0467] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0468] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0469] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0470] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0471] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0472] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0473] This invention provides an online communal dining service designed to alleviate feelings of loneliness when users eat alone and address the problem of social isolation. This system is primarily implemented using terminals, servers, and communication technologies between them.
[0474] Users launch a dedicated application on their smartphones or PCs. They take pictures of the food they are eating with their device's camera and send them to the server via the application. At this time, users can also set conversation topics and the personality of their avatar character.
[0475] The server uses image recognition technology to analyze the received food images. This identifies the type and characteristics of the food and generates an avatar eating a meal similar to the one the user is eating. Image analysis is performed based on information such as color, shape, and arrangement.
[0476] Once an avatar is generated, the server controls its animation to match the user's eating pace. Furthermore, the server uses natural language processing to create conversation scenarios to support interaction with the user and determines what the avatar should say at the appropriate time.
[0477] To enable natural dialogue, the terminal receives the user's voice input, converts it into text data, and sends it to the server. The server analyzes the conversation based on the text data, generates the avatar's next response, and sends it to the terminal. The terminal uses speech synthesis technology to have the avatar speak that response aloud, which is then played for the user to hear.
[0478] As a concrete example, consider a scenario where a user is eating a salad for dinner. The user uploads an image of the salad, and the server uses image recognition technology to identify the salad and generates an avatar eating it. The avatar then speaks to the user, saying things like, "That salad looks delicious, what dressing are you using?" and responds to the user's answer with something like, "That sounds delicious, I'd like to try it too."
[0479] In this way, users can enjoy meals with their avatars without feeling lonely, even when dining alone. This system aims not only to alleviate feelings of loneliness but also to provide users with an enjoyable conversational experience.
[0480] The following describes the processing flow.
[0481] Step 1:
[0482] Users launch the shared meal app on their device, take photos of their meals, and upload them to the app.
[0483] Step 2:
[0484] The device sends images taken by the user directly to the server. It also simultaneously sends information about the conversation topic and avatar's personality that the user has set.
[0485] Step 3:
[0486] The server executes an image recognition algorithm to analyze the received food images. This process identifies the type and characteristics of the food (e.g., color and shape) from the image.
[0487] Step 4:
[0488] Based on the identified meal information, the server determines an avatar meal similar to the one the user is eating, and then generates the avatar based on that information.
[0489] Step 5:
[0490] The server constructs conversation scenarios based on the user's settings and prepares avatar responses using a natural language processing system. The avatar's movements are configured to adapt to the user's eating speed.
[0491] Step 6:
[0492] The server sends the generated avatar data and the corresponding conversation scenario to the terminal.
[0493] Step 7:
[0494] The device displays an avatar and begins to act in accordance with the user's eating speed. The avatar also speaks to the user based on a conversation scenario received from the server.
[0495] Step 8:
[0496] The user initiates a conversation by responding to a question from the avatar. The device receives the voice response.
[0497] Step 9:
[0498] The terminal converts the user's voice input into text using speech recognition and sends it to the server.
[0499] Step 10:
[0500] The server processes the received text and generates an appropriate response to the user's reply. The response is then sent back to the terminal.
[0501] Step 11:
[0502] The device uses speech synthesis technology to convert responses from the server into speech, which is then presented to the user through an avatar.
[0503] Step 12:
[0504] This conversation continues until the user finishes, at which point the user either closes the app or ends the conversation.
[0505] Step 13:
[0506] The server saves session data, activates the feedback function, and prepares to collect additional input and comments from the user.
[0507] (Example 1)
[0508] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0509] In modern society, changes in individual lifestyles and work styles are leading to an increase in people experiencing feelings of loneliness and social isolation. These feelings are particularly pronounced in situations where people frequently eat alone. There is a need for systems that alleviate the associated psychological burden and provide opportunities for people to enjoy richer social interactions.
[0510] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0511] In this invention, the server includes an image analysis means for receiving images of food taken by the user through an information processing device and identifying the type of food from the images; a virtual character generation means for generating a virtual character that consumes similar food based on the identified type of food; and an action adjustment means for the generated virtual character to move in accordance with the user's eating speed. This makes it possible for the user to enjoy a meal with a virtual character even when eating alone, reducing feelings of loneliness and providing an enjoyable conversational experience.
[0512] An "information processing device" is a device that processes image data and input information captured by a user and has the function of communicating with a server.
[0513] "Image analysis means" refers to technical methods and processes for recognizing objects based on specific features from received image data and analyzing that data.
[0514] A "virtual character generation method" refers to a technology or system that creates an animated character that interacts with the user digitally and mimics their actions, based on analyzed data.
[0515] "Motion adjustment means" refers to technology that adapts the movements and behavior of the generated virtual character in real time to the user's actions and pace.
[0516] "Natural language processing means" refers to computer programs and algorithms that understand the context of user utterances and dialogues and generate appropriate responses.
[0517] A "generative AI model" is a machine learning model used to generate new information or content based on data.
[0518] "Conversation control means" refers to technologies that manage the flow and content of conversations so that virtual characters provide consistent responses.
[0519] This invention is an online communal dining system aimed at making meals more enjoyable and reducing feelings of loneliness for users. Users launch a dedicated application using a smartphone or PC, which is an information processing device. This application has a camera function, allowing users to take pictures of their meals. The captured images are temporarily stored on the device and sent to the server along with the conversation topic selected by the user and the personality settings of their avatar.
[0520] The server uses image analysis techniques to process the received image data. Specifically, it utilizes machine learning models to analyze the types and characteristics of food from the images. For example, a convolutional neural network (CNN) can be used. Based on the analysis results, the server generates a virtual character that consumes food similar to the user's meal. This character is adjusted to operate in real time in accordance with the user's eating pace and movements.
[0521] Interaction between the virtual character and the user is achieved using natural language processing technology. The server creates conversation scenarios using a generated AI model and controls the virtual character to ensure consistent speech. Additionally, voice input from the user is converted into text data by the terminal and sent to the server. Based on this text data, the server generates appropriate responses for the user and uses speech synthesis technology to vocalize what the virtual character would say.
[0522] As a concrete example, consider a scenario where a user is eating a salad for dinner. The user takes a picture of the salad with their smartphone and selects "healthy eating habits" as the conversation topic. The server identifies the salad from the image and generates a virtual character eating the salad. Based on the user's selection, the character starts a conversation such as, "That's a healthy meal, are you trying any special recipes?" and provides dialogue that responds to the user's answers.
[0523] An example of a prompt message is, "Generate a dialogue scenario for a virtual character, assuming the user is eating a healthy salad." In this way, the system alleviates feelings of loneliness and provides the user with an enjoyable dining experience.
[0524] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0525] Step 1:
[0526] The user launches a dedicated application on the information processing device and takes a picture of their meal. The captured image is temporarily stored on the terminal. The input is image data of the meal acquired by the camera, and the output is image data ready to be sent to the server. Specifically, the user operates the camera function using the user interface to capture and save a still image.
[0527] Step 2:
[0528] The device sends captured image data, along with user-defined conversation topics and avatar personality information, to the server. Input consists of user-provided images and settings, while output is a data packet containing this information. Specifically, the data is packetized and transmitted using a secure communication protocol.
[0529] Step 3:
[0530] The server receives the image data and uses image analysis tools to identify the type and characteristics of the food. The input is image data and user configuration information, and the output is information about the type and characteristics of the food. Here, an image analysis algorithm (e.g., a convolutional neural network) is applied to analyze the patterns of color and shape.
[0531] Step 4:
[0532] The server generates a virtual character to consume similar foods based on identified food information. The input is identified food information, and the output is model data for the virtual character. This includes specific actions to create the character's appearance and movement patterns using the generated AI model.
[0533] Step 5:
[0534] The server adjusts the virtual character's movements to match the user's eating pace and generates conversation scenarios using natural language processing. Inputs include user-provided settings and food information, while outputs consist of dynamically controlled animation sequences and conversation scenario text. Specific operations include adjusting the character's movement speed and creating dialogue for appropriate timing.
[0535] Step 6:
[0536] The device receives voice input from the user and converts it into text data. The input is the user's voice data, and the output is text data. Specifically, it uses a microphone to acquire voice and speech recognition software (e.g., a voice-to-text API) to convert it into text.
[0537] Step 7:
[0538] The server generates a virtual character response to the user based on text data and sends it to the terminal. The input is the user's text data, and the output is the virtual character's response text. The specific operation here is the process of generating a context-appropriate response using a generative AI model.
[0539] Step 8:
[0540] The terminal uses speech synthesis technology to convert the received response text into speech, and a virtual character speaks it aloud. The input is response text data, and the output is audio data. Specifically, it uses a text-to-speech engine to convert the text into speech and plays it back through the speaker.
[0541] (Application Example 1)
[0542] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0543] This invention aims to provide a system that reduces feelings of loneliness and makes the dining experience more enjoyable, given the increasing number of people who eat alone in modern society. In particular, it seeks to help users maintain their interest in food and increase social interaction through interaction with virtual characters and virtual store experiences in augmented reality spaces.
[0544] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0545] In this invention, the server includes: an image analysis means for receiving images of meals acquired by the user through an information processing device and identifying the type of meal from the images; a virtual character generation means for generating a virtual character eating a similar meal based on the identified type of meal; an action control means for the generated virtual character to move in accordance with the user's eating pace; and a means for displaying a virtual space on the user's visual device using augmented reality technology and conducting a conversation within that virtual space. This enables the user to enjoy a realistic and interactive dining experience by utilizing the virtual space.
[0546] "Image analysis means" refers to a technology that receives images of meals acquired by a user through an information processing device and has the function of identifying the type of meal from those images.
[0547] "Virtual character generation means" refers to a technology for generating virtual characters that eat similar meals based on a specified type of meal.
[0548] "Motion control means" refers to a technology that adjusts the generated virtual character to move in accordance with the user's eating pace.
[0549] "Natural language processing means" refers to technologies that perform language analysis and generation to enable natural conversations between users and virtual characters.
[0550] Augmented reality technology is a technology that displays a virtual space on the user's visual device and allows them to converse with virtual characters within that virtual space.
[0551] An "information processing device" is an electronic device used by a user to acquire images and process the received data.
[0552] A "visual device" is a device that uses augmented reality technology to display a virtual space to the user.
[0553] A "virtual space" is a digital environment constructed within the user's visual environment using augmented reality technology.
[0554] This invention is implemented using an information processing device and related technologies to improve the experience of users dining alone. The system primarily utilizes image analysis means, virtual character generation means, motion control means, natural language processing means, and augmented reality technology.
[0555] The server receives images of meals taken by users using information processing devices such as smartphones and personal computers. These images undergo data analysis using the Google Cloud Vision API to identify the type of meal. At this stage, features such as color and shape are analyzed, and information on similar meals is retrieved from the database.
[0556] Once the type of meal is identified, the server uses a generative AI model to generate a virtual character common to the user's meals. This virtual character is animated and controlled on platforms such as Unity or Unreal Engine, expressing movements that correspond to the user's eating pace.
[0557] Furthermore, the server generates conversation content through natural language processing technology. In this process, OpenAI's GPT model is utilized to determine the flow of the conversation and the appropriate timing for what to say. Dialogue scenarios with the user are generated based on prompt sentences such as, "This looks delicious! What's the secret ingredient in this dish?"
[0558] Furthermore, augmented reality technology plays a role in facilitating interaction by displaying a virtual space on the user's visual devices (such as smart glasses or head-mounted displays). Virtual characters provide a realistic dining experience, allowing users to enjoy interactive conversations even when alone.
[0559] For example, if a user is trying a new cheese fondue, a virtual character might say, "This cheese has a unique flavor, doesn't it? Do you have any particular preferences when choosing bread?" enriching the dining experience. In this way, an environment is created where users can immerse themselves in their meal without feeling isolated.
[0560] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0561] Step 1:
[0562] Users take pictures of their meals using the camera on their devices, such as smartphones or personal computers. The captured images are sent to a server via a dedicated application on the device. In this input process, the user's identification information is added along with the transmitted image data.
[0563] Step 2:
[0564] The server uses the received meal image data to perform image analysis using the Google Cloud Vision API. As a result of the analysis, data regarding the type and characteristics of the meal is extracted. Specifically, the color, shape, and arrangement of ingredients are analyzed and classified into specific categories. The output is information regarding the identified type of meal.
[0565] Step 3:
[0566] The server generates a virtual character using a generative AI model based on the image analysis results. During the generation process, it utilizes OpenAI's GPT model to generate character information related to the user's meal. This virtual character includes conversation topics and personality settings related to the type of meal. The output consists of the digital model of this virtual character and conversation prompts.
[0567] Step 4:
[0568] The server uses Unity or Unreal Engine to control the generated virtual character's movements. During this process, a sequence of movements is designed based on the user's eating pace, and the virtual character's animations are adjusted accordingly. The input is the user's eating pace information, and the output is the character's movements synchronized with the user's pace.
[0569] Step 5:
[0570] The server generates conversational content for virtual characters using natural language processing. It constructs a concrete conversational flow based on prompts obtained from a generating AI model, taking nuances and timing into consideration. The input is user-generated text data, and the output is the virtual character's response data based on that input.
[0571] Step 6:
[0572] The device displays a virtual space generated using augmented reality technology on the user's visual device. Through smart glasses or a head-mounted display, the user experiences real-time interaction with a virtual character. The output consists of a visual virtual space and synthesized speech from the character.
[0573] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0574] This invention provides an online shared dining service designed to alleviate feelings of loneliness and provide emotional support to users when they are eating alone. This invention achieves more personalized interaction by combining an emotion recognition engine. The system primarily utilizes terminals, servers, and communication technologies between them.
[0575] Users launch a dedicated application on a device such as a smartphone or PC. First, users can take a picture of their meal with their camera and send the image to the server via the app. Furthermore, users can also set conversation topics and the personality of their avatar.
[0576] The server uses advanced image recognition algorithms to analyze food images. This identifies the type and characteristics of the food and generates an avatar that eats with the user. The generated avatar can move in sync with the user's eating speed.
[0577] Furthermore, the system incorporates an emotion recognition engine. The server analyzes the user's voice tone and facial expressions, evaluating the user's emotional state in real time. This information is reflected in the avatar's responses and actions, providing conversations appropriate to the user's emotions. For example, if the user appears depressed, the avatar can continue the conversation with words of encouragement or humor.
[0578] As a concrete example, consider a scenario where a user is eating a salad for dinner. When the user uploads an image of the salad, the server recognizes the salad and generates a suitable avatar. When the user says, "I'm tired today," and the emotion engine detects a downcast tone, the server sends a response such as, "That must have been tough, let's talk about something relaxing together." In this way, users can enjoy a shared dining experience that is empathetic and doesn't make them feel lonely.
[0579] This system aims to create a more fulfilling mealtime experience by alleviating users' feelings of loneliness and providing appropriate responses tailored to their individual emotions.
[0580] The following describes the processing flow.
[0581] Step 1:
[0582] Users launch the shared meal app on their device, take photos of their meal before starting, and upload them to the app.
[0583] Step 2:
[0584] The device sends the captured image of the meal, along with the user's chosen conversation topics and avatar personality, to the server.
[0585] Step 3:
[0586] The server analyzes the transmitted food images using image recognition technology to identify the type and characteristics of the food.
[0587] Step 4:
[0588] The server generates avatars that eat similar meals based on the identified meal type and prepares the avatar data. At this time, it configures the avatars to operate in accordance with the user's eating speed.
[0589] Step 5:
[0590] The emotion recognition engine analyzes the user's voice data and, if necessary, facial expression data from camera footage in real time to determine the user's emotional state.
[0591] Step 6:
[0592] The server generates conversational content suitable for the user based on the emotional state obtained from the emotion recognition engine, and constructs response scenarios for the avatar.
[0593] Step 7:
[0594] The server sends the generated avatar and response scenario to the terminal.
[0595] Step 8:
[0596] The device displays an avatar on the screen, begins performing actions that match the user's eating speed, and the avatar speaks to the user based on a pre-set scenario.
[0597] Step 9:
[0598] The user communicates with the avatar using voice, and the device receives that audio.
[0599] Step 10:
[0600] The terminal converts the user's voice input into text data and sends it to the server.
[0601] Step 11:
[0602] The server analyzes the text data and generates an appropriate response that takes into account the user's reply. In doing so, it also creates a response that reflects the user's emotional state.
[0603] Step 12:
[0604] The server sends a response to the terminal, and the terminal uses speech synthesis to have an avatar speak the content to the user.
[0605] Step 13:
[0606] When the user finishes their meal and ends the session, the app displays a screen requesting feedback.
[0607] Step 14:
[0608] The server collects user feedback and stores it in appropriate data storage to help improve future services.
[0609] (Example 2)
[0610] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0611] Traditionally, eating alone has often been associated with feelings of loneliness, making it difficult to satisfy users' emotional needs. Furthermore, there has been a need for ways to enrich and personalize the general food and beverage consumption experience and provide emotional support.
[0612] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0613] In this invention, the server includes: image recognition means for receiving images of food and beverages acquired by the user through an information processing device and identifying the type of food and beverage from the images; character generation means for generating a virtual character that consumes similar food and beverages based on the identified type of food and beverage; action adjustment means for the generated virtual character to move in accordance with the user's food and beverage consumption rate; emotion analysis means for analyzing the user's voice tone and facial expressions and recognizing the user's emotions; and natural language processing means for the virtual character to engage in natural dialogue with the user in accordance with the analyzed emotions. As a result, even when eating alone, the user can receive emotional support from the virtual character, experience personalized interaction, and reduce feelings of loneliness.
[0614] An "information processing device" is a device that allows users to input digital data and process or communicate that information, and includes smartphones and personal computers.
[0615] "Image recognition means" refers to technologies and algorithms for analyzing received image data and identifying its content and characteristics.
[0616] A "virtual character" is a digital agent created within a digital environment and designed to interact with users.
[0617] "Character generation means" refers to technologies and algorithms for constructing virtual characters based on user information.
[0618] "Motion adjustment means" refers to technology for adjusting the movements of a virtual character to match the user's actions and speed.
[0619] "Emotional analysis means" refers to technology for analyzing and evaluating a user's emotional state from their voice and facial expressions.
[0620] "Natural language processing means" refers to technologies that process information necessary for users and systems to engage in natural language conversations.
[0621] This invention provides a system that allows users to alleviate feelings of loneliness and receive emotional support while eating alone. Specifically, users can utilize this system by using a dedicated application on an information processing device.
[0622] The user takes a picture of their food or drink using their device's camera and sends the image to the server via the application. The server uses image recognition technology to process the received image. This involves using software such as TensorFlow or OpenCV to identify the type of food or drink from the image. Based on this identified information, the server generates a virtual character. 3D modeling software such as Blender can be used for this generation.
[0623] The generated virtual character is adjusted to move in accordance with the user's food and drink consumption rate. The server also captures the user's voice and facial expressions through the terminal's microphone and camera, and analyzes the user's emotional state using emotion analysis tools. This process utilizes PyTorch for voice analysis and Dlib for facial expression analysis.
[0624] Based on these analysis results, the server uses natural language processing technology to generate a response appropriate to the user's emotions. The generated response is then delivered to the user in real time through a virtual character.
[0625] As a concrete example of its use, if a user says "I'm tired today" during a meal, the server can analyze this and a virtual character can respond with encouraging words such as "You've worked hard, let's take a break." In this way, users can enjoy an emotionally supportive dining experience even when dining alone.
[0626] Examples of prompts to input into the generation AI model include, "Please generate a virtual character that can provide emotional support when a user is eating alone," and "Please analyze images of food and drinks and suggest a character that suits the user."
[0627] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0628] Step 1:
[0629] The user launches a dedicated application on the information processing device and takes pictures of food and beverages using the device's camera. The food and beverage images, as input, are sent to the server via the app. Specifically, the user taps a button within the app, activates the camera, captures the image, and uploads it.
[0630] Step 2:
[0631] The server processes the received images using image recognition technology. It performs data analysis on input images of food and beverages using TensorFlow and OpenCV to identify the type and characteristics of the food and beverages. The output provides information about the identified food and beverages. Specifically, the server analyzes the image pixel data and performs matching by comparing it with a known database of food and beverages.
[0632] Step 3:
[0633] The server generates virtual characters based on identified food and beverage information. Using the food and beverage information as input, it generates 3D characters using modeling tools such as Blender. The output is a model of a virtual character associated with the food and beverage. In terms of specific operation, the character's appearance is adjusted according to the selected template, creating synergy with the food and beverage.
[0634] Step 4:
[0635] The generated virtual character's movements are adjusted to match the user's eating speed. The user's eating speed, as input, is analyzed from the user's operation history and video data, and the character's animation is adjusted using a movement adjustment mechanism. The output is character movements tailored to the user. Specifically, the user's movement patterns are analyzed, and the character mimics table manners.
[0636] Step 5:
[0637] The device acquires the user's voice and facial expressions through the microphone and camera and transmits them to the server. It collects the user's voice and video as input data and transfers it to the server in real time. Specifically, the microphone's voice recording function and the camera's video recording function work in conjunction.
[0638] Step 6:
[0639] The server uses this data to perform emotion analysis. It processes audio and facial expression data as input using PyTorch or Dlib to identify the user's emotional state. The output is the result of the user's emotion analysis. Specifically, the server performs audio waveform analysis and facial feature point detection to calculate an emotional index.
[0640] Step 7:
[0641] The server considers the user's emotional state and generates appropriate responses using natural language processing. It utilizes a generative AI model based on the emotional analysis results as input to generate conversational content tailored to the user. The output is a text response to the user. In its specific operation, the server uses an AI-powered conversation generation engine to create voice or text responses through a virtual character.
[0642] (Application Example 2)
[0643] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0644] In modern society, the number of people dining alone is increasing, and the loneliness they experience in such situations is becoming a mental burden. There is a need for ways to provide emotional satisfaction and social connection even when customers visit restaurants alone. In particular, for customers who regularly dine alone, the challenge is to provide an interactive dining experience in physical restaurants that they can enjoy without feeling lonely.
[0645] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0646] In this invention, the server includes an image recognition means for receiving images of food and beverages taken by the user through an information processing device and identifying the type of food and beverage from the image; a virtual character generation means for generating a virtual character representing a similar food and beverage based on the identified type of food and beverage; and an emotion recognition means for analyzing voice data and nonverbal signs obtained from the user and evaluating the user's emotional state. This makes it possible to alleviate feelings of loneliness and provide an enjoyable dining experience to customers who visit a restaurant and dine alone through natural conversation using a virtual character.
[0647] An "information processing device" generally refers to a device such as a computer, tablet, or smartphone that inputs, processes, and outputs information.
[0648] "Food and beverages" refers to food and beverages that have been prepared or supplied for human consumption.
[0649] "Image recognition means" refers to a means that uses computer vision technology to analyze images and recognize specific objects or features from them.
[0650] A "virtual character generation method" is a means of generating a character that mimics the form of a human being in a digital environment based on specific conditions, using computer technology.
[0651] "Operation adjustment means" refers to means for adjusting the appropriate timing and operation according to the user's actions and circumstances.
[0652] An "emotion recognition tool" is a tool that analyzes a user's voice, facial expressions, and actions to evaluate their emotional state.
[0653] A "generative AI model" is an artificial intelligence technology that learns from large amounts of data and generates appropriate responses and content in response to various inputs.
[0654] "Natural language processing methods" refer to methods that utilize technologies for computers to understand, process, and generate human language.
[0655] A "dialogue provision method" refers to a technology or system for providing dialogue through interaction with the user.
[0656] To implement this invention, a terminal acting as an information processing device is first required. The user uses this terminal to take pictures of food and beverages and transmits them to the server. The server, acting as an image recognition means, analyzes the received images via computer vision technology to identify the type of food and beverage. Then, a virtual person generation means generates a virtual person to eat and drink with the user based on the identified food and beverage.
[0657] This virtual character moves in accordance with the user's eating and drinking speed using motion adjustment means. Furthermore, it analyzes the user's voice and facial expressions through emotion recognition means to evaluate the user's emotional state in real time. Natural language processing means incorporating a generative AI model generates appropriate responses according to the user's emotional state, providing natural and emotionally resonant dialogue.
[0658] As a concrete example, consider a customer who visits a restaurant alone and orders a dessert. The customer takes a picture of the dessert with their device's camera and sends it to the server. The server recognizes the dessert using image recognition and generates a virtual person. When the customer says in a cheerful voice, "Today was a good day!", emotion recognition captures the positive tone. Based on this information, the generative AI model generates a response such as, "That's wonderful! What happened?" and continues the conversation.
[0659] The specific technologies used include Google Cloud Vision API for image recognition, Google Cloud Speech-to-Text for speech-to-text conversion, Azure Emotion API for emotion recognition, and AI models such as GPT-3 for natural language generation.
[0660] As an example of a prompt for a generative AI model, in a situation where a customer says, "I want to have a good time today," the following would be used:
[0661] "The customer is speaking in a cheerful tone and says they want to have a good time today. Please bring up some lighthearted topics."
[0662] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0663] Step 1:
[0664] The user uses a device to take pictures of food and drinks. The input is image data from the device's camera. The output is the transmission of the image data to the server. The device prepares the captured images to be sent to the server in the appropriate format.
[0665] Step 2:
[0666] The server receives image data and uses image recognition to identify the type of food or beverage. The input is image data, and the output is the recognized type of food or beverage information. The server uses the Google Cloud Vision API for image recognition, analyzing objects in the image to classify the food or beverage.
[0667] Step 3:
[0668] The server generates a virtual character using a virtual character generation mechanism based on the recognized type of food and beverage. The input is information about the type of food and beverage, and the output is the generated virtual character data. The server designs a virtual character for an appropriate eating and drinking situation and sets visual and behavioral parameters.
[0669] Step 4:
[0670] The user sends voice and facial expressions to the server via their device. Input consists of the user's voice and facial expression data. Output is the transmission of data to the server for emotion analysis. The user's device captures this data in real time and sends it to the server for analysis.
[0671] Step 5:
[0672] The server uses emotion recognition to evaluate the user's emotional state from their voice and facial expressions. The input consists of voice data and facial expression data, and the output is information about the user's emotional state. The server uses the Azure Emotion API to identify emotions from the input data and record the state.
[0673] Step 6:
[0674] The server uses a generative AI model to generate appropriate responses based on the user's emotional state and food / drink consumption. The input is emotional state information and virtual character data, and the output is the generated response message. The server uses a GPT-3 model to construct natural-sounding conversational content based on the prompt sentences.
[0675] Step 7:
[0676] The terminal receives response messages from the server and provides dialogue to the user through a virtual character. The input is the response message from the server, and the output is the dialogue content displayed to the user. The terminal reproduces the obtained dialogue content on the screen as actions of the virtual character and conveys it to the user.
[0677] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0678] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0679] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0680] [Fourth Embodiment]
[0681] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0682] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0683] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0684] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0685] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0686] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0687] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0688] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0689] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0690] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0691] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0692] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0693] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0694] This invention provides an online communal dining service designed to alleviate feelings of loneliness when users eat alone and address the problem of social isolation. This system is primarily implemented using terminals, servers, and communication technologies between them.
[0695] Users launch a dedicated application on their smartphones or PCs. They take pictures of the food they are eating with their device's camera and send them to the server via the application. At this time, users can also set conversation topics and the personality of their avatar character.
[0696] The server uses image recognition technology to analyze the received food images. This identifies the type and characteristics of the food and generates an avatar eating a meal similar to the one the user is eating. Image analysis is performed based on information such as color, shape, and arrangement.
[0697] Once an avatar is generated, the server controls its animation to match the user's eating pace. Furthermore, the server uses natural language processing to create conversation scenarios to support interaction with the user and determines what the avatar should say at the appropriate time.
[0698] To enable natural dialogue, the terminal receives the user's voice input, converts it into text data, and sends it to the server. The server analyzes the conversation based on the text data, generates the avatar's next response, and sends it to the terminal. The terminal uses speech synthesis technology to have the avatar speak that response aloud, which is then played for the user to hear.
[0699] As a concrete example, consider a scenario where a user is eating a salad for dinner. The user uploads an image of the salad, and the server uses image recognition technology to identify the salad and generates an avatar eating it. The avatar then speaks to the user, saying things like, "That salad looks delicious, what dressing are you using?" and responds to the user's answer with something like, "That sounds delicious, I'd like to try it too."
[0700] In this way, users can enjoy meals with their avatars without feeling lonely, even when dining alone. This system aims not only to alleviate feelings of loneliness but also to provide users with an enjoyable conversational experience.
[0701] The following describes the processing flow.
[0702] Step 1:
[0703] Users launch the shared meal app on their device, take photos of their meals, and upload them to the app.
[0704] Step 2:
[0705] The device sends images taken by the user directly to the server. It also simultaneously sends information about the conversation topic and avatar's personality that the user has set.
[0706] Step 3:
[0707] The server executes an image recognition algorithm to analyze the received food images. This process identifies the type and characteristics of the food (e.g., color and shape) from the image.
[0708] Step 4:
[0709] Based on the identified meal information, the server determines an avatar meal similar to the one the user is eating, and then generates the avatar based on that information.
[0710] Step 5:
[0711] The server constructs conversation scenarios based on the user's settings and prepares avatar responses using a natural language processing system. The avatar's movements are configured to adapt to the user's eating speed.
[0712] Step 6:
[0713] The server sends the generated avatar data and the corresponding conversation scenario to the terminal.
[0714] Step 7:
[0715] The device displays an avatar and begins to act in accordance with the user's eating speed. The avatar also speaks to the user based on a conversation scenario received from the server.
[0716] Step 8:
[0717] The user initiates a conversation by responding to a question from the avatar. The device receives the voice response.
[0718] Step 9:
[0719] The terminal converts the user's voice input into text using speech recognition and sends it to the server.
[0720] Step 10:
[0721] The server processes the received text and generates an appropriate response to the user's reply. The response is then sent back to the terminal.
[0722] Step 11:
[0723] The device uses speech synthesis technology to convert responses from the server into speech, which is then presented to the user through an avatar.
[0724] Step 12:
[0725] This conversation continues until the user finishes, at which point the user either closes the app or ends the conversation.
[0726] Step 13:
[0727] The server saves session data, activates the feedback function, and prepares to collect additional input and comments from the user.
[0728] (Example 1)
[0729] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0730] In modern society, changes in individual lifestyles and work styles are leading to an increase in people experiencing feelings of loneliness and social isolation. These feelings are particularly pronounced in situations where people frequently eat alone. There is a need for systems that alleviate the associated psychological burden and provide opportunities for people to enjoy richer social interactions.
[0731] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0732] In this invention, the server includes an image analysis means for receiving images of food taken by the user through an information processing device and identifying the type of food from the images; a virtual character generation means for generating a virtual character that consumes similar food based on the identified type of food; and an action adjustment means for the generated virtual character to move in accordance with the user's eating speed. This makes it possible for the user to enjoy a meal with a virtual character even when eating alone, reducing feelings of loneliness and providing an enjoyable conversational experience.
[0733] An "information processing device" is a device that processes image data and input information captured by a user and has the function of communicating with a server.
[0734] "Image analysis means" refers to technical methods and processes for recognizing objects based on specific features from received image data and analyzing that data.
[0735] A "virtual character generation method" refers to a technology or system that creates an animated character that interacts with the user digitally and mimics their actions, based on analyzed data.
[0736] "Motion adjustment means" refers to technology that adapts the movements and behavior of the generated virtual character in real time to the user's actions and pace.
[0737] "Natural language processing means" refers to computer programs and algorithms that understand the context of user utterances and dialogues and generate appropriate responses.
[0738] A "generative AI model" is a machine learning model used to generate new information or content based on data.
[0739] "Conversation control means" refers to technologies that manage the flow and content of conversations so that virtual characters provide consistent responses.
[0740] This invention is an online communal dining system aimed at making meals more enjoyable and reducing feelings of loneliness for users. Users launch a dedicated application using a smartphone or PC, which is an information processing device. This application has a camera function, allowing users to take pictures of their meals. The captured images are temporarily stored on the device and sent to the server along with the conversation topic selected by the user and the personality settings of their avatar.
[0741] The server uses image analysis techniques to process the received image data. Specifically, it utilizes machine learning models to analyze the types and characteristics of food from the images. For example, a convolutional neural network (CNN) can be used. Based on the analysis results, the server generates a virtual character that consumes food similar to the user's meal. This character is adjusted to operate in real time in accordance with the user's eating pace and movements.
[0742] Interaction between the virtual character and the user is achieved using natural language processing technology. The server creates conversation scenarios using a generated AI model and controls the virtual character to ensure consistent speech. Additionally, voice input from the user is converted into text data by the terminal and sent to the server. Based on this text data, the server generates appropriate responses for the user and uses speech synthesis technology to vocalize what the virtual character would say.
[0743] As a concrete example, consider a scenario where a user is eating a salad for dinner. The user takes a picture of the salad with their smartphone and selects "healthy eating habits" as the conversation topic. The server identifies the salad from the image and generates a virtual character eating the salad. Based on the user's selection, the character starts a conversation such as, "That's a healthy meal, are you trying any special recipes?" and provides dialogue that responds to the user's answers.
[0744] An example of a prompt message is, "Generate a dialogue scenario for a virtual character, assuming the user is eating a healthy salad." In this way, the system alleviates feelings of loneliness and provides the user with an enjoyable dining experience.
[0745] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0746] Step 1:
[0747] The user launches a dedicated application on the information processing device and takes a picture of their meal. The captured image is temporarily stored on the terminal. The input is image data of the meal acquired by the camera, and the output is image data ready to be sent to the server. Specifically, the user operates the camera function using the user interface to capture and save a still image.
[0748] Step 2:
[0749] The device sends captured image data, along with user-defined conversation topics and avatar personality information, to the server. Input consists of user-provided images and settings, while output is a data packet containing this information. Specifically, the data is packetized and transmitted using a secure communication protocol.
[0750] Step 3:
[0751] The server receives the image data and uses image analysis tools to identify the type and characteristics of the food. The input is image data and user configuration information, and the output is information about the type and characteristics of the food. Here, an image analysis algorithm (e.g., a convolutional neural network) is applied to analyze the patterns of color and shape.
[0752] Step 4:
[0753] The server generates a virtual character to consume similar foods based on identified food information. The input is identified food information, and the output is model data for the virtual character. This includes specific actions to create the character's appearance and movement patterns using the generated AI model.
[0754] Step 5:
[0755] The server adjusts the virtual character's movements to match the user's eating pace and generates conversation scenarios using natural language processing. Inputs include user-provided settings and food information, while outputs consist of dynamically controlled animation sequences and conversation scenario text. Specific operations include adjusting the character's movement speed and creating dialogue for appropriate timing.
[0756] Step 6:
[0757] The device receives voice input from the user and converts it into text data. The input is the user's voice data, and the output is text data. Specifically, it uses a microphone to acquire voice and speech recognition software (e.g., a voice-to-text API) to convert it into text.
[0758] Step 7:
[0759] The server generates a virtual character response to the user based on text data and sends it to the terminal. The input is the user's text data, and the output is the virtual character's response text. The specific operation here is the process of generating a context-appropriate response using a generative AI model.
[0760] Step 8:
[0761] The terminal uses speech synthesis technology to convert the received response text into speech, and a virtual character speaks it aloud. The input is response text data, and the output is audio data. Specifically, it uses a text-to-speech engine to convert the text into speech and plays it back through the speaker.
[0762] (Application Example 1)
[0763] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0764] This invention aims to provide a system that reduces feelings of loneliness and makes the dining experience more enjoyable, given the increasing number of people who eat alone in modern society. In particular, it seeks to help users maintain their interest in food and increase social interaction through interaction with virtual characters and virtual store experiences in augmented reality spaces.
[0765] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0766] In this invention, the server includes: an image analysis means for receiving images of meals acquired by the user through an information processing device and identifying the type of meal from the images; a virtual character generation means for generating a virtual character eating a similar meal based on the identified type of meal; an action control means for the generated virtual character to move in accordance with the user's eating pace; and a means for displaying a virtual space on the user's visual device using augmented reality technology and conducting a conversation within that virtual space. This enables the user to enjoy a realistic and interactive dining experience by utilizing the virtual space.
[0767] "Image analysis means" refers to a technology that receives images of meals acquired by a user through an information processing device and has the function of identifying the type of meal from those images.
[0768] "Virtual character generation means" refers to a technology for generating virtual characters that eat similar meals based on a specified type of meal.
[0769] "Motion control means" refers to a technology that adjusts the generated virtual character to move in accordance with the user's eating pace.
[0770] "Natural language processing means" refers to technologies that perform language analysis and generation to enable natural conversations between users and virtual characters.
[0771] Augmented reality technology is a technology that displays a virtual space on the user's visual device and allows them to converse with virtual characters within that virtual space.
[0772] An "information processing device" is an electronic device used by a user to acquire images and process the received data.
[0773] A "visual device" is a device that uses augmented reality technology to display a virtual space to the user.
[0774] A "virtual space" is a digital environment constructed within the user's visual environment using augmented reality technology.
[0775] This invention is implemented using an information processing device and related technologies to improve the experience of users dining alone. The system primarily utilizes image analysis means, virtual character generation means, motion control means, natural language processing means, and augmented reality technology.
[0776] The server receives images of meals taken by users using information processing devices such as smartphones and personal computers. These images undergo data analysis using the Google Cloud Vision API to identify the type of meal. At this stage, features such as color and shape are analyzed, and information on similar meals is retrieved from the database.
[0777] Once the type of meal is identified, the server uses a generative AI model to generate a virtual character common to the user's meals. This virtual character is animated and controlled on platforms such as Unity or Unreal Engine, expressing movements that correspond to the user's eating pace.
[0778] Furthermore, the server generates conversation content through natural language processing technology. In this process, OpenAI's GPT model is utilized to determine the flow of the conversation and the appropriate timing for what to say. Dialogue scenarios with the user are generated based on prompt sentences such as, "This looks delicious! What's the secret ingredient in this dish?"
[0779] Furthermore, augmented reality technology plays a role in facilitating interaction by displaying a virtual space on the user's visual devices (such as smart glasses or head-mounted displays). Virtual characters provide a realistic dining experience, allowing users to enjoy interactive conversations even when alone.
[0780] For example, if a user is trying a new cheese fondue, a virtual character might say, "This cheese has a unique flavor, doesn't it? Do you have any particular preferences when choosing bread?" enriching the dining experience. In this way, an environment is created where users can immerse themselves in their meal without feeling isolated.
[0781] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0782] Step 1:
[0783] Users take pictures of their meals using the camera on their devices, such as smartphones or personal computers. The captured images are sent to a server via a dedicated application on the device. In this input process, the user's identification information is added along with the transmitted image data.
[0784] Step 2:
[0785] The server uses the received meal image data to perform image analysis using the Google Cloud Vision API. As a result of the analysis, data regarding the type and characteristics of the meal is extracted. Specifically, the color, shape, and arrangement of ingredients are analyzed and classified into specific categories. The output is information regarding the identified type of meal.
[0786] Step 3:
[0787] The server generates a virtual character using a generative AI model based on the image analysis results. During the generation process, it utilizes OpenAI's GPT model to generate character information related to the user's meal. This virtual character includes conversation topics and personality settings related to the type of meal. The output consists of the digital model of this virtual character and conversation prompts.
[0788] Step 4:
[0789] The server uses Unity or Unreal Engine to control the generated virtual character's movements. During this process, a sequence of movements is designed based on the user's eating pace, and the virtual character's animations are adjusted accordingly. The input is the user's eating pace information, and the output is the character's movements synchronized with the user's pace.
[0790] Step 5:
[0791] The server generates conversational content for virtual characters using natural language processing. It constructs a concrete conversational flow based on prompts obtained from a generating AI model, taking nuances and timing into consideration. The input is user-generated text data, and the output is the virtual character's response data based on that input.
[0792] Step 6:
[0793] The device displays a virtual space generated using augmented reality technology on the user's visual device. Through smart glasses or a head-mounted display, the user experiences real-time interaction with a virtual character. The output consists of a visual virtual space and synthesized speech from the character.
[0794] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0795] This invention provides an online shared dining service designed to alleviate feelings of loneliness and provide emotional support to users when they are eating alone. This invention achieves more personalized interaction by combining an emotion recognition engine. The system primarily utilizes terminals, servers, and communication technologies between them.
[0796] Users launch a dedicated application on a device such as a smartphone or PC. First, users can take a picture of their meal with their camera and send the image to the server via the app. Furthermore, users can also set conversation topics and the personality of their avatar.
[0797] The server uses advanced image recognition algorithms to analyze food images. This identifies the type and characteristics of the food and generates an avatar that eats with the user. The generated avatar can move in sync with the user's eating speed.
[0798] Furthermore, the system incorporates an emotion recognition engine. The server analyzes the user's voice tone and facial expressions, evaluating the user's emotional state in real time. This information is reflected in the avatar's responses and actions, providing conversations appropriate to the user's emotions. For example, if the user appears depressed, the avatar can continue the conversation with words of encouragement or humor.
[0799] As a concrete example, consider a scenario where a user is eating a salad for dinner. When the user uploads an image of the salad, the server recognizes the salad and generates a suitable avatar. When the user says, "I'm tired today," and the emotion engine detects a downcast tone, the server sends a response such as, "That must have been tough, let's talk about something relaxing together." In this way, users can enjoy a shared dining experience that is empathetic and doesn't make them feel lonely.
[0800] This system aims to create a more fulfilling mealtime experience by alleviating users' feelings of loneliness and providing appropriate responses tailored to their individual emotions.
[0801] The following describes the processing flow.
[0802] Step 1:
[0803] Users launch the shared meal app on their device, take photos of their meal before starting, and upload them to the app.
[0804] Step 2:
[0805] The device sends the captured image of the meal, along with the user's chosen conversation topics and avatar personality, to the server.
[0806] Step 3:
[0807] The server analyzes the transmitted food images using image recognition technology to identify the type and characteristics of the food.
[0808] Step 4:
[0809] The server generates avatars that eat similar meals based on the identified meal type and prepares the avatar data. At this time, it configures the avatars to operate in accordance with the user's eating speed.
[0810] Step 5:
[0811] The emotion recognition engine analyzes the user's voice data and, if necessary, facial expression data from camera footage in real time to determine the user's emotional state.
[0812] Step 6:
[0813] The server generates conversational content suitable for the user based on the emotional state obtained from the emotion recognition engine, and constructs response scenarios for the avatar.
[0814] Step 7:
[0815] The server sends the generated avatar and response scenario to the terminal.
[0816] Step 8:
[0817] The device displays an avatar on the screen, begins performing actions that match the user's eating speed, and the avatar speaks to the user based on a pre-set scenario.
[0818] Step 9:
[0819] The user communicates with the avatar using voice, and the device receives that audio.
[0820] Step 10:
[0821] The terminal converts the user's voice input into text data and sends it to the server.
[0822] Step 11:
[0823] The server analyzes the text data and generates an appropriate response that takes into account the user's reply. In doing so, it also creates a response that reflects the user's emotional state.
[0824] Step 12:
[0825] The server sends a response to the terminal, and the terminal uses speech synthesis to have an avatar speak the content to the user.
[0826] Step 13:
[0827] When the user finishes their meal and ends the session, the app displays a screen requesting feedback.
[0828] Step 14:
[0829] The server collects user feedback and stores it in appropriate data storage to help improve future services.
[0830] (Example 2)
[0831] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0832] Traditionally, eating alone has often been associated with feelings of loneliness, making it difficult to satisfy users' emotional needs. Furthermore, there has been a need for ways to enrich and personalize the general food and beverage consumption experience and provide emotional support.
[0833] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0834] In this invention, the server includes: image recognition means for receiving images of food and beverages acquired by the user through an information processing device and identifying the type of food and beverage from the images; character generation means for generating a virtual character that consumes similar food and beverages based on the identified type of food and beverage; action adjustment means for the generated virtual character to move in accordance with the user's food and beverage consumption rate; emotion analysis means for analyzing the user's voice tone and facial expressions and recognizing the user's emotions; and natural language processing means for the virtual character to engage in natural dialogue with the user in accordance with the analyzed emotions. As a result, even when eating alone, the user can receive emotional support from the virtual character, experience personalized interaction, and reduce feelings of loneliness.
[0835] An "information processing device" is a device that allows users to input digital data and process or communicate that information, and includes smartphones and personal computers.
[0836] "Image recognition means" refers to technologies and algorithms for analyzing received image data and identifying its content and characteristics.
[0837] A "virtual character" is a digital agent created within a digital environment and designed to interact with users.
[0838] "Character generation means" refers to technologies and algorithms for constructing virtual characters based on user information.
[0839] "Motion adjustment means" refers to technology for adjusting the movements of a virtual character to match the user's actions and speed.
[0840] "Emotional analysis means" refers to technology for analyzing and evaluating a user's emotional state from their voice and facial expressions.
[0841] "Natural language processing means" refers to technologies that process information necessary for users and systems to engage in natural language conversations.
[0842] This invention provides a system that allows users to alleviate feelings of loneliness and receive emotional support while eating alone. Specifically, users can utilize this system by using a dedicated application on an information processing device.
[0843] The user takes a picture of their food or drink using their device's camera and sends the image to the server via the application. The server uses image recognition technology to process the received image. This involves using software such as TensorFlow or OpenCV to identify the type of food or drink from the image. Based on this identified information, the server generates a virtual character. 3D modeling software such as Blender can be used for this generation.
[0844] The generated virtual character is adjusted to move in accordance with the user's food and drink consumption rate. The server also captures the user's voice and facial expressions through the terminal's microphone and camera, and analyzes the user's emotional state using emotion analysis tools. This process utilizes PyTorch for voice analysis and Dlib for facial expression analysis.
[0845] Based on these analysis results, the server uses natural language processing technology to generate a response appropriate to the user's emotions. The generated response is then delivered to the user in real time through a virtual character.
[0846] As a concrete example of its use, if a user says "I'm tired today" during a meal, the server can analyze this and a virtual character can respond with encouraging words such as "You've worked hard, let's take a break." In this way, users can enjoy an emotionally supportive dining experience even when dining alone.
[0847] Examples of prompts to input into the generation AI model include, "Please generate a virtual character that can provide emotional support when a user is eating alone," and "Please analyze images of food and drinks and suggest a character that suits the user."
[0848] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0849] Step 1:
[0850] The user launches a dedicated application on the information processing device and takes pictures of food and beverages using the device's camera. The food and beverage images, as input, are sent to the server via the app. Specifically, the user taps a button within the app, activates the camera, captures the image, and uploads it.
[0851] Step 2:
[0852] The server processes the received images using image recognition technology. It performs data analysis on input images of food and beverages using TensorFlow and OpenCV to identify the type and characteristics of the food and beverages. The output provides information about the identified food and beverages. Specifically, the server analyzes the image pixel data and performs matching by comparing it with a known database of food and beverages.
[0853] Step 3:
[0854] The server generates virtual characters based on identified food and beverage information. Using the food and beverage information as input, it generates 3D characters using modeling tools such as Blender. The output is a model of a virtual character associated with the food and beverage. In terms of specific operation, the character's appearance is adjusted according to the selected template, creating synergy with the food and beverage.
[0855] Step 4:
[0856] The generated virtual character's movements are adjusted to match the user's eating speed. The user's eating speed, as input, is analyzed from the user's operation history and video data, and the character's animation is adjusted using a movement adjustment mechanism. The output is character movements tailored to the user. Specifically, the user's movement patterns are analyzed, and the character mimics table manners.
[0857] Step 5:
[0858] The device acquires the user's voice and facial expressions through the microphone and camera and transmits them to the server. It collects the user's voice and video as input data and transfers it to the server in real time. Specifically, the microphone's voice recording function and the camera's video recording function work in conjunction.
[0859] Step 6:
[0860] The server uses this data to perform emotion analysis. It processes audio and facial expression data as input using PyTorch or Dlib to identify the user's emotional state. The output is the result of the user's emotion analysis. Specifically, the server performs audio waveform analysis and facial feature point detection to calculate an emotional index.
[0861] Step 7:
[0862] The server considers the user's emotional state and generates appropriate responses using natural language processing. It utilizes a generative AI model based on the emotional analysis results as input to generate conversational content tailored to the user. The output is a text response to the user. In its specific operation, the server uses an AI-powered conversation generation engine to create voice or text responses through a virtual character.
[0863] (Application Example 2)
[0864] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0865] In modern society, the number of people dining alone is increasing, and the loneliness they experience in such situations is becoming a mental burden. There is a need for ways to provide emotional satisfaction and social connection even when customers visit restaurants alone. In particular, for customers who regularly dine alone, the challenge is to provide an interactive dining experience in physical restaurants that they can enjoy without feeling lonely.
[0866] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0867] In this invention, the server includes an image recognition means for receiving images of food and beverages taken by the user through an information processing device and identifying the type of food and beverage from the image; a virtual character generation means for generating a virtual character representing a similar food and beverage based on the identified type of food and beverage; and an emotion recognition means for analyzing voice data and nonverbal signs obtained from the user and evaluating the user's emotional state. This makes it possible to alleviate feelings of loneliness and provide an enjoyable dining experience to customers who visit a restaurant and dine alone through natural conversation using a virtual character.
[0868] An "information processing device" generally refers to a device such as a computer, tablet, or smartphone that inputs, processes, and outputs information.
[0869] "Food and beverages" refers to food and beverages that have been prepared or supplied for human consumption.
[0870] "Image recognition means" refers to a means that uses computer vision technology to analyze images and recognize specific objects or features from them.
[0871] A "virtual character generation method" is a means of generating a character that mimics the form of a human being in a digital environment based on specific conditions, using computer technology.
[0872] "Operation adjustment means" refers to means for adjusting the appropriate timing and operation according to the user's actions and circumstances.
[0873] An "emotion recognition tool" is a tool that analyzes a user's voice, facial expressions, and actions to evaluate their emotional state.
[0874] A "generative AI model" is an artificial intelligence technology that learns from large amounts of data and generates appropriate responses and content in response to various inputs.
[0875] "Natural language processing methods" refer to methods that utilize technologies for computers to understand, process, and generate human language.
[0876] A "dialogue provision method" refers to a technology or system for providing dialogue through interaction with the user.
[0877] To implement this invention, a terminal acting as an information processing device is first required. The user uses this terminal to take pictures of food and beverages and transmits them to the server. The server, acting as an image recognition means, analyzes the received images via computer vision technology to identify the type of food and beverage. Then, a virtual person generation means generates a virtual person to eat and drink with the user based on the identified food and beverage.
[0878] This virtual character moves in accordance with the user's eating and drinking speed using motion adjustment means. Furthermore, it analyzes the user's voice and facial expressions through emotion recognition means to evaluate the user's emotional state in real time. Natural language processing means incorporating a generative AI model generates appropriate responses according to the user's emotional state, providing natural and emotionally resonant dialogue.
[0879] As a concrete example, consider a customer who visits a restaurant alone and orders a dessert. The customer takes a picture of the dessert with their device's camera and sends it to the server. The server recognizes the dessert using image recognition and generates a virtual person. When the customer says in a cheerful voice, "Today was a good day!", emotion recognition captures the positive tone. Based on this information, the generative AI model generates a response such as, "That's wonderful! What happened?" and continues the conversation.
[0880] The specific technologies used include Google Cloud Vision API for image recognition, Google Cloud Speech-to-Text for speech-to-text conversion, Azure Emotion API for emotion recognition, and AI models such as GPT-3 for natural language generation.
[0881] As an example of a prompt for a generative AI model, in a situation where a customer says, "I want to have a good time today," the following would be used:
[0882] "The customer is speaking in a cheerful tone and says they want to have a good time today. Please bring up some lighthearted topics."
[0883] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0884] Step 1:
[0885] The user uses a device to take pictures of food and drinks. The input is image data from the device's camera. The output is the transmission of the image data to the server. The device prepares the captured images to be sent to the server in the appropriate format.
[0886] Step 2:
[0887] The server receives image data and uses image recognition to identify the type of food or beverage. The input is image data, and the output is the recognized type of food or beverage information. The server uses the Google Cloud Vision API for image recognition, analyzing objects in the image to classify the food or beverage.
[0888] Step 3:
[0889] The server generates a virtual character using a virtual character generation mechanism based on the recognized type of food and beverage. The input is information about the type of food and beverage, and the output is the generated virtual character data. The server designs a virtual character for an appropriate eating and drinking situation and sets visual and behavioral parameters.
[0890] Step 4:
[0891] The user sends voice and facial expressions to the server via their device. Input consists of the user's voice and facial expression data. Output is the transmission of data to the server for emotion analysis. The user's device captures this data in real time and sends it to the server for analysis.
[0892] Step 5:
[0893] The server uses emotion recognition to evaluate the user's emotional state from their voice and facial expressions. The input consists of voice data and facial expression data, and the output is information about the user's emotional state. The server uses the Azure Emotion API to identify emotions from the input data and record the state.
[0894] Step 6:
[0895] The server uses a generative AI model to generate appropriate responses based on the user's emotional state and food / drink consumption. The input is emotional state information and virtual character data, and the output is the generated response message. The server uses a GPT-3 model to construct natural-sounding conversational content based on the prompt sentences.
[0896] Step 7:
[0897] The terminal receives response messages from the server and provides dialogue to the user through a virtual character. The input is the response message from the server, and the output is the dialogue content displayed to the user. The terminal reproduces the obtained dialogue content on the screen as actions of the virtual character and conveys it to the user.
[0898] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0899] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0900] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0901] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0902] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0903] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0904] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0905] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0906] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0907] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0908] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0909] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0910] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0911] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0912] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0913] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0914] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0915] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0916] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0917] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0918] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0919] The following is further disclosed regarding the embodiments described above.
[0920] (Claim 1)
[0921] An image recognition means for receiving images of meals taken by the user through a terminal and identifying the type of meal from the image,
[0922] An avatar generation means that generates avatars that eat similar meals based on the type of meal that was identified,
[0923] A means for adjusting the movement of the generated avatar to match the user's eating speed,
[0924] A natural language processing method for realizing natural dialogue between the user and the avatar,
[0925] A system that includes this.
[0926] (Claim 2)
[0927] The system according to claim 1, characterized in that the avatar converts audio data received from the user into text data, generates an appropriate response based on the text data, and transmits it to the terminal.
[0928] (Claim 3)
[0929] The system according to claim 1, characterized in that the system receives user feedback when the user finishes their meal, and uses that feedback to improve the system.
[0930] "Example 1"
[0931] (Claim 1)
[0932] An image analysis means for receiving images of food taken by a user through an information processing device and identifying the type of food from the image,
[0933] A virtual character generation means that generates a virtual character that consumes similar foods based on the type of food identified,
[0934] A means for adjusting the movement of the generated virtual character to match the user's eating speed,
[0935] A natural language processing method for realizing natural dialogue between a user and a virtual character,
[0936] A conversation control method that generates conversation scenarios using a generative AI model and causes a virtual character to make consistent statements,
[0937] A system that includes this.
[0938] (Claim 2)
[0939] The system according to claim 1, characterized in that the virtual character converts audio data received from the user into text data, generates an appropriate response based on the text data, and transmits it to an information processing device.
[0940] (Claim 3)
[0941] The system according to claim 1, characterized in that the system receives user feedback when the user finishes their meal and uses that feedback to improve the system.
[0942] "Application Example 1"
[0943] (Claim 1)
[0944] An image analysis means for receiving an image of a meal acquired by a user through an information processing device and identifying the type of meal from the image,
[0945] A virtual character generation means that generates a virtual character that eats a similar meal based on a specified type of meal,
[0946] A motion control means for the generated virtual character to move in accordance with the user's eating pace,
[0947] A natural language processing method for realizing natural conversation between a user and a virtual character,
[0948] A means of displaying a virtual space on a user's visual device using augmented reality technology and conducting a conversation within that virtual space,
[0949] A system that includes this.
[0950] (Claim 2)
[0951] The system according to claim 1, characterized in that the virtual character converts audio data received from the user into text data, generates an appropriate response based on the text data, and transmits it to an information processing device.
[0952] (Claim 3)
[0953] The system according to claim 1, characterized in that the system receives user evaluation information when the user finishes their meal activity and uses that evaluation information to improve the system.
[0954] "Example 2 of combining an emotion engine"
[0955] (Claim 1)
[0956] A video recognition means for receiving images of food and beverages acquired by a user through an information processing device and identifying the type of food and beverage from the image,
[0957] A character generation means that generates a virtual character that consumes similar food and beverages based on the type of food and beverage identified,
[0958] A means for adjusting the behavior of the generated virtual character to match the user's food and beverage consumption rate,
[0959] An emotion analysis means that analyzes the user's voice tone and facial expressions to recognize the user's emotions,
[0960] A natural language processing method for enabling a virtual character to have a natural conversation with the user in response to analyzed emotions,
[0961] A system that includes this.
[0962] (Claim 2)
[0963] The system according to claim 1, characterized in that the virtual character converts audio information received from the user into text information, generates an appropriate response based on the text information, and transmits it to an information processing device.
[0964] (Claim 3)
[0965] The system according to claim 1, characterized in that the system receives user feedback when the user finishes their food and beverage consumption activity and uses that feedback to improve the system.
[0966] "Application example 2 of combining emotional engines"
[0967] (Claim 1)
[0968] An image recognition means for receiving images of food and beverages taken by a user through an information processing device and identifying the type of food and beverage from the image,
[0969] A virtual character generation means that generates a virtual character representing a similar food or drink based on the type of food or drink identified,
[0970] A means for adjusting the movement of the generated virtual character to match the user's eating and drinking speed,
[0971] An emotion recognition means that analyzes voice data and nonverbal signs obtained from the user to evaluate the user's emotional state,
[0972] A natural language processing method using a generative AI model for generating appropriate responses based on evaluated emotional states,
[0973] A means of providing dialogue to allow a customer to interact with a virtual dining partner on an electronic device while they are dining alone,
[0974] A system that includes this.
[0975] (Claim 2)
[0976] The system according to claim 1, characterized in that the virtual person converts spoken data received from a user into text information, utilizes a generative AI model that generates an appropriate response based on the text information and evaluated emotions, and transmits it to an information processing device.
[0977] (Claim 3)
[0978] The system according to claim 1, characterized in that the system receives a user's response when the user finishes eating and drinking, and utilizes the response to improve the generated response and the behavior of the virtual character. [Explanation of Symbols]
[0979] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An image recognition means for receiving images of meals taken by the user through a terminal and identifying the type of meal from the image, An avatar generation means that generates avatars that eat similar meals based on the type of meal that was identified, A means for adjusting the movement of the generated avatar to match the user's eating speed, A natural language processing method for realizing natural dialogue between the user and the avatar, A system that includes this.
2. The system according to claim 1, characterized in that the avatar converts audio data received from the user into text data, generates an appropriate response based on the text data, and transmits it to the terminal.
3. The system according to claim 1, characterized in that the system receives user feedback when the user finishes their meal and uses that feedback to improve the system.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A