system
The integration of extended reality and generative AI with a video acquisition device and virtual character addresses user resistance by offering personalized health advice, enhancing social acceptance and promoting a healthy lifestyle through continuous improvement.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-21
- Publication Date
- 2026-05-07
AI Technical Summary
Users experience resistance to using extended reality devices due to insufficient evaluation of practicality, convenience, and clear benefits, hindering social acceptance and daily health promotion.
A system integrating extended reality technology with generative AI, utilizing a video acquisition device, machine learning algorithms, and a virtual character for interactive dialogue to provide health-related advice and continuously improve through user feedback.
Enhances user acceptance and promotes a healthy lifestyle by providing personalized health advice and improving the system's accuracy over time.
Smart Images

Figure 2026074907000001_ABST
Abstract
Description
Technical Field
[0004] , ,
[0005] , , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, devices using extended reality technology have evolved, but there are users who feel resistance when using these devices in daily life. This sense of resistance is due to the fact that the practicality, convenience, or clear benefits to users of the devices have not been sufficiently evaluated. The present invention aims to reduce the resistance to use and enhance social acceptance by integrating extended reality technology and generative AI to provide a system that enables users to naturally achieve a healthy lifestyle while wearing the device.
Means for Solving the Problems
[0005] This invention includes a video acquisition device that acquires environmental data, uses a machine learning algorithm to analyze the data to identify a specific object, and generates health information related to that object. Furthermore, the generated health information is visually presented to the user through a display device, and a virtual character engages in interactive dialogue with the user to provide advice on exercise and diet. This allows the user to naturally take health-promoting actions in their daily life, and the system can continuously improve its machine learning algorithm using the feedback information.
[0006] A "video acquisition device" is a device used to capture environmental data around the user, and mainly includes cameras and sensors.
[0007] "Environmental data" refers to information that digitally represents the physical elements present in the user's surroundings, and can take various forms such as images and audio.
[0008] A "machine learning algorithm" is a general term for algorithms that use methods for computers to learn patterns and rules from data, thereby automatically improving performance.
[0009] "Target" refers to individual objects or phenomena within the environmental data captured by the video acquisition device, which are the objects of identification and analysis.
[0010] "Health information" refers to data and advice related to the user's physical and mental health status, generated based on the identified subject.
[0011] A "display device" is a device used to visually present generated information to a user, and includes AR glasses, among others.
[0012] A "virtual character" is a digital character created for interactive dialogue with users, and it operates using generative AI.
[0013] "Interactive dialogue" refers to two-way communication between a user and a digital system, in which responses are dynamically generated in response to user input.
[0014] "Exercise and dietary advice" refers to specific action suggestions to promote the user's health, and is personalized based on data analyzed by the generating AI.
[0015] "Feedback information" refers to data about user experience and responses provided to the system, which is used for the continuous improvement of the system. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10]Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] The system of this invention is designed to help users promote a healthy lifestyle in their daily lives. This system operates on a platform that utilizes augmented reality technology.
[0038] First, the device uses a video acquisition device to capture data about the user's everyday environment. For example, when the user is eating, it collects images of the food and drinks in the surrounding area in real time. This allows for the acquisition of data based on the user's lifestyle.
[0039] Next, the server analyzes these captured image data. Using machine learning algorithms, it identifies specific objects or foods and generates associated nutritional and calorie information. For example, if the recognized object is "pizza," it identifies and generates its calorie value and nutritional information.
[0040] Next, the device provides the user with generated health information. Through the display, it visually presents the user with calorie information about meals and healthy options. Furthermore, a virtual character appears on the screen, interacts with the user, and provides specific advice regarding exercise and diet. For example, it might suggest, "This pizza is 700 kcal. We recommend a 20-minute jog."
[0041] This system can also receive feedback from users. Users can take action based on the information provided and input the results and their impressions as feedback. The server then uses this feedback data to refine its machine learning algorithms, continuously improving the accuracy and appropriateness of the information it generates.
[0042] For example, when elderly people use the system, advice and action suggestions that are easier to understand are provided through a virtual character, allowing the system to be flexibly customized to suit the user. In this way, the present invention comprehensively supports the user's health and enables natural health promotion in daily life.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The device activates a video acquisition device to capture image data in real time, capturing the environment within the user's field of view. The camera continuously acquires images, recording objects and activities present around the user.
[0046] Step 2:
[0047] The device compresses the captured image data for processing and sends it to a server in the cloud using a secure communication protocol. Encryption technology is used to ensure that data transfer is fast and secure.
[0048] Step 3:
[0049] The server inputs the received image data into a machine learning algorithm to identify specific objects, particularly food and everyday items. This process is carried out by a trained model, which analyzes attributes associated with the identified objects (e.g., calories, ingredients).
[0050] Step 4:
[0051] The server generates health information based on the identified subject and creates a detailed report that includes nutritional advice regarding diet and recommendations regarding exercise. This information also takes into account the user's past behavioral history and health status.
[0052] Step 5:
[0053] The server sends the generated health information to the terminal, which then presents it to the user. The information is displayed visually on an AR display and also conveyed audibly through a virtual character.
[0054] Step 6:
[0055] Based on the health information and advice presented, users make decisions about actions such as diet and exercise. For example, they might approve an exercise plan and record its implementation on their device.
[0056] Step 7:
[0057] The device then collects user feedback and actions taken, and sends them back to the server. This feedback is used to learn from and improve the system, influencing future data processing.
[0058] Step 8:
[0059] The server retrains its machine learning model based on the collected feedback information, improving the accuracy and adaptability of the algorithm. This enhances the quality of the advice provided, enabling it to continue effectively supporting users' health management.
[0060] (Example 1)
[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0062] In modern society, establishing a healthy lifestyle is a crucial issue for many people. However, it is difficult to properly manage one's own health and diet in daily life. In particular, obtaining appropriate information and taking action based on that information is not easy, and specific advice tailored to individual circumstances is needed. To solve this problem, the provision of real-time health information and practical advice based on that information is necessary.
[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0064] In this invention, the server includes means for acquiring information about the user's surroundings using a video acquisition system, means for identifying a specific object using a data analysis algorithm to analyze the acquired information and generate health-related information related to that object, and means for visually presenting the generated health-related information to the user through a display system and providing interaction with the user. This enables the user to intuitively and effectively manage their health in their daily life.
[0065] A "video acquisition system" is a device or technology that can collect information about the user's surroundings, and is a means of acquiring environmental data in real time.
[0066] "Users" refer to individuals who use the system, and are the target audience for the provision and management of health information in their daily lives.
[0067] A "data analysis algorithm" refers to a mathematical processing method used to analyze acquired information and identify specific objects, and is particularly a means of generating health-related information.
[0068] "Health-related information" refers to specific data related to the user's health management, such as calorie values and nutrient information for individual foods and exercises.
[0069] A "display system" is a device or technology for visually communicating generated health-related information to users, and is a means of enabling interaction.
[0070] A "virtual character" is a person or character represented on the screen by computer generation, and it serves to provide users with advice on exercise and nutrition management.
[0071] This invention is a comprehensive system designed to support users' healthy lifestyles. It provides users with specific means to monitor and improve their health in their daily lives.
[0072] First, the device uses a video acquisition system to capture the user's daily living environment. For example, it uses a dedicated augmented reality-enabled device to capture information in real time while the user is eating. This device is equipped with a camera and sensors to collect information about food and drinks that come into the user's field of vision.
[0073] Next, the server analyzes the acquired information using data analysis algorithms such as TENSORFLOW® and PyTorch. This allows it to identify specific foods from image data and generate related health information, such as calorie and nutrient data. A pre-trained food database is used for identification.
[0074] The terminal then uses a display system to visually present the generated health-related information to the user. A virtual character is displayed on the screen along with the generated information, offering advice to the user on specific exercise and nutritional management. This advice is individually customized for the user and may be presented in the form of, for example, "This dessert is 350kcal. If you want to be even healthier, we recommend a 15-minute walk."
[0075] Furthermore, users can provide feedback via voice input or touch interfaces. The server analyzes this feedback information and improves its data analysis algorithms, enabling it to provide more accurate and user-friendly health-related information.
[0076] This invention provides advanced support for users' health management and offers practical guidance for daily life. In this way, users can improve their health more efficiently and effectively.
[0077] Example of a prompt:
[0078] "Please explain how augmented reality technology can be used to provide calorie information about food when a user is eating."
[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0080] Step 1:
[0081] The device uses a video acquisition system to capture the user's daily living environment. Specifically, the device's camera captures food and drinks that come into the user's field of view. The acquired data, as input, is real-time image data. As output, this image data is sent to a server in the cloud.
[0082] Step 2:
[0083] The server acquires image data sent from the terminal and performs analysis using data analysis algorithms. Specifically, it uses TensorFlow and PyTorch to identify objects within the image. This process involves comparing the image data with a pre-built food database to identify specific foods. The input is image data received from the terminal. The output generates health-related information such as calories and nutrients associated with the identified foods.
[0084] Step 3:
[0085] The terminal receives health-related information transmitted from the server and uses a display system to provide it to the user. Specifically, calorie and nutrient information is visually presented on the display screen. Furthermore, a virtual character appears and provides health advice to the user. The input is health-related information sent from the server. The output includes information visually presented to the user and accompanying specific action suggestions.
[0086] Step 4:
[0087] Users act based on the provided health information and input the results and their impressions into the device. Input interfaces include voice input and touchscreens. The input consists of user feedback. As output, this feedback is sent to a server and used for subsequent data analysis.
[0088] Step 5:
[0089] The server receives user feedback and uses it to improve the data analysis algorithm. Specifically, it analyzes the feedback data and uses it to improve the accuracy of the model. This process enables the generated AI model to have higher accuracy and adaptability in future generation of health-related information. The input is user feedback data. The output is a state where the algorithm's accuracy has been improved.
[0090] (Application Example 1)
[0091] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0092] In modern society, a lack of information to help consumers make healthy food choices is a significant challenge. In particular, the restaurant industry often lacks clear nutritional and calorie information for each menu item, making it difficult for consumers to make health-conscious choices. Furthermore, even when health information is available, there is a lack of mechanisms to propose concrete action plans based on that information.
[0093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0094] In this invention, the server includes means for acquiring environmental information around the user using a video acquisition device, means for identifying a specific object using a learning algorithm to analyze the environmental information and generate nutritional information related to that object, and means for presenting the generated nutritional information to the user through a display device and providing interaction with the user. This makes it possible for users to instantly know the nutritional information and calories of each menu item, even when eating out, and to receive specific advice to make choices that contribute to their health.
[0095] A "video acquisition device" is a device that acquires information about the user's surrounding environment in real time using cameras and sensors.
[0096] "Environmental information" refers to data about objects and conditions present in or around the user's location.
[0097] A "learning algorithm" is an artificial intelligence technique used to analyze large amounts of data and identify specific patterns or features.
[0098] "Nutritional information" refers to health-related data such as calories, vitamins, and minerals contained in specific meals or foods.
[0099] A "display device" is a digital display or screen that allows users to visually confirm information.
[0100] "Interaction" refers to the process and reactions in which a user interacts with and interacts with a system.
[0101] "Additional information" refers to supplementary data provided alongside the main information.
[0102] "Visual devices" are devices used to display images or videos to users, and usually refer to glasses-type devices or monitors.
[0103] A "virtual character" is an interactive entity, such as a fictional person or animal, created in a digital environment.
[0104] This invention provides a system to support healthy food choices in the restaurant industry. The server has the function of acquiring images of food in a restaurant when the user wears smart glasses. Specifically, the image acquisition device is a camera mounted on the smart glasses, which captures images of the food in real time. The server uses a learning algorithm operated on the cloud to analyze the acquired image data, perform image recognition, and identify specific dishes and ingredients. Machine learning frameworks such as TensorFlow are applied to the algorithms used.
[0105] The server extracts the relevant nutritional information from the database based on the identification result and displays the generated nutritional information on the smart glasses' display. This display device is the visual device of the smart glasses. A virtual character appears on the screen and provides interaction with the user. The virtual character suggests dishes that contribute to health, proposes exercises, and provides specific advice based on the user's choices.
[0106] For example, when a user points their camera at a hamburger steak, a virtual character will advise, "This hamburger steak is approximately 500 kcal. It would be ideal to eat it with this salad." Another example of a prompt message is, "Based on image analysis, display the nutritional information and calories of the recognized dish and provide healthy advice for that dish." This system allows users to obtain appropriate nutritional information and health-based choices in real time.
[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0108] Step 1:
[0109] The smart glasses, acting as the terminal, use a camera to capture images of food within the user's field of view. The input is real-time video data, and the output is still image data of that food. This image is then prepared to be sent to a server for recognition.
[0110] Step 2:
[0111] The server analyzes the received image data and identifies dishes using machine learning algorithms (e.g., TensorFlow). Here, it takes still image data as input and outputs specific dish names and ingredient information through data calculations. This identified data is used for accessing the server's internal database.
[0112] Step 3:
[0113] The server extracts nutritional information related to identified dishes from a database. The input is the dish identification data, and the output is the nutritional and calorie information for that dish. Database queries are used to collect relevant health information and compile it in a user-friendly format.
[0114] Step 4:
[0115] The smart glasses, acting as the terminal, generate a virtual character along with the acquired nutritional information and visually present the information on the display. The input is nutritional information and a data format for display, while the output is the visual presentation to the user. Here, the character makes healthy suggestions to the user.
[0116] Step 5:
[0117] Users make meal choices based on the provided nutritional information and advice from virtual characters. They can input feedback based on their choices and experiences into the terminal. The input is user feedback data, and the output is feedback data sent to the server for future analysis.
[0118] Step 6:
[0119] The server incorporates user feedback into its machine learning model, continuously updating and optimizing the algorithm. Here, user feedback is used as input, and data calculations such as algorithm parameter tuning are performed to improve the accuracy of subsequent recognitions. The output is the updated learning model, which improves the overall system performance.
[0120] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0121] The system of this invention is designed to support users' daily lives and promote health management. This system operates on a platform that integrates augmented reality technology and emotion recognition technology.
[0122] First, the device uses a video acquisition device to capture the environment within the user's field of view, acquiring image data in real time. This data includes the user's surroundings and facial expressions. Furthermore, the device also acquires the user's voice using a microphone, preparing it for analysis by the emotion engine.
[0123] Next, the server receives this data and performs analysis using machine learning algorithms. Through the analysis of image data, specific objects (e.g., food) are identified, and related health information (e.g., calories and ingredients) is generated. Simultaneously, the emotion engine analyzes the user's facial expressions and voice to identify their emotional state. This enables the generation of health information and advice that takes the user's current emotions into account.
[0124] The generated information is presented to the user by the terminal. Through the display device, the user not only sees health information and lifestyle advice visually, but also engages in interactive dialogue using a virtual character. The character changes its response based on the user's emotional state; for example, if the user is tired, it will offer encouraging advice in a gentle tone.
[0125] Furthermore, users can use this information to make choices about their diet and exercise. The selected actions are input into the system as feedback. The device then sends this feedback information back to the server, where it is used to refine the emotion engine and machine learning algorithms. This process allows the system to continuously provide personalized health support that is more adapted to the user.
[0126] As a concrete example, consider a scenario where a user is choosing lunch. In this case, the system recognizes from the user's facial expression that they are relaxed and offers healthy and nutritious options. If the user responds to the choice with "That's a good choice," the system learns from the positive response and strengthens its advice position for future occasions. In total, this invention supports comprehensive health management that takes the user's emotions into account, making it easier to make healthy choices in daily life.
[0127] The following describes the processing flow.
[0128] Step 1:
[0129] The device uses a video acquisition device and microphone to capture the user's surroundings as well as the user's face and voice in real time. By capturing images of the food the user is looking at and what they are saying, it simultaneously collects environmental data and audio data.
[0130] Step 2:
[0131] The device compresses and encrypts the captured environmental and audio data before sending it to a server in the cloud. Communication protocols are used to ensure that the data is transmitted securely and efficiently.
[0132] Step 3:
[0133] The server uses the received data to execute an image recognition algorithm and identify specific objects (such as food) from the environmental data. It then generates health information, such as calorie and nutritional information, for the identified objects.
[0134] Step 4:
[0135] The server utilizes an emotion engine to analyze voice and facial expression data. It identifies the user's emotional state (for example, whether they are feeling stressed) and determines how this will influence the health information generated based on that state.
[0136] Step 5:
[0137] The server optimizes health advice and creates customized information based on identified objects and the user's emotional state. For example, if the user is tired, it will suggest easy-to-prepare healthy meals.
[0138] Step 6:
[0139] The terminal presents health information and advice from the server to the user visually and audibly through a display device. A virtual character offers suggestions for diet and exercise in a tone that matches the user's emotions.
[0140] Step 7:
[0141] Users act based on the advice they receive. For example, they might decide to cook a dish using a suggested recipe and provide feedback by entering the result into their device.
[0142] Step 8:
[0143] The device collects user feedback and forwards it back to the server. The server uses this feedback to update its machine learning model and sentiment engine, improving the accuracy of the information. This will enable the provision of more appropriate health management advice in the future.
[0144] (Example 2)
[0145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0146] In modern society, providing personalized health management and support for daily life is becoming increasingly important. However, conventional health management systems often fail to adequately consider users' emotions and real-time circumstances, thus failing to maximize the user experience. Therefore, there is a need to realize a system that provides dynamically customized, individualized health information and advice in response to users' emotions and circumstances, thereby increasing user satisfaction and effectiveness.
[0147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0148] In this invention, the server includes a visual acquisition means that operates to acquire objects and facial expressions around the user; an analysis means that analyzes the acquired visual and audio data and uses machine learning means to identify the object and the user's emotional state; and a visual display means that presents health information and advice that takes into account the user's emotions, based on the analysis results. This enables highly accurate health support tailored to the user's individual circumstances and emotions.
[0149] "Visual acquisition means" refers to devices and sensors used to capture the user's surrounding environment and facial expressions, thereby enabling the acquisition of real-time visual data.
[0150] "Machine learning methods" refer to algorithms and their execution environments used to analyze data obtained from users and identify the objects contained in that data or the emotional state of the user.
[0151] "Analysis means" refers to functions that process acquired visual and audio data to identify objects and analyze user emotions.
[0152] "Visual display means" refers to devices or systems that present information generated based on analysis results to the user, thereby enabling effective information sharing with the user.
[0153] A "virtual dialogue target" refers to a digital character or virtual agent created to facilitate interaction with the user and provide emotion-based advice.
[0154] The system of the present invention aims to manage the user's health and support their daily life, and is implemented using a device that integrates augmented reality technology and emotion recognition technology. First, the terminal uses a camera function as a visual acquisition means to capture the environment within the user's field of view. This camera acquires the environment around the user and the user's own facial expressions in real time. In addition, a microphone is used to collect data on the tone of the user's voice, which indicates their speech and emotions.
[0155] Next, the server receives visual and audio data sent from the terminal. In this scenario, the server has machine learning-based analysis capabilities, using libraries such as TensorFlow to identify specific objects from image data and determine the user's emotions from audio data. Based on the analysis results, the server generates health information and advice tailored to the user's current situation.
[0156] The generated information is communicated to the user through the device's visual display. Specifically, the AR glasses display is used to visually present the information to the user, and a virtual dialogue partner engages in an interactive conversation with the user. The dialogue partner adopts an approach tailored to the user's emotional state; for example, it might offer healthy eating advice in a gentle voice to a relaxed user.
[0157] As a concrete example, let's consider a scenario where a user is choosing lunch. In this case, the system determines from the user's facial expressions and voice that they are relaxed and provides a meal option with appropriate calories. When the user provides feedback such as "That's a good choice," the server learns from this positive response and improves the accuracy of future advice.
[0158] An example of a prompt message is, "When a user is choosing lunch and has a relaxed expression, what healthy meal options should be suggested?" In this way, the present invention can provide users with individually tailored health support and facilitate healthy choices in their daily lives.
[0159] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0160] Step 1:
[0161] The device uses a camera to capture the user's surroundings and acquire visual data. The input is the camera's video feed, and by extracting data on surrounding objects and facial expressions from this video feed, image data is obtained as output. This includes specific objects within the user's field of vision and the user's facial expressions.
[0162] Step 2:
[0163] The device uses a microphone to capture ambient sounds and collects audio data. The input is sound from the environment, and the output is audio data obtained by digitally converting that sound. This audio data allows the user's speech and voice tone to be obtained.
[0164] Step 3:
[0165] The terminal transmits the acquired image and audio data to the server. The input is the image and audio data collected in the previous step, which are transferred via the wireless network. As output, the data reaches the server.
[0166] Step 4:
[0167] The server analyzes received image data and uses machine learning algorithms to identify specific objects. The input is image data, and the analysis generates object information and related health information as output. For example, it can recognize food and calculate its calorie information.
[0168] Step 5:
[0169] The server uses an emotion recognition engine to analyze voice data and determine the user's emotional state. The input is voice data, which is analyzed to identify the user's emotional state (e.g., relaxed, stressed). Based on this output, the way information is provided changes.
[0170] Step 6:
[0171] The server generates health information and advice based on the analysis results. The input consists of object information and emotional states, which are combined to generate information best suited to the user's situation. The output is customized health information presented to the user.
[0172] Step 7:
[0173] The device presents information received from the server to the user through a visual display. The input is health information and advice sent from the server, and by displaying this information on the AR glasses' screen, the output is visualized data delivered to the user.
[0174] Step 8:
[0175] The user selects an action based on the information presented and inputs the result as feedback into the terminal. The input consists of the user's actions and comments, and this feedback is sent to the system as output.
[0176] Step 9:
[0177] The terminal sends user feedback back to the server, which is then used to refine the machine learning algorithm. The input is user feedback information, which the server uses to generate output that improves the system's accuracy.
[0178] (Application Example 2)
[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0180] In recent years, the importance of optimizing health and daily life choices and providing personalized information tailored to each user has increased. However, conventional systems often failed to consider the user's emotional state and were limited to providing general information. Furthermore, mechanisms for optimizing the system using user feedback were insufficient, making it difficult to provide advice appropriate to individual users.
[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0182] In this invention, the server includes means for capturing environmental data around the user using a video acquisition device, means for identifying specific objects and generating health information using a machine learning algorithm to analyze the environmental data, and means for presenting the generated health information or product information in accordance with the user's psychological state based on their emotional state. This makes it possible to provide optimal advice and information in real time that is tailored to the user's emotions and environment.
[0183] A "video acquisition device" is a device that captures environmental data around the user.
[0184] "Environmental data" refers to data that includes various information that can be perceived within the user's field of vision.
[0185] A "machine learning algorithm" is a mathematical procedure used to analyze data and identify specific objects.
[0186] "Health information" refers to health-related data associated with a specific subject.
[0187] A "display device" is a device that visually presents generated information to the user.
[0188] "Means of providing interaction" refers to methods that enable two-way communication between the user and the system.
[0189] "Voice input for recognizing the user's emotional state" refers to a method of understanding a user's psychological state through the voice they produce.
[0190] "Feedback information" refers to data about responses and reactions received from users.
[0191] A "virtual character" is a digitized entity created by a computer and used to interact with users.
[0192] "Product information" refers to data about attributes and characteristics related to a product.
[0193] The system designed to realize this application aims to improve the in-store shopping experience and health management for users of smart glasses. The server, terminal, and user work together as follows:
[0194] First, a camera built into the smart glasses captures the user's field of view in real time, and this video data is sent to a server via the device. The server uses machine learning libraries such as TensorFlow and OpenCV to identify objects in the video and extract related health and product information. Furthermore, it analyzes the user's psychological state using voice input and generates data corresponding to the user's emotions. At this time, the analyzed video data and emotion data are processed integrally by an emotion recognition engine on the same server.
[0195] Information about products the user sees, including details and potential health effects, is visually displayed on the user's smart glasses using augmented reality technology. The system also generates virtual characters that interact with the user, offering advice and recommendations based on their emotional state. For example, if the system analyzes the user as relaxed, the character will offer more friendly and proactive recommendations.
[0196] The generated information is continuously improved by machine learning algorithms on the server based on user feedback, and optimized to provide more appropriate and advanced information in the future. Furthermore, the feedback data is input into an emotion recognition engine, enabling the provision of information that is even more tailored to the user's emotions.
[0197] For example, if a user picks up a snack at a supermarket, the smart glasses will display information about the health effects of that snack and alternative options. When the user comments on it, the server immediately uses that feedback information to learn from the system.
[0198] Examples of prompt statements include:
[0199] "Provide health information about the products that users see."
[0200] "Propose products that align with the user's emotions."
[0201] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0202] Step 1:
[0203] The device uses a video acquisition device to capture the user's field of view. It acquires video data from the user's field of view as input and sends it to the server as output. Specifically, the smart glasses' camera collects data in real time and prepares it to be sent to the server in its original format.
[0204] Step 2:
[0205] The server analyzes received video data using machine learning algorithms. It receives video data sent from a terminal as input and generates a list of identified objects and related health or product information as output. This process utilizes TensorFlow and OpenCV to recognize objects in the video and retrieve relevant information from a database.
[0206] Step 3:
[0207] The server analyzes the user's voice input using an emotion recognition engine. It receives voice data sent from the terminal as input and identifies the user's emotional state as output. Here, the emotion recognition engine extracts voice features, analyzes them with a generative AI model, and determines the current emotion.
[0208] Step 4:
[0209] The server integrates the video analysis results and emotion analysis results to generate information to be visually presented to the user. It uses the outputs from steps 2 and 3 as input and sends information display instructions to the terminal as output. In this step, the server selects the most appropriate advice or product suggestions based on the information obtained from both analyses.
[0210] Step 5:
[0211] The device presents comprehensive information to the user through a display device. It receives information display instructions from a server as input and provides health information, product options, and interactive advice from a virtual character on the smart glasses' display as output. Specifically, it uses augmented reality technology to overlay information onto the user's field of view.
[0212] Step 6:
[0213] The user provides feedback to the system. As input, the user provides feedback through their device, including selected information and voice comments, which are then returned to the server. As output, this feedback data exists on the server and is used for subsequent analyses.
[0214] Step 7:
[0215] The server optimizes the machine learning algorithm based on feedback information. It receives user feedback data as input and generates an updated model of the algorithm as output. This enables the server to provide more adaptive and accurate information to users in the future.
[0216] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0217] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0219] [Second Embodiment]
[0220] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0221] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0223] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0227] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0228] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0229] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0230] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0231] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0232] The system of this invention is designed to help users promote a healthy lifestyle in their daily lives. This system operates on a platform that utilizes augmented reality technology.
[0233] First, the device uses a video acquisition device to capture data about the user's everyday environment. For example, when the user is eating, it collects images of the food and drinks in the surrounding area in real time. This allows for the acquisition of data based on the user's lifestyle.
[0234] Next, the server analyzes these captured image data. Using machine learning algorithms, it identifies specific objects or foods and generates associated nutritional and calorie information. For example, if the recognized object is "pizza," it identifies and generates its calorie value and nutritional information.
[0235] Next, the device provides the user with generated health information. Through the display, it visually presents the user with calorie information about meals and healthy options. Furthermore, a virtual character appears on the screen, interacts with the user, and provides specific advice regarding exercise and diet. For example, it might suggest, "This pizza is 700 kcal. We recommend a 20-minute jog."
[0236] This system can also receive feedback from users. Users can take action based on the information provided and input the results and their impressions as feedback. The server then uses this feedback data to refine its machine learning algorithms, continuously improving the accuracy and appropriateness of the information it generates.
[0237] For example, when elderly people use the system, advice and action suggestions that are easier to understand are provided through a virtual character, allowing the system to be flexibly customized to suit the user. In this way, the present invention comprehensively supports the user's health and enables natural health promotion in daily life.
[0238] The following describes the processing flow.
[0239] Step 1:
[0240] The device activates a video acquisition device to capture image data in real time, capturing the environment within the user's field of view. The camera continuously acquires images, recording objects and activities present around the user.
[0241] Step 2:
[0242] The device compresses the captured image data for processing and sends it to a server in the cloud using a secure communication protocol. Encryption technology is used to ensure that data transfer is fast and secure.
[0243] Step 3:
[0244] The server inputs the received image data into a machine learning algorithm to identify specific objects, particularly food and everyday items. This process is carried out by a trained model, which analyzes attributes associated with the identified objects (e.g., calories, ingredients).
[0245] Step 4:
[0246] The server generates health information based on the identified subject and creates a detailed report that includes nutritional advice regarding diet and recommendations regarding exercise. This information also takes into account the user's past behavioral history and health status.
[0247] Step 5:
[0248] The server sends the generated health information to the terminal, which then presents it to the user. The information is displayed visually on an AR display and also conveyed audibly through a virtual character.
[0249] Step 6:
[0250] Based on the health information and advice presented, users make decisions about actions such as diet and exercise. For example, they might approve an exercise plan and record its implementation on their device.
[0251] Step 7:
[0252] The device then collects user feedback and actions taken, and sends them back to the server. This feedback is used to learn from and improve the system, influencing future data processing.
[0253] Step 8:
[0254] The server retrains its machine learning model based on the collected feedback information, improving the accuracy and adaptability of the algorithm. This enhances the quality of the advice provided, enabling it to continue effectively supporting users' health management.
[0255] (Example 1)
[0256] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0257] In modern society, establishing a healthy lifestyle is a crucial issue for many people. However, it is difficult to properly manage one's own health and diet in daily life. In particular, obtaining appropriate information and taking action based on that information is not easy, and specific advice tailored to individual circumstances is needed. To solve this problem, the provision of real-time health information and practical advice based on that information is necessary.
[0258] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0259] In this invention, the server includes means for acquiring information about the user's surroundings using a video acquisition system, means for identifying a specific object using a data analysis algorithm to analyze the acquired information and generate health-related information related to that object, and means for visually presenting the generated health-related information to the user through a display system and providing interaction with the user. This enables the user to intuitively and effectively manage their health in their daily life.
[0260] A "video acquisition system" is a device or technology that can collect information about the user's surroundings, and is a means of acquiring environmental data in real time.
[0261] "Users" refer to individuals who use the system, and are the target audience for the provision and management of health information in their daily lives.
[0262] A "data analysis algorithm" refers to a mathematical processing method used to analyze acquired information and identify specific objects, and is particularly a means of generating health-related information.
[0263] "Health-related information" refers to specific data related to the user's health management, such as calorie values and nutrient information for individual foods and exercises.
[0264] A "display system" is a device or technology for visually communicating generated health-related information to users, and is a means of enabling interaction.
[0265] A "virtual character" is a person or character represented on the screen by computer generation, and it serves to provide users with advice on exercise and nutrition management.
[0266] This invention is a comprehensive system designed to support users' healthy lifestyles. It provides users with specific means to monitor and improve their health in their daily lives.
[0267] First, the device uses a video acquisition system to capture the user's daily living environment. For example, it uses a dedicated augmented reality-enabled device to capture information in real time while the user is eating. This device is equipped with a camera and sensors to collect information about food and drinks that come into the user's field of vision.
[0268] Next, the server analyzes the acquired information using data analysis algorithms such as TensorFlow and PyTorch. This allows it to identify specific foods from image data and generate related health information, such as calorie and nutrient data. A pre-trained food database is used for identification.
[0269] The terminal then uses a display system to visually present the generated health-related information to the user. A virtual character is displayed on the screen along with the generated information, offering advice to the user on specific exercise and nutritional management. This advice is individually customized for the user and may be presented in the form of, for example, "This dessert is 350kcal. If you want to be even healthier, we recommend a 15-minute walk."
[0270] Furthermore, users can provide feedback via voice input or touch interfaces. The server analyzes this feedback information and improves its data analysis algorithms, enabling it to provide more accurate and user-friendly health-related information.
[0271] This invention provides advanced support for users' health management and offers practical guidance for daily life. In this way, users can improve their health more efficiently and effectively.
[0272] Example of a prompt:
[0273] "Please explain how augmented reality technology can be used to provide calorie information about food when a user is eating."
[0274] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0275] Step 1:
[0276] The device uses a video acquisition system to capture the user's daily living environment. Specifically, the device's camera captures food and drinks that come into the user's field of view. The acquired data, as input, is real-time image data. As output, this image data is sent to a server in the cloud.
[0277] Step 2:
[0278] The server acquires image data sent from the terminal and performs analysis using data analysis algorithms. Specifically, it uses TensorFlow and PyTorch to identify objects within the image. This process involves comparing the image data with a pre-built food database to identify specific foods. The input is image data received from the terminal. The output generates health-related information such as calories and nutrients associated with the identified foods.
[0279] Step 3:
[0280] The terminal receives health-related information transmitted from the server and uses a display system to provide it to the user. Specifically, calorie and nutrient information is visually presented on the display screen. Furthermore, a virtual character appears and provides health advice to the user. The input is health-related information sent from the server. The output includes information visually presented to the user and accompanying specific action suggestions.
[0281] Step 4:
[0282] Users act based on the provided health information and input the results and their impressions into the device. Input interfaces include voice input and touchscreens. The input consists of user feedback. As output, this feedback is sent to a server and used for subsequent data analysis.
[0283] Step 5:
[0284] The server receives feedback from users and uses it to improve data analysis algorithms. As a specific operation, it analyzes the feedback data and utilizes it to improve the accuracy of the model. Through this process, the generative AI model can be made more accurate and adaptable in generating future health-related information. The input is the user's feedback data, and the output is the state where the accuracy of the algorithm is improved.
[0285] (Application Example 1)
[0286] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as "server", and the smart glasses 214 are referred to as "terminal".
[0287] In modern society, it is an issue that there is insufficient information for making healthy dietary choices. Especially in the food service industry, the nutritional information and calories of each menu are unclear, making it difficult for consumers to make choices that contribute to their health. Also, even if health information is obtained, there is a lack of a mechanism to propose a specific action plan according to it.
[0288] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0289] In this invention, the server includes means for acquiring environmental information around the user using an image acquisition device, means for identifying a specific object using a learning algorithm for analyzing the environmental information and generating nutritional information related to the object, and means for presenting the generated nutritional information to the user through a display device and providing interaction with the user. Thereby, even when the user dines out, it is possible to immediately know the nutritional information and calories of each menu and receive specific advice for making health-beneficial choices.
[0290] The "image acquisition device" is a device that acquires environmental information around the user in real time using a camera or a sensor.
[0291] "Environmental information" refers to data about objects and conditions present in or around the user's location.
[0292] A "learning algorithm" is an artificial intelligence technique used to analyze large amounts of data and identify specific patterns or features.
[0293] "Nutritional information" refers to health-related data such as calories, vitamins, and minerals contained in specific meals or foods.
[0294] A "display device" is a digital display or screen that allows users to visually confirm information.
[0295] "Interaction" refers to the process and reactions in which a user interacts with and interacts with a system.
[0296] "Additional information" refers to supplementary data provided alongside the main information.
[0297] "Visual devices" are devices used to display images or videos to users, and usually refer to glasses-type devices or monitors.
[0298] A "virtual character" is an interactive entity, such as a fictional person or animal, created in a digital environment.
[0299] This invention provides a system to support healthy food choices in the restaurant industry. The server has the function of acquiring images of food in a restaurant when the user wears smart glasses. Specifically, the image acquisition device is a camera mounted on the smart glasses, which captures images of the food in real time. The server uses a learning algorithm operated on the cloud to analyze the acquired image data, perform image recognition, and identify specific dishes and ingredients. Machine learning frameworks such as TensorFlow are applied to the algorithms used.
[0300] The server extracts the relevant nutritional information from the database based on the identification result and displays the generated nutritional information on the smart glasses' display. This display device is the visual device of the smart glasses. A virtual character appears on the screen and provides interaction with the user. The virtual character suggests dishes that contribute to health, proposes exercises, and provides specific advice based on the user's choices.
[0301] For example, when a user points their camera at a hamburger steak, a virtual character will advise, "This hamburger steak is approximately 500 kcal. It would be ideal to eat it with this salad." Another example of a prompt message is, "Based on image analysis, display the nutritional information and calories of the recognized dish and provide healthy advice for that dish." This system allows users to obtain appropriate nutritional information and health-based choices in real time.
[0302] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0303] Step 1:
[0304] The smart glasses, acting as the terminal, use a camera to capture images of food within the user's field of view. The input is real-time video data, and the output is still image data of that food. This image is then prepared to be sent to a server for recognition.
[0305] Step 2:
[0306] The server analyzes the received image data and identifies dishes using machine learning algorithms (e.g., TensorFlow). Here, it takes still image data as input and outputs specific dish names and ingredient information through data calculations. This identified data is used for accessing the server's internal database.
[0307] Step 3:
[0308] The server extracts the nutritional information related to the identified dish from the database. The input is the identification data of the dish, and the output is the nutritional information and calorie information of that dish. Using a database query, collect the relevant health information and summarize it in a user-friendly format.
[0309] Step 4:
[0310] The smart glasses, which are the terminal, generate a virtual character together with the acquired nutritional information and visually present the information on the display. The input is the nutritional information and the data format for display, and the output is the visual presentation content for the user. Here, the character makes healthy suggestions to the user together.
[0311] Step 5:
[0312] The user makes a meal choice based on the provided nutritional information and advice from the virtual character. Feedback based on the user's choices and experiences can be input into the terminal. The input is the user's feedback data, and the output is the feedback data sent to the server for future analysis.
[0313] Step 6:
[0314] The server incorporates the user's feedback into the machine learning model and continuously updates and optimizes the algorithm. Here, the user's feedback is used as the input, and through data operations such as adjusting the parameters of the algorithm, the goal is to improve the identification accuracy hereafter. The output is the updated learning model, which improves the performance of the entire system.
[0315] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform specific processing using the user's emotion.
[0316] The system of this invention is designed to support users' daily lives and promote health management. This system operates on a platform that integrates augmented reality technology and emotion recognition technology.
[0317] First, the device uses a video acquisition device to capture the environment within the user's field of view, acquiring image data in real time. This data includes the user's surroundings and facial expressions. Furthermore, the device also acquires the user's voice using a microphone, preparing it for analysis by the emotion engine.
[0318] Next, the server receives this data and performs analysis using machine learning algorithms. Through the analysis of image data, specific objects (e.g., food) are identified, and related health information (e.g., calories and ingredients) is generated. Simultaneously, the emotion engine analyzes the user's facial expressions and voice to identify their emotional state. This enables the generation of health information and advice that takes the user's current emotions into account.
[0319] The generated information is presented to the user by the terminal. Through the display device, the user not only sees health information and lifestyle advice visually, but also engages in interactive dialogue using a virtual character. The character changes its response based on the user's emotional state; for example, if the user is tired, it will offer encouraging advice in a gentle tone.
[0320] Furthermore, users can use this information to make choices about their diet and exercise. The selected actions are input into the system as feedback. The device then sends this feedback information back to the server, where it is used to refine the emotion engine and machine learning algorithms. This process allows the system to continuously provide personalized health support that is more adapted to the user.
[0321] As a concrete example, consider a scenario where a user is choosing lunch. In this case, the system recognizes from the user's facial expression that they are relaxed and offers healthy and nutritious options. If the user responds to the choice with "That's a good choice," the system learns from the positive response and strengthens its advice position for future occasions. In total, this invention supports comprehensive health management that takes the user's emotions into account, making it easier to make healthy choices in daily life.
[0322] The following describes the processing flow.
[0323] Step 1:
[0324] The device uses a video acquisition device and microphone to capture the user's surroundings as well as the user's face and voice in real time. By capturing images of the food the user is looking at and what they are saying, it simultaneously collects environmental data and audio data.
[0325] Step 2:
[0326] The device compresses and encrypts the captured environmental and audio data before sending it to a server in the cloud. Communication protocols are used to ensure that the data is transmitted securely and efficiently.
[0327] Step 3:
[0328] The server uses the received data to execute an image recognition algorithm and identify specific objects (such as food) from the environmental data. It then generates health information, such as calorie and nutritional information, for the identified objects.
[0329] Step 4:
[0330] The server utilizes an emotion engine to analyze voice and facial expression data. It identifies the user's emotional state (for example, whether they are feeling stressed) and determines how this will influence the health information generated based on that state.
[0331] Step 5:
[0332] The server optimizes health advice and creates customized information based on identified objects and the user's emotional state. For example, if the user is tired, it will suggest easy-to-prepare healthy meals.
[0333] Step 6:
[0334] The terminal presents health information and advice from the server to the user visually and audibly through a display device. A virtual character offers suggestions for diet and exercise in a tone that matches the user's emotions.
[0335] Step 7:
[0336] Users act based on the advice they receive. For example, they might decide to cook a dish using a suggested recipe and provide feedback by entering the result into their device.
[0337] Step 8:
[0338] The device collects user feedback and forwards it back to the server. The server uses this feedback to update its machine learning model and sentiment engine, improving the accuracy of the information. This will enable the provision of more appropriate health management advice in the future.
[0339] (Example 2)
[0340] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0341] In modern society, providing personalized health management and support for daily life is becoming increasingly important. However, conventional health management systems often fail to adequately consider users' emotions and real-time circumstances, thus failing to maximize the user experience. Therefore, there is a need to realize a system that provides dynamically customized, individualized health information and advice in response to users' emotions and circumstances, thereby increasing user satisfaction and effectiveness.
[0342] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0343] In this invention, the server includes a visual acquisition means that operates to acquire objects and facial expressions around the user; an analysis means that analyzes the acquired visual and audio data and uses machine learning means to identify the object and the user's emotional state; and a visual display means that presents health information and advice that takes into account the user's emotions, based on the analysis results. This enables highly accurate health support tailored to the user's individual circumstances and emotions.
[0344] "Visual acquisition means" refers to devices and sensors used to capture the user's surrounding environment and facial expressions, thereby enabling the acquisition of real-time visual data.
[0345] "Machine learning methods" refer to algorithms and their execution environments used to analyze data obtained from users and identify the objects contained in that data or the emotional state of the user.
[0346] "Analysis means" refers to functions that process acquired visual and audio data to identify objects and analyze user emotions.
[0347] "Visual display means" refers to devices or systems that present information generated based on analysis results to the user, thereby enabling effective information sharing with the user.
[0348] A "virtual dialogue target" refers to a digital character or virtual agent created to facilitate interaction with the user and provide emotion-based advice.
[0349] The system of the present invention aims to manage the user's health and support their daily life, and is implemented using a device that integrates augmented reality technology and emotion recognition technology. First, the terminal uses a camera function as a visual acquisition means to capture the environment within the user's field of view. This camera acquires the environment around the user and the user's own facial expressions in real time. In addition, a microphone is used to collect data on the tone of the user's voice, which indicates their speech and emotions.
[0350] Next, the server receives visual and audio data sent from the terminal. In this scenario, the server has machine learning-based analysis capabilities, using libraries such as TensorFlow to identify specific objects from image data and determine the user's emotions from audio data. Based on the analysis results, the server generates health information and advice tailored to the user's current situation.
[0351] The generated information is communicated to the user through the device's visual display. Specifically, the AR glasses display is used to visually present the information to the user, and a virtual dialogue partner engages in an interactive conversation with the user. The dialogue partner adopts an approach tailored to the user's emotional state; for example, it might offer healthy eating advice in a gentle voice to a relaxed user.
[0352] As a concrete example, let's consider a scenario where a user is choosing lunch. In this case, the system determines from the user's facial expressions and voice that they are relaxed and provides a meal option with appropriate calories. When the user provides feedback such as "That's a good choice," the server learns from this positive response and improves the accuracy of future advice.
[0353] An example of a prompt message is, "When a user is choosing lunch and has a relaxed expression, what healthy meal options should be suggested?" In this way, the present invention can provide users with individually tailored health support and facilitate healthy choices in their daily lives.
[0354] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0355] Step 1:
[0356] The device uses a camera to capture the user's surroundings and acquire visual data. The input is the camera's video feed, and by extracting data on surrounding objects and facial expressions from this video feed, image data is obtained as output. This includes specific objects within the user's field of vision and the user's facial expressions.
[0357] Step 2:
[0358] The device uses a microphone to capture ambient sounds and collects audio data. The input is sound from the environment, and the output is audio data obtained by digitally converting that sound. This audio data allows the user's speech and voice tone to be obtained.
[0359] Step 3:
[0360] The terminal transmits the acquired image and audio data to the server. The input is the image and audio data collected in the previous step, which are transferred via the wireless network. As output, the data reaches the server.
[0361] Step 4:
[0362] The server analyzes received image data and uses machine learning algorithms to identify specific objects. The input is image data, and the analysis generates object information and related health information as output. For example, it can recognize food and calculate its calorie information.
[0363] Step 5:
[0364] The server uses an emotion recognition engine to analyze voice data and determine the user's emotional state. The input is voice data, which is analyzed to identify the user's emotional state (e.g., relaxed, stressed). Based on this output, the way information is provided changes.
[0365] Step 6:
[0366] The server generates health information and advice based on the analysis results. The input consists of object information and emotional states, which are combined to generate information best suited to the user's situation. The output is customized health information presented to the user.
[0367] Step 7:
[0368] The device presents information received from the server to the user through a visual display. The input is health information and advice sent from the server, and by displaying this information on the AR glasses' screen, the output is visualized data delivered to the user.
[0369] Step 8:
[0370] The user selects an action based on the information presented and inputs the result as feedback into the terminal. The input consists of the user's actions and comments, and this feedback is sent to the system as output.
[0371] Step 9:
[0372] The terminal sends user feedback back to the server, which is then used to refine the machine learning algorithm. The input is user feedback information, which the server uses to generate output that improves the system's accuracy.
[0373] (Application Example 2)
[0374] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0375] In recent years, the importance of optimizing health and daily life choices and providing personalized information tailored to each user has increased. However, conventional systems often failed to consider the user's emotional state and were limited to providing general information. Furthermore, mechanisms for optimizing the system using user feedback were insufficient, making it difficult to provide advice appropriate to individual users.
[0376] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0377] In this invention, the server includes means for capturing environmental data around the user using a video acquisition device, means for identifying specific objects and generating health information using a machine learning algorithm to analyze the environmental data, and means for presenting the generated health information or product information in accordance with the user's psychological state based on their emotional state. This makes it possible to provide optimal advice and information in real time that is tailored to the user's emotions and environment.
[0378] A "video acquisition device" is a device that captures environmental data around the user.
[0379] "Environmental data" refers to data that includes various information that can be perceived within the user's field of vision.
[0380] A "machine learning algorithm" is a mathematical procedure used to analyze data and identify specific objects.
[0381] "Health information" refers to health-related data associated with a specific subject.
[0382] A "display device" is a device that visually presents generated information to the user.
[0383] "Means of providing interaction" refers to methods that enable two-way communication between the user and the system.
[0384] "Voice input for recognizing the user's emotional state" refers to a method of understanding a user's psychological state through the voice they produce.
[0385] "Feedback information" refers to data about responses and reactions received from users.
[0386] A "virtual character" is a digitized entity created by a computer and used to interact with users.
[0387] "Product information" refers to data about attributes and characteristics related to a product.
[0388] The system designed to realize this application aims to improve the in-store shopping experience and health management for users of smart glasses. The server, terminal, and user work together as follows:
[0389] First, a camera built into the smart glasses captures the user's field of view in real time, and this video data is sent to a server via the device. The server uses machine learning libraries such as TensorFlow and OpenCV to identify objects in the video and extract related health and product information. Furthermore, it analyzes the user's psychological state using voice input and generates data corresponding to the user's emotions. At this time, the analyzed video data and emotion data are processed integrally by an emotion recognition engine on the same server.
[0390] Information about products the user sees, including details and potential health effects, is visually displayed on the user's smart glasses using augmented reality technology. The system also generates virtual characters that interact with the user, offering advice and recommendations based on their emotional state. For example, if the system analyzes the user as relaxed, the character will offer more friendly and proactive recommendations.
[0391] The generated information is continuously improved by machine learning algorithms on the server based on user feedback, and optimized to provide more appropriate and advanced information in the future. Furthermore, the feedback data is input into an emotion recognition engine, enabling the provision of information that is even more tailored to the user's emotions.
[0392] For example, if a user picks up a snack at a supermarket, the smart glasses will display information about the health effects of that snack and alternative options. When the user comments on it, the server immediately uses that feedback information to learn from the system.
[0393] Examples of prompt statements include:
[0394] "Provide health information about the products that users see."
[0395] "Propose products that align with the user's emotions."
[0396] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0397] Step 1:
[0398] The device uses a video acquisition device to capture the user's field of view. It acquires video data from the user's field of view as input and sends it to the server as output. Specifically, the smart glasses' camera collects data in real time and prepares it to be sent to the server in its original format.
[0399] Step 2:
[0400] The server analyzes received video data using machine learning algorithms. It receives video data sent from a terminal as input and generates a list of identified objects and related health or product information as output. This process utilizes TensorFlow and OpenCV to recognize objects in the video and retrieve relevant information from a database.
[0401] Step 3:
[0402] The server analyzes the user's voice input using an emotion recognition engine. It receives voice data sent from the terminal as input and identifies the user's emotional state as output. Here, the emotion recognition engine extracts voice features, analyzes them with a generative AI model, and determines the current emotion.
[0403] Step 4:
[0404] The server integrates the video analysis results and emotion analysis results to generate information to be visually presented to the user. It uses the outputs from steps 2 and 3 as input and sends information display instructions to the terminal as output. In this step, the server selects the most appropriate advice or product suggestions based on the information obtained from both analyses.
[0405] Step 5:
[0406] The device presents comprehensive information to the user through a display device. It receives information display instructions from a server as input and provides health information, product options, and interactive advice from a virtual character on the smart glasses' display as output. Specifically, it uses augmented reality technology to overlay information onto the user's field of view.
[0407] Step 6:
[0408] The user provides feedback to the system. As input, the user provides feedback through their device, including selected information and voice comments, which are then returned to the server. As output, this feedback data exists on the server and is used for subsequent analyses.
[0409] Step 7:
[0410] The server optimizes the machine learning algorithm based on feedback information. It receives user feedback data as input and generates an updated model of the algorithm as output. This enables the server to provide more adaptive and accurate information to users in the future.
[0411] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0412] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0413] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0414] [Third Embodiment]
[0415] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0416] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0417] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0418] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0419] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0420] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0421] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0422] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0423] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0424] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0425] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0426] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0427] The system of this invention is designed to help users promote a healthy lifestyle in their daily lives. This system operates on a platform that utilizes augmented reality technology.
[0428] First, the device uses a video acquisition device to capture data about the user's everyday environment. For example, when the user is eating, it collects images of the food and drinks in the surrounding area in real time. This allows for the acquisition of data based on the user's lifestyle.
[0429] Next, the server analyzes these captured image data. Using machine learning algorithms, it identifies specific objects or foods and generates associated nutritional and calorie information. For example, if the recognized object is "pizza," it identifies and generates its calorie value and nutritional information.
[0430] Next, the device provides the user with generated health information. Through the display, it visually presents the user with calorie information about meals and healthy options. Furthermore, a virtual character appears on the screen, interacts with the user, and provides specific advice regarding exercise and diet. For example, it might suggest, "This pizza is 700 kcal. We recommend a 20-minute jog."
[0431] This system can also receive feedback from users. Users can take action based on the information provided and input the results and their impressions as feedback. The server then uses this feedback data to refine its machine learning algorithms, continuously improving the accuracy and appropriateness of the information it generates.
[0432] For example, when elderly people use the system, advice and action suggestions that are easier to understand are provided through a virtual character, allowing the system to be flexibly customized to suit the user. In this way, the present invention comprehensively supports the user's health and enables natural health promotion in daily life.
[0433] The following describes the processing flow.
[0434] Step 1:
[0435] The device activates a video acquisition device to capture image data in real time, capturing the environment within the user's field of view. The camera continuously acquires images, recording objects and activities present around the user.
[0436] Step 2:
[0437] The device compresses the captured image data for processing and sends it to a server in the cloud using a secure communication protocol. Encryption technology is used to ensure that data transfer is fast and secure.
[0438] Step 3:
[0439] The server inputs the received image data into a machine learning algorithm to identify specific objects, particularly food and everyday items. This process is carried out by a trained model, which analyzes attributes associated with the identified objects (e.g., calories, ingredients).
[0440] Step 4:
[0441] The server generates health information based on the identified subject and creates a detailed report that includes nutritional advice regarding diet and recommendations regarding exercise. This information also takes into account the user's past behavioral history and health status.
[0442] Step 5:
[0443] The server sends the generated health information to the terminal, which then presents it to the user. The information is displayed visually on an AR display and also conveyed audibly through a virtual character.
[0444] Step 6:
[0445] Based on the health information and advice presented, users make decisions about actions such as diet and exercise. For example, they might approve an exercise plan and record its implementation on their device.
[0446] Step 7:
[0447] The device then collects user feedback and actions taken, and sends them back to the server. This feedback is used to learn from and improve the system, influencing future data processing.
[0448] Step 8:
[0449] The server retrains its machine learning model based on the collected feedback information, improving the accuracy and adaptability of the algorithm. This enhances the quality of the advice provided, enabling it to continue effectively supporting users' health management.
[0450] (Example 1)
[0451] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0452] In modern society, establishing a healthy lifestyle is a crucial issue for many people. However, it is difficult to properly manage one's own health and diet in daily life. In particular, obtaining appropriate information and taking action based on that information is not easy, and specific advice tailored to individual circumstances is needed. To solve this problem, the provision of real-time health information and practical advice based on that information is necessary.
[0453] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0454] In this invention, the server includes means for acquiring information about the user's surroundings using a video acquisition system, means for identifying a specific object using a data analysis algorithm to analyze the acquired information and generate health-related information related to that object, and means for visually presenting the generated health-related information to the user through a display system and providing interaction with the user. This enables the user to intuitively and effectively manage their health in their daily life.
[0455] A "video acquisition system" is a device or technology that can collect information about the user's surroundings, and is a means of acquiring environmental data in real time.
[0456] "Users" refer to individuals who use the system, and are the target audience for the provision and management of health information in their daily lives.
[0457] A "data analysis algorithm" refers to a mathematical processing method used to analyze acquired information and identify specific objects, and is particularly a means of generating health-related information.
[0458] "Health-related information" refers to specific data related to the user's health management, such as calorie values and nutrient information for individual foods and exercises.
[0459] A "display system" is a device or technology for visually communicating generated health-related information to users, and is a means of enabling interaction.
[0460] A "virtual character" is a person or character represented on the screen by computer generation, and it serves to provide users with advice on exercise and nutrition management.
[0461] This invention is a comprehensive system designed to support users' healthy lifestyles. It provides users with specific means to monitor and improve their health in their daily lives.
[0462] First, the device uses a video acquisition system to capture the user's daily living environment. For example, it uses a dedicated augmented reality-enabled device to capture information in real time while the user is eating. This device is equipped with a camera and sensors to collect information about food and drinks that come into the user's field of vision.
[0463] Next, the server analyzes the acquired information using data analysis algorithms such as TensorFlow and PyTorch. This allows it to identify specific foods from image data and generate related health information, such as calorie and nutrient data. A pre-trained food database is used for identification.
[0464] The terminal then uses a display system to visually present the generated health-related information to the user. A virtual character is displayed on the screen along with the generated information, offering advice to the user on specific exercise and nutritional management. This advice is individually customized for the user and may be presented in the form of, for example, "This dessert is 350kcal. If you want to be even healthier, we recommend a 15-minute walk."
[0465] Furthermore, users can provide feedback via voice input or touch interfaces. The server analyzes this feedback information and improves its data analysis algorithms, enabling it to provide more accurate and user-friendly health-related information.
[0466] This invention provides advanced support for users' health management and offers practical guidance for daily life. In this way, users can improve their health more efficiently and effectively.
[0467] Example of a prompt:
[0468] "Please explain how augmented reality technology can be used to provide calorie information about food when a user is eating."
[0469] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0470] Step 1:
[0471] The device uses a video acquisition system to capture the user's daily living environment. Specifically, the device's camera captures food and drinks that come into the user's field of view. The acquired data, as input, is real-time image data. As output, this image data is sent to a server in the cloud.
[0472] Step 2:
[0473] The server acquires image data sent from the terminal and performs analysis using data analysis algorithms. Specifically, it uses TensorFlow and PyTorch to identify objects within the image. This process involves comparing the image data with a pre-built food database to identify specific foods. The input is image data received from the terminal. The output generates health-related information such as calories and nutrients associated with the identified foods.
[0474] Step 3:
[0475] The terminal receives health-related information transmitted from the server and uses a display system to provide it to the user. Specifically, calorie and nutrient information is visually presented on the display screen. Furthermore, a virtual character appears and provides health advice to the user. The input is health-related information sent from the server. The output includes information visually presented to the user and accompanying specific action suggestions.
[0476] Step 4:
[0477] Users act based on the provided health information and input the results and their impressions into the device. Input interfaces include voice input and touchscreens. The input consists of user feedback. As output, this feedback is sent to a server and used for subsequent data analysis.
[0478] Step 5:
[0479] The server receives user feedback and uses it to improve the data analysis algorithm. Specifically, it analyzes the feedback data and uses it to improve the accuracy of the model. This process enables the generated AI model to have higher accuracy and adaptability in future generation of health-related information. The input is user feedback data. The output is a state where the algorithm's accuracy has been improved.
[0480] (Application Example 1)
[0481] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0482] In modern society, a lack of information to help consumers make healthy food choices is a significant challenge. In particular, the restaurant industry often lacks clear nutritional and calorie information for each menu item, making it difficult for consumers to make health-conscious choices. Furthermore, even when health information is available, there is a lack of mechanisms to propose concrete action plans based on that information.
[0483] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0484] In this invention, the server includes means for acquiring environmental information around the user using a video acquisition device, means for identifying a specific object using a learning algorithm to analyze the environmental information and generate nutritional information related to that object, and means for presenting the generated nutritional information to the user through a display device and providing interaction with the user. This makes it possible for users to instantly know the nutritional information and calories of each menu item, even when eating out, and to receive specific advice to make choices that contribute to their health.
[0485] A "video acquisition device" is a device that acquires information about the user's surrounding environment in real time using cameras and sensors.
[0486] "Environmental information" refers to data about objects and conditions present in or around the user's location.
[0487] A "learning algorithm" is an artificial intelligence technique used to analyze large amounts of data and identify specific patterns or features.
[0488] "Nutritional information" refers to health-related data such as calories, vitamins, and minerals contained in specific meals or foods.
[0489] A "display device" is a digital display or screen that allows users to visually confirm information.
[0490] "Interaction" refers to the process and reactions in which a user interacts with and interacts with a system.
[0491] "Additional information" refers to supplementary data provided alongside the main information.
[0492] "Visual devices" are devices used to display images or videos to users, and usually refer to glasses-type devices or monitors.
[0493] A "virtual character" is an interactive entity, such as a fictional person or animal, created in a digital environment.
[0494] This invention provides a system to support healthy food choices in the restaurant industry. The server has the function of acquiring images of food in a restaurant when the user wears smart glasses. Specifically, the image acquisition device is a camera mounted on the smart glasses, which captures images of the food in real time. The server uses a learning algorithm operated on the cloud to analyze the acquired image data, perform image recognition, and identify specific dishes and ingredients. Machine learning frameworks such as TensorFlow are applied to the algorithms used.
[0495] The server extracts the relevant nutritional information from the database based on the identification result and displays the generated nutritional information on the smart glasses' display. This display device is the visual device of the smart glasses. A virtual character appears on the screen and provides interaction with the user. The virtual character suggests dishes that contribute to health, proposes exercises, and provides specific advice based on the user's choices.
[0496] For example, when a user points their camera at a hamburger steak, a virtual character will advise, "This hamburger steak is approximately 500 kcal. It would be ideal to eat it with this salad." Another example of a prompt message is, "Based on image analysis, display the nutritional information and calories of the recognized dish and provide healthy advice for that dish." This system allows users to obtain appropriate nutritional information and health-based choices in real time.
[0497] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0498] Step 1:
[0499] The smart glasses, acting as the terminal, use a camera to capture images of food within the user's field of view. The input is real-time video data, and the output is still image data of that food. This image is then prepared to be sent to a server for recognition.
[0500] Step 2:
[0501] The server analyzes the received image data and identifies dishes using machine learning algorithms (e.g., TensorFlow). Here, it takes still image data as input and outputs specific dish names and ingredient information through data calculations. This identified data is used for accessing the server's internal database.
[0502] Step 3:
[0503] The server extracts nutritional information related to identified dishes from a database. The input is the dish identification data, and the output is the nutritional and calorie information for that dish. Database queries are used to collect relevant health information and compile it in a user-friendly format.
[0504] Step 4:
[0505] The smart glasses, acting as the terminal, generate a virtual character along with the acquired nutritional information and visually present the information on the display. The input is nutritional information and a data format for display, while the output is the visual presentation to the user. Here, the character makes healthy suggestions to the user.
[0506] Step 5:
[0507] Users make meal choices based on the provided nutritional information and advice from virtual characters. They can input feedback based on their choices and experiences into the terminal. The input is user feedback data, and the output is feedback data sent to the server for future analysis.
[0508] Step 6:
[0509] The server incorporates user feedback into its machine learning model, continuously updating and optimizing the algorithm. Here, user feedback is used as input, and data calculations such as algorithm parameter tuning are performed to improve the accuracy of subsequent recognitions. The output is the updated learning model, which improves the overall system performance.
[0510] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0511] The system of this invention is designed to support users' daily lives and promote health management. This system operates on a platform that integrates augmented reality technology and emotion recognition technology.
[0512] First, the device uses a video acquisition device to capture the environment within the user's field of view, acquiring image data in real time. This data includes the user's surroundings and facial expressions. Furthermore, the device also acquires the user's voice using a microphone, preparing it for analysis by the emotion engine.
[0513] Next, the server receives this data and performs analysis using machine learning algorithms. Through the analysis of image data, specific objects (e.g., food) are identified, and related health information (e.g., calories and ingredients) is generated. Simultaneously, the emotion engine analyzes the user's facial expressions and voice to identify their emotional state. This enables the generation of health information and advice that takes the user's current emotions into account.
[0514] The generated information is presented to the user by the terminal. Through the display device, the user not only sees health information and lifestyle advice visually, but also engages in interactive dialogue using a virtual character. The character changes its response based on the user's emotional state; for example, if the user is tired, it will offer encouraging advice in a gentle tone.
[0515] Furthermore, users can use this information to make choices about their diet and exercise. The selected actions are input into the system as feedback. The device then sends this feedback information back to the server, where it is used to refine the emotion engine and machine learning algorithms. This process allows the system to continuously provide personalized health support that is more adapted to the user.
[0516] As a concrete example, consider a scenario where a user is choosing lunch. In this case, the system recognizes from the user's facial expression that they are relaxed and offers healthy and nutritious options. If the user responds to the choice with "That's a good choice," the system learns from the positive response and strengthens its advice position for future occasions. In total, this invention supports comprehensive health management that takes the user's emotions into account, making it easier to make healthy choices in daily life.
[0517] The following describes the processing flow.
[0518] Step 1:
[0519] The device uses a video acquisition device and microphone to capture the user's surroundings as well as the user's face and voice in real time. By capturing images of the food the user is looking at and what they are saying, it simultaneously collects environmental data and audio data.
[0520] Step 2:
[0521] The device compresses and encrypts the captured environmental and audio data before sending it to a server in the cloud. Communication protocols are used to ensure that the data is transmitted securely and efficiently.
[0522] Step 3:
[0523] The server uses the received data to execute an image recognition algorithm and identify specific objects (such as food) from the environmental data. It then generates health information, such as calorie and nutritional information, for the identified objects.
[0524] Step 4:
[0525] The server utilizes an emotion engine to analyze voice and facial expression data. It identifies the user's emotional state (for example, whether they are feeling stressed) and determines how this will influence the health information generated based on that state.
[0526] Step 5:
[0527] The server optimizes health advice and creates customized information based on identified objects and the user's emotional state. For example, if the user is tired, it will suggest easy-to-prepare healthy meals.
[0528] Step 6:
[0529] The terminal presents health information and advice from the server to the user visually and audibly through a display device. A virtual character offers suggestions for diet and exercise in a tone that matches the user's emotions.
[0530] Step 7:
[0531] Users act based on the advice they receive. For example, they might decide to cook a dish using a suggested recipe and provide feedback by entering the result into their device.
[0532] Step 8:
[0533] The device collects user feedback and forwards it back to the server. The server uses this feedback to update its machine learning model and sentiment engine, improving the accuracy of the information. This will enable the provision of more appropriate health management advice in the future.
[0534] (Example 2)
[0535] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0536] In modern society, providing personalized health management and support for daily life is becoming increasingly important. However, conventional health management systems often fail to adequately consider users' emotions and real-time circumstances, thus failing to maximize the user experience. Therefore, there is a need to realize a system that provides dynamically customized, individualized health information and advice in response to users' emotions and circumstances, thereby increasing user satisfaction and effectiveness.
[0537] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0538] In this invention, the server includes a visual acquisition means that operates to acquire objects and facial expressions around the user; an analysis means that analyzes the acquired visual and audio data and uses machine learning means to identify the object and the user's emotional state; and a visual display means that presents health information and advice that takes into account the user's emotions, based on the analysis results. This enables highly accurate health support tailored to the user's individual circumstances and emotions.
[0539] "Visual acquisition means" refers to devices and sensors used to capture the user's surrounding environment and facial expressions, thereby enabling the acquisition of real-time visual data.
[0540] "Machine learning methods" refer to algorithms and their execution environments used to analyze data obtained from users and identify the objects contained in that data or the emotional state of the user.
[0541] "Analysis means" refers to functions that process acquired visual and audio data to identify objects and analyze user emotions.
[0542] "Visual display means" refers to devices or systems that present information generated based on analysis results to the user, thereby enabling effective information sharing with the user.
[0543] A "virtual dialogue target" refers to a digital character or virtual agent created to facilitate interaction with the user and provide emotion-based advice.
[0544] The system of the present invention aims to manage the user's health and support their daily life, and is implemented using a device that integrates augmented reality technology and emotion recognition technology. First, the terminal uses a camera function as a visual acquisition means to capture the environment within the user's field of view. This camera acquires the environment around the user and the user's own facial expressions in real time. In addition, a microphone is used to collect data on the tone of the user's voice, which indicates their speech and emotions.
[0545] Next, the server receives visual and audio data sent from the terminal. In this scenario, the server has machine learning-based analysis capabilities, using libraries such as TensorFlow to identify specific objects from image data and determine the user's emotions from audio data. Based on the analysis results, the server generates health information and advice tailored to the user's current situation.
[0546] The generated information is communicated to the user through the device's visual display. Specifically, the AR glasses display is used to visually present the information to the user, and a virtual dialogue partner engages in an interactive conversation with the user. The dialogue partner adopts an approach tailored to the user's emotional state; for example, it might offer healthy eating advice in a gentle voice to a relaxed user.
[0547] As a concrete example, let's consider a scenario where a user is choosing lunch. In this case, the system determines from the user's facial expressions and voice that they are relaxed and provides a meal option with appropriate calories. When the user provides feedback such as "That's a good choice," the server learns from this positive response and improves the accuracy of future advice.
[0548] An example of a prompt message is, "When a user is choosing lunch and has a relaxed expression, what healthy meal options should be suggested?" In this way, the present invention can provide users with individually tailored health support and facilitate healthy choices in their daily lives.
[0549] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0550] Step 1:
[0551] The device uses a camera to capture the user's surroundings and acquire visual data. The input is the camera's video feed, and by extracting data on surrounding objects and facial expressions from this video feed, image data is obtained as output. This includes specific objects within the user's field of vision and the user's facial expressions.
[0552] Step 2:
[0553] The device uses a microphone to capture ambient sounds and collects audio data. The input is sound from the environment, and the output is audio data obtained by digitally converting that sound. This audio data allows the user's speech and voice tone to be obtained.
[0554] Step 3:
[0555] The terminal transmits the acquired image and audio data to the server. The input is the image and audio data collected in the previous step, which are transferred via the wireless network. As output, the data reaches the server.
[0556] Step 4:
[0557] The server analyzes received image data and uses machine learning algorithms to identify specific objects. The input is image data, and the analysis generates object information and related health information as output. For example, it can recognize food and calculate its calorie information.
[0558] Step 5:
[0559] The server uses an emotion recognition engine to analyze voice data and determine the user's emotional state. The input is voice data, which is analyzed to identify the user's emotional state (e.g., relaxed, stressed). Based on this output, the way information is provided changes.
[0560] Step 6:
[0561] The server generates health information and advice based on the analysis results. The input consists of object information and emotional states, which are combined to generate information best suited to the user's situation. The output is customized health information presented to the user.
[0562] Step 7:
[0563] The device presents information received from the server to the user through a visual display. The input is health information and advice sent from the server, and by displaying this information on the AR glasses' screen, the output is visualized data delivered to the user.
[0564] Step 8:
[0565] The user selects an action based on the information presented and inputs the result as feedback into the terminal. The input consists of the user's actions and comments, and this feedback is sent to the system as output.
[0566] Step 9:
[0567] The terminal sends user feedback back to the server, which is then used to refine the machine learning algorithm. The input is user feedback information, which the server uses to generate output that improves the system's accuracy.
[0568] (Application Example 2)
[0569] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0570] In recent years, the importance of optimizing health and daily life choices and providing personalized information tailored to each user has increased. However, conventional systems often failed to consider the user's emotional state and were limited to providing general information. Furthermore, mechanisms for optimizing the system using user feedback were insufficient, making it difficult to provide advice appropriate to individual users.
[0571] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0572] In this invention, the server includes means for capturing environmental data around the user using a video acquisition device, means for identifying specific objects and generating health information using a machine learning algorithm to analyze the environmental data, and means for presenting the generated health information or product information in accordance with the user's psychological state based on their emotional state. This makes it possible to provide optimal advice and information in real time that is tailored to the user's emotions and environment.
[0573] A "video acquisition device" is a device that captures environmental data around the user.
[0574] "Environmental data" refers to data that includes various information that can be perceived within the user's field of vision.
[0575] A "machine learning algorithm" is a mathematical procedure used to analyze data and identify specific objects.
[0576] "Health information" refers to health-related data associated with a specific subject.
[0577] A "display device" is a device that visually presents generated information to the user.
[0578] "Means of providing interaction" refers to methods that enable two-way communication between the user and the system.
[0579] "Voice input for recognizing the user's emotional state" refers to a method of understanding a user's psychological state through the voice they produce.
[0580] "Feedback information" refers to data about responses and reactions received from users.
[0581] A "virtual character" is a digitized entity created by a computer and used to interact with users.
[0582] "Product information" refers to data about attributes and characteristics related to a product.
[0583] The system designed to realize this application aims to improve the in-store shopping experience and health management for users of smart glasses. The server, terminal, and user work together as follows:
[0584] First, a camera built into the smart glasses captures the user's field of view in real time, and this video data is sent to a server via the device. The server uses machine learning libraries such as TensorFlow and OpenCV to identify objects in the video and extract related health and product information. Furthermore, it analyzes the user's psychological state using voice input and generates data corresponding to the user's emotions. At this time, the analyzed video data and emotion data are processed integrally by an emotion recognition engine on the same server.
[0585] Information about products the user sees, including details and potential health effects, is visually displayed on the user's smart glasses using augmented reality technology. The system also generates virtual characters that interact with the user, offering advice and recommendations based on their emotional state. For example, if the system analyzes the user as relaxed, the character will offer more friendly and proactive recommendations.
[0586] The generated information is continuously improved by machine learning algorithms on the server based on user feedback, and optimized to provide more appropriate and advanced information in the future. Furthermore, the feedback data is input into an emotion recognition engine, enabling the provision of information that is even more tailored to the user's emotions.
[0587] For example, if a user picks up a snack at a supermarket, the smart glasses will display information about the health effects of that snack and alternative options. When the user comments on it, the server immediately uses that feedback information to learn from the system.
[0588] Examples of prompt statements include:
[0589] "Provide health information about the products that users see."
[0590] "Propose products that align with the user's emotions."
[0591] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0592] Step 1:
[0593] The device uses a video acquisition device to capture the user's field of view. It acquires video data from the user's field of view as input and sends it to the server as output. Specifically, the smart glasses' camera collects data in real time and prepares it to be sent to the server in its original format.
[0594] Step 2:
[0595] The server analyzes received video data using machine learning algorithms. It receives video data sent from a terminal as input and generates a list of identified objects and related health or product information as output. This process utilizes TensorFlow and OpenCV to recognize objects in the video and retrieve relevant information from a database.
[0596] Step 3:
[0597] The server analyzes the user's voice input using an emotion recognition engine. It receives voice data sent from the terminal as input and identifies the user's emotional state as output. Here, the emotion recognition engine extracts voice features, analyzes them with a generative AI model, and determines the current emotion.
[0598] Step 4:
[0599] The server integrates the video analysis results and emotion analysis results to generate information to be visually presented to the user. It uses the outputs from steps 2 and 3 as input and sends information display instructions to the terminal as output. In this step, the server selects the most appropriate advice or product suggestions based on the information obtained from both analyses.
[0600] Step 5:
[0601] The device presents comprehensive information to the user through a display device. It receives information display instructions from a server as input and provides health information, product options, and interactive advice from a virtual character on the smart glasses' display as output. Specifically, it uses augmented reality technology to overlay information onto the user's field of view.
[0602] Step 6:
[0603] The user provides feedback to the system. As input, the user provides feedback through their device, including selected information and voice comments, which are then returned to the server. As output, this feedback data exists on the server and is used for subsequent analyses.
[0604] Step 7:
[0605] The server optimizes the machine learning algorithm based on feedback information. It receives user feedback data as input and generates an updated model of the algorithm as output. This enables the server to provide more adaptive and accurate information to users in the future.
[0606] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0607] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0608] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0609] [Fourth Embodiment]
[0610] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0611] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0612] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0613] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0614] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0615] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0616] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0617] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0618] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0619] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0620] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0621] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0622] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0623] The system of this invention is designed to help users promote a healthy lifestyle in their daily lives. This system operates on a platform that utilizes augmented reality technology.
[0624] First, the device uses a video acquisition device to capture data about the user's everyday environment. For example, when the user is eating, it collects images of the food and drinks in the surrounding area in real time. This allows for the acquisition of data based on the user's lifestyle.
[0625] Next, the server analyzes these captured image data. Using machine learning algorithms, it identifies specific objects or foods and generates associated nutritional and calorie information. For example, if the recognized object is "pizza," it identifies and generates its calorie value and nutritional information.
[0626] Next, the device provides the user with generated health information. Through the display, it visually presents the user with calorie information about meals and healthy options. Furthermore, a virtual character appears on the screen, interacts with the user, and provides specific advice regarding exercise and diet. For example, it might suggest, "This pizza is 700 kcal. We recommend a 20-minute jog."
[0627] This system can also receive feedback from users. Users can take action based on the information provided and input the results and their impressions as feedback. The server then uses this feedback data to refine its machine learning algorithms, continuously improving the accuracy and appropriateness of the information it generates.
[0628] For example, when elderly people use the system, advice and action suggestions that are easier to understand are provided through a virtual character, allowing the system to be flexibly customized to suit the user. In this way, the present invention comprehensively supports the user's health and enables natural health promotion in daily life.
[0629] The following describes the processing flow.
[0630] Step 1:
[0631] The device activates a video acquisition device to capture image data in real time, capturing the environment within the user's field of view. The camera continuously acquires images, recording objects and activities present around the user.
[0632] Step 2:
[0633] The device compresses the captured image data for processing and sends it to a server in the cloud using a secure communication protocol. Encryption technology is used to ensure that data transfer is fast and secure.
[0634] Step 3:
[0635] The server inputs the received image data into a machine learning algorithm to identify specific objects, particularly food and everyday items. This process is carried out by a trained model, which analyzes attributes associated with the identified objects (e.g., calories, ingredients).
[0636] Step 4:
[0637] The server generates health information based on the identified subject and creates a detailed report that includes nutritional advice regarding diet and recommendations regarding exercise. This information also takes into account the user's past behavioral history and health status.
[0638] Step 5:
[0639] The server sends the generated health information to the terminal, which then presents it to the user. The information is displayed visually on an AR display and also conveyed audibly through a virtual character.
[0640] Step 6:
[0641] Based on the health information and advice presented, users make decisions about actions such as diet and exercise. For example, they might approve an exercise plan and record its implementation on their device.
[0642] Step 7:
[0643] The device then collects user feedback and actions taken, and sends them back to the server. This feedback is used to learn from and improve the system, influencing future data processing.
[0644] Step 8:
[0645] The server retrains its machine learning model based on the collected feedback information, improving the accuracy and adaptability of the algorithm. This enhances the quality of the advice provided, enabling it to continue effectively supporting users' health management.
[0646] (Example 1)
[0647] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0648] In modern society, establishing a healthy lifestyle is a crucial issue for many people. However, it is difficult to properly manage one's own health and diet in daily life. In particular, obtaining appropriate information and taking action based on that information is not easy, and specific advice tailored to individual circumstances is needed. To solve this problem, the provision of real-time health information and practical advice based on that information is necessary.
[0649] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0650] In this invention, the server includes means for acquiring information about the user's surroundings using a video acquisition system, means for identifying a specific object using a data analysis algorithm to analyze the acquired information and generate health-related information related to that object, and means for visually presenting the generated health-related information to the user through a display system and providing interaction with the user. This enables the user to intuitively and effectively manage their health in their daily life.
[0651] A "video acquisition system" is a device or technology that can collect information about the user's surroundings, and is a means of acquiring environmental data in real time.
[0652] "Users" refer to individuals who use the system, and are the target audience for the provision and management of health information in their daily lives.
[0653] A "data analysis algorithm" refers to a mathematical processing method used to analyze acquired information and identify specific objects, and is particularly a means of generating health-related information.
[0654] "Health-related information" refers to specific data related to the user's health management, such as calorie values and nutrient information for individual foods and exercises.
[0655] A "display system" is a device or technology for visually communicating generated health-related information to users, and is a means of enabling interaction.
[0656] A "virtual character" is a person or character represented on the screen by computer generation, and it serves to provide users with advice on exercise and nutrition management.
[0657] This invention is a comprehensive system designed to support users' healthy lifestyles. It provides users with specific means to monitor and improve their health in their daily lives.
[0658] First, the device uses a video acquisition system to capture the user's daily living environment. For example, it uses a dedicated augmented reality-enabled device to capture information in real time while the user is eating. This device is equipped with a camera and sensors to collect information about food and drinks that come into the user's field of vision.
[0659] Next, the server analyzes the acquired information using data analysis algorithms such as TensorFlow and PyTorch. This allows it to identify specific foods from image data and generate related health information, such as calorie and nutrient data. A pre-trained food database is used for identification.
[0660] The terminal then uses a display system to visually present the generated health-related information to the user. A virtual character is displayed on the screen along with the generated information, offering advice to the user on specific exercise and nutritional management. This advice is individually customized for the user and may be presented in the form of, for example, "This dessert is 350kcal. If you want to be even healthier, we recommend a 15-minute walk."
[0661] Furthermore, users can provide feedback via voice input or touch interfaces. The server analyzes this feedback information and improves its data analysis algorithms, enabling it to provide more accurate and user-friendly health-related information.
[0662] This invention provides advanced support for users' health management and offers practical guidance for daily life. In this way, users can improve their health more efficiently and effectively.
[0663] Example of a prompt:
[0664] "Please explain how augmented reality technology can be used to provide calorie information about food when a user is eating."
[0665] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0666] Step 1:
[0667] The device uses a video acquisition system to capture the user's daily living environment. Specifically, the device's camera captures food and drinks that come into the user's field of view. The acquired data, as input, is real-time image data. As output, this image data is sent to a server in the cloud.
[0668] Step 2:
[0669] The server acquires image data sent from the terminal and performs analysis using data analysis algorithms. Specifically, it uses TensorFlow and PyTorch to identify objects within the image. This process involves comparing the image data with a pre-built food database to identify specific foods. The input is image data received from the terminal. The output generates health-related information such as calories and nutrients associated with the identified foods.
[0670] Step 3:
[0671] The terminal receives health-related information transmitted from the server and uses a display system to provide it to the user. Specifically, calorie and nutrient information is visually presented on the display screen. Furthermore, a virtual character appears and provides health advice to the user. The input is health-related information sent from the server. The output includes information visually presented to the user and accompanying specific action suggestions.
[0672] Step 4:
[0673] Users act based on the provided health information and input the results and their impressions into the device. Input interfaces include voice input and touchscreens. The input consists of user feedback. As output, this feedback is sent to a server and used for subsequent data analysis.
[0674] Step 5:
[0675] The server receives user feedback and uses it to improve the data analysis algorithm. Specifically, it analyzes the feedback data and uses it to improve the accuracy of the model. This process enables the generated AI model to have higher accuracy and adaptability in future generation of health-related information. The input is user feedback data. The output is a state where the algorithm's accuracy has been improved.
[0676] (Application Example 1)
[0677] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0678] In modern society, a lack of information to help consumers make healthy food choices is a significant challenge. In particular, the restaurant industry often lacks clear nutritional and calorie information for each menu item, making it difficult for consumers to make health-conscious choices. Furthermore, even when health information is available, there is a lack of mechanisms to propose concrete action plans based on that information.
[0679] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0680] In this invention, the server includes means for acquiring environmental information around the user using a video acquisition device, means for identifying a specific object using a learning algorithm to analyze the environmental information and generate nutritional information related to that object, and means for presenting the generated nutritional information to the user through a display device and providing interaction with the user. This makes it possible for users to instantly know the nutritional information and calories of each menu item, even when eating out, and to receive specific advice to make choices that contribute to their health.
[0681] A "video acquisition device" is a device that acquires information about the user's surrounding environment in real time using cameras and sensors.
[0682] "Environmental information" refers to data about objects and conditions present in or around the user's location.
[0683] A "learning algorithm" is an artificial intelligence technique used to analyze large amounts of data and identify specific patterns or features.
[0684] "Nutritional information" refers to health-related data such as calories, vitamins, and minerals contained in specific meals or foods.
[0685] A "display device" is a digital display or screen that allows users to visually confirm information.
[0686] "Interaction" refers to the process and reactions in which a user interacts with and interacts with a system.
[0687] "Additional information" refers to supplementary data provided alongside the main information.
[0688] "Visual devices" are devices used to display images or videos to users, and usually refer to glasses-type devices or monitors.
[0689] A "virtual character" is an interactive entity, such as a fictional person or animal, created in a digital environment.
[0690] This invention provides a system to support healthy food choices in the restaurant industry. The server has the function of acquiring images of food in a restaurant when the user wears smart glasses. Specifically, the image acquisition device is a camera mounted on the smart glasses, which captures images of the food in real time. The server uses a learning algorithm operated on the cloud to analyze the acquired image data, perform image recognition, and identify specific dishes and ingredients. Machine learning frameworks such as TensorFlow are applied to the algorithms used.
[0691] The server extracts the relevant nutritional information from the database based on the identification result and displays the generated nutritional information on the smart glasses' display. This display device is the visual device of the smart glasses. A virtual character appears on the screen and provides interaction with the user. The virtual character suggests dishes that contribute to health, proposes exercises, and provides specific advice based on the user's choices.
[0692] For example, when a user points their camera at a hamburger steak, a virtual character will advise, "This hamburger steak is approximately 500 kcal. It would be ideal to eat it with this salad." Another example of a prompt message is, "Based on image analysis, display the nutritional information and calories of the recognized dish and provide healthy advice for that dish." This system allows users to obtain appropriate nutritional information and health-based choices in real time.
[0693] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0694] Step 1:
[0695] The smart glasses, acting as the terminal, use a camera to capture images of food within the user's field of view. The input is real-time video data, and the output is still image data of that food. This image is then prepared to be sent to a server for recognition.
[0696] Step 2:
[0697] The server analyzes the received image data and identifies dishes using machine learning algorithms (e.g., TensorFlow). Here, it takes still image data as input and outputs specific dish names and ingredient information through data calculations. This identified data is used for accessing the server's internal database.
[0698] Step 3:
[0699] The server extracts nutritional information related to identified dishes from a database. The input is the dish identification data, and the output is the nutritional and calorie information for that dish. Database queries are used to collect relevant health information and compile it in a user-friendly format.
[0700] Step 4:
[0701] The smart glasses, acting as the terminal, generate a virtual character along with the acquired nutritional information and visually present the information on the display. The input is nutritional information and a data format for display, while the output is the visual presentation to the user. Here, the character makes healthy suggestions to the user.
[0702] Step 5:
[0703] Users make meal choices based on the provided nutritional information and advice from virtual characters. They can input feedback based on their choices and experiences into the terminal. The input is user feedback data, and the output is feedback data sent to the server for future analysis.
[0704] Step 6:
[0705] The server incorporates user feedback into its machine learning model, continuously updating and optimizing the algorithm. Here, user feedback is used as input, and data calculations such as algorithm parameter tuning are performed to improve the accuracy of subsequent recognitions. The output is the updated learning model, which improves the overall system performance.
[0706] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0707] The system of this invention is designed to support users' daily lives and promote health management. This system operates on a platform that integrates augmented reality technology and emotion recognition technology.
[0708] First, the device uses a video acquisition device to capture the environment within the user's field of view, acquiring image data in real time. This data includes the user's surroundings and facial expressions. Furthermore, the device also acquires the user's voice using a microphone, preparing it for analysis by the emotion engine.
[0709] Next, the server receives this data and performs analysis using machine learning algorithms. Through the analysis of image data, specific objects (e.g., food) are identified, and related health information (e.g., calories and ingredients) is generated. Simultaneously, the emotion engine analyzes the user's facial expressions and voice to identify their emotional state. This enables the generation of health information and advice that takes the user's current emotions into account.
[0710] The generated information is presented to the user by the terminal. Through the display device, the user not only sees health information and lifestyle advice visually, but also engages in interactive dialogue using a virtual character. The character changes its response based on the user's emotional state; for example, if the user is tired, it will offer encouraging advice in a gentle tone.
[0711] Furthermore, users can use this information to make choices about their diet and exercise. The selected actions are input into the system as feedback. The device then sends this feedback information back to the server, where it is used to refine the emotion engine and machine learning algorithms. This process allows the system to continuously provide personalized health support that is more adapted to the user.
[0712] As a concrete example, consider a scenario where a user is choosing lunch. In this case, the system recognizes from the user's facial expression that they are relaxed and offers healthy and nutritious options. If the user responds to the choice with "That's a good choice," the system learns from the positive response and strengthens its advice position for future occasions. In total, this invention supports comprehensive health management that takes the user's emotions into account, making it easier to make healthy choices in daily life.
[0713] The following describes the processing flow.
[0714] Step 1:
[0715] The device uses a video acquisition device and microphone to capture the user's surroundings as well as the user's face and voice in real time. By capturing images of the food the user is looking at and what they are saying, it simultaneously collects environmental data and audio data.
[0716] Step 2:
[0717] The device compresses and encrypts the captured environmental and audio data before sending it to a server in the cloud. Communication protocols are used to ensure that the data is transmitted securely and efficiently.
[0718] Step 3:
[0719] The server uses the received data to execute an image recognition algorithm and identify specific objects (such as food) from the environmental data. It then generates health information, such as calorie and nutritional information, for the identified objects.
[0720] Step 4:
[0721] The server utilizes an emotion engine to analyze voice and facial expression data. It identifies the user's emotional state (for example, whether they are feeling stressed) and determines how this will influence the health information generated based on that state.
[0722] Step 5:
[0723] The server optimizes health advice and creates customized information based on identified objects and the user's emotional state. For example, if the user is tired, it will suggest easy-to-prepare healthy meals.
[0724] Step 6:
[0725] The terminal presents health information and advice from the server to the user visually and audibly through a display device. A virtual character offers suggestions for diet and exercise in a tone that matches the user's emotions.
[0726] Step 7:
[0727] Users act based on the advice they receive. For example, they might decide to cook a dish using a suggested recipe and provide feedback by entering the result into their device.
[0728] Step 8:
[0729] The device collects user feedback and forwards it back to the server. The server uses this feedback to update its machine learning model and sentiment engine, improving the accuracy of the information. This will enable the provision of more appropriate health management advice in the future.
[0730] (Example 2)
[0731] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0732] In modern society, providing personalized health management and support for daily life is becoming increasingly important. However, conventional health management systems often fail to adequately consider users' emotions and real-time circumstances, thus failing to maximize the user experience. Therefore, there is a need to realize a system that provides dynamically customized, individualized health information and advice in response to users' emotions and circumstances, thereby increasing user satisfaction and effectiveness.
[0733] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0734] In this invention, the server includes a visual acquisition means that operates to acquire objects and facial expressions around the user; an analysis means that analyzes the acquired visual and audio data and uses machine learning means to identify the object and the user's emotional state; and a visual display means that presents health information and advice that takes into account the user's emotions, based on the analysis results. This enables highly accurate health support tailored to the user's individual circumstances and emotions.
[0735] "Visual acquisition means" refers to devices and sensors used to capture the user's surrounding environment and facial expressions, thereby enabling the acquisition of real-time visual data.
[0736] "Machine learning methods" refer to algorithms and their execution environments used to analyze data obtained from users and identify the objects contained in that data or the emotional state of the user.
[0737] "Analysis means" refers to functions that process acquired visual and audio data to identify objects and analyze user emotions.
[0738] "Visual display means" refers to devices or systems that present information generated based on analysis results to the user, thereby enabling effective information sharing with the user.
[0739] A "virtual dialogue target" refers to a digital character or virtual agent created to facilitate interaction with the user and provide emotion-based advice.
[0740] The system of the present invention aims to manage the user's health and support their daily life, and is implemented using a device that integrates augmented reality technology and emotion recognition technology. First, the terminal uses a camera function as a visual acquisition means to capture the environment within the user's field of view. This camera acquires the environment around the user and the user's own facial expressions in real time. In addition, a microphone is used to collect data on the tone of the user's voice, which indicates their speech and emotions.
[0741] Next, the server receives visual and audio data sent from the terminal. In this scenario, the server has machine learning-based analysis capabilities, using libraries such as TensorFlow to identify specific objects from image data and determine the user's emotions from audio data. Based on the analysis results, the server generates health information and advice tailored to the user's current situation.
[0742] The generated information is communicated to the user through the device's visual display. Specifically, the AR glasses display is used to visually present the information to the user, and a virtual dialogue partner engages in an interactive conversation with the user. The dialogue partner adopts an approach tailored to the user's emotional state; for example, it might offer healthy eating advice in a gentle voice to a relaxed user.
[0743] As a concrete example, let's consider a scenario where a user is choosing lunch. In this case, the system determines from the user's facial expressions and voice that they are relaxed and provides a meal option with appropriate calories. When the user provides feedback such as "That's a good choice," the server learns from this positive response and improves the accuracy of future advice.
[0744] An example of a prompt message is, "When a user is choosing lunch and has a relaxed expression, what healthy meal options should be suggested?" In this way, the present invention can provide users with individually tailored health support and facilitate healthy choices in their daily lives.
[0745] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0746] Step 1:
[0747] The device uses a camera to capture the user's surroundings and acquire visual data. The input is the camera's video feed, and by extracting data on surrounding objects and facial expressions from this video feed, image data is obtained as output. This includes specific objects within the user's field of vision and the user's facial expressions.
[0748] Step 2:
[0749] The device uses a microphone to capture ambient sounds and collects audio data. The input is sound from the environment, and the output is audio data obtained by digitally converting that sound. This audio data allows the user's speech and voice tone to be obtained.
[0750] Step 3:
[0751] The terminal transmits the acquired image and audio data to the server. The input is the image and audio data collected in the previous step, which are transferred via the wireless network. As output, the data reaches the server.
[0752] Step 4:
[0753] The server analyzes received image data and uses machine learning algorithms to identify specific objects. The input is image data, and the analysis generates object information and related health information as output. For example, it can recognize food and calculate its calorie information.
[0754] Step 5:
[0755] The server uses an emotion recognition engine to analyze voice data and determine the user's emotional state. The input is voice data, which is analyzed to identify the user's emotional state (e.g., relaxed, stressed). Based on this output, the way information is provided changes.
[0756] Step 6:
[0757] The server generates health information and advice based on the analysis results. The input consists of object information and emotional states, which are combined to generate information best suited to the user's situation. The output is customized health information presented to the user.
[0758] Step 7:
[0759] The device presents information received from the server to the user through a visual display. The input is health information and advice sent from the server, and by displaying this information on the AR glasses' screen, the output is visualized data delivered to the user.
[0760] Step 8:
[0761] The user selects an action based on the information presented and inputs the result as feedback into the terminal. The input consists of the user's actions and comments, and this feedback is sent to the system as output.
[0762] Step 9:
[0763] The terminal sends user feedback back to the server, which is then used to refine the machine learning algorithm. The input is user feedback information, which the server uses to generate output that improves the system's accuracy.
[0764] (Application Example 2)
[0765] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0766] In recent years, the importance of optimizing health and daily life choices and providing personalized information tailored to each user has increased. However, conventional systems often failed to consider the user's emotional state and were limited to providing general information. Furthermore, mechanisms for optimizing the system using user feedback were insufficient, making it difficult to provide advice appropriate to individual users.
[0767] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0768] In this invention, the server includes means for capturing environmental data around the user using a video acquisition device, means for identifying specific objects and generating health information using a machine learning algorithm to analyze the environmental data, and means for presenting the generated health information or product information in accordance with the user's psychological state based on their emotional state. This makes it possible to provide optimal advice and information in real time that is tailored to the user's emotions and environment.
[0769] A "video acquisition device" is a device that captures environmental data around the user.
[0770] "Environmental data" refers to data that includes various information that can be perceived within the user's field of vision.
[0771] A "machine learning algorithm" is a mathematical procedure used to analyze data and identify specific objects.
[0772] "Health information" refers to health-related data associated with a specific subject.
[0773] A "display device" is a device that visually presents generated information to the user.
[0774] "Means of providing interaction" refers to methods that enable two-way communication between the user and the system.
[0775] "Voice input for recognizing the user's emotional state" refers to a method of understanding a user's psychological state through the voice they produce.
[0776] "Feedback information" refers to data about responses and reactions received from users.
[0777] A "virtual character" is a digitized entity created by a computer and used to interact with users.
[0778] "Product information" refers to data about attributes and characteristics related to a product.
[0779] The system designed to realize this application aims to improve the in-store shopping experience and health management for users of smart glasses. The server, terminal, and user work together as follows:
[0780] First, a camera built into the smart glasses captures the user's field of view in real time, and this video data is sent to a server via the device. The server uses machine learning libraries such as TensorFlow and OpenCV to identify objects in the video and extract related health and product information. Furthermore, it analyzes the user's psychological state using voice input and generates data corresponding to the user's emotions. At this time, the analyzed video data and emotion data are processed integrally by an emotion recognition engine on the same server.
[0781] Information about products the user sees, including details and potential health effects, is visually displayed on the user's smart glasses using augmented reality technology. The system also generates virtual characters that interact with the user, offering advice and recommendations based on their emotional state. For example, if the system analyzes the user as relaxed, the character will offer more friendly and proactive recommendations.
[0782] The generated information is continuously improved by machine learning algorithms on the server based on user feedback, and optimized to provide more appropriate and advanced information in the future. Furthermore, the feedback data is input into an emotion recognition engine, enabling the provision of information that is even more tailored to the user's emotions.
[0783] For example, if a user picks up a snack at a supermarket, the smart glasses will display information about the health effects of that snack and alternative options. When the user comments on it, the server immediately uses that feedback information to learn from the system.
[0784] Examples of prompt statements include:
[0785] "Provide health information about the products that users see."
[0786] "Propose products that align with the user's emotions."
[0787] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0788] Step 1:
[0789] The device uses a video acquisition device to capture the user's field of view. It acquires video data from the user's field of view as input and sends it to the server as output. Specifically, the smart glasses' camera collects data in real time and prepares it to be sent to the server in its original format.
[0790] Step 2:
[0791] The server analyzes received video data using machine learning algorithms. It receives video data sent from a terminal as input and generates a list of identified objects and related health or product information as output. This process utilizes TensorFlow and OpenCV to recognize objects in the video and retrieve relevant information from a database.
[0792] Step 3:
[0793] The server analyzes the user's voice input using an emotion recognition engine. It receives voice data sent from the terminal as input and identifies the user's emotional state as output. Here, the emotion recognition engine extracts voice features, analyzes them with a generative AI model, and determines the current emotion.
[0794] Step 4:
[0795] The server integrates the video analysis results and emotion analysis results to generate information to be visually presented to the user. It uses the outputs from steps 2 and 3 as input and sends information display instructions to the terminal as output. In this step, the server selects the most appropriate advice or product suggestions based on the information obtained from both analyses.
[0796] Step 5:
[0797] The device presents comprehensive information to the user through a display device. It receives information display instructions from a server as input and provides health information, product options, and interactive advice from a virtual character on the smart glasses' display as output. Specifically, it uses augmented reality technology to overlay information onto the user's field of view.
[0798] Step 6:
[0799] The user provides feedback to the system. As input, the user provides feedback through their device, including selected information and voice comments, which are then returned to the server. As output, this feedback data exists on the server and is used for subsequent analyses.
[0800] Step 7:
[0801] The server optimizes the machine learning algorithm based on feedback information. It receives user feedback data as input and generates an updated model of the algorithm as output. This enables the server to provide more adaptive and accurate information to users in the future.
[0802] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0803] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0804] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0805] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0806] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0807] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0808] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0809] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0810] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0811] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0812] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0813] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0814] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0815] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0816] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0817] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0818] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0819] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0820] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0821] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0822] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0823] The following is further disclosed regarding the embodiments described above.
[0824] (Claim 1)
[0825] A means for capturing environmental data around the user using a video acquisition device,
[0826] A means for using a machine learning algorithm to analyze the aforementioned environmental data, to identify a specific target, and to generate health information related to that target,
[0827] A means of visually presenting generated health information to the user through a display device and providing interaction with the user,
[0828] A system that includes this.
[0829] (Claim 2)
[0830] The system according to claim 1, further comprising means for continuously improving the machine learning algorithm using user feedback information and generating health information that has been optimized again.
[0831] (Claim 3)
[0832] The system according to claim 1, further comprising means for generating a virtual character using the display device and for the character to perform an interactive dialogue in which the character provides advice to the user regarding exercise and diet.
[0833] "Example 1"
[0834] (Claim 1)
[0835] A means of acquiring information about the user's surroundings using a video acquisition system,
[0836] A means for using a data analysis algorithm to identify a specific target in order to analyze the acquired information and to generate health-related information related to that target,
[0837] A means of visually presenting generated health-related information to users through a display system and providing interaction with users,
[0838] A means of providing assistance in making proposals that are suitable for the user,
[0839] A system that includes this.
[0840] (Claim 2)
[0841] The system according to claim 1, further comprising means for continuously improving the data analysis algorithm and regenerating optimized health-related information using evaluation information from users.
[0842] (Claim 3)
[0843] The system according to claim 1, further comprising means for generating a virtual character using the display system and for the character to perform an interactive exchange in which the character provides advice to the user regarding exercise and nutrition management.
[0844] "Application Example 1"
[0845] (Claim 1)
[0846] A means of acquiring environmental information around the user using a video acquisition device,
[0847] A means for using a learning algorithm to identify a specific target in order to analyze the aforementioned environmental information and to generate nutritional information related to that target,
[0848] A means of presenting generated nutritional information to the user through a display device and providing interaction with the user,
[0849] A means for presenting additional information in real time on a visual device using the aforementioned nutritional information,
[0850] A system that includes this.
[0851] (Claim 2)
[0852] The system according to claim 1, further comprising means for continuously improving the learning algorithm using user feedback information and regenerating optimized nutrition information.
[0853] (Claim 3)
[0854] The system according to claim 1, further comprising means for generating a virtual character using the display device and for the character to perform a dialogue in which the character provides advice to the user regarding exercise and diet.
[0855] "Example 2 of combining an emotion engine"
[0856] (Claim 1)
[0857] A visual acquisition means that operates to acquire objects and facial expressions around the user,
[0858] An analysis means that uses machine learning means to analyze acquired visual and audio data and identify the emotional state of the object and the user,
[0859] A visual display means that presents health information generated based on the analysis results and advice that takes into account the user's emotions,
[0860] A system that includes this.
[0861] (Claim 2)
[0862] The system according to claim 1, further comprising a learning means for receiving user feedback data, continuously optimizing the analysis means, and generating personalized health information.
[0863] (Claim 3)
[0864] The system according to claim 1, further comprising a dialogue means for displaying a virtual dialogue target using a visual display means, and the dialogue target providing advice on exercise and diet in accordance with the user's emotions.
[0865] "Application example 2 of combining emotional engines"
[0866] (Claim 1)
[0867] A means for capturing environmental data around the user using a video acquisition device,
[0868] A means for using a machine learning algorithm to analyze the aforementioned environmental data, to identify a specific target, and to generate health information related to that target,
[0869] A means of visually presenting generated health information to the user through a display device and providing interaction with the user,
[0870] A means for analyzing voice input to recognize the user's emotional state and providing health information or product information tailored to the user's psychological state,
[0871] A system that includes this.
[0872] (Claim 2)
[0873] The system according to claim 1, further comprising means for continuously improving the machine learning algorithm using user feedback information and generating further optimized health information or product information.
[0874] (Claim 3)
[0875] The system according to claim 1, further comprising means for generating a virtual character using the display device and for the character to perform an interactive dialogue in which the character provides advice to the user regarding exercise, diet and product selection. [Explanation of symbols]
[0876] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for capturing environmental data around the user using a video acquisition device, A means for using a machine learning algorithm to analyze the aforementioned environmental data, to identify a specific target, and to generate health information related to that target, A means of visually presenting generated health information to the user through a display device and providing interaction with the user, A system that includes this.
2. The system according to claim 1, further comprising means for continuously improving the machine learning algorithm using user feedback information and generating health information that has been optimized again.
3. The system according to claim 1, further comprising means for generating a virtual character using the display device and for the character to perform an interactive dialogue in which the character provides advice to the user regarding exercise and diet.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A