System
The system addresses the complexity and inefficiency of conventional life log recording by using a camera-equipped glasses to automatically capture and analyze images, generating personalized lifestyle suggestions based on user behavior, thus enhancing lifestyle management.
Patent Information
- Application Number
- JP2024116504
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Conventional life log recording methods and lifestyle improvement systems require dedicated devices, are complicated to operate, and suffer from low accuracy and efficiency due to manual recording and analysis.
A system that uses a built-in camera in glasses to capture real-time images, preprocess them to remove noise, send them to a server in batch format, analyze the images using AI models to identify objects and behaviors, record them as life logs, and generate lifestyle improvement suggestions, which are then transmitted to a terminal for user notification.
Enables automated lifestyle improvements and accurate life log recording without special user operations, providing personalized suggestions for enhancing lifestyle habits.
Smart Images

Figure 2026015030000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional life log recording methods and improvement suggestion systems require dedicated devices and are complicated to operate, which has led to problems for many users, preventing them from continuing to improve their lifestyles or record their logs. Furthermore, manual recording and analysis have led to problems with low accuracy and efficiency. This invention aims to solve these problems by automating users' lifestyle improvements and recording their life logs, and simplifying operation. [Means for solving the problem]
[0005] The present invention includes a means for capturing a user's field of view in real time using a built-in camera, a means for preprocessing the captured image to remove noise, and a means for sending the preprocessed image to a server in batch format. The server provides a means for analyzing the received image data to identify important objects and categorizing the user's behavior based on the analysis results. The server also includes a means for recording behavioral data as text information in a life log and analyzing the recorded life log to generate lifestyle improvement suggestions. The system further includes a means for transmitting the generated suggestions and life log to a terminal and notifying the user of them. This system allows users to automate lifestyle improvements and life log recording in a simplified environment.
[0006] An "integrated camera" is a camera device built into the glasses that captures your field of vision in real time.
[0007] "Preprocessing" refers to processing to remove noise from the captured image and improve image clarity.
[0008] "Batch format" is a method of processing and transmitting multiple image data at once at regular intervals.
[0009] A "server" is a remote computer system that receives data sent from a terminal, analyzes it, and performs the necessary processing.
[0010] "Image analysis" is the process of analyzing captured images using AI models to identify objects and scenes within the images.
[0011] An "object" is a significant element (e.g., food, person, place) that can be recognized or identified in an image.
[0012] "Behavioral classification" is the process of assigning user behavior to specific categories (e.g., eating, exercise, travel) based on the analysis results.
[0013] A "life log" is data that records a user's daily actions and events.
[0014] "Improvement suggestions" are advice or suggestions for improving the user's lifestyle habits based on the analysis results of the life log.
[0015] A "terminal" is a device that is used as glasses and provides information to the user.
[0016] "Notification" is the act of conveying information from a terminal to a user.
[0017] "Analysis results" refers to data and information obtained through image analysis. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] This invention is a system that uses glasses-type devices equipped with a built-in camera and image processing functions to automatically record a user's daily activities, analyze those activities via a server, and generate lifestyle improvement suggestions.With this system, users can efficiently obtain life logs without performing any special operations and use them to improve their lifestyles.
[0040] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view in real time. This capture operation occurs every second, and the obtained image data is temporarily stored in the device's internal memory. After capturing the images, the device performs preprocessing on them to remove noise and improve clarity. This preprocessing uses image processing techniques such as Gaussian filters and sharpening algorithms.
[0041] Next, the device sends the pre-processed image data to the server in batch format at regular intervals (e.g., every hour). A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety.
[0042] The server inputs the image data received from the device into a generative AI model to analyze the image, using deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks).
[0043] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, it identifies "salad," "steak," and "rice" from an image captured during a meal and records it as "lunch." The server records this behavior data as text information in a life log and saves it in a database.
[0044] The server then analyzes the recorded life log and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You're not getting enough vegetables. Eat more vegetables next time."
[0045] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: salad, steak, rice. Eat more vegetables."
[0046] Users can receive this notification, refer to their life log and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lifestyle and manage their health.
[0047] In this way, the present invention realizes a new system that combines a built-in camera, image processing technology, and analysis algorithms to automatically record users' behavior and provide specific suggestions that will help improve their lifestyle.
[0048] The processing flow will be explained below.
[0049] Step 1:
[0050] The device (glasses-type device) uses a built-in camera to capture the user's field of view at one-second intervals.
[0051] Specifically, the camera automatically captures an image and temporarily stores it in its internal memory.
[0052] Step 2:
[0053] The device preprocesses the captured image to remove noise and improve image clarity.
[0054] Specifically, the image quality is improved by applying a Gaussian filter or sharpening algorithm.
[0055] Step 3:
[0056] The terminal converts the preprocessed image data into a batch format at regular time intervals (e.g., every hour).
[0057] Specifically, the image data is compressed into one file and prepared for transmission.
[0058] Step 4:
[0059] The terminal transmits the image data in a batch format to the server using a secure communication protocol (e.g., HTTPS).
[0060] Specifically, a session for transmitting image data is established, and data uploading begins.
[0061] Step 5:
[0062] The server inputs the image data received from the terminal into the generated AI model and analyzes the image.
[0063] Specifically, the AI model processes the received image data and identifies important objects (e.g., food, people, landmarks).
[0064] Step 6:
[0065] The server classifies the user's behavior into categories (e.g., eating, exercise, travel) based on the image analysis results.
[0066] Specifically, the analysis results are assigned to preset categories and organized according to each category.
[0067] Step 7:
[0068] The server records the behavioral data organized by category as text information in a life log.
[0069] Specifically, text information such as "2023-10-05 12:30 Lunch: salad, steak, rice" is saved in the database.
[0070] Step 8:
[0071] The server analyzes the recorded life log and generates suggestions for improving lifestyle habits.
[0072] Specifically, it compares and analyzes the user's past behavioral logs and generates specific advice such as "eat more vegetables."
[0073] Step 9:
[0074] The server sends the generated improvement suggestions and detailed life logs to the device.
[0075] Specifically, the system compresses the improvement suggestions and life log information and sends them to the device using a secure communication protocol.
[0076] Step 10:
[0077] The device notifies the user of the received suggestions and life logs so that they can be viewed.
[0078] Specifically, the device's display and voice notification functions are used to notify the user, displaying "Lunch record: salad, steak, rice. Eat more vegetables."
[0079] Step 11:
[0080] Users receive notifications from their devices and view life logs and improvement suggestions.
[0081] Specifically, the user checks the suggestions on the device, reviews their lifestyle habits based on them, and decides on their next course of action (e.g., eating more vegetables at the next meal).
[0082] Example 1
[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0084] Conventional life log management systems have had problems with the effort required for users to record their own data and the accuracy of the records. In addition, the lifestyle improvement suggestions they offer are often not specialized enough, and their effectiveness in improving users' health is limited.
[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0086] In this invention, the server includes means for capturing the field of view in real time using a built-in camera, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for inputting the images into a generative AI model to identify important objects, means for categorizing the user's behavior based on the analysis results, means for recording the behavioral data as text information in a life log, means for analyzing the recorded life log with an algorithm based on expert knowledge to generate lifestyle improvement suggestions, means for transmitting the generated suggestions and life log to a terminal, and means for notifying the user of the suggestions and life log. This allows the user to easily obtain a highly accurate life log and receive expert improvement suggestions.
[0087] A "built-in camera" is a camera built into a device that captures the external field of view in real time.
[0088] "Means for capturing the field of view in real time" is a function that uses the built-in camera to continuously capture images that are within the user's field of view.
[0089] The "means for pre-processing the captured image" is a function for performing processing on the obtained image data to remove noise and improve clarity.
[0090] The "means for transmitting to the server in batch format at specific time intervals" is a function for collecting preprocessed image data at regular time intervals and transmitting the collected data to the server using a secure communication protocol.
[0091] A "generative AI model" refers to an algorithm that uses artificial intelligence to analyze data and perform specific tasks.
[0092] "Means for identifying significant objects" refers to the ability to use generative AI models to recognize and analyze objects, people, landmarks, etc. in images.
[0093] "Means for categorizing user behavior" is a function that records user behavior into categories such as food, exercise, and travel based on the analysis results.
[0094] A "life log" is data that records a user's daily activities and behaviors in chronological order.
[0095] An "algorithm based on expert knowledge" is a computational method for performing analysis that incorporates specialized knowledge in a specific field.
[0096] The "means for generating lifestyle improvement suggestions" is a function that generates advice for improving the user's lifestyle using an algorithm based on the recorded life log and expert knowledge.
[0097] The "means for transmitting to the terminal" is a function for transferring the generated proposals and life logs to the terminal.
[0098] The "means of notifying the user of suggestions and life logs" is a function that notifies the user of the contents of suggestions and life logs through the device's built-in display or audio output.
[0099] MODE FOR CARRYING OUT THE INVENTION
[0100] This invention is a system that uses glasses-type devices equipped with a built-in camera and image processing functions to automatically record a user's daily activities, analyze those activities via a server, and generate lifestyle improvement suggestions.With this system, users can efficiently obtain life logs without performing any special operations and use them to improve their lifestyles.
[0101] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view in real time. This capture operation occurs every second, and the obtained image data is temporarily stored in the device's internal memory. After capturing the images, the device performs preprocessing on them to remove noise and improve clarity. This preprocessing uses image processing techniques such as Gaussian filters and sharpening algorithms.
[0102] Next, the device sends the pre-processed image data to the server in batch format at regular intervals (e.g., every hour). A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety.
[0103] The server inputs the image data received from the device into a generative AI model to analyze the image, using deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks).
[0104] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, it identifies "salad," "steak," and "rice" from an image captured during a meal and records it as "lunch." The server records this behavior data as text information in a life log and saves it in a database.
[0105] The server then analyzes the recorded life log and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You're not getting enough vegetables. Eat more vegetables next time."
[0106] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: salad, steak, rice. Eat more vegetables."
[0107] Users can receive this notification, refer to their life log and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lifestyle and manage their health.
[0108] An example of a specific prompt is, "Please identify whether there is food in this image and identify the specific type." Based on this prompt, the generative AI model recognizes and classifies the objects in the image.
[0109] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0110] Step 1:
[0111] The device uses a built-in camera to capture the user's field of view in real time.
[0112] Input data: user's view
[0113] How it works: The device captures images within its field of view every second and stores them in its internal memory, recording the user's daily activities as video.
[0114] Output data: Captured image data
[0115] Step 2:
[0116] The device pre-processes the captured image to remove noise and improve clarity.
[0117] Input data: Captured image data
[0118] What it does: The device uses a Gaussian filter to remove noise from the image and a sharpening algorithm to improve clarity, allowing it to accurately identify important objects during analysis.
[0119] Output data: Preprocessed image data
[0120] Step 3:
[0121] The terminal sends the preprocessed image data in batch format to the server at regular time intervals (e.g., every hour).
[0122] Input data: Preprocessed image data
[0123] Specific operation: The device periodically compiles image data into a batch format and sends it securely to the server using HTTPS.
[0124] Output data: batch image data sent to the server
[0125] Step 4:
[0126] The server inputs the received image data into the generated AI model and analyzes the image.
[0127] Input data: batch image data
[0128] How it works: The server uses a generative AI model (e.g., YOLO, RetinaNet, etc.) to analyze the input image and recognize and identify important objects (e.g., food, people, landmarks). For example, it can recognize food items such as "salad" and "steak" from an image of a meal.
[0129] Output data: Analysis results (recognized objects)
[0130] Step 5:
[0131] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel).
[0132] Input data: Analysis results (recognized objects)
[0133] Specific actions: Based on the recognized object information, actions are classified into categories such as "meal," "exercise," and "travel," and recorded as a life log. For example, if "salad," "steak," and "rice" are detected, they are classified into the category of "lunch."
[0134] Output data: Behavioral data categorized by category
[0135] Step 6:
[0136] The server records the behavioral data as text information in a life log and stores it in a database.
[0137] Input data: Categorized behavioral data
[0138] Specific operation: The classified behavioral data is converted into text format and saved in a database, thereby saving the user's past behavior as a history.
[0139] Output data: Saved life logs
[0140] Step 7:
[0141] The server analyzes the recorded life log and generates suggestions for improving lifestyle habits.
[0142] Input data: Saved life logs
[0143] Specific operation: Using an algorithm based on expert knowledge, the system generates lifestyle improvement suggestions based on the user's behavioral history. For example, it generates specific advice such as, "You're not getting enough vegetables. Next time, eat more vegetables."
[0144] Output data: Suggestions for improving lifestyle habits
[0145] Step 8:
[0146] The server then sends the generated improvement suggestions and detailed life logs back to the device.
[0147] Input data: Generated lifestyle improvement suggestions, detailed life log
[0148] Specific operation: A secure communication protocol is used to send the generated lifestyle improvement suggestions and detailed life logs to the device.
[0149] Output data: Improvement suggestions and life logs sent to the device
[0150] Step 9:
[0151] The device receives improvement suggestions and life logs and notifies the user.
[0152] Input data: improvement proposals, life logs
[0153] Specific operation: The device notifies the user of improvement suggestions and the contents of the life log via the device's built-in display and voice output. For example, it displays "Lunch record: salad, steak, rice. Eat more vegetables."
[0154] Output data: Improvement suggestions notified to the user and lifelog
[0155] Step 10:
[0156] Users receive notifications, refer to their own life logs and improvement suggestions, and reflect them in their next actions.
[0157] Input data: Notified improvement suggestions, lifelog
[0158] Specific actions: Refer to the notification and reflect it in your next actions. For example, take action to improve your lifestyle, such as consciously eating more vegetables at your next meal.
[0159] Output data: Actions that reflect improvement suggestions
[0160] (Application example 1)
[0161] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0162] Conventional lifestyle recording systems and in-store customer behavior analysis systems require users and customers to record and operate the systems themselves, which is time-consuming and makes it difficult to obtain accurate data.In addition, it is not possible to grasp in real time what products or areas in the store customers are interested in, which makes it difficult to provide efficient customer service and marketing.
[0163] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0164] In this invention, the server includes means for capturing the field of view in real time using a built-in image capture device, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for analyzing the received image data to identify important objects, means for classifying the user's behavior based on the analysis results, means for saving the behavior data as text information in a lifestyle log, means for analyzing the recorded lifestyle log to generate lifestyle improvement suggestions, means for sending the generated suggestions and lifestyle log to a terminal, means for notifying the user of the suggestions and lifestyle log, means for measuring the level of customer interest in specific products or areas in the store, and means for notifying store managers and staff of the analysis results. This reduces the effort required for users and customers, enables accurate data acquisition and analysis, and enables efficient customer service and marketing.
[0165] An "integrated imaging device" is a device, such as a camera or image sensor, built into a device that captures the surrounding field of view in real time.
[0166] "Preprocessing" refers to a series of steps performed on captured image data to remove noise and improve image clarity.
[0167] "Batch format" is a method of processing and sending data for a certain period of time all at once.
[0168] "Important objects" are objects or areas identified by the analysis that should be of interest to the user or the system.
[0169] "Behavior" refers to the activity or pattern of a user or customer at a particular time.
[0170] "Classification" refers to organizing user behavior into categories based on the analysis results.
[0171] "Life Record" is a database that stores user behavior data as text information.
[0172] "Lifestyle improvement suggestions" are specific advice for improving your health and quality of life that is created by analyzing your recorded lifestyle records.
[0173] A "terminal" is a device worn by a user and has various functions including a built-in image capture device.
[0174] "Notification" refers to information transmission activities that inform users of generated suggestions and life logs via their devices.
[0175] A "customer" is a user who visits a store.
[0176] "In-store" means the physical premises where products are displayed and sold.
[0177] "Interested" means that a customer is taking actions that show interest, such as focusing on a particular product or area or staying there for a long time.
[0178] "Store Manager" means a person responsible for the operation and management of a store.
[0179] "Staff" refers to employees who deal with customers and manage merchandise within the store.
[0180] The system that realizes this application example is composed of a terminal equipped with a built-in image capture device, a server that uses a communication protocol, and a system that notifies users based on the analysis results. This system operates as follows.
[0181] The device captures the user's or customer's field of view in real time using a built-in image capture device, taking images every second. The captured images are then pre-processed to remove noise and improve image clarity, using OpenCV's Gaussian filter and sharpening algorithms, for example.
[0182] The preprocessed image data is sent to the server in batch format at regular intervals. This data transmission uses a secure communication protocol (e.g., HTTPS) to ensure data security. The server then analyzes the received image data and identifies important objects using a generative AI model (e.g., YOLO or RetinaNet). This analysis identifies user behavior categories and the products and areas in the store that customers are interested in.
[0183] The server classifies user behavior based on the analysis results and saves them as text information in a daily log. It also notifies store managers and staff in real time based on the analysis results of customer behavior in the store. Notifications are sent via smart glasses or smartphones, and specific alerts, such as "A customer is interested in the wine section," are displayed.
[0184] The recorded lifestyle records and behavioral data are further analyzed using deep learning algorithms to generate lifestyle improvement suggestions. For example, advice such as "You're not getting enough vegetables. Eat more vegetables next time" may be generated. The generated suggestions are resent to the device and notified to the user. This notification is sent via the smart glasses display or the smartphone notification function.
[0185] For example, if a customer stares at the wine section of a store for a long time, that information is sent to the server and analyzed as "the customer is interested in the wine section." As a result, sales staff are notified that "they should attend to the customer in the wine section," thereby enabling personalized customer service.
[0186] An example of a prompt for a generative AI model is as follows:
[0187] Suggest the best way to serve customers when they are interested in the wine section.
[0188] This system reduces the workload for users and stores, while enabling efficient and effective customer service and data-based marketing strategies.
[0189] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0190] Step 1:
[0191] The device captures the field of view in real time using a built-in image capture device. Specifically, the device acquires an image every second. At this point, the input is the captured image data, and the output is the image data before preprocessing.
[0192] Step 2:
[0193] The device preprocesses the captured image to remove noise and improve image clarity, using OpenCV's Gaussian filter and sharpening algorithms. The input to this process is the captured image data, and the output is image data with noise removed and improved clarity.
[0194] Step 3:
[0195] The terminal transmits the preprocessed image data to the server in batch format at regular intervals. Specifically, the data is stored in a fixed buffer and transmitted, for example, every hour. The input in this process is a plurality of preprocessed image data, and the output is batch data transmitted to the server.
[0196] Step 4:
[0197] The server analyzes the received image data. Specifically, it uses a generative AI model (e.g., YOLO or RetinaNet) to identify important objects in the image. The input to this process is a batch of image data, and the output is a list of objects as the analysis result.
[0198] Step 5:
[0199] The server categorizes the user's behavior based on the analysis results. Specifically, it sets categories such as eating, exercise, and shopping, and categorizes the data based on these. The input for this process is a list of objects, and the output is behavioral data categorized by category.
[0200] Step 6:
[0201] The server saves the behavioral data as text information in a life log. Specifically, it saves it in a database. The input for this process is behavioral data categorized by category, and the output is life log data as text information.
[0202] Step 7:
[0203] The server analyzes the recorded life log and generates lifestyle improvement suggestions. Specifically, it uses a deep learning algorithm to generate advice based on past logs and specialized knowledge. The input in this process is life log data, and the output is improvement suggestions.
[0204] Step 8:
[0205] The server transmits the generated suggestions and lifelog data to the device using a secure communication protocol. The inputs to this process are the improvement suggestions and lifelog data, and the output is the data transmitted to the device.
[0206] Step 9:
[0207] The device notifies the user of the suggestions and lifelogs by displaying a message on the display or by issuing a voice notification. The input to this process is the suggestions and lifelog data received from the server, and the output is the notification delivered to the user.
[0208] Step 10:
[0209] The server analyzes the level of interest customers have in specific products or areas in the store and notifies store managers and staff of the results. Specifically, it analyzes customer gaze data and notifies the results in real time. The input to this process is gaze data, and the output is a notification message as the analysis result.
[0210] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0211] This invention is a system that uses a glasses-type device with a built-in camera and image processing function to automatically record a user's daily activities, and further recognizes and analyzes the user's emotions using an emotion engine to generate lifestyle improvement suggestions.With this system, users can efficiently collect life logs and receive advice on lifestyle improvements based on their emotions without performing any special operations.
[0212] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view at one-second intervals. This capture operation is performed automatically, and the obtained image data is temporarily stored in the device's internal memory. At the same time, the user's face is also periodically captured by the camera and analyzed by the emotion engine.
[0213] The device then pre-processes the captured image to remove noise and improve clarity, using image processing techniques such as Gaussian filters and sharpening algorithms.
[0214] The device sends the preprocessed image data in batches at regular intervals (e.g., every hour) to the server. A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety. Emotion data is also sent to the server in the same way.
[0215] The server inputs the image data received from the device into a generative AI model to analyze the image. This image analysis utilizes deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks). The emotion engine also analyzes the user's face to identify their emotional state (e.g., joy, sadness, anger, etc.).
[0216] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, from an image captured during a meal, "salad," "steak," and "rice" are identified and recorded as "lunch." The server also uses facial recognition technology to classify the user's emotions and adds them to the behavioral data. For example, if "happiness" is detected on the user's face during lunch, it is recorded as "lunch (happiness)." The server records this behavioral data as text information in a life log and saves it in a database.
[0217] The server then analyzes the recorded life log and emotional data to generate lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You haven't been eating enough vegetables, so I recommend you eat more next time. Also, your recent meals have been associated with pleasure, so maintain this pattern to reduce stress."
[0218] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern."
[0219] Users can receive notifications, refer to their life logs and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lives and manage their health based on their emotions.
[0220] In this way, the present invention combines a built-in camera, image processing technology, and an emotion engine to realize a new system that automatically records a user's behavior and emotions and provides specific lifestyle improvement suggestions based on that information.
[0221] The processing flow will be explained below.
[0222] Step 1:
[0223] The device (glasses-type device) uses a built-in camera to capture the user's field of view at one-second intervals.
[0224] Specifically, the camera automatically captures an image and temporarily stores it in its internal memory.
[0225] Step 2:
[0226] The device periodically captures the user's face using a camera and analyzes it using an emotion engine.
[0227] Specifically, it uses a facial recognition algorithm to identify the user's face and extract their emotional state (e.g., joy, sadness, anger, etc.).
[0228] Step 3:
[0229] The device pre-processes the captured field of view image to remove noise and improve image clarity.
[0230] Specifically, the image quality is improved by applying a Gaussian filter or sharpening algorithm.
[0231] Step 4:
[0232] The terminal collectively converts the preprocessed image data and emotion data into a batch format at regular time intervals (e.g., every hour).
[0233] Specifically, the image data and emotion data are compressed into a single file and preparations for transmission are made.
[0234] Step 5:
[0235] The terminal transmits the batched data to the server using a secure communication protocol (e.g., HTTPS).
[0236] Specifically, the operation is to establish a session for data transmission and start uploading data.
[0237] Step 6:
[0238] The server inputs the data received from the device into the generative AI model and analyzes the field of view image.
[0239] Specifically, the AI model processes the received image data and identifies important objects (e.g., food, people, landmarks).
[0240] Step 7:
[0241] The server associates the emotion data analyzed by the emotion engine with the behavioral data.
[0242] Specifically, behavioral data is tagged with emotion tags such as "joy" or "sadness" and stored in a database.
[0243] Step 8:
[0244] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel).
[0245] Specifically, the analysis results are assigned to preset categories and organized according to each category.
[0246] Step 9:
[0247] The server records behavioral data and emotional data organized by category as text information in a life log.
[0248] Specifically, the system stores text information such as "2023-10-05 12:30 Lunch: Salad, steak, rice (joy)" in a database.
[0249] Step 10:
[0250] The server analyzes the recorded life log and emotional data and generates suggestions for improving lifestyle habits.
[0251] Specifically, it generates specific advice such as "Eat more vegetables and maintain your current eating patterns" based on the user's past behavioral logs and emotional data.
[0252] Step 11:
[0253] The server sends the generated improvement suggestions and detailed life logs to the device.
[0254] Specifically, the system compresses the improvement suggestions and life log information and sends them to the device using a secure communication protocol.
[0255] Step 12:
[0256] The device notifies the user of the received suggestions and life logs so that they can be viewed.
[0257] Specifically, the device's display and voice notification function will notify the user: "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern."
[0258] Step 13:
[0259] Users receive notifications from their devices and view life logs and improvement suggestions.
[0260] Specifically, the user checks the suggestions on the device and decides on the next action based on them (e.g., eating more vegetables at the next meal).
[0261] Example 2
[0262] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0263] Conventional lifelog acquisition systems require manual user operation, making it difficult to efficiently record detailed daily activities and analyze emotional states. It is also difficult to generate specific lifestyle improvement suggestions based on acquired lifelog data. Furthermore, data security and privacy protection are insufficient, creating a need for a system that users can use with confidence.
[0264] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0265] In this invention, the server includes: a means for capturing the field of view in real time using a built-in camera; a means for preprocessing the captured images to remove noise; a means for sending the preprocessed images to the server in batch format at specific time intervals; a means for analyzing the received image data to identify important objects; a means for categorizing the user's behavior based on the analysis results; a means for analyzing the user's emotions using facial recognition technology; a means for recording the behavioral data and emotional data as text information in a life log; a means for analyzing the recorded life log and generating lifestyle improvement suggestions; a means for transmitting the generated suggestions and life log to a terminal; and a means for notifying the user of the suggestions and life log. This allows the user to obtain a detailed life log without performing any special operations and receive specific lifestyle improvement suggestions based on their emotions. Furthermore, the use of a secure communication protocol ensures data security and privacy protection.
[0266] The "built-in camera" is a camera built into the eyeglass-type terminal, and is a device that has the function of capturing the user's field of view in real time.
[0267] "Pre-processing" refers to a general range of image processing techniques used to remove noise and improve clarity of captured image data, including Gaussian filters and sharpening algorithms.
[0268] "Batch format" refers to a method of processing data collectively at regular intervals, and is used to efficiently transmit and process large amounts of data.
[0269] A "secure communication protocol" is a communication protocol used to ensure data security and privacy, such as HTTPS.
[0270] "Image analysis" refers to the process of using generative AI models and deep learning algorithms to recognize and identify important objects in captured image data.
[0271] "Emotion analysis" refers to the process of using facial recognition technology to identify emotional states (e.g., joy, sadness, anger, etc.) from a user's face.
[0272] "Behavior categorization" refers to the process of classifying user behavior into multiple categories (e.g., eating, exercise, travel) based on the analysis results.
[0273] A "life log" refers to a collection of data that records a user's daily activities and emotional data as text information.
[0274] "Lifestyle improvement suggestions" refer to suggestions that include specific advice for improving the user's lifestyle based on the recorded life log and emotional data.
[0275] "Notification" refers to a method of informing the user of the generated suggestions and details of the life log, and is done through display or audio output.
[0276] The present invention is a system that uses a glasses-type device equipped with a built-in camera and image processing functions to automatically record a user's daily activities, and further recognizes and analyzes the user's emotions using an emotion engine, thereby generating lifestyle improvement suggestions. Detailed embodiments of the present invention will be described below.
[0277] Hardware and software used
[0278] Glasses-type device: Equipped with a built-in camera, internal memory, display, and audio output functions, it captures and temporarily stores data.
[0279] Server: Runs generative AI models, emotion engines, and databases for data analysis and storage. Performs data analysis and proposal generation.
[0280] Communication protocol: HTTPS is used to ensure secure communication.
[0281] Program processing
[0282] The device automatically captures the user's field of view using the built-in camera at one-second intervals, and the user's face is also captured periodically and analyzed by the emotion engine.
[0283] The device pre-processes the captured image data to remove noise and improve clarity, using image processing techniques such as Gaussian filters and sharpening algorithms.
[0284] The terminal transmits the preprocessed image data and emotion data to the server in batch format at regular time intervals using a secure communication protocol (HTTPS).
[0285] The server inputs the received image data into a generative AI model (e.g., YOLO, RetinaNet) for image analysis. This image analysis recognizes and identifies important objects (e.g., food, people, landmarks). The emotion engine also analyzes the user's face to identify their emotional state (e.g., joy, sadness, anger, etc.).
[0286] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, from an image captured during a meal, "salad," "steak," and "rice" are identified and recorded as "lunch." Furthermore, facial recognition technology is used to classify the user's emotions and add them to the behavioral data. For example, if "happiness" is detected on the user's face during lunch, it is categorized as "lunch (happiness)."
[0287] The server analyzes the recorded life log and emotional data and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, the server might generate advice such as, "You're not eating enough vegetables, so we recommend you eat more next time. Also, your recent meals have been associated with pleasure, so maintain this pattern to reduce stress."
[0288] Examples of concrete examples and prompts
[0289] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or voice output. For example, the device's display might say, "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern," and a voice output might say, "Improvement suggestions have been received."
[0290] Prompt Sentence Examples
[0291] "Show me your lunch log and provide relevant sentiment data and improvement suggestions."
[0292] The present invention allows users to effortlessly improve their lifestyle and manage their health based on their emotions. In this way, the present invention realizes a new system that combines a built-in camera, image processing technology, and an emotion engine to automatically record a user's behavior and emotions and provide specific lifestyle improvement suggestions based on the recorded information.
[0293] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0294] Step 1:
[0295] Data Capture
[0296] The device uses a built-in camera to capture the user's field of view at one-second intervals. The captured image data is temporarily stored in the internal memory. At the same time, the device periodically captures the user's face, which is then analyzed by the emotion engine.
[0297] Input: Images of the user's field of view and face
[0298] Output: Temporarily saved captured image data and facial image data
[0299] Specific operation: The device will start the camera, a shutter sound will be heard every second, and the field of view and face will be captured. These image data will be stored in the internal memory.
[0300] Step 2:
[0301] Image preprocessing
[0302] The device pre-processes the captured image data to remove noise and improve clarity, using Gaussian filters and sharpening algorithms.
[0303] Input: Temporarily saved captured image data
[0304] Output: Preprocessed image data
[0305] What happens: The device's image processing algorithm runs, and the progress is displayed on the screen.
[0306] Step 3:
[0307] Data transmission
[0308] The device sends the preprocessed image data and emotion data to the server in batch format at regular intervals using a secure communication protocol (HTTPS).
[0309] Input: Preprocessed image data and emotion data
[0310] Output: The dataset sent to the server
[0311] Specific operation: While data transmission is in progress, the display will show "Sending data...", and once transmission is complete, the display will notify you that "Data transmitted successfully."
[0312] Step 4:
[0313] Data analysis
[0314] The server inputs the received image data into a generative AI model (e.g., YOLO, RetinaNet) to recognize and identify important objects, and also uses an emotion engine to analyze emotions from the facial image data.
[0315] Input: Dataset sent to the server (image data and emotion data)
[0316] Output: Analysis results (object recognition and emotional state)
[0317] Specific operation: The server's processing unit performs image analysis and emotion analysis, and displays the progress on the operation console. After the analysis is complete, a list of recognized objects and emotional states is displayed.
[0318] Step 5:
[0319] Classification of behaviors and emotions
[0320] Based on the analysis results, the server categorizes the user's behavior (e.g., eating, exercise, travel) and also associates emotional data with the behavior.
[0321] Input: Analysis results (object recognition and emotional state)
[0322] Output: Classified behavioral and emotional data
[0323] Specific operation: The server stores the action and emotion classification results in a database in the form of, for example, "Lunch (salad, steak, rice, joy)".
[0324] Step 6:
[0325] Generate improvement suggestions
[0326] The server analyzes the recorded life log and emotional data and generates lifestyle improvement suggestions using an algorithm based on past behavioral logs and expert knowledge.
[0327] Input: Classified behavioral and emotional data
[0328] Output: Generated lifestyle improvement suggestions
[0329] Specific operation: The server's algorithm runs and generates improvement suggestions. The console screen displays a list of the generated suggestions.
[0330] Step 7:
[0331] User Notification
[0332] The server sends the generated improvement suggestions and detailed life logs to the device, which receives them and notifies the user through a display or audio output.
[0333] Input: Generated improvement suggestions and detailed lifelog
[0334] Output: Suggestions and lifelog information notified to the user
[0335] Specific operation: The device display will show "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern," and a voice output will announce, "Improvement suggestions have been received."
[0336] (Application example 2)
[0337] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0338] In modern society, it is difficult for users to easily collect life logs to maintain their health and receive lifestyle improvement suggestions based on those logs in their busy daily lives. It is also difficult to provide personalized dietary suggestions that take the user's emotions into account. Conventional systems require users to manually enter data, which is time-consuming and makes it difficult to collect accurate data. To solve these issues, a system is needed that utilizes a built-in camera, image processing technology, and an emotion engine to automatically collect life logs and provide dietary suggestions based on emotions.
[0339] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the field of view in real time using a built-in camera, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for analyzing the received image data to identify important objects, means for categorizing the user's behavior based on the analysis results, means for recording the behavioral data as text information in a life log, means for analyzing the recorded life log to generate lifestyle improvement suggestions, means for transmitting the generated suggestions and life log to the terminal, means for notifying the user of the suggestions and life log, means for analyzing the user's emotional state and generating personalized meal suggestions based on the emotional state, and means for notifying the user of the generated meal suggestions. This allows the user to receive personalized lifestyle improvement suggestions and meal suggestions based on their emotions without performing any special operations.
[0340] The "built-in camera" is a camera device built into the eyeglass-type terminal to capture the user's field of view in real time.
[0341] "Preprocessing" refers to initial image processing operations to remove noise and improve clarity from a captured image.
[0342] The "batch format" is a data transfer method in which preprocessed image data is sent in batches at regular time intervals.
[0343] A "server" is a remote computer system that receives, analyzes, and stores data sent from user terminals, and generates suggestions based on the results of the analysis.
[0344] "Analysis" is the process of identifying significant objects from the received image data and discerning the user's behavior and emotional state.
[0345] The "behavior category" is a category for classifying the user's daily behavior into specific activity types based on the analysis results.
[0346] A "life log" is a digital history that records and saves a user's daily behavioral and emotional data as text information.
[0347] "Improvement suggestions" are specific advice for improving the user's lifestyle habits that are generated from the analysis results of the life log.
[0348] "Emotional state" refers to the type of emotion (e.g., joy, sadness, anger) analyzed from the user's facial image.
[0349] "Meal Suggestions" are personalized meal suggestions generated based on the user's emotional state and past meal data.
[0350] "Notification" is the process of sending messages to communicate generated suggestions and lifelogs to users.
[0351] The system of the present invention includes a camera built into an eyeglass-type terminal, a server, and a user's terminal (a smartphone or smart glasses).
[0352] First, the device's built-in camera captures the user's field of view in real time, capturing images every second. These captured images are then pre-processed using the device's image processing capabilities. This pre-processing uses image processing techniques such as the OpenCV library to remove noise and improve image clarity.
[0353] The preprocessed image data is sent to the server in batches at regular intervals. A secure communication protocol (HTTPS) is used for transmission to ensure data security. Emotion data is also sent to the server.
[0354] The server analyzes the received image data using a deep learning framework (such as TensorFlow or YOLO) to identify important objects. Based on the analysis results, it classifies the user's behavior into categories (e.g., eating, exercise, travel). It also uses facial recognition technology to analyze the user's emotional state and adds it to the behavioral data.
[0355] For example, if an image of a user eating a salad is captured and joy is identified from the user's emotional state, the behavioral data is recorded in the life log as "Lunch (Joy)." The server analyzes this behavioral data and emotional data and generates lifestyle improvement suggestions. This analysis uses algorithms based on the user's past behavioral logs and expert knowledge.
[0356] The server then generates personalized meal suggestions, taking into account the user's emotional state, using a generative AI model (e.g., GPT) and prompting the user with the following sentences:
[0357] "The user is having salad, soup, and bread for lunch. Their facial expressions indicate they are enjoying the meal. Please suggest a meal for this user's next meal."
[0358] The generated improvement suggestions and meal suggestions are sent back to the device and notified to the user. Notification methods include displaying the information on the device's display or using audio output. This allows the user to receive their own life log, improvement suggestions, and personalized meal suggestions and reflect them in their next actions without performing any special operations.
[0359] For example, if a user eats a salad at lunch and the emotion of joy is detected, the device will notify them with a suggestion: "We recommend that you incorporate more vegetables into your next meal to improve balance and maintain your current eating pattern." In this way, users can receive practical advice based on their emotions and behavioral data.
[0360] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0361] Step 1:
[0362] The device uses the built-in camera to capture the user's field of view at one-second intervals. At this time, the camera API operates and generates image data obtained from the user's field of view. The input is raw data from the camera, and the output is the captured image file.
[0363] Step 2:
[0364] The device pre-processes the captured image to remove noise and improve clarity. This process uses an image processing library (OpenCV). The input is the captured image file, and the output is a pre-processed, clear image file.
[0365] Step 3:
[0366] The terminal sends the preprocessed image data to the server in batch format at specific time intervals. A secure communication protocol (HTTPS) is used for data transmission to ensure data security. The input is the preprocessed image data, and the output is the data transfer to the server.
[0367] Step 4:
[0368] The server analyzes the received image data using a deep learning framework (TensorFlow, YOLO) to identify important objects. This analysis step recognizes and classifies objects in the image (e.g., food, people, landmarks, etc.). The input is the transmitted image data, and the output is the analyzed object information.
[0369] Step 5:
[0370] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). It also analyzes the user's emotional state using facial recognition technology and adds it to the behavioral data. For example, the data is recorded as "lunch (joy)." The input is the analyzed object information and emotional data, and the output is behavioral data classified by category.
[0371] Step 6:
[0372] The server analyzes this behavioral and emotional data and generates lifestyle improvement suggestions. The algorithm utilizes the user's past behavioral logs and expert knowledge. The input is behavioral and emotional data categorized by category, and the output is the generated improvement suggestions.
[0373] Step 7:
[0374] The server generates personalized meal suggestions using a generative AI model (e.g., GPT) based on the user's emotional state. For example, you can get suggestions by entering a prompt like this:
[0375] "The user is having salad, soup, and bread for lunch. Their facial expressions indicate they are enjoying the meal. Please suggest a meal for this user's next meal."
[0376] The input is a prompt based on the user's behavioral and emotional data, and the output is personalized meal suggestions.
[0377] Step 8:
[0378] The server sends the generated improvement suggestions and meal suggestions to the terminal, which receives them and notifies the user. The input is the generated suggestions and notification data, and the output is the notification to the user.
[0379] Step 9:
[0380] The user receives notifications from the device, refers to their own life log, improvement suggestions, and personalized meal suggestions, and reflects them in their next actions. The input is the notification information from the device, and the output is the action selected by the user.
[0381] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0382] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0383] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0384] [Second embodiment]
[0385] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0386] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0387] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0388] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0389] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0390] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0391] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0392] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0393] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0394] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0395] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0396] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0397] This invention is a system that uses glasses-type devices equipped with a built-in camera and image processing functions to automatically record a user's daily activities, analyze those activities via a server, and generate lifestyle improvement suggestions.With this system, users can efficiently obtain life logs without performing any special operations and use them to improve their lifestyles.
[0398] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view in real time. This capture operation occurs every second, and the obtained image data is temporarily stored in the device's internal memory. After capturing the images, the device performs preprocessing on them to remove noise and improve clarity. This preprocessing uses image processing techniques such as Gaussian filters and sharpening algorithms.
[0399] Next, the device sends the pre-processed image data to the server in batch format at regular intervals (e.g., every hour). A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety.
[0400] The server inputs the image data received from the device into a generative AI model to analyze the image, using deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks).
[0401] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, it identifies "salad," "steak," and "rice" from an image captured during a meal and records it as "lunch." The server records this behavior data as text information in a life log and saves it in a database.
[0402] The server then analyzes the recorded life log and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You're not getting enough vegetables. Eat more vegetables next time."
[0403] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: salad, steak, rice. Eat more vegetables."
[0404] Users can receive this notification, refer to their life log and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lifestyle and manage their health.
[0405] In this way, the present invention realizes a new system that combines a built-in camera, image processing technology, and analysis algorithms to automatically record users' behavior and provide specific suggestions that will help improve their lifestyle.
[0406] The processing flow will be explained below.
[0407] Step 1:
[0408] The device (glasses-type device) uses a built-in camera to capture the user's field of view at one-second intervals.
[0409] Specifically, the camera automatically captures an image and temporarily stores it in its internal memory.
[0410] Step 2:
[0411] The device preprocesses the captured image to remove noise and improve image clarity.
[0412] Specifically, the image quality is improved by applying a Gaussian filter or sharpening algorithm.
[0413] Step 3:
[0414] The terminal converts the preprocessed image data into a batch format at regular time intervals (e.g., every hour).
[0415] Specifically, the image data is compressed into one file and prepared for transmission.
[0416] Step 4:
[0417] The terminal transmits the image data in a batch format to the server using a secure communication protocol (e.g., HTTPS).
[0418] Specifically, a session for transmitting image data is established, and data uploading begins.
[0419] Step 5:
[0420] The server inputs the image data received from the terminal into the generated AI model and analyzes the image.
[0421] Specifically, the AI model processes the received image data and identifies important objects (e.g., food, people, landmarks).
[0422] Step 6:
[0423] The server classifies the user's behavior into categories (e.g., eating, exercise, travel) based on the image analysis results.
[0424] Specifically, the analysis results are assigned to preset categories and organized according to each category.
[0425] Step 7:
[0426] The server records the behavioral data organized by category as text information in a life log.
[0427] Specifically, text information such as "2023-10-05 12:30 Lunch: salad, steak, rice" is saved in the database.
[0428] Step 8:
[0429] The server analyzes the recorded life log and generates suggestions for improving lifestyle habits.
[0430] Specifically, it compares and analyzes the user's past behavioral logs and generates specific advice such as "eat more vegetables."
[0431] Step 9:
[0432] The server sends the generated improvement suggestions and detailed life logs to the device.
[0433] Specifically, the system compresses the improvement suggestions and life log information and sends them to the device using a secure communication protocol.
[0434] Step 10:
[0435] The device notifies the user of the received suggestions and life logs so that they can be viewed.
[0436] Specifically, the device's display and voice notification functions are used to notify the user, displaying "Lunch record: salad, steak, rice. Eat more vegetables."
[0437] Step 11:
[0438] Users receive notifications from their devices and view life logs and improvement suggestions.
[0439] Specifically, the user checks the suggestions on the device, reviews their lifestyle habits based on them, and decides on their next course of action (e.g., eating more vegetables at the next meal).
[0440] Example 1
[0441] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0442] Conventional life log management systems have had problems with the effort required for users to record their own data and the accuracy of the records. In addition, the lifestyle improvement suggestions they offer are often not specialized enough, and their effectiveness in improving users' health is limited.
[0443] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0444] In this invention, the server includes means for capturing the field of view in real time using a built-in camera, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for inputting the images into a generative AI model to identify important objects, means for categorizing the user's behavior based on the analysis results, means for recording the behavioral data as text information in a life log, means for analyzing the recorded life log with an algorithm based on expert knowledge to generate lifestyle improvement suggestions, means for transmitting the generated suggestions and life log to a terminal, and means for notifying the user of the suggestions and life log. This allows the user to easily obtain a highly accurate life log and receive expert improvement suggestions.
[0445] A "built-in camera" is a camera built into a device that captures the external field of view in real time.
[0446] "Means for capturing the field of view in real time" is a function that uses the built-in camera to continuously capture images that are within the user's field of view.
[0447] The "means for pre-processing the captured image" is a function for performing processing on the obtained image data to remove noise and improve clarity.
[0448] The "means for transmitting to the server in batch format at specific time intervals" is a function for collecting preprocessed image data at regular time intervals and transmitting the collected data to the server using a secure communication protocol.
[0449] A "generative AI model" refers to an algorithm that uses artificial intelligence to analyze data and perform specific tasks.
[0450] "Means for identifying significant objects" refers to the ability to use generative AI models to recognize and analyze objects, people, landmarks, etc. in images.
[0451] "Means for categorizing user behavior" is a function that records user behavior into categories such as food, exercise, and travel based on the analysis results.
[0452] A "life log" is data that records a user's daily activities and behaviors in chronological order.
[0453] An "algorithm based on expert knowledge" is a computational method for performing analysis that incorporates specialized knowledge in a specific field.
[0454] The "means for generating lifestyle improvement suggestions" is a function that generates advice for improving the user's lifestyle using an algorithm based on the recorded life log and expert knowledge.
[0455] The "means for transmitting to the terminal" is a function for transferring the generated proposals and life logs to the terminal.
[0456] The "means of notifying the user of suggestions and life logs" is a function that notifies the user of the contents of suggestions and life logs through the device's built-in display or audio output.
[0457] MODE FOR CARRYING OUT THE INVENTION
[0458] This invention is a system that uses glasses-type devices equipped with a built-in camera and image processing functions to automatically record a user's daily activities, analyze those activities via a server, and generate lifestyle improvement suggestions.With this system, users can efficiently obtain life logs without performing any special operations and use them to improve their lifestyles.
[0459] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view in real time. This capture operation occurs every second, and the obtained image data is temporarily stored in the device's internal memory. After capturing the images, the device performs preprocessing on them to remove noise and improve clarity. This preprocessing uses image processing techniques such as Gaussian filters and sharpening algorithms.
[0460] Next, the device sends the pre-processed image data to the server in batch format at regular intervals (e.g., every hour). A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety.
[0461] The server inputs the image data received from the device into a generative AI model to analyze the image, using deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks).
[0462] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, it identifies "salad," "steak," and "rice" from an image captured during a meal and records it as "lunch." The server records this behavior data as text information in a life log and saves it in a database.
[0463] The server then analyzes the recorded life log and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You're not getting enough vegetables. Eat more vegetables next time."
[0464] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: salad, steak, rice. Eat more vegetables."
[0465] Users can receive this notification, refer to their life log and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lifestyle and manage their health.
[0466] An example of a specific prompt is, "Please identify whether there is food in this image and identify the specific type." Based on this prompt, the generative AI model recognizes and classifies the objects in the image.
[0467] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0468] Step 1:
[0469] The device uses a built-in camera to capture the user's field of view in real time.
[0470] Input data: user's view
[0471] How it works: The device captures images within its field of view every second and stores them in its internal memory, recording the user's daily activities as video.
[0472] Output data: Captured image data
[0473] Step 2:
[0474] The device pre-processes the captured image to remove noise and improve clarity.
[0475] Input data: Captured image data
[0476] What it does: The device uses a Gaussian filter to remove noise from the image and a sharpening algorithm to improve clarity, allowing it to accurately identify important objects during analysis.
[0477] Output data: Preprocessed image data
[0478] Step 3:
[0479] The terminal sends the preprocessed image data in batch format to the server at regular time intervals (e.g., every hour).
[0480] Input data: Preprocessed image data
[0481] Specific operation: The device periodically compiles image data into a batch format and sends it securely to the server using HTTPS.
[0482] Output data: batch image data sent to the server
[0483] Step 4:
[0484] The server inputs the received image data into the generated AI model and analyzes the image.
[0485] Input data: batch image data
[0486] How it works: The server uses a generative AI model (e.g., YOLO, RetinaNet, etc.) to analyze the input image and recognize and identify important objects (e.g., food, people, landmarks). For example, it can recognize food items such as "salad" and "steak" from an image of a meal.
[0487] Output data: Analysis results (recognized objects)
[0488] Step 5:
[0489] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel).
[0490] Input data: Analysis results (recognized objects)
[0491] Specific actions: Based on the recognized object information, actions are classified into categories such as "meal," "exercise," and "travel," and recorded as a life log. For example, if "salad," "steak," and "rice" are detected, they are classified into the category of "lunch."
[0492] Output data: Behavioral data categorized by category
[0493] Step 6:
[0494] The server records the behavioral data as text information in a life log and stores it in a database.
[0495] Input data: Categorized behavioral data
[0496] Specific operation: The classified behavioral data is converted into text format and saved in a database, thereby saving the user's past behavior as a history.
[0497] Output data: Saved life logs
[0498] Step 7:
[0499] The server analyzes the recorded life log and generates suggestions for improving lifestyle habits.
[0500] Input data: Saved life logs
[0501] Specific operation: Using an algorithm based on expert knowledge, the system generates lifestyle improvement suggestions based on the user's behavioral history. For example, it generates specific advice such as, "You're not getting enough vegetables. Next time, eat more vegetables."
[0502] Output data: Suggestions for improving lifestyle habits
[0503] Step 8:
[0504] The server then sends the generated improvement suggestions and detailed life logs back to the device.
[0505] Input data: Generated lifestyle improvement suggestions, detailed life log
[0506] Specific operation: A secure communication protocol is used to send the generated lifestyle improvement suggestions and detailed life logs to the device.
[0507] Output data: Improvement suggestions and life logs sent to the device
[0508] Step 9:
[0509] The device receives improvement suggestions and life logs and notifies the user.
[0510] Input data: improvement proposals, life logs
[0511] Specific operation: The device notifies the user of improvement suggestions and the contents of the life log via the device's built-in display and voice output. For example, it displays "Lunch record: salad, steak, rice. Eat more vegetables."
[0512] Output data: Improvement suggestions notified to the user and lifelog
[0513] Step 10:
[0514] Users receive notifications, refer to their own life logs and improvement suggestions, and reflect them in their next actions.
[0515] Input data: Notified improvement suggestions, lifelog
[0516] Specific actions: Refer to the notification and reflect it in your next actions. For example, take action to improve your lifestyle, such as consciously eating more vegetables at your next meal.
[0517] Output data: Actions that reflect improvement suggestions
[0518] (Application example 1)
[0519] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0520] Conventional lifestyle recording systems and in-store customer behavior analysis systems require users and customers to record and operate the systems themselves, which is time-consuming and makes it difficult to obtain accurate data.In addition, it is not possible to grasp in real time what products or areas in the store customers are interested in, which makes it difficult to provide efficient customer service and marketing.
[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0522] In this invention, the server includes means for capturing the field of view in real time using a built-in image capture device, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for analyzing the received image data to identify important objects, means for classifying the user's behavior based on the analysis results, means for saving the behavior data as text information in a lifestyle log, means for analyzing the recorded lifestyle log to generate lifestyle improvement suggestions, means for sending the generated suggestions and lifestyle log to a terminal, means for notifying the user of the suggestions and lifestyle log, means for measuring the level of customer interest in specific products or areas in the store, and means for notifying store managers and staff of the analysis results. This reduces the effort required for users and customers, enables accurate data acquisition and analysis, and enables efficient customer service and marketing.
[0523] An "integrated imaging device" is a device, such as a camera or image sensor, built into a device that captures the surrounding field of view in real time.
[0524] "Preprocessing" refers to a series of steps performed on captured image data to remove noise and improve image clarity.
[0525] "Batch format" is a method of processing and sending data for a certain period of time all at once.
[0526] "Important objects" are objects or areas identified by the analysis that should be of interest to the user or the system.
[0527] "Behavior" refers to the activity or pattern of a user or customer at a particular time.
[0528] "Classification" refers to organizing user behavior into categories based on the analysis results.
[0529] "Life Record" is a database that stores user behavior data as text information.
[0530] "Lifestyle improvement suggestions" are specific advice for improving your health and quality of life that is created by analyzing your recorded lifestyle records.
[0531] A "terminal" is a device worn by a user and has various functions including a built-in image capture device.
[0532] "Notification" refers to information transmission activities that inform users of generated suggestions and life logs via their devices.
[0533] A "customer" is a user who visits a store.
[0534] "In-store" means the physical premises where products are displayed and sold.
[0535] "Interested" means that a customer is taking actions that show interest, such as focusing on a particular product or area or staying there for a long time.
[0536] "Store Manager" means a person responsible for the operation and management of a store.
[0537] "Staff" refers to employees who deal with customers and manage merchandise within the store.
[0538] The system that realizes this application example is composed of a terminal equipped with a built-in image capture device, a server that uses a communication protocol, and a system that notifies users based on the analysis results. This system operates as follows.
[0539] The device captures the user's or customer's field of view in real time using a built-in image capture device, taking images every second. The captured images are then pre-processed to remove noise and improve image clarity, using OpenCV's Gaussian filter and sharpening algorithms, for example.
[0540] The preprocessed image data is sent to the server in batch format at regular intervals. This data transmission uses a secure communication protocol (e.g., HTTPS) to ensure data security. The server then analyzes the received image data and identifies important objects using a generative AI model (e.g., YOLO or RetinaNet). This analysis identifies user behavior categories and the products and areas in the store that customers are interested in.
[0541] The server classifies user behavior based on the analysis results and saves them as text information in a daily log. It also notifies store managers and staff in real time based on the analysis results of customer behavior in the store. Notifications are sent via smart glasses or smartphones, and specific alerts, such as "A customer is interested in the wine section," are displayed.
[0542] The recorded lifestyle records and behavioral data are further analyzed using deep learning algorithms to generate lifestyle improvement suggestions. For example, advice such as "You're not getting enough vegetables. Eat more vegetables next time" may be generated. The generated suggestions are resent to the device and notified to the user. This notification is sent via the smart glasses display or the smartphone notification function.
[0543] For example, if a customer stares at the wine section of a store for a long time, that information is sent to the server and analyzed as "the customer is interested in the wine section." As a result, sales staff are notified that "they should attend to the customer in the wine section," thereby enabling personalized customer service.
[0544] An example of a prompt for a generative AI model is as follows:
[0545] Suggest the best way to serve customers when they are interested in the wine section.
[0546] This system reduces the workload for users and stores, while enabling efficient and effective customer service and data-based marketing strategies.
[0547] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0548] Step 1:
[0549] The device captures the field of view in real time using a built-in image capture device. Specifically, the device acquires an image every second. At this point, the input is the captured image data, and the output is the image data before preprocessing.
[0550] Step 2:
[0551] The device preprocesses the captured image to remove noise and improve image clarity, using OpenCV's Gaussian filter and sharpening algorithms. The input to this process is the captured image data, and the output is image data with noise removed and improved clarity.
[0552] Step 3:
[0553] The terminal transmits the preprocessed image data to the server in batch format at regular intervals. Specifically, the data is stored in a fixed buffer and transmitted, for example, every hour. The input in this process is a plurality of preprocessed image data, and the output is batch data transmitted to the server.
[0554] Step 4:
[0555] The server analyzes the received image data. Specifically, it uses a generative AI model (e.g., YOLO or RetinaNet) to identify important objects in the image. The input to this process is a batch of image data, and the output is a list of objects as the analysis result.
[0556] Step 5:
[0557] The server categorizes the user's behavior based on the analysis results. Specifically, it sets categories such as eating, exercise, and shopping, and categorizes the data based on these. The input for this process is a list of objects, and the output is behavioral data categorized by category.
[0558] Step 6:
[0559] The server saves the behavioral data as text information in a life log. Specifically, it saves it in a database. The input for this process is behavioral data categorized by category, and the output is life log data as text information.
[0560] Step 7:
[0561] The server analyzes the recorded life log and generates lifestyle improvement suggestions. Specifically, it uses a deep learning algorithm to generate advice based on past logs and specialized knowledge. The input in this process is life log data, and the output is improvement suggestions.
[0562] Step 8:
[0563] The server transmits the generated suggestions and lifelog data to the device using a secure communication protocol. The inputs to this process are the improvement suggestions and lifelog data, and the output is the data transmitted to the device.
[0564] Step 9:
[0565] The device notifies the user of the suggestions and lifelogs by displaying a message on the display or by issuing a voice notification. The input to this process is the suggestions and lifelog data received from the server, and the output is the notification delivered to the user.
[0566] Step 10:
[0567] The server analyzes the level of interest customers have in specific products or areas in the store and notifies store managers and staff of the results. Specifically, it analyzes customer gaze data and notifies the results in real time. The input to this process is gaze data, and the output is a notification message as the analysis result.
[0568] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0569] This invention is a system that uses a glasses-type device with a built-in camera and image processing function to automatically record a user's daily activities, and further recognizes and analyzes the user's emotions using an emotion engine to generate lifestyle improvement suggestions.With this system, users can efficiently collect life logs and receive advice on lifestyle improvements based on their emotions without performing any special operations.
[0570] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view at one-second intervals. This capture operation is performed automatically, and the obtained image data is temporarily stored in the device's internal memory. At the same time, the user's face is also periodically captured by the camera and analyzed by the emotion engine.
[0571] The device then pre-processes the captured image to remove noise and improve clarity, using image processing techniques such as Gaussian filters and sharpening algorithms.
[0572] The device sends the preprocessed image data in batches at regular intervals (e.g., every hour) to the server. A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety. Emotion data is also sent to the server in the same way.
[0573] The server inputs the image data received from the device into a generative AI model to analyze the image. This image analysis utilizes deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks). The emotion engine also analyzes the user's face to identify their emotional state (e.g., joy, sadness, anger, etc.).
[0574] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, from an image captured during a meal, "salad," "steak," and "rice" are identified and recorded as "lunch." The server also uses facial recognition technology to classify the user's emotions and adds them to the behavioral data. For example, if "happiness" is detected on the user's face during lunch, it is recorded as "lunch (happiness)." The server records this behavioral data as text information in a life log and saves it in a database.
[0575] The server then analyzes the recorded life log and emotional data to generate lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You haven't been eating enough vegetables, so I recommend you eat more next time. Also, your recent meals have been associated with pleasure, so maintain this pattern to reduce stress."
[0576] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern."
[0577] Users can receive notifications, refer to their life logs and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lives and manage their health based on their emotions.
[0578] In this way, the present invention combines a built-in camera, image processing technology, and an emotion engine to realize a new system that automatically records a user's behavior and emotions and provides specific lifestyle improvement suggestions based on that information.
[0579] The processing flow will be explained below.
[0580] Step 1:
[0581] The device (glasses-type device) uses a built-in camera to capture the user's field of view at one-second intervals.
[0582] Specifically, the camera automatically captures an image and temporarily stores it in its internal memory.
[0583] Step 2:
[0584] The device periodically captures the user's face using a camera and analyzes it using an emotion engine.
[0585] Specifically, it uses a facial recognition algorithm to identify the user's face and extract their emotional state (e.g., joy, sadness, anger, etc.).
[0586] Step 3:
[0587] The device pre-processes the captured field of view image to remove noise and improve image clarity.
[0588] Specifically, the image quality is improved by applying a Gaussian filter or sharpening algorithm.
[0589] Step 4:
[0590] The terminal collectively converts the preprocessed image data and emotion data into a batch format at regular time intervals (e.g., every hour).
[0591] Specifically, the image data and emotion data are compressed into a single file and preparations for transmission are made.
[0592] Step 5:
[0593] The terminal transmits the batched data to the server using a secure communication protocol (e.g., HTTPS).
[0594] Specifically, the operation is to establish a session for data transmission and start uploading data.
[0595] Step 6:
[0596] The server inputs the data received from the device into the generative AI model and analyzes the field of view image.
[0597] Specifically, the AI model processes the received image data and identifies important objects (e.g., food, people, landmarks).
[0598] Step 7:
[0599] The server associates the emotion data analyzed by the emotion engine with the behavioral data.
[0600] Specifically, behavioral data is tagged with emotion tags such as "joy" or "sadness" and stored in a database.
[0601] Step 8:
[0602] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel).
[0603] Specifically, the analysis results are assigned to preset categories and organized according to each category.
[0604] Step 9:
[0605] The server records behavioral data and emotional data organized by category as text information in a life log.
[0606] Specifically, the system stores text information such as "2023-10-05 12:30 Lunch: Salad, steak, rice (joy)" in a database.
[0607] Step 10:
[0608] The server analyzes the recorded life log and emotional data and generates suggestions for improving lifestyle habits.
[0609] Specifically, it generates specific advice such as "Eat more vegetables and maintain your current eating patterns" based on the user's past behavioral logs and emotional data.
[0610] Step 11:
[0611] The server sends the generated improvement suggestions and detailed life logs to the device.
[0612] Specifically, the system compresses the improvement suggestions and life log information and sends them to the device using a secure communication protocol.
[0613] Step 12:
[0614] The device notifies the user of the received suggestions and life logs so that they can be viewed.
[0615] Specifically, the device's display and voice notification function will notify the user: "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern."
[0616] Step 13:
[0617] Users receive notifications from their devices and view life logs and improvement suggestions.
[0618] Specifically, the user checks the suggestions on the device and decides on the next action based on them (e.g., eating more vegetables at the next meal).
[0619] Example 2
[0620] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0621] Conventional lifelog acquisition systems require manual user operation, making it difficult to efficiently record detailed daily activities and analyze emotional states. It is also difficult to generate specific lifestyle improvement suggestions based on acquired lifelog data. Furthermore, data security and privacy protection are insufficient, creating a need for a system that users can use with confidence.
[0622] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0623] In this invention, the server includes: a means for capturing the field of view in real time using a built-in camera; a means for preprocessing the captured images to remove noise; a means for sending the preprocessed images to the server in batch format at specific time intervals; a means for analyzing the received image data to identify important objects; a means for categorizing the user's behavior based on the analysis results; a means for analyzing the user's emotions using facial recognition technology; a means for recording the behavioral data and emotional data as text information in a life log; a means for analyzing the recorded life log and generating lifestyle improvement suggestions; a means for transmitting the generated suggestions and life log to a terminal; and a means for notifying the user of the suggestions and life log. This allows the user to obtain a detailed life log without performing any special operations and receive specific lifestyle improvement suggestions based on their emotions. Furthermore, the use of a secure communication protocol ensures data security and privacy protection.
[0624] The "built-in camera" is a camera built into the eyeglass-type terminal, and is a device that has the function of capturing the user's field of view in real time.
[0625] "Pre-processing" refers to a general range of image processing techniques used to remove noise and improve clarity of captured image data, including Gaussian filters and sharpening algorithms.
[0626] "Batch format" refers to a method of processing data collectively at regular intervals, and is used to efficiently transmit and process large amounts of data.
[0627] A "secure communication protocol" is a communication protocol used to ensure data security and privacy, such as HTTPS.
[0628] "Image analysis" refers to the process of using generative AI models and deep learning algorithms to recognize and identify important objects in captured image data.
[0629] "Emotion analysis" refers to the process of using facial recognition technology to identify emotional states (e.g., joy, sadness, anger, etc.) from a user's face.
[0630] "Behavior categorization" refers to the process of classifying user behavior into multiple categories (e.g., eating, exercise, travel) based on the analysis results.
[0631] A "life log" refers to a collection of data that records a user's daily activities and emotional data as text information.
[0632] "Lifestyle improvement suggestions" refer to suggestions that include specific advice for improving the user's lifestyle based on the recorded life log and emotional data.
[0633] "Notification" refers to a method of informing the user of the generated suggestions and details of the life log, and is done through display or audio output.
[0634] The present invention is a system that uses a glasses-type device equipped with a built-in camera and image processing functions to automatically record a user's daily activities, and further recognizes and analyzes the user's emotions using an emotion engine, thereby generating lifestyle improvement suggestions. Detailed embodiments of the present invention will be described below.
[0635] Hardware and software used
[0636] Glasses-type device: Equipped with a built-in camera, internal memory, display, and audio output functions, it captures and temporarily stores data.
[0637] Server: Runs generative AI models, emotion engines, and databases for data analysis and storage. Performs data analysis and proposal generation.
[0638] Communication protocol: HTTPS is used to ensure secure communication.
[0639] Program processing
[0640] The device automatically captures the user's field of view using the built-in camera at one-second intervals, and the user's face is also captured periodically and analyzed by the emotion engine.
[0641] The device pre-processes the captured image data to remove noise and improve clarity, using image processing techniques such as Gaussian filters and sharpening algorithms.
[0642] The terminal transmits the preprocessed image data and emotion data to the server in batch format at regular time intervals using a secure communication protocol (HTTPS).
[0643] The server inputs the received image data into a generative AI model (e.g., YOLO, RetinaNet) for image analysis. This image analysis recognizes and identifies important objects (e.g., food, people, landmarks). The emotion engine also analyzes the user's face to identify their emotional state (e.g., joy, sadness, anger, etc.).
[0644] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, from an image captured during a meal, "salad," "steak," and "rice" are identified and recorded as "lunch." Furthermore, facial recognition technology is used to classify the user's emotions and add them to the behavioral data. For example, if "happiness" is detected on the user's face during lunch, it is categorized as "lunch (happiness)."
[0645] The server analyzes the recorded life log and emotional data and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, the server might generate advice such as, "You're not eating enough vegetables, so we recommend you eat more next time. Also, your recent meals have been associated with pleasure, so maintain this pattern to reduce stress."
[0646] Examples of concrete examples and prompts
[0647] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or voice output. For example, the device's display might say, "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern," and a voice output might say, "Improvement suggestions have been received."
[0648] Prompt Sentence Examples
[0649] "Show me your lunch log and provide relevant sentiment data and improvement suggestions."
[0650] The present invention allows users to effortlessly improve their lifestyle and manage their health based on their emotions. In this way, the present invention realizes a new system that combines a built-in camera, image processing technology, and an emotion engine to automatically record a user's behavior and emotions and provide specific lifestyle improvement suggestions based on the recorded information.
[0651] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0652] Step 1:
[0653] Data Capture
[0654] The device uses a built-in camera to capture the user's field of view at one-second intervals. The captured image data is temporarily stored in the internal memory. At the same time, the device periodically captures the user's face, which is then analyzed by the emotion engine.
[0655] Input: Images of the user's field of view and face
[0656] Output: Temporarily saved captured image data and facial image data
[0657] Specific operation: The device will start the camera, a shutter sound will be heard every second, and the field of view and face will be captured. These image data will be stored in the internal memory.
[0658] Step 2:
[0659] Image preprocessing
[0660] The device pre-processes the captured image data to remove noise and improve clarity, using Gaussian filters and sharpening algorithms.
[0661] Input: Temporarily saved captured image data
[0662] Output: Preprocessed image data
[0663] What happens: The device's image processing algorithm runs, and the progress is displayed on the screen.
[0664] Step 3:
[0665] Data transmission
[0666] The device sends the preprocessed image data and emotion data to the server in batch format at regular intervals using a secure communication protocol (HTTPS).
[0667] Input: Preprocessed image data and emotion data
[0668] Output: The dataset sent to the server
[0669] Specific operation: While data transmission is in progress, the display will show "Sending data...", and once transmission is complete, the display will notify you that "Data transmitted successfully."
[0670] Step 4:
[0671] Data analysis
[0672] The server inputs the received image data into a generative AI model (e.g., YOLO, RetinaNet) to recognize and identify important objects, and also uses an emotion engine to analyze emotions from the facial image data.
[0673] Input: Dataset sent to the server (image data and emotion data)
[0674] Output: Analysis results (object recognition and emotional state)
[0675] Specific operation: The server's processing unit performs image analysis and emotion analysis, and displays the progress on the operation console. After the analysis is complete, a list of recognized objects and emotional states is displayed.
[0676] Step 5:
[0677] Classification of behaviors and emotions
[0678] Based on the analysis results, the server categorizes the user's behavior (e.g., eating, exercise, travel) and also associates emotional data with the behavior.
[0679] Input: Analysis results (object recognition and emotional state)
[0680] Output: Classified behavioral and emotional data
[0681] Specific operation: The server stores the action and emotion classification results in a database in the form of, for example, "Lunch (salad, steak, rice, joy)".
[0682] Step 6:
[0683] Generate improvement suggestions
[0684] The server analyzes the recorded life log and emotional data and generates lifestyle improvement suggestions using an algorithm based on past behavioral logs and expert knowledge.
[0685] Input: Classified behavioral and emotional data
[0686] Output: Generated lifestyle improvement suggestions
[0687] Specific operation: The server's algorithm runs and generates improvement suggestions. The console screen displays a list of the generated suggestions.
[0688] Step 7:
[0689] User Notification
[0690] The server sends the generated improvement suggestions and detailed life logs to the device, which receives them and notifies the user through a display or audio output.
[0691] Input: Generated improvement suggestions and detailed lifelog
[0692] Output: Suggestions and lifelog information notified to the user
[0693] Specific operation: The device display will show "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern," and a voice output will announce, "Improvement suggestions have been received."
[0694] (Application example 2)
[0695] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0696] In modern society, it is difficult for users to easily collect life logs to maintain their health and receive lifestyle improvement suggestions based on those logs in their busy daily lives. It is also difficult to provide personalized dietary suggestions that take the user's emotions into account. Conventional systems require users to manually enter data, which is time-consuming and makes it difficult to collect accurate data. To solve these issues, a system is needed that utilizes a built-in camera, image processing technology, and an emotion engine to automatically collect life logs and provide dietary suggestions based on emotions.
[0697] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the field of view in real time using a built-in camera, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for analyzing the received image data to identify important objects, means for categorizing the user's behavior based on the analysis results, means for recording the behavioral data as text information in a life log, means for analyzing the recorded life log to generate lifestyle improvement suggestions, means for transmitting the generated suggestions and life log to the terminal, means for notifying the user of the suggestions and life log, means for analyzing the user's emotional state and generating personalized meal suggestions based on the emotional state, and means for notifying the user of the generated meal suggestions. This allows the user to receive personalized lifestyle improvement suggestions and meal suggestions based on their emotions without performing any special operations.
[0698] The "built-in camera" is a camera device built into the eyeglass-type terminal to capture the user's field of view in real time.
[0699] "Preprocessing" refers to initial image processing operations to remove noise and improve clarity from a captured image.
[0700] The "batch format" is a data transfer method in which preprocessed image data is sent in batches at regular time intervals.
[0701] A "server" is a remote computer system that receives, analyzes, and stores data sent from user terminals, and generates suggestions based on the results of the analysis.
[0702] "Analysis" is the process of identifying significant objects from the received image data and discerning the user's behavior and emotional state.
[0703] The "behavior category" is a category for classifying the user's daily behavior into specific activity types based on the analysis results.
[0704] A "life log" is a digital history that records and saves a user's daily behavioral and emotional data as text information.
[0705] "Improvement suggestions" are specific advice for improving the user's lifestyle habits that are generated from the analysis results of the life log.
[0706] "Emotional state" refers to the type of emotion (e.g., joy, sadness, anger) analyzed from the user's facial image.
[0707] "Meal Suggestions" are personalized meal suggestions generated based on the user's emotional state and past meal data.
[0708] "Notification" is the process of sending messages to communicate generated suggestions and lifelogs to users.
[0709] The system of the present invention includes a camera built into an eyeglass-type terminal, a server, and a user's terminal (a smartphone or smart glasses).
[0710] First, the device's built-in camera captures the user's field of view in real time, capturing images every second. These captured images are then pre-processed using the device's image processing capabilities. This pre-processing uses image processing techniques such as the OpenCV library to remove noise and improve image clarity.
[0711] The preprocessed image data is sent to the server in batches at regular intervals. A secure communication protocol (HTTPS) is used for transmission to ensure data security. Emotion data is also sent to the server.
[0712] The server analyzes the received image data using a deep learning framework (such as TensorFlow or YOLO) to identify important objects. Based on the analysis results, it classifies the user's behavior into categories (e.g., eating, exercise, travel). It also uses facial recognition technology to analyze the user's emotional state and adds it to the behavioral data.
[0713] For example, if an image of a user eating a salad is captured and joy is identified from the user's emotional state, the behavioral data is recorded in the life log as "Lunch (Joy)." The server analyzes this behavioral data and emotional data and generates lifestyle improvement suggestions. This analysis uses algorithms based on the user's past behavioral logs and expert knowledge.
[0714] The server then generates personalized meal suggestions, taking into account the user's emotional state, using a generative AI model (e.g., GPT) and prompting the user with the following sentences:
[0715] "The user is having salad, soup, and bread for lunch. Their facial expressions indicate they are enjoying the meal. Please suggest a meal for this user's next meal."
[0716] The generated improvement suggestions and meal suggestions are sent back to the device and notified to the user. Notification methods include displaying the information on the device's display or using audio output. This allows the user to receive their own life log, improvement suggestions, and personalized meal suggestions and reflect them in their next actions without performing any special operations.
[0717] For example, if a user eats a salad at lunch and the emotion of joy is detected, the device will notify them with a suggestion: "We recommend that you incorporate more vegetables into your next meal to improve balance and maintain your current eating pattern." In this way, users can receive practical advice based on their emotions and behavioral data.
[0718] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0719] Step 1:
[0720] The device uses the built-in camera to capture the user's field of view at one-second intervals. At this time, the camera API operates and generates image data obtained from the user's field of view. The input is raw data from the camera, and the output is the captured image file.
[0721] Step 2:
[0722] The device pre-processes the captured image to remove noise and improve clarity. This process uses an image processing library (OpenCV). The input is the captured image file, and the output is a pre-processed, clear image file.
[0723] Step 3:
[0724] The terminal sends the preprocessed image data to the server in batch format at specific time intervals. A secure communication protocol (HTTPS) is used for data transmission to ensure data security. The input is the preprocessed image data, and the output is the data transfer to the server.
[0725] Step 4:
[0726] The server analyzes the received image data using a deep learning framework (TensorFlow, YOLO) to identify important objects. This analysis step recognizes and classifies objects in the image (e.g., food, people, landmarks, etc.). The input is the transmitted image data, and the output is the analyzed object information.
[0727] Step 5:
[0728] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). It also analyzes the user's emotional state using facial recognition technology and adds it to the behavioral data. For example, the data is recorded as "lunch (joy)." The input is the analyzed object information and emotional data, and the output is behavioral data classified by category.
[0729] Step 6:
[0730] The server analyzes this behavioral and emotional data and generates lifestyle improvement suggestions. The algorithm utilizes the user's past behavioral logs and expert knowledge. The input is behavioral and emotional data categorized by category, and the output is the generated improvement suggestions.
[0731] Step 7:
[0732] The server generates personalized meal suggestions using a generative AI model (e.g., GPT) based on the user's emotional state. For example, you can get suggestions by entering a prompt like this:
[0733] "The user is having salad, soup, and bread for lunch. Their facial expressions indicate they are enjoying the meal. Please suggest a meal for this user's next meal."
[0734] The input is a prompt based on the user's behavioral and emotional data, and the output is personalized meal suggestions.
[0735] Step 8:
[0736] The server sends the generated improvement suggestions and meal suggestions to the terminal, which receives them and notifies the user. The input is the generated suggestions and notification data, and the output is the notification to the user.
[0737] Step 9:
[0738] The user receives notifications from the device, refers to their own life log, improvement suggestions, and personalized meal suggestions, and reflects them in their next actions. The input is the notification information from the device, and the output is the action selected by the user.
[0739] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0740] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0741] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0742] [Third embodiment]
[0743] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0744] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0745] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0746] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0747] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0748] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0749] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0750] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0751] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0752] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0753] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0754] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0755] This invention is a system that uses glasses-type devices equipped with a built-in camera and image processing functions to automatically record a user's daily activities, analyze those activities via a server, and generate lifestyle improvement suggestions.With this system, users can efficiently obtain life logs without performing any special operations and use them to improve their lifestyles.
[0756] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view in real time. This capture operation occurs every second, and the obtained image data is temporarily stored in the device's internal memory. After capturing the images, the device performs preprocessing on them to remove noise and improve clarity. This preprocessing uses image processing techniques such as Gaussian filters and sharpening algorithms.
[0757] Next, the device sends the pre-processed image data to the server in batch format at regular intervals (e.g., every hour). A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety.
[0758] The server inputs the image data received from the device into a generative AI model to analyze the image, using deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks).
[0759] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, it identifies "salad," "steak," and "rice" from an image captured during a meal and records it as "lunch." The server records this behavior data as text information in a life log and saves it in a database.
[0760] The server then analyzes the recorded life log and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You're not getting enough vegetables. Eat more vegetables next time."
[0761] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: salad, steak, rice. Eat more vegetables."
[0762] Users can receive this notification, refer to their life log and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lifestyle and manage their health.
[0763] In this way, the present invention realizes a new system that combines a built-in camera, image processing technology, and analysis algorithms to automatically record users' behavior and provide specific suggestions that will help improve their lifestyle.
[0764] The processing flow will be explained below.
[0765] Step 1:
[0766] The device (glasses-type device) uses a built-in camera to capture the user's field of view at one-second intervals.
[0767] Specifically, the camera automatically captures an image and temporarily stores it in its internal memory.
[0768] Step 2:
[0769] The device preprocesses the captured image to remove noise and improve image clarity.
[0770] Specifically, the image quality is improved by applying a Gaussian filter or sharpening algorithm.
[0771] Step 3:
[0772] The terminal converts the preprocessed image data into a batch format at regular time intervals (e.g., every hour).
[0773] Specifically, the image data is compressed into one file and prepared for transmission.
[0774] Step 4:
[0775] The terminal transmits the image data in a batch format to the server using a secure communication protocol (e.g., HTTPS).
[0776] Specifically, a session for transmitting image data is established, and data uploading begins.
[0777] Step 5:
[0778] The server inputs the image data received from the terminal into the generated AI model and analyzes the image.
[0779] Specifically, the AI model processes the received image data and identifies important objects (e.g., food, people, landmarks).
[0780] Step 6:
[0781] The server classifies the user's behavior into categories (e.g., eating, exercise, travel) based on the image analysis results.
[0782] Specifically, the analysis results are assigned to preset categories and organized according to each category.
[0783] Step 7:
[0784] The server records the behavioral data organized by category as text information in a life log.
[0785] Specifically, text information such as "2023-10-05 12:30 Lunch: salad, steak, rice" is saved in the database.
[0786] Step 8:
[0787] The server analyzes the recorded life log and generates suggestions for improving lifestyle habits.
[0788] Specifically, it compares and analyzes the user's past behavioral logs and generates specific advice such as "eat more vegetables."
[0789] Step 9:
[0790] The server sends the generated improvement suggestions and detailed life logs to the device.
[0791] Specifically, the system compresses the improvement suggestions and life log information and sends them to the device using a secure communication protocol.
[0792] Step 10:
[0793] The device notifies the user of the received suggestions and life logs so that they can be viewed.
[0794] Specifically, the device's display and voice notification functions are used to notify the user, displaying "Lunch record: salad, steak, rice. Eat more vegetables."
[0795] Step 11:
[0796] Users receive notifications from their devices and view life logs and improvement suggestions.
[0797] Specifically, the user checks the suggestions on the device, reviews their lifestyle habits based on them, and decides on their next course of action (e.g., eating more vegetables at the next meal).
[0798] Example 1
[0799] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0800] Conventional life log management systems have had problems with the effort required for users to record their own data and the accuracy of the records. In addition, the lifestyle improvement suggestions they offer are often not specialized enough, and their effectiveness in improving users' health is limited.
[0801] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0802] In this invention, the server includes means for capturing the field of view in real time using a built-in camera, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for inputting the images into a generative AI model to identify important objects, means for categorizing the user's behavior based on the analysis results, means for recording the behavioral data as text information in a life log, means for analyzing the recorded life log with an algorithm based on expert knowledge to generate lifestyle improvement suggestions, means for transmitting the generated suggestions and life log to a terminal, and means for notifying the user of the suggestions and life log. This allows the user to easily obtain a highly accurate life log and receive expert improvement suggestions.
[0803] A "built-in camera" is a camera built into a device that captures the external field of view in real time.
[0804] "Means for capturing the field of view in real time" is a function that uses the built-in camera to continuously capture images that are within the user's field of view.
[0805] The "means for pre-processing the captured image" is a function for performing processing on the obtained image data to remove noise and improve clarity.
[0806] The "means for transmitting to the server in batch format at specific time intervals" is a function for collecting preprocessed image data at regular time intervals and transmitting the collected data to the server using a secure communication protocol.
[0807] A "generative AI model" refers to an algorithm that uses artificial intelligence to analyze data and perform specific tasks.
[0808] "Means for identifying significant objects" refers to the ability to use generative AI models to recognize and analyze objects, people, landmarks, etc. in images.
[0809] "Means for categorizing user behavior" is a function that records user behavior into categories such as food, exercise, and travel based on the analysis results.
[0810] A "life log" is data that records a user's daily activities and behaviors in chronological order.
[0811] An "algorithm based on expert knowledge" is a computational method for performing analysis that incorporates specialized knowledge in a specific field.
[0812] The "means for generating lifestyle improvement suggestions" is a function that generates advice for improving the user's lifestyle using an algorithm based on the recorded life log and expert knowledge.
[0813] The "means for transmitting to the terminal" is a function for transferring the generated proposals and life logs to the terminal.
[0814] The "means of notifying the user of suggestions and life logs" is a function that notifies the user of the contents of suggestions and life logs through the device's built-in display or audio output.
[0815] MODE FOR CARRYING OUT THE INVENTION
[0816] This invention is a system that uses glasses-type devices equipped with a built-in camera and image processing functions to automatically record a user's daily activities, analyze those activities via a server, and generate lifestyle improvement suggestions.With this system, users can efficiently obtain life logs without performing any special operations and use them to improve their lifestyles.
[0817] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view in real time. This capture operation occurs every second, and the obtained image data is temporarily stored in the device's internal memory. After capturing the images, the device performs preprocessing on them to remove noise and improve clarity. This preprocessing uses image processing techniques such as Gaussian filters and sharpening algorithms.
[0818] Next, the device sends the pre-processed image data to the server in batch format at regular intervals (e.g., every hour). A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety.
[0819] The server inputs the image data received from the device into a generative AI model to analyze the image, using deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks).
[0820] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, it identifies "salad," "steak," and "rice" from an image captured during a meal and records it as "lunch." The server records this behavior data as text information in a life log and saves it in a database.
[0821] The server then analyzes the recorded life log and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You're not getting enough vegetables. Eat more vegetables next time."
[0822] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: salad, steak, rice. Eat more vegetables."
[0823] Users can receive this notification, refer to their life log and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lifestyle and manage their health.
[0824] An example of a specific prompt is, "Please identify whether there is food in this image and identify the specific type." Based on this prompt, the generative AI model recognizes and classifies the objects in the image.
[0825] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0826] Step 1:
[0827] The device uses a built-in camera to capture the user's field of view in real time.
[0828] Input data: user's view
[0829] How it works: The device captures images within its field of view every second and stores them in its internal memory, recording the user's daily activities as video.
[0830] Output data: Captured image data
[0831] Step 2:
[0832] The device pre-processes the captured image to remove noise and improve clarity.
[0833] Input data: Captured image data
[0834] What it does: The device uses a Gaussian filter to remove noise from the image and a sharpening algorithm to improve clarity, allowing it to accurately identify important objects during analysis.
[0835] Output data: Preprocessed image data
[0836] Step 3:
[0837] The terminal sends the preprocessed image data in batch format to the server at regular time intervals (e.g., every hour).
[0838] Input data: Preprocessed image data
[0839] Specific operation: The device periodically compiles image data into a batch format and sends it securely to the server using HTTPS.
[0840] Output data: batch image data sent to the server
[0841] Step 4:
[0842] The server inputs the received image data into the generated AI model and analyzes the image.
[0843] Input data: batch image data
[0844] How it works: The server uses a generative AI model (e.g., YOLO, RetinaNet, etc.) to analyze the input image and recognize and identify important objects (e.g., food, people, landmarks). For example, it can recognize food items such as "salad" and "steak" from an image of a meal.
[0845] Output data: Analysis results (recognized objects)
[0846] Step 5:
[0847] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel).
[0848] Input data: Analysis results (recognized objects)
[0849] Specific actions: Based on the recognized object information, actions are classified into categories such as "meal," "exercise," and "travel," and recorded as a life log. For example, if "salad," "steak," and "rice" are detected, they are classified into the category of "lunch."
[0850] Output data: Behavioral data categorized by category
[0851] Step 6:
[0852] The server records the behavioral data as text information in a life log and stores it in a database.
[0853] Input data: Categorized behavioral data
[0854] Specific operation: The classified behavioral data is converted into text format and saved in a database, thereby saving the user's past behavior as a history.
[0855] Output data: Saved life logs
[0856] Step 7:
[0857] The server analyzes the recorded life log and generates suggestions for improving lifestyle habits.
[0858] Input data: Saved life logs
[0859] Specific operation: Using an algorithm based on expert knowledge, the system generates lifestyle improvement suggestions based on the user's behavioral history. For example, it generates specific advice such as, "You're not getting enough vegetables. Next time, eat more vegetables."
[0860] Output data: Suggestions for improving lifestyle habits
[0861] Step 8:
[0862] The server then sends the generated improvement suggestions and detailed life logs back to the device.
[0863] Input data: Generated lifestyle improvement suggestions, detailed life log
[0864] Specific operation: A secure communication protocol is used to send the generated lifestyle improvement suggestions and detailed life logs to the device.
[0865] Output data: Improvement suggestions and life logs sent to the device
[0866] Step 9:
[0867] The device receives improvement suggestions and life logs and notifies the user.
[0868] Input data: improvement proposals, life logs
[0869] Specific operation: The device notifies the user of improvement suggestions and the contents of the life log via the device's built-in display and voice output. For example, it displays "Lunch record: salad, steak, rice. Eat more vegetables."
[0870] Output data: Improvement suggestions notified to the user and lifelog
[0871] Step 10:
[0872] Users receive notifications, refer to their own life logs and improvement suggestions, and reflect them in their next actions.
[0873] Input data: Notified improvement suggestions, lifelog
[0874] Specific actions: Refer to the notification and reflect it in your next actions. For example, take action to improve your lifestyle, such as consciously eating more vegetables at your next meal.
[0875] Output data: Actions that reflect improvement suggestions
[0876] (Application example 1)
[0877] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0878] Conventional lifestyle recording systems and in-store customer behavior analysis systems require users and customers to record and operate the systems themselves, which is time-consuming and makes it difficult to obtain accurate data.In addition, it is not possible to grasp in real time what products or areas in the store customers are interested in, which makes it difficult to provide efficient customer service and marketing.
[0879] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0880] In this invention, the server includes means for capturing the field of view in real time using a built-in image capture device, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for analyzing the received image data to identify important objects, means for classifying the user's behavior based on the analysis results, means for saving the behavior data as text information in a lifestyle log, means for analyzing the recorded lifestyle log to generate lifestyle improvement suggestions, means for sending the generated suggestions and lifestyle log to a terminal, means for notifying the user of the suggestions and lifestyle log, means for measuring the level of customer interest in specific products or areas in the store, and means for notifying store managers and staff of the analysis results. This reduces the effort required for users and customers, enables accurate data acquisition and analysis, and enables efficient customer service and marketing.
[0881] An "integrated imaging device" is a device, such as a camera or image sensor, built into a device that captures the surrounding field of view in real time.
[0882] "Preprocessing" refers to a series of steps performed on captured image data to remove noise and improve image clarity.
[0883] "Batch format" is a method of processing and sending data for a certain period of time all at once.
[0884] "Important objects" are objects or areas identified by the analysis that should be of interest to the user or the system.
[0885] "Behavior" refers to the activity or pattern of a user or customer at a particular time.
[0886] "Classification" refers to organizing user behavior into categories based on the analysis results.
[0887] "Life Record" is a database that stores user behavior data as text information.
[0888] "Lifestyle improvement suggestions" are specific advice for improving your health and quality of life that is created by analyzing your recorded lifestyle records.
[0889] A "terminal" is a device worn by a user and has various functions including a built-in image capture device.
[0890] "Notification" refers to information transmission activities that inform users of generated suggestions and life logs via their devices.
[0891] A "customer" is a user who visits a store.
[0892] "In-store" means the physical premises where products are displayed and sold.
[0893] "Interested" means that a customer is taking actions that show interest, such as focusing on a particular product or area or staying there for a long time.
[0894] "Store Manager" means a person responsible for the operation and management of a store.
[0895] "Staff" refers to employees who deal with customers and manage merchandise within the store.
[0896] The system that realizes this application example is composed of a terminal equipped with a built-in image capture device, a server that uses a communication protocol, and a system that notifies users based on the analysis results. This system operates as follows.
[0897] The device captures the user's or customer's field of view in real time using a built-in image capture device, taking images every second. The captured images are then pre-processed to remove noise and improve image clarity, using OpenCV's Gaussian filter and sharpening algorithms, for example.
[0898] The preprocessed image data is sent to the server in batch format at regular intervals. This data transmission uses a secure communication protocol (e.g., HTTPS) to ensure data security. The server then analyzes the received image data and identifies important objects using a generative AI model (e.g., YOLO or RetinaNet). This analysis identifies user behavior categories and the products and areas in the store that customers are interested in.
[0899] The server classifies user behavior based on the analysis results and saves them as text information in a daily log. It also notifies store managers and staff in real time based on the analysis results of customer behavior in the store. Notifications are sent via smart glasses or smartphones, and specific alerts, such as "A customer is interested in the wine section," are displayed.
[0900] The recorded lifestyle records and behavioral data are further analyzed using deep learning algorithms to generate lifestyle improvement suggestions. For example, advice such as "You're not getting enough vegetables. Eat more vegetables next time" may be generated. The generated suggestions are resent to the device and notified to the user. This notification is sent via the smart glasses display or the smartphone notification function.
[0901] For example, if a customer stares at the wine section of a store for a long time, that information is sent to the server and analyzed as "the customer is interested in the wine section." As a result, sales staff are notified that "they should attend to the customer in the wine section," thereby enabling personalized customer service.
[0902] An example of a prompt for a generative AI model is as follows:
[0903] Suggest the best way to serve customers when they are interested in the wine section.
[0904] This system reduces the workload for users and stores, while enabling efficient and effective customer service and data-based marketing strategies.
[0905] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0906] Step 1:
[0907] The device captures the field of view in real time using a built-in image capture device. Specifically, the device acquires an image every second. At this point, the input is the captured image data, and the output is the image data before preprocessing.
[0908] Step 2:
[0909] The device preprocesses the captured image to remove noise and improve image clarity, using OpenCV's Gaussian filter and sharpening algorithms. The input to this process is the captured image data, and the output is image data with noise removed and improved clarity.
[0910] Step 3:
[0911] The terminal transmits the preprocessed image data to the server in batch format at regular intervals. Specifically, the data is stored in a fixed buffer and transmitted, for example, every hour. The input in this process is a plurality of preprocessed image data, and the output is batch data transmitted to the server.
[0912] Step 4:
[0913] The server analyzes the received image data. Specifically, it uses a generative AI model (e.g., YOLO or RetinaNet) to identify important objects in the image. The input to this process is a batch of image data, and the output is a list of objects as the analysis result.
[0914] Step 5:
[0915] The server categorizes the user's behavior based on the analysis results. Specifically, it sets categories such as eating, exercise, and shopping, and categorizes the data based on these. The input for this process is a list of objects, and the output is behavioral data categorized by category.
[0916] Step 6:
[0917] The server saves the behavioral data as text information in a life log. Specifically, it saves it in a database. The input for this process is behavioral data categorized by category, and the output is life log data as text information.
[0918] Step 7:
[0919] The server analyzes the recorded life log and generates lifestyle improvement suggestions. Specifically, it uses a deep learning algorithm to generate advice based on past logs and specialized knowledge. The input in this process is life log data, and the output is improvement suggestions.
[0920] Step 8:
[0921] The server transmits the generated suggestions and lifelog data to the device using a secure communication protocol. The inputs to this process are the improvement suggestions and lifelog data, and the output is the data transmitted to the device.
[0922] Step 9:
[0923] The device notifies the user of the suggestions and lifelogs by displaying a message on the display or by issuing a voice notification. The input to this process is the suggestions and lifelog data received from the server, and the output is the notification delivered to the user.
[0924] Step 10:
[0925] The server analyzes the level of interest customers have in specific products or areas in the store and notifies store managers and staff of the results. Specifically, it analyzes customer gaze data and notifies the results in real time. The input to this process is gaze data, and the output is a notification message as the analysis result.
[0926] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0927] This invention is a system that uses a glasses-type device with a built-in camera and image processing function to automatically record a user's daily activities, and further recognizes and analyzes the user's emotions using an emotion engine to generate lifestyle improvement suggestions.With this system, users can efficiently collect life logs and receive advice on lifestyle improvements based on their emotions without performing any special operations.
[0928] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view at one-second intervals. This capture operation is performed automatically, and the obtained image data is temporarily stored in the device's internal memory. At the same time, the user's face is also periodically captured by the camera and analyzed by the emotion engine.
[0929] The device then pre-processes the captured image to remove noise and improve clarity, using image processing techniques such as Gaussian filters and sharpening algorithms.
[0930] The device sends the preprocessed image data in batches at regular intervals (e.g., every hour) to the server. A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety. Emotion data is also sent to the server in the same way.
[0931] The server inputs the image data received from the device into a generative AI model to analyze the image. This image analysis utilizes deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks). The emotion engine also analyzes the user's face to identify their emotional state (e.g., joy, sadness, anger, etc.).
[0932] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, from an image captured during a meal, "salad," "steak," and "rice" are identified and recorded as "lunch." The server also uses facial recognition technology to classify the user's emotions and adds them to the behavioral data. For example, if "happiness" is detected on the user's face during lunch, it is recorded as "lunch (happiness)." The server records this behavioral data as text information in a life log and saves it in a database.
[0933] The server then analyzes the recorded life log and emotional data to generate lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You haven't been eating enough vegetables, so I recommend you eat more next time. Also, your recent meals have been associated with pleasure, so maintain this pattern to reduce stress."
[0934] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern."
[0935] Users can receive notifications, refer to their life logs and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lives and manage their health based on their emotions.
[0936] In this way, the present invention combines a built-in camera, image processing technology, and an emotion engine to realize a new system that automatically records a user's behavior and emotions and provides specific lifestyle improvement suggestions based on that information.
[0937] The processing flow will be explained below.
[0938] Step 1:
[0939] The device (glasses-type device) uses a built-in camera to capture the user's field of view at one-second intervals.
[0940] Specifically, the camera automatically captures an image and temporarily stores it in its internal memory.
[0941] Step 2:
[0942] The device periodically captures the user's face using a camera and analyzes it using an emotion engine.
[0943] Specifically, it uses a facial recognition algorithm to identify the user's face and extract their emotional state (e.g., joy, sadness, anger, etc.).
[0944] Step 3:
[0945] The device pre-processes the captured field of view image to remove noise and improve image clarity.
[0946] Specifically, the image quality is improved by applying a Gaussian filter or sharpening algorithm.
[0947] Step 4:
[0948] The terminal collectively converts the preprocessed image data and emotion data into a batch format at regular time intervals (e.g., every hour).
[0949] Specifically, the image data and emotion data are compressed into a single file and preparations for transmission are made.
[0950] Step 5:
[0951] The terminal transmits the batched data to the server using a secure communication protocol (e.g., HTTPS).
[0952] Specifically, the operation is to establish a session for data transmission and start uploading data.
[0953] Step 6:
[0954] The server inputs the data received from the device into the generative AI model and analyzes the field of view image.
[0955] Specifically, the AI model processes the received image data and identifies important objects (e.g., food, people, landmarks).
[0956] Step 7:
[0957] The server associates the emotion data analyzed by the emotion engine with the behavioral data.
[0958] Specifically, behavioral data is tagged with emotion tags such as "joy" or "sadness" and stored in a database.
[0959] Step 8:
[0960] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel).
[0961] Specifically, the analysis results are assigned to preset categories and organized according to each category.
[0962] Step 9:
[0963] The server records behavioral data and emotional data organized by category as text information in a life log.
[0964] Specifically, the system stores text information such as "2023-10-05 12:30 Lunch: Salad, steak, rice (joy)" in a database.
[0965] Step 10:
[0966] The server analyzes the recorded life log and emotional data and generates suggestions for improving lifestyle habits.
[0967] Specifically, it generates specific advice such as "Eat more vegetables and maintain your current eating patterns" based on the user's past behavioral logs and emotional data.
[0968] Step 11:
[0969] The server sends the generated improvement suggestions and detailed life logs to the device.
[0970] Specifically, the system compresses the improvement suggestions and life log information and sends them to the device using a secure communication protocol.
[0971] Step 12:
[0972] The device notifies the user of the received suggestions and life logs so that they can be viewed.
[0973] Specifically, the device's display and voice notification function will notify the user: "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern."
[0974] Step 13:
[0975] Users receive notifications from their devices and view life logs and improvement suggestions.
[0976] Specifically, the user checks the suggestions on the device and decides on the next action based on them (e.g., eating more vegetables at the next meal).
[0977] Example 2
[0978] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0979] Conventional lifelog acquisition systems require manual user operation, making it difficult to efficiently record detailed daily activities and analyze emotional states. It is also difficult to generate specific lifestyle improvement suggestions based on acquired lifelog data. Furthermore, data security and privacy protection are insufficient, creating a need for a system that users can use with confidence.
[0980] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0981] In this invention, the server includes: a means for capturing the field of view in real time using a built-in camera; a means for preprocessing the captured images to remove noise; a means for sending the preprocessed images to the server in batch format at specific time intervals; a means for analyzing the received image data to identify important objects; a means for categorizing the user's behavior based on the analysis results; a means for analyzing the user's emotions using facial recognition technology; a means for recording the behavioral data and emotional data as text information in a life log; a means for analyzing the recorded life log and generating lifestyle improvement suggestions; a means for transmitting the generated suggestions and life log to a terminal; and a means for notifying the user of the suggestions and life log. This allows the user to obtain a detailed life log without performing any special operations and receive specific lifestyle improvement suggestions based on their emotions. Furthermore, the use of a secure communication protocol ensures data security and privacy protection.
[0982] The "built-in camera" is a camera built into the eyeglass-type terminal, and is a device that has the function of capturing the user's field of view in real time.
[0983] "Pre-processing" refers to a general range of image processing techniques used to remove noise and improve clarity of captured image data, including Gaussian filters and sharpening algorithms.
[0984] "Batch format" refers to a method of processing data collectively at regular intervals, and is used to efficiently transmit and process large amounts of data.
[0985] A "secure communication protocol" is a communication protocol used to ensure data security and privacy, such as HTTPS.
[0986] "Image analysis" refers to the process of using generative AI models and deep learning algorithms to recognize and identify important objects in captured image data.
[0987] "Emotion analysis" refers to the process of using facial recognition technology to identify emotional states (e.g., joy, sadness, anger, etc.) from a user's face.
[0988] "Behavior categorization" refers to the process of classifying user behavior into multiple categories (e.g., eating, exercise, travel) based on the analysis results.
[0989] A "life log" refers to a collection of data that records a user's daily activities and emotional data as text information.
[0990] "Lifestyle improvement suggestions" refer to suggestions that include specific advice for improving the user's lifestyle based on the recorded life log and emotional data.
[0991] "Notification" refers to a method of informing the user of the generated suggestions and details of the life log, and is done through display or audio output.
[0992] The present invention is a system that uses a glasses-type device equipped with a built-in camera and image processing functions to automatically record a user's daily activities, and further recognizes and analyzes the user's emotions using an emotion engine, thereby generating lifestyle improvement suggestions. Detailed embodiments of the present invention will be described below.
[0993] Hardware and software used
[0994] Glasses-type device: Equipped with a built-in camera, internal memory, display, and audio output functions, it captures and temporarily stores data.
[0995] Server: Runs generative AI models, emotion engines, and databases for data analysis and storage. Performs data analysis and proposal generation.
[0996] Communication protocol: HTTPS is used to ensure secure communication.
[0997] Program processing
[0998] The device automatically captures the user's field of view using the built-in camera at one-second intervals, and the user's face is also captured periodically and analyzed by the emotion engine.
[0999] The device pre-processes the captured image data to remove noise and improve clarity, using image processing techniques such as Gaussian filters and sharpening algorithms.
[1000] The terminal transmits the preprocessed image data and emotion data to the server in batch format at regular time intervals using a secure communication protocol (HTTPS).
[1001] The server inputs the received image data into a generative AI model (e.g., YOLO, RetinaNet) for image analysis. This image analysis recognizes and identifies important objects (e.g., food, people, landmarks). The emotion engine also analyzes the user's face to identify their emotional state (e.g., joy, sadness, anger, etc.).
[1002] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, from an image captured during a meal, "salad," "steak," and "rice" are identified and recorded as "lunch." Furthermore, facial recognition technology is used to classify the user's emotions and add them to the behavioral data. For example, if "happiness" is detected on the user's face during lunch, it is categorized as "lunch (happiness)."
[1003] The server analyzes the recorded life log and emotional data and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, the server might generate advice such as, "You're not eating enough vegetables, so we recommend you eat more next time. Also, your recent meals have been associated with pleasure, so maintain this pattern to reduce stress."
[1004] Examples of concrete examples and prompts
[1005] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or voice output. For example, the device's display might say, "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern," and a voice output might say, "Improvement suggestions have been received."
[1006] Prompt Sentence Examples
[1007] "Show me your lunch log and provide relevant sentiment data and improvement suggestions."
[1008] The present invention allows users to effortlessly improve their lifestyle and manage their health based on their emotions. In this way, the present invention realizes a new system that combines a built-in camera, image processing technology, and an emotion engine to automatically record a user's behavior and emotions and provide specific lifestyle improvement suggestions based on the recorded information.
[1009] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1010] Step 1:
[1011] Data Capture
[1012] The device uses a built-in camera to capture the user's field of view at one-second intervals. The captured image data is temporarily stored in the internal memory. At the same time, the device periodically captures the user's face, which is then analyzed by the emotion engine.
[1013] Input: Images of the user's field of view and face
[1014] Output: Temporarily saved captured image data and facial image data
[1015] Specific operation: The device will start the camera, a shutter sound will be heard every second, and the field of view and face will be captured. These image data will be stored in the internal memory.
[1016] Step 2:
[1017] Image preprocessing
[1018] The device pre-processes the captured image data to remove noise and improve clarity, using Gaussian filters and sharpening algorithms.
[1019] Input: Temporarily saved captured image data
[1020] Output: Preprocessed image data
[1021] What happens: The device's image processing algorithm runs, and the progress is displayed on the screen.
[1022] Step 3:
[1023] Data transmission
[1024] The device sends the preprocessed image data and emotion data to the server in batch format at regular intervals using a secure communication protocol (HTTPS).
[1025] Input: Preprocessed image data and emotion data
[1026] Output: The dataset sent to the server
[1027] Specific operation: While data transmission is in progress, the display will show "Sending data...", and once transmission is complete, the display will notify you that "Data transmitted successfully."
[1028] Step 4:
[1029] Data analysis
[1030] The server inputs the received image data into a generative AI model (e.g., YOLO, RetinaNet) to recognize and identify important objects, and also uses an emotion engine to analyze emotions from the facial image data.
[1031] Input: Dataset sent to the server (image data and emotion data)
[1032] Output: Analysis results (object recognition and emotional state)
[1033] Specific operation: The server's processing unit performs image analysis and emotion analysis, and displays the progress on the operation console. After the analysis is complete, a list of recognized objects and emotional states is displayed.
[1034] Step 5:
[1035] Classification of behaviors and emotions
[1036] Based on the analysis results, the server categorizes the user's behavior (e.g., eating, exercise, travel) and also associates emotional data with the behavior.
[1037] Input: Analysis results (object recognition and emotional state)
[1038] Output: Classified behavioral and emotional data
[1039] Specific operation: The server stores the action and emotion classification results in a database in the form of, for example, "Lunch (salad, steak, rice, joy)".
[1040] Step 6:
[1041] Generate improvement suggestions
[1042] The server analyzes the recorded life log and emotional data and generates lifestyle improvement suggestions using an algorithm based on past behavioral logs and expert knowledge.
[1043] Input: Classified behavioral and emotional data
[1044] Output: Generated lifestyle improvement suggestions
[1045] Specific operation: The server's algorithm runs and generates improvement suggestions. The console screen displays a list of the generated suggestions.
[1046] Step 7:
[1047] User Notification
[1048] The server sends the generated improvement suggestions and detailed life logs to the device, which receives them and notifies the user through a display or audio output.
[1049] Input: Generated improvement suggestions and detailed lifelog
[1050] Output: Suggestions and lifelog information notified to the user
[1051] Specific operation: The device display will show "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern," and a voice output will announce, "Improvement suggestions have been received."
[1052] (Application example 2)
[1053] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1054] In modern society, it is difficult for users to easily collect life logs to maintain their health and receive lifestyle improvement suggestions based on those logs in their busy daily lives. It is also difficult to provide personalized dietary suggestions that take the user's emotions into account. Conventional systems require users to manually enter data, which is time-consuming and makes it difficult to collect accurate data. To solve these issues, a system is needed that utilizes a built-in camera, image processing technology, and an emotion engine to automatically collect life logs and provide dietary suggestions based on emotions.
[1055] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the field of view in real time using a built-in camera, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for analyzing the received image data to identify important objects, means for categorizing the user's behavior based on the analysis results, means for recording the behavioral data as text information in a life log, means for analyzing the recorded life log to generate lifestyle improvement suggestions, means for transmitting the generated suggestions and life log to the terminal, means for notifying the user of the suggestions and life log, means for analyzing the user's emotional state and generating personalized meal suggestions based on the emotional state, and means for notifying the user of the generated meal suggestions. This allows the user to receive personalized lifestyle improvement suggestions and meal suggestions based on their emotions without performing any special operations.
[1056] The "built-in camera" is a camera device built into the eyeglass-type terminal to capture the user's field of view in real time.
[1057] "Preprocessing" refers to initial image processing operations to remove noise and improve clarity from a captured image.
[1058] The "batch format" is a data transfer method in which preprocessed image data is sent in batches at regular time intervals.
[1059] A "server" is a remote computer system that receives, analyzes, and stores data sent from user terminals, and generates suggestions based on the results of the analysis.
[1060] "Analysis" is the process of identifying significant objects from the received image data and discerning the user's behavior and emotional state.
[1061] The "behavior category" is a category for classifying the user's daily behavior into specific activity types based on the analysis results.
[1062] A "life log" is a digital history that records and saves a user's daily behavioral and emotional data as text information.
[1063] "Improvement suggestions" are specific advice for improving the user's lifestyle habits that are generated from the analysis results of the life log.
[1064] "Emotional state" refers to the type of emotion (e.g., joy, sadness, anger) analyzed from the user's facial image.
[1065] "Meal Suggestions" are personalized meal suggestions generated based on the user's emotional state and past meal data.
[1066] "Notification" is the process of sending messages to communicate generated suggestions and lifelogs to users.
[1067] The system of the present invention includes a camera built into an eyeglass-type terminal, a server, and a user's terminal (a smartphone or smart glasses).
[1068] First, the device's built-in camera captures the user's field of view in real time, capturing images every second. These captured images are then pre-processed using the device's image processing capabilities. This pre-processing uses image processing techniques such as the OpenCV library to remove noise and improve image clarity.
[1069] The preprocessed image data is sent to the server in batches at regular intervals. A secure communication protocol (HTTPS) is used for transmission to ensure data security. Emotion data is also sent to the server.
[1070] The server analyzes the received image data using a deep learning framework (such as TensorFlow or YOLO) to identify important objects. Based on the analysis results, it classifies the user's behavior into categories (e.g., eating, exercise, travel). It also uses facial recognition technology to analyze the user's emotional state and adds it to the behavioral data.
[1071] For example, if an image of a user eating a salad is captured and joy is identified from the user's emotional state, the behavioral data is recorded in the life log as "Lunch (Joy)." The server analyzes this behavioral data and emotional data and generates lifestyle improvement suggestions. This analysis uses algorithms based on the user's past behavioral logs and expert knowledge.
[1072] The server then generates personalized meal suggestions, taking into account the user's emotional state, using a generative AI model (e.g., GPT) and prompting the user with the following sentences:
[1073] "The user is having salad, soup, and bread for lunch. Their facial expressions indicate they are enjoying the meal. Please suggest a meal for this user's next meal."
[1074] The generated improvement suggestions and meal suggestions are sent back to the device and notified to the user. Notification methods include displaying the information on the device's display or using audio output. This allows the user to receive their own life log, improvement suggestions, and personalized meal suggestions and reflect them in their next actions without performing any special operations.
[1075] For example, if a user eats a salad at lunch and the emotion of joy is detected, the device will notify them with a suggestion: "We recommend that you incorporate more vegetables into your next meal to improve balance and maintain your current eating pattern." In this way, users can receive practical advice based on their emotions and behavioral data.
[1076] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1077] Step 1:
[1078] The device uses the built-in camera to capture the user's field of view at one-second intervals. At this time, the camera API operates and generates image data obtained from the user's field of view. The input is raw data from the camera, and the output is the captured image file.
[1079] Step 2:
[1080] The device pre-processes the captured image to remove noise and improve clarity. This process uses an image processing library (OpenCV). The input is the captured image file, and the output is a pre-processed, clear image file.
[1081] Step 3:
[1082] The terminal sends the preprocessed image data to the server in batch format at specific time intervals. A secure communication protocol (HTTPS) is used for data transmission to ensure data security. The input is the preprocessed image data, and the output is the data transfer to the server.
[1083] Step 4:
[1084] The server analyzes the received image data using a deep learning framework (TensorFlow, YOLO) to identify important objects. This analysis step recognizes and classifies objects in the image (e.g., food, people, landmarks, etc.). The input is the transmitted image data, and the output is the analyzed object information.
[1085] Step 5:
[1086] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). It also analyzes the user's emotional state using facial recognition technology and adds it to the behavioral data. For example, the data is recorded as "lunch (joy)." The input is the analyzed object information and emotional data, and the output is behavioral data classified by category.
[1087] Step 6:
[1088] The server analyzes this behavioral and emotional data and generates lifestyle improvement suggestions. The algorithm utilizes the user's past behavioral logs and expert knowledge. The input is behavioral and emotional data categorized by category, and the output is the generated improvement suggestions.
[1089] Step 7:
[1090] The server generates personalized meal suggestions using a generative AI model (e.g., GPT) based on the user's emotional state. For example, you can get suggestions by entering a prompt like this:
[1091] "The user is having salad, soup, and bread for lunch. Their facial expressions indicate they are enjoying the meal. Please suggest a meal for this user's next meal."
[1092] The input is a prompt based on the user's behavioral and emotional data, and the output is personalized meal suggestions.
[1093] Step 8:
[1094] The server sends the generated improvement suggestions and meal suggestions to the terminal, which receives them and notifies the user. The input is the generated suggestions and notification data, and the output is the notification to the user.
[1095] Step 9:
[1096] The user receives notifications from the device, refers to their own life log, improvement suggestions, and personalized meal suggestions, and reflects them in their next actions. The input is the notification information from the device, and the output is the action selected by the user.
[1097] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1098] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1099] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1100] [Fourth embodiment]
[1101] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1102] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1103] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1104] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1105] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1106] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1107] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1108] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1109] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1110] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1111] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1112] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1113] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1114] This invention is a system that uses glasses-type devices equipped with a built-in camera and image processing functions to automatically record a user's daily activities, analyze those activities via a server, and generate lifestyle improvement suggestions.With this system, users can efficiently obtain life logs without performing any special operations and use them to improve their lifestyles.
[1115] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view in real time. This capture operation occurs every second, and the obtained image data is temporarily stored in the device's internal memory. After capturing the images, the device performs preprocessing on them to remove noise and improve clarity. This preprocessing uses image processing techniques such as Gaussian filters and sharpening algorithms.
[1116] Next, the device sends the pre-processed image data to the server in batch format at regular intervals (e.g., every hour). A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety.
[1117] The server inputs the image data received from the device into a generative AI model to analyze the image, using deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks).
[1118] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, it identifies "salad," "steak," and "rice" from an image captured during a meal and records it as "lunch." The server records this behavior data as text information in a life log and saves it in a database.
[1119] The server then analyzes the recorded life log and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You're not getting enough vegetables. Eat more vegetables next time."
[1120] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: salad, steak, rice. Eat more vegetables."
[1121] Users can receive this notification, refer to their life log and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lifestyle and manage their health.
[1122] In this way, the present invention realizes a new system that combines a built-in camera, image processing technology, and analysis algorithms to automatically record users' behavior and provide specific suggestions that will help improve their lifestyle.
[1123] The processing flow will be explained below.
[1124] Step 1:
[1125] The device (glasses-type device) uses a built-in camera to capture the user's field of view at one-second intervals.
[1126] Specifically, the camera automatically captures an image and temporarily stores it in its internal memory.
[1127] Step 2:
[1128] The device preprocesses the captured image to remove noise and improve image clarity.
[1129] Specifically, the image quality is improved by applying a Gaussian filter or sharpening algorithm.
[1130] Step 3:
[1131] The terminal converts the preprocessed image data into a batch format at regular time intervals (e.g., every hour).
[1132] Specifically, the image data is compressed into one file and prepared for transmission.
[1133] Step 4:
[1134] The terminal transmits the image data in a batch format to the server using a secure communication protocol (e.g., HTTPS).
[1135] Specifically, a session for transmitting image data is established, and data uploading begins.
[1136] Step 5:
[1137] The server inputs the image data received from the terminal into the generated AI model and analyzes the image.
[1138] Specifically, the AI model processes the received image data and identifies important objects (e.g., food, people, landmarks).
[1139] Step 6:
[1140] The server classifies the user's behavior into categories (e.g., eating, exercise, travel) based on the image analysis results.
[1141] Specifically, the analysis results are assigned to preset categories and organized according to each category.
[1142] Step 7:
[1143] The server records the behavioral data organized by category as text information in a life log.
[1144] Specifically, text information such as "2023-10-05 12:30 Lunch: salad, steak, rice" is saved in the database.
[1145] Step 8:
[1146] The server analyzes the recorded life log and generates suggestions for improving lifestyle habits.
[1147] Specifically, it compares and analyzes the user's past behavioral logs and generates specific advice such as "eat more vegetables."
[1148] Step 9:
[1149] The server sends the generated improvement suggestions and detailed life logs to the device.
[1150] Specifically, the system compresses the improvement suggestions and life log information and sends them to the device using a secure communication protocol.
[1151] Step 10:
[1152] The device notifies the user of the received suggestions and life logs so that they can be viewed.
[1153] Specifically, the device's display and voice notification functions are used to notify the user, displaying "Lunch record: salad, steak, rice. Eat more vegetables."
[1154] Step 11:
[1155] Users receive notifications from their devices and view life logs and improvement suggestions.
[1156] Specifically, the user checks the suggestions on the device, reviews their lifestyle habits based on them, and decides on their next course of action (e.g., eating more vegetables at the next meal).
[1157] Example 1
[1158] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1159] Conventional life log management systems have had problems with the effort required for users to record their own data and the accuracy of the records. In addition, the lifestyle improvement suggestions they offer are often not specialized enough, and their effectiveness in improving users' health is limited.
[1160] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1161] In this invention, the server includes means for capturing the field of view in real time using a built-in camera, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for inputting the images into a generative AI model to identify important objects, means for categorizing the user's behavior based on the analysis results, means for recording the behavioral data as text information in a life log, means for analyzing the recorded life log with an algorithm based on expert knowledge to generate lifestyle improvement suggestions, means for transmitting the generated suggestions and life log to a terminal, and means for notifying the user of the suggestions and life log. This allows the user to easily obtain a highly accurate life log and receive expert improvement suggestions.
[1162] A "built-in camera" is a camera built into a device that captures the external field of view in real time.
[1163] "Means for capturing the field of view in real time" is a function that uses the built-in camera to continuously capture images that are within the user's field of view.
[1164] The "means for pre-processing the captured image" is a function for performing processing on the obtained image data to remove noise and improve clarity.
[1165] The "means for transmitting to the server in batch format at specific time intervals" is a function for collecting preprocessed image data at regular time intervals and transmitting the collected data to the server using a secure communication protocol.
[1166] A "generative AI model" refers to an algorithm that uses artificial intelligence to analyze data and perform specific tasks.
[1167] "Means for identifying significant objects" refers to the ability to use generative AI models to recognize and analyze objects, people, landmarks, etc. in images.
[1168] "Means for categorizing user behavior" is a function that records user behavior into categories such as food, exercise, and travel based on the analysis results.
[1169] A "life log" is data that records a user's daily activities and behaviors in chronological order.
[1170] An "algorithm based on expert knowledge" is a computational method for performing analysis that incorporates specialized knowledge in a specific field.
[1171] The "means for generating lifestyle improvement suggestions" is a function that generates advice for improving the user's lifestyle using an algorithm based on the recorded life log and expert knowledge.
[1172] The "means for transmitting to the terminal" is a function for transferring the generated proposals and life logs to the terminal.
[1173] The "means of notifying the user of suggestions and life logs" is a function that notifies the user of the contents of suggestions and life logs through the device's built-in display or audio output.
[1174] MODE FOR CARRYING OUT THE INVENTION
[1175] This invention is a system that uses glasses-type devices equipped with a built-in camera and image processing functions to automatically record a user's daily activities, analyze those activities via a server, and generate lifestyle improvement suggestions.With this system, users can efficiently obtain life logs without performing any special operations and use them to improve their lifestyles.
[1176] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view in real time. This capture operation occurs every second, and the obtained image data is temporarily stored in the device's internal memory. After capturing the images, the device performs preprocessing on them to remove noise and improve clarity. This preprocessing uses image processing techniques such as Gaussian filters and sharpening algorithms.
[1177] Next, the device sends the pre-processed image data to the server in batch format at regular intervals (e.g., every hour). A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety.
[1178] The server inputs the image data received from the device into a generative AI model to analyze the image, using deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks).
[1179] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, it identifies "salad," "steak," and "rice" from an image captured during a meal and records it as "lunch." The server records this behavior data as text information in a life log and saves it in a database.
[1180] The server then analyzes the recorded life log and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You're not getting enough vegetables. Eat more vegetables next time."
[1181] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: salad, steak, rice. Eat more vegetables."
[1182] Users can receive this notification, refer to their life log and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lifestyle and manage their health.
[1183] An example of a specific prompt is, "Please identify whether there is food in this image and identify the specific type." Based on this prompt, the generative AI model recognizes and classifies the objects in the image.
[1184] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1185] Step 1:
[1186] The device uses a built-in camera to capture the user's field of view in real time.
[1187] Input data: user's view
[1188] How it works: The device captures images within its field of view every second and stores them in its internal memory, recording the user's daily activities as video.
[1189] Output data: Captured image data
[1190] Step 2:
[1191] The device pre-processes the captured image to remove noise and improve clarity.
[1192] Input data: Captured image data
[1193] What it does: The device uses a Gaussian filter to remove noise from the image and a sharpening algorithm to improve clarity, allowing it to accurately identify important objects during analysis.
[1194] Output data: Preprocessed image data
[1195] Step 3:
[1196] The terminal sends the preprocessed image data in batch format to the server at regular time intervals (e.g., every hour).
[1197] Input data: Preprocessed image data
[1198] Specific operation: The device periodically compiles image data into a batch format and sends it securely to the server using HTTPS.
[1199] Output data: batch image data sent to the server
[1200] Step 4:
[1201] The server inputs the received image data into the generated AI model and analyzes the image.
[1202] Input data: batch image data
[1203] How it works: The server uses a generative AI model (e.g., YOLO, RetinaNet, etc.) to analyze the input image and recognize and identify important objects (e.g., food, people, landmarks). For example, it can recognize food items such as "salad" and "steak" from an image of a meal.
[1204] Output data: Analysis results (recognized objects)
[1205] Step 5:
[1206] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel).
[1207] Input data: Analysis results (recognized objects)
[1208] Specific actions: Based on the recognized object information, actions are classified into categories such as "meal," "exercise," and "travel," and recorded as a life log. For example, if "salad," "steak," and "rice" are detected, they are classified into the category of "lunch."
[1209] Output data: Behavioral data categorized by category
[1210] Step 6:
[1211] The server records the behavioral data as text information in a life log and stores it in a database.
[1212] Input data: Categorized behavioral data
[1213] Specific operation: The classified behavioral data is converted into text format and saved in a database, thereby saving the user's past behavior as a history.
[1214] Output data: Saved life logs
[1215] Step 7:
[1216] The server analyzes the recorded life log and generates suggestions for improving lifestyle habits.
[1217] Input data: Saved life logs
[1218] Specific operation: Using an algorithm based on expert knowledge, the system generates lifestyle improvement suggestions based on the user's behavioral history. For example, it generates specific advice such as, "You're not getting enough vegetables. Next time, eat more vegetables."
[1219] Output data: Suggestions for improving lifestyle habits
[1220] Step 8:
[1221] The server then sends the generated improvement suggestions and detailed life logs back to the device.
[1222] Input data: Generated lifestyle improvement suggestions, detailed life log
[1223] Specific operation: A secure communication protocol is used to send the generated lifestyle improvement suggestions and detailed life logs to the device.
[1224] Output data: Improvement suggestions and life logs sent to the device
[1225] Step 9:
[1226] The device receives improvement suggestions and life logs and notifies the user.
[1227] Input data: improvement proposals, life logs
[1228] Specific operation: The device notifies the user of improvement suggestions and the contents of the life log via the device's built-in display and voice output. For example, it displays "Lunch record: salad, steak, rice. Eat more vegetables."
[1229] Output data: Improvement suggestions notified to the user and lifelog
[1230] Step 10:
[1231] Users receive notifications, refer to their own life logs and improvement suggestions, and reflect them in their next actions.
[1232] Input data: Notified improvement suggestions, lifelog
[1233] Specific actions: Refer to the notification and reflect it in your next actions. For example, take action to improve your lifestyle, such as consciously eating more vegetables at your next meal.
[1234] Output data: Actions that reflect improvement suggestions
[1235] (Application example 1)
[1236] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1237] Conventional lifestyle recording systems and in-store customer behavior analysis systems require users and customers to record and operate the systems themselves, which is time-consuming and makes it difficult to obtain accurate data.In addition, it is not possible to grasp in real time what products or areas in the store customers are interested in, which makes it difficult to provide efficient customer service and marketing.
[1238] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1239] In this invention, the server includes means for capturing the field of view in real time using a built-in image capture device, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for analyzing the received image data to identify important objects, means for classifying the user's behavior based on the analysis results, means for saving the behavior data as text information in a lifestyle log, means for analyzing the recorded lifestyle log to generate lifestyle improvement suggestions, means for sending the generated suggestions and lifestyle log to a terminal, means for notifying the user of the suggestions and lifestyle log, means for measuring the level of customer interest in specific products or areas in the store, and means for notifying store managers and staff of the analysis results. This reduces the effort required for users and customers, enables accurate data acquisition and analysis, and enables efficient customer service and marketing.
[1240] An "integrated imaging device" is a device, such as a camera or image sensor, built into a device that captures the surrounding field of view in real time.
[1241] "Preprocessing" refers to a series of steps performed on captured image data to remove noise and improve image clarity.
[1242] "Batch format" is a method of processing and sending data for a certain period of time all at once.
[1243] "Important objects" are objects or areas identified by the analysis that should be of interest to the user or the system.
[1244] "Behavior" refers to the activity or pattern of a user or customer at a particular time.
[1245] "Classification" refers to organizing user behavior into categories based on the analysis results.
[1246] "Life Record" is a database that stores user behavior data as text information.
[1247] "Lifestyle improvement suggestions" are specific advice for improving your health and quality of life that is created by analyzing your recorded lifestyle records.
[1248] A "terminal" is a device worn by a user and has various functions including a built-in image capture device.
[1249] "Notification" refers to information transmission activities that inform users of generated suggestions and life logs via their devices.
[1250] A "customer" is a user who visits a store.
[1251] "In-store" means the physical premises where products are displayed and sold.
[1252] "Interested" means that a customer is taking actions that show interest, such as focusing on a particular product or area or staying there for a long time.
[1253] "Store Manager" means a person responsible for the operation and management of a store.
[1254] "Staff" refers to employees who deal with customers and manage merchandise within the store.
[1255] The system that realizes this application example is composed of a terminal equipped with a built-in image capture device, a server that uses a communication protocol, and a system that notifies users based on the analysis results. This system operates as follows.
[1256] The device captures the user's or customer's field of view in real time using a built-in image capture device, taking images every second. The captured images are then pre-processed to remove noise and improve image clarity, using OpenCV's Gaussian filter and sharpening algorithms, for example.
[1257] The preprocessed image data is sent to the server in batch format at regular intervals. This data transmission uses a secure communication protocol (e.g., HTTPS) to ensure data security. The server then analyzes the received image data and identifies important objects using a generative AI model (e.g., YOLO or RetinaNet). This analysis identifies user behavior categories and the products and areas in the store that customers are interested in.
[1258] The server classifies user behavior based on the analysis results and saves them as text information in a daily log. It also notifies store managers and staff in real time based on the analysis results of customer behavior in the store. Notifications are sent via smart glasses or smartphones, and specific alerts, such as "A customer is interested in the wine section," are displayed.
[1259] The recorded lifestyle records and behavioral data are further analyzed using deep learning algorithms to generate lifestyle improvement suggestions. For example, advice such as "You're not getting enough vegetables. Eat more vegetables next time" may be generated. The generated suggestions are resent to the device and notified to the user. This notification is sent via the smart glasses display or the smartphone notification function.
[1260] For example, if a customer stares at the wine section of a store for a long time, that information is sent to the server and analyzed as "the customer is interested in the wine section." As a result, sales staff are notified that "they should attend to the customer in the wine section," thereby enabling personalized customer service.
[1261] An example of a prompt for a generative AI model is as follows:
[1262] Suggest the best way to serve customers when they are interested in the wine section.
[1263] This system reduces the workload for users and stores, while enabling efficient and effective customer service and data-based marketing strategies.
[1264] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1265] Step 1:
[1266] The device captures the field of view in real time using a built-in image capture device. Specifically, the device acquires an image every second. At this point, the input is the captured image data, and the output is the image data before preprocessing.
[1267] Step 2:
[1268] The device preprocesses the captured image to remove noise and improve image clarity, using OpenCV's Gaussian filter and sharpening algorithms. The input to this process is the captured image data, and the output is image data with noise removed and improved clarity.
[1269] Step 3:
[1270] The terminal transmits the preprocessed image data to the server in batch format at regular intervals. Specifically, the data is stored in a fixed buffer and transmitted, for example, every hour. The input in this process is a plurality of preprocessed image data, and the output is batch data transmitted to the server.
[1271] Step 4:
[1272] The server analyzes the received image data. Specifically, it uses a generative AI model (e.g., YOLO or RetinaNet) to identify important objects in the image. The input to this process is a batch of image data, and the output is a list of objects as the analysis result.
[1273] Step 5:
[1274] The server categorizes the user's behavior based on the analysis results. Specifically, it sets categories such as eating, exercise, and shopping, and categorizes the data based on these. The input for this process is a list of objects, and the output is behavioral data categorized by category.
[1275] Step 6:
[1276] The server saves the behavioral data as text information in a life log. Specifically, it saves it in a database. The input for this process is behavioral data categorized by category, and the output is life log data as text information.
[1277] Step 7:
[1278] The server analyzes the recorded life log and generates lifestyle improvement suggestions. Specifically, it uses a deep learning algorithm to generate advice based on past logs and specialized knowledge. The input in this process is life log data, and the output is improvement suggestions.
[1279] Step 8:
[1280] The server transmits the generated suggestions and lifelog data to the device using a secure communication protocol. The inputs to this process are the improvement suggestions and lifelog data, and the output is the data transmitted to the device.
[1281] Step 9:
[1282] The device notifies the user of the suggestions and lifelogs by displaying a message on the display or by issuing a voice notification. The input to this process is the suggestions and lifelog data received from the server, and the output is the notification delivered to the user.
[1283] Step 10:
[1284] The server analyzes the level of interest customers have in specific products or areas in the store and notifies store managers and staff of the results. Specifically, it analyzes customer gaze data and notifies the results in real time. The input to this process is gaze data, and the output is a notification message as the analysis result.
[1285] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1286] This invention is a system that uses a glasses-type device with a built-in camera and image processing function to automatically record a user's daily activities, and further recognizes and analyzes the user's emotions using an emotion engine to generate lifestyle improvement suggestions.With this system, users can efficiently collect life logs and receive advice on lifestyle improvements based on their emotions without performing any special operations.
[1287] First, the device (glasses-type device) uses its built-in camera to capture the user's field of view at one-second intervals. This capture operation is performed automatically, and the obtained image data is temporarily stored in the device's internal memory. At the same time, the user's face is also periodically captured by the camera and analyzed by the emotion engine.
[1288] The device then pre-processes the captured image to remove noise and improve clarity, using image processing techniques such as Gaussian filters and sharpening algorithms.
[1289] The device sends the preprocessed image data in batches at regular intervals (e.g., every hour) to the server. A secure communication protocol (e.g., HTTPS) is used to transmit the data, ensuring its safety. Emotion data is also sent to the server in the same way.
[1290] The server inputs the image data received from the device into a generative AI model to analyze the image. This image analysis utilizes deep learning algorithms (e.g., YOLO, RetinaNet, etc.) to recognize and identify important objects (e.g., food, people, landmarks). The emotion engine also analyzes the user's face to identify their emotional state (e.g., joy, sadness, anger, etc.).
[1291] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, from an image captured during a meal, "salad," "steak," and "rice" are identified and recorded as "lunch." The server also uses facial recognition technology to classify the user's emotions and adds them to the behavioral data. For example, if "happiness" is detected on the user's face during lunch, it is recorded as "lunch (happiness)." The server records this behavioral data as text information in a life log and saves it in a database.
[1292] The server then analyzes the recorded life log and emotional data to generate lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, it provides specific advice such as, "You haven't been eating enough vegetables, so I recommend you eat more next time. Also, your recent meals have been associated with pleasure, so maintain this pattern to reduce stress."
[1293] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or audio output. For example, the device's display might say, "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern."
[1294] Users can receive notifications, refer to their life logs and improvement suggestions, and reflect them in their next actions.This system allows users to easily improve their lives and manage their health based on their emotions.
[1295] In this way, the present invention combines a built-in camera, image processing technology, and an emotion engine to realize a new system that automatically records a user's behavior and emotions and provides specific lifestyle improvement suggestions based on that information.
[1296] The processing flow will be explained below.
[1297] Step 1:
[1298] The device (glasses-type device) uses a built-in camera to capture the user's field of view at one-second intervals.
[1299] Specifically, the camera automatically captures an image and temporarily stores it in its internal memory.
[1300] Step 2:
[1301] The device periodically captures the user's face using a camera and analyzes it using an emotion engine.
[1302] Specifically, it uses a facial recognition algorithm to identify the user's face and extract their emotional state (e.g., joy, sadness, anger, etc.).
[1303] Step 3:
[1304] The device pre-processes the captured field of view image to remove noise and improve image clarity.
[1305] Specifically, the image quality is improved by applying a Gaussian filter or sharpening algorithm.
[1306] Step 4:
[1307] The terminal collectively converts the preprocessed image data and emotion data into a batch format at regular time intervals (e.g., every hour).
[1308] Specifically, the image data and emotion data are compressed into a single file and preparations for transmission are made.
[1309] Step 5:
[1310] The terminal transmits the batched data to the server using a secure communication protocol (e.g., HTTPS).
[1311] Specifically, the operation is to establish a session for data transmission and start uploading data.
[1312] Step 6:
[1313] The server inputs the data received from the device into the generative AI model and analyzes the field of view image.
[1314] Specifically, the AI model processes the received image data and identifies important objects (e.g., food, people, landmarks).
[1315] Step 7:
[1316] The server associates the emotion data analyzed by the emotion engine with the behavioral data.
[1317] Specifically, behavioral data is tagged with emotion tags such as "joy" or "sadness" and stored in a database.
[1318] Step 8:
[1319] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel).
[1320] Specifically, the analysis results are assigned to preset categories and organized according to each category.
[1321] Step 9:
[1322] The server records behavioral data and emotional data organized by category as text information in a life log.
[1323] Specifically, the system stores text information such as "2023-10-05 12:30 Lunch: Salad, steak, rice (joy)" in a database.
[1324] Step 10:
[1325] The server analyzes the recorded life log and emotional data and generates suggestions for improving lifestyle habits.
[1326] Specifically, it generates specific advice such as "Eat more vegetables and maintain your current eating patterns" based on the user's past behavioral logs and emotional data.
[1327] Step 11:
[1328] The server sends the generated improvement suggestions and detailed life logs to the device.
[1329] Specifically, the system compresses the improvement suggestions and life log information and sends them to the device using a secure communication protocol.
[1330] Step 12:
[1331] The device notifies the user of the received suggestions and life logs so that they can be viewed.
[1332] Specifically, the device's display and voice notification function will notify the user: "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern."
[1333] Step 13:
[1334] Users receive notifications from their devices and view life logs and improvement suggestions.
[1335] Specifically, the user checks the suggestions on the device and decides on the next action based on them (e.g., eating more vegetables at the next meal).
[1336] Example 2
[1337] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1338] Conventional lifelog acquisition systems require manual user operation, making it difficult to efficiently record detailed daily activities and analyze emotional states. It is also difficult to generate specific lifestyle improvement suggestions based on acquired lifelog data. Furthermore, data security and privacy protection are insufficient, creating a need for a system that users can use with confidence.
[1339] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1340] In this invention, the server includes: a means for capturing the field of view in real time using a built-in camera; a means for preprocessing the captured images to remove noise; a means for sending the preprocessed images to the server in batch format at specific time intervals; a means for analyzing the received image data to identify important objects; a means for categorizing the user's behavior based on the analysis results; a means for analyzing the user's emotions using facial recognition technology; a means for recording the behavioral data and emotional data as text information in a life log; a means for analyzing the recorded life log and generating lifestyle improvement suggestions; a means for transmitting the generated suggestions and life log to a terminal; and a means for notifying the user of the suggestions and life log. This allows the user to obtain a detailed life log without performing any special operations and receive specific lifestyle improvement suggestions based on their emotions. Furthermore, the use of a secure communication protocol ensures data security and privacy protection.
[1341] The "built-in camera" is a camera built into the eyeglass-type terminal, and is a device that has the function of capturing the user's field of view in real time.
[1342] "Pre-processing" refers to a general range of image processing techniques used to remove noise and improve clarity of captured image data, including Gaussian filters and sharpening algorithms.
[1343] "Batch format" refers to a method of processing data collectively at regular intervals, and is used to efficiently transmit and process large amounts of data.
[1344] A "secure communication protocol" is a communication protocol used to ensure data security and privacy, such as HTTPS.
[1345] "Image analysis" refers to the process of using generative AI models and deep learning algorithms to recognize and identify important objects in captured image data.
[1346] "Emotion analysis" refers to the process of using facial recognition technology to identify emotional states (e.g., joy, sadness, anger, etc.) from a user's face.
[1347] "Behavior categorization" refers to the process of classifying user behavior into multiple categories (e.g., eating, exercise, travel) based on the analysis results.
[1348] A "life log" refers to a collection of data that records a user's daily activities and emotional data as text information.
[1349] "Lifestyle improvement suggestions" refer to suggestions that include specific advice for improving the user's lifestyle based on the recorded life log and emotional data.
[1350] "Notification" refers to a method of informing the user of the generated suggestions and details of the life log, and is done through display or audio output.
[1351] The present invention is a system that uses a glasses-type device equipped with a built-in camera and image processing functions to automatically record a user's daily activities, and further recognizes and analyzes the user's emotions using an emotion engine, thereby generating lifestyle improvement suggestions. Detailed embodiments of the present invention will be described below.
[1352] Hardware and software used
[1353] Glasses-type device: Equipped with a built-in camera, internal memory, display, and audio output functions, it captures and temporarily stores data.
[1354] Server: Runs generative AI models, emotion engines, and databases for data analysis and storage. Performs data analysis and proposal generation.
[1355] Communication protocol: HTTPS is used to ensure secure communication.
[1356] Program processing
[1357] The device automatically captures the user's field of view using the built-in camera at one-second intervals, and the user's face is also captured periodically and analyzed by the emotion engine.
[1358] The device pre-processes the captured image data to remove noise and improve clarity, using image processing techniques such as Gaussian filters and sharpening algorithms.
[1359] The terminal transmits the preprocessed image data and emotion data to the server in batch format at regular time intervals using a secure communication protocol (HTTPS).
[1360] The server inputs the received image data into a generative AI model (e.g., YOLO, RetinaNet) for image analysis. This image analysis recognizes and identifies important objects (e.g., food, people, landmarks). The emotion engine also analyzes the user's face to identify their emotional state (e.g., joy, sadness, anger, etc.).
[1361] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). For example, from an image captured during a meal, "salad," "steak," and "rice" are identified and recorded as "lunch." Furthermore, facial recognition technology is used to classify the user's emotions and add them to the behavioral data. For example, if "happiness" is detected on the user's face during lunch, it is categorized as "lunch (happiness)."
[1362] The server analyzes the recorded life log and emotional data and generates lifestyle improvement suggestions. This analysis is performed using an algorithm based on the user's past behavioral logs and expert knowledge. For example, the server might generate advice such as, "You're not eating enough vegetables, so we recommend you eat more next time. Also, your recent meals have been associated with pleasure, so maintain this pattern to reduce stress."
[1363] Examples of concrete examples and prompts
[1364] The generated improvement suggestions and detailed life logs are then sent back to the device. The device receives them and notifies the user. Notifications are given via a display or voice output. For example, the device's display might say, "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern," and a voice output might say, "Improvement suggestions have been received."
[1365] Prompt Sentence Examples
[1366] "Show me your lunch log and provide relevant sentiment data and improvement suggestions."
[1367] The present invention allows users to effortlessly improve their lifestyle and manage their health based on their emotions. In this way, the present invention realizes a new system that combines a built-in camera, image processing technology, and an emotion engine to automatically record a user's behavior and emotions and provide specific lifestyle improvement suggestions based on the recorded information.
[1368] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1369] Step 1:
[1370] Data Capture
[1371] The device uses a built-in camera to capture the user's field of view at one-second intervals. The captured image data is temporarily stored in the internal memory. At the same time, the device periodically captures the user's face, which is then analyzed by the emotion engine.
[1372] Input: Images of the user's field of view and face
[1373] Output: Temporarily saved captured image data and facial image data
[1374] Specific operation: The device will start the camera, a shutter sound will be heard every second, and the field of view and face will be captured. These image data will be stored in the internal memory.
[1375] Step 2:
[1376] Image preprocessing
[1377] The device pre-processes the captured image data to remove noise and improve clarity, using Gaussian filters and sharpening algorithms.
[1378] Input: Temporarily saved captured image data
[1379] Output: Preprocessed image data
[1380] What happens: The device's image processing algorithm runs, and the progress is displayed on the screen.
[1381] Step 3:
[1382] Data transmission
[1383] The device sends the preprocessed image data and emotion data to the server in batch format at regular intervals using a secure communication protocol (HTTPS).
[1384] Input: Preprocessed image data and emotion data
[1385] Output: The dataset sent to the server
[1386] Specific operation: While data transmission is in progress, the display will show "Sending data...", and once transmission is complete, the display will notify you that "Data transmitted successfully."
[1387] Step 4:
[1388] Data analysis
[1389] The server inputs the received image data into a generative AI model (e.g., YOLO, RetinaNet) to recognize and identify important objects, and also uses an emotion engine to analyze emotions from the facial image data.
[1390] Input: Dataset sent to the server (image data and emotion data)
[1391] Output: Analysis results (object recognition and emotional state)
[1392] Specific operation: The server's processing unit performs image analysis and emotion analysis, and displays the progress on the operation console. After the analysis is complete, a list of recognized objects and emotional states is displayed.
[1393] Step 5:
[1394] Classification of behaviors and emotions
[1395] Based on the analysis results, the server categorizes the user's behavior (e.g., eating, exercise, travel) and also associates emotional data with the behavior.
[1396] Input: Analysis results (object recognition and emotional state)
[1397] Output: Classified behavioral and emotional data
[1398] Specific operation: The server stores the action and emotion classification results in a database in the form of, for example, "Lunch (salad, steak, rice, joy)".
[1399] Step 6:
[1400] Generate improvement suggestions
[1401] The server analyzes the recorded life log and emotional data and generates lifestyle improvement suggestions using an algorithm based on past behavioral logs and expert knowledge.
[1402] Input: Classified behavioral and emotional data
[1403] Output: Generated lifestyle improvement suggestions
[1404] Specific operation: The server's algorithm runs and generates improvement suggestions. The console screen displays a list of the generated suggestions.
[1405] Step 7:
[1406] User Notification
[1407] The server sends the generated improvement suggestions and detailed life logs to the device, which receives them and notifies the user through a display or audio output.
[1408] Input: Generated improvement suggestions and detailed lifelog
[1409] Output: Suggestions and lifelog information notified to the user
[1410] Specific operation: The device display will show "Lunch record: Salad, steak, rice (joy). Eat more vegetables and maintain your current eating pattern," and a voice output will announce, "Improvement suggestions have been received."
[1411] (Application example 2)
[1412] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1413] In modern society, it is difficult for users to easily collect life logs to maintain their health and receive lifestyle improvement suggestions based on those logs in their busy daily lives. It is also difficult to provide personalized dietary suggestions that take the user's emotions into account. Conventional systems require users to manually enter data, which is time-consuming and makes it difficult to collect accurate data. To solve these issues, a system is needed that utilizes a built-in camera, image processing technology, and an emotion engine to automatically collect life logs and provide dietary suggestions based on emotions.
[1414] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing the field of view in real time using a built-in camera, means for preprocessing the captured images to remove noise, means for sending the preprocessed images to the server in batch format at specific time intervals, means for analyzing the received image data to identify important objects, means for categorizing the user's behavior based on the analysis results, means for recording the behavioral data as text information in a life log, means for analyzing the recorded life log to generate lifestyle improvement suggestions, means for transmitting the generated suggestions and life log to the terminal, means for notifying the user of the suggestions and life log, means for analyzing the user's emotional state and generating personalized meal suggestions based on the emotional state, and means for notifying the user of the generated meal suggestions. This allows the user to receive personalized lifestyle improvement suggestions and meal suggestions based on their emotions without performing any special operations.
[1415] The "built-in camera" is a camera device built into the eyeglass-type terminal to capture the user's field of view in real time.
[1416] "Preprocessing" refers to initial image processing operations to remove noise and improve clarity from a captured image.
[1417] The "batch format" is a data transfer method in which preprocessed image data is sent in batches at regular time intervals.
[1418] A "server" is a remote computer system that receives, analyzes, and stores data sent from user terminals, and generates suggestions based on the results of the analysis.
[1419] "Analysis" is the process of identifying significant objects from the received image data and discerning the user's behavior and emotional state.
[1420] The "behavior category" is a category for classifying the user's daily behavior into specific activity types based on the analysis results.
[1421] A "life log" is a digital history that records and saves a user's daily behavioral and emotional data as text information.
[1422] "Improvement suggestions" are specific advice for improving the user's lifestyle habits that are generated from the analysis results of the life log.
[1423] "Emotional state" refers to the type of emotion (e.g., joy, sadness, anger) analyzed from the user's facial image.
[1424] "Meal Suggestions" are personalized meal suggestions generated based on the user's emotional state and past meal data.
[1425] "Notification" is the process of sending messages to communicate generated suggestions and lifelogs to users.
[1426] The system of the present invention includes a camera built into an eyeglass-type terminal, a server, and a user's terminal (a smartphone or smart glasses).
[1427] First, the device's built-in camera captures the user's field of view in real time, capturing images every second. These captured images are then pre-processed using the device's image processing capabilities. This pre-processing uses image processing techniques such as the OpenCV library to remove noise and improve image clarity.
[1428] The preprocessed image data is sent to the server in batches at regular intervals. A secure communication protocol (HTTPS) is used for transmission to ensure data security. Emotion data is also sent to the server.
[1429] The server analyzes the received image data using a deep learning framework (such as TensorFlow or YOLO) to identify important objects. Based on the analysis results, it classifies the user's behavior into categories (e.g., eating, exercise, travel). It also uses facial recognition technology to analyze the user's emotional state and adds it to the behavioral data.
[1430] For example, if an image of a user eating a salad is captured and joy is identified from the user's emotional state, the behavioral data is recorded in the life log as "Lunch (Joy)." The server analyzes this behavioral data and emotional data and generates lifestyle improvement suggestions. This analysis uses algorithms based on the user's past behavioral logs and expert knowledge.
[1431] The server then generates personalized meal suggestions, taking into account the user's emotional state, using a generative AI model (e.g., GPT) and prompting the user with the following sentences:
[1432] "The user is having salad, soup, and bread for lunch. Their facial expressions indicate they are enjoying the meal. Please suggest a meal for this user's next meal."
[1433] The generated improvement suggestions and meal suggestions are sent back to the device and notified to the user. Notification methods include displaying the information on the device's display or using audio output. This allows the user to receive their own life log, improvement suggestions, and personalized meal suggestions and reflect them in their next actions without performing any special operations.
[1434] For example, if a user eats a salad at lunch and the emotion of joy is detected, the device will notify them with a suggestion: "We recommend that you incorporate more vegetables into your next meal to improve balance and maintain your current eating pattern." In this way, users can receive practical advice based on their emotions and behavioral data.
[1435] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1436] Step 1:
[1437] The device uses the built-in camera to capture the user's field of view at one-second intervals. At this time, the camera API operates and generates image data obtained from the user's field of view. The input is raw data from the camera, and the output is the captured image file.
[1438] Step 2:
[1439] The device pre-processes the captured image to remove noise and improve clarity. This process uses an image processing library (OpenCV). The input is the captured image file, and the output is a pre-processed, clear image file.
[1440] Step 3:
[1441] The terminal sends the preprocessed image data to the server in batch format at specific time intervals. A secure communication protocol (HTTPS) is used for data transmission to ensure data security. The input is the preprocessed image data, and the output is the data transfer to the server.
[1442] Step 4:
[1443] The server analyzes the received image data using a deep learning framework (TensorFlow, YOLO) to identify important objects. This analysis step recognizes and classifies objects in the image (e.g., food, people, landmarks, etc.). The input is the transmitted image data, and the output is the analyzed object information.
[1444] Step 5:
[1445] Based on the analysis results, the server classifies the user's behavior into categories (e.g., eating, exercise, travel). It also analyzes the user's emotional state using facial recognition technology and adds it to the behavioral data. For example, the data is recorded as "lunch (joy)." The input is the analyzed object information and emotional data, and the output is behavioral data classified by category.
[1446] Step 6:
[1447] The server analyzes this behavioral and emotional data and generates lifestyle improvement suggestions. The algorithm utilizes the user's past behavioral logs and expert knowledge. The input is behavioral and emotional data categorized by category, and the output is the generated improvement suggestions.
[1448] Step 7:
[1449] The server generates personalized meal suggestions using a generative AI model (e.g., GPT) based on the user's emotional state. For example, you can get suggestions by entering a prompt like this:
[1450] "The user is having salad, soup, and bread for lunch. Their facial expressions indicate they are enjoying the meal. Please suggest a meal for this user's next meal."
[1451] The input is a prompt based on the user's behavioral and emotional data, and the output is personalized meal suggestions.
[1452] Step 8:
[1453] The server sends the generated improvement suggestions and meal suggestions to the terminal, which receives them and notifies the user. The input is the generated suggestions and notification data, and the output is the notification to the user.
[1454] Step 9:
[1455] The user receives notifications from the device, refers to their own life log, improvement suggestions, and personalized meal suggestions, and reflects them in their next actions. The input is the notification information from the device, and the output is the action selected by the user.
[1456] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1457] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1458] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1459] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1460] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1461] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1462] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1463] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1464] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1465] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1466] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1467] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1468] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1469] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1470] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1471] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1472] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1473] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1474] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1475] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1476] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1477] The following is further disclosed regarding the above embodiment.
[1478] (Claim 1)
[1479] A built-in camera captures the field of view in real time,
[1480] A means of pre-processing the captured image to remove noise;
[1481] means for transmitting the preprocessed images in batches to a server at specific time intervals;
[1482] means for analyzing the received image data to identify objects of interest;
[1483] A means of categorizing user behavior based on the analysis results;
[1484] A means for recording behavioral data as text information in a life log;
[1485] A means for analyzing the recorded life log and generating lifestyle improvement suggestions;
[1486] A means for transmitting the generated suggestions and life logs to the terminal;
[1487] A means of notifying users of suggestions and life logs;
[1488] A system including:
[1489] (Claim 2)
[1490] 10. The system of claim 1, wherein the on-board camera captures an image of the field of view every second.
[1491] (Claim 3)
[1492] 2. The system according to claim 1, wherein lifestyle improvement suggestions are generated based on expert knowledge.
[1493] "Example 1"
[1494] (Claim 1)
[1495] A built-in camera captures the field of view in real time,
[1496] A means of pre-processing the captured image to remove noise;
[1497] means for transmitting the preprocessed images in batches to a server at specific time intervals;
[1498] a means for inputting the received image data into a generative AI model to identify objects of interest;
[1499] A means of categorizing user behavior based on the analysis results;
[1500] A means for recording behavioral data as text information in a life log;
[1501] A means for analyzing recorded life logs using an algorithm based on specialized knowledge to generate lifestyle improvement suggestions;
[1502] A means for transmitting the generated suggestions and life logs to the terminal;
[1503] A means of notifying users of suggestions and life logs;
[1504] A system including:
[1505] (Claim 2)
[1506] 10. The system of claim 1, wherein the built-in camera captures images of the field of view at specific time intervals (every second).
[1507] (Claim 3)
[1508] The system of claim 1, wherein lifestyle improvement suggestions are generated using an algorithm based on expert knowledge for analyzing the recorded life log.
[1509] "Application Example 1"
[1510] (Claim 1)
[1511] a means for capturing a field of view in real time by an internal imaging device;
[1512] A means of pre-processing the captured image to remove noise;
[1513] means for transmitting the preprocessed images in batches to a server at specific time intervals;
[1514] means for analyzing the received image data to identify objects of interest;
[1515] A means for classifying user behavior based on the analysis results;
[1516] A means for storing behavioral data as text information in a life record;
[1517] A means for analyzing the recorded life log and generating lifestyle improvement suggestions;
[1518] A means for transmitting the generated suggestions and life records to a terminal;
[1519] A means for notifying users of suggestions and life records;
[1520] A way to measure customer interest in specific products or areas of the store,
[1521] A means of notifying store managers and staff of the analysis results,
[1522] A system including:
[1523] (Claim 2)
[1524] 10. The system of claim 1, wherein the on-board imaging device captures an image of the field of view every second.
[1525] (Claim 3)
[1526] 2. The system according to claim 1, wherein lifestyle improvement suggestions are generated based on expert knowledge.
[1527] "Example 2: Combining Emotion Engines"
[1528] (Claim 1)
[1529] A built-in camera captures the field of view in real time,
[1530] A means of pre-processing the captured image to remove noise;
[1531] means for transmitting the preprocessed images in batches to a server at specific time intervals;
[1532] means for analyzing the received image data to identify objects of interest;
[1533] A means of categorizing user behavior based on the analysis results;
[1534] A means for analyzing a user's emotions using facial recognition technology;
[1535] a means for recording behavioral data and emotional data as character information in a life log;
[1536] A means for analyzing the recorded life log and generating lifestyle improvement suggestions;
[1537] A means for transmitting the generated suggestions and life logs to the terminal;
[1538] A means of notifying users of suggestions and life logs;
[1539] A system including:
[1540] (Claim 2)
[1541] 10. The system of claim 1, wherein the on-board camera captures an image of the field of view every second.
[1542] (Claim 3)
[1543] 2. The system according to claim 1, wherein lifestyle improvement suggestions are generated based on expert knowledge and past behavior logs.
[1544] "Application example 2 when combining emotion engines"
[1545] (Claim 1)
[1546] A built-in camera captures the field of view in real time,
[1547] A means of pre-processing the captured image to remove noise;
[1548] means for transmitting the preprocessed images in batches to a server at specific time intervals;
[1549] means for analyzing the received image data to identify objects of interest;
[1550] A means of categorizing user behavior based on the analysis results;
[1551] A means for recording behavioral data as text information in a life log;
[1552] A means for analyzing the recorded life log and generating lifestyle improvement suggestions;
[1553] A means for transmitting the generated suggestions and life logs to the terminal;
[1554] A means of notifying users of suggestions and life logs;
[1555] means for analyzing a user's emotional state and generating personalized meal suggestions based thereon;
[1556] means for notifying a user of the generated meal suggestions;
[1557] A system including:
[1558] (Claim 2)
[1559] 10. The system of claim 1, wherein the on-board camera captures an image of the field of view every second.
[1560] (Claim 3)
[1561] 2. The system according to claim 1, wherein lifestyle improvement suggestions and dietary suggestions are generated based on expert knowledge. [Explanation of symbols]
[1562] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A built-in camera captures the field of view in real time, A means of pre-processing the captured image to remove noise; means for transmitting the preprocessed images in batches to a server at specific time intervals; means for analyzing the received image data to identify objects of interest; A means of categorizing user behavior based on the analysis results; A means for recording behavioral data as text information in a life log; A means for analyzing the recorded life log and generating lifestyle improvement suggestions; A means for transmitting the generated suggestions and life logs to the terminal; A means of notifying users of suggestions and life logs; A system including:
2. 10. The system of claim 1, wherein the on-board camera captures an image of the field of view every second.
3. 2. The system according to claim 1, wherein the lifestyle improvement suggestions are generated based on specialized knowledge.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A