System

A wearable device with image recognition and a server-based analysis system addresses the challenges of manual recording and privacy concerns, enabling efficient lifestyle management and personalized advice.

JP2026022471APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024123988
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing methods for recording lifestyle habits and activities are time-consuming, laborious, and pose privacy concerns, making it difficult for users to maintain self-management and receive accurate feedback on their improvements.

Method used

A system comprising a wearable device with image recognition capabilities that captures and analyzes user behavior, a server for data analysis, and a notification mechanism to protect privacy, providing advice based on the analysis results.

Benefits of technology

Enables efficient and hassle-free lifestyle management by automatically recording and analyzing user behavior, offering personalized advice while ensuring privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022471000001_ABST
    Figure 2026022471000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: a device with image recognition capabilities; means for transmitting image data captured by the device to a server; means for analyzing the image data at the server to identify user behavior; and means for providing advice to a user based on the analysis.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Many people today want to improve their lifestyle habits and record their daily activities, but the time-consuming and laborious process makes it difficult to maintain self-management. Manually recording trips and special events is also tedious and poses the risk of missing events. Meanwhile, image recording devices raise privacy and security concerns, creating significant barriers for users. This invention aims to solve these problems and provide a system for efficiently and safely recording lifestyle habits and activities. [Means for solving the problem]

[0005] The present invention provides a system including a device with an image recognition function, means for transmitting image data acquired by the device to a server, means for analyzing the image data in the server and identifying user behavior, and means for providing advice to the user based on the analysis.

[0006] Specifically, the device includes a means for analyzing the dietary content contained in the image data and providing advice on nutritional balance, a means for analyzing the exercise status contained in the image data and providing advice on improving lifestyle habits, a means for identifying and automatically recording scenery during travel from the image data, and a means for notifying the user that the device is recording images, thereby resolving privacy issues.

[0007] Furthermore, by using the image data stored in the server, the system can efficiently manage data by categorizing the user's behavioral data. In this way, the system provides a user with a hassle-free and efficient way to improve their lifestyle habits and record their behavior.

[0008] A "device with image recognition functionality" is a wearable device that has a built-in camera and image analysis functionality and is capable of capturing and recognizing the surrounding environment and the user's behavior.

[0009] "Image data" is digital data containing visual information acquired by a device with image recognition capabilities.

[0010] The "server" is a central processing unit that receives, stores, and analyzes acquired image data, generates advice based on the analysis results, and transmits the advice to the user terminal.

[0011] "Analysis" is the process of recognizing elements contained in image data and classifying and evaluating them.

[0012] "User behavior" refers to actions and events related to the user's daily activities and lifestyle habits.

[0013] "Advice" refers to specific suggestions or instructions for improving the user's behavior or lifestyle based on the analysis results.

[0014] "Nutritional balance" refers to the appropriate distribution and balance of nutrients in the diet taken by the user.

[0015] "Exercise status" refers to observation results regarding the type, frequency, and intensity of physical activity performed by the user.

[0016] "Scenery during travel" refers to visual information such as scenery and tourist spots that the user sees during the trip.

[0017] A "means for notifying the user that an image is being recorded" is a visual or audio interface that notifies the user while the device is taking an image.

[0018] "Categorizing" is the process of organizing analyzed behavioral data according to specific themes or purposes and storing them in a database. [Brief explanation of the drawings]

[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0027] [First embodiment]

[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0040] This invention proposes a system that uses a device and a server equipped with image recognition functionality to automatically record user behavior and provide advice based on the analysis results. Hereinafter, an embodiment of the invention will be described in detail.

[0041] System configuration

[0042] 1. Devices with image recognition capabilities

[0043] Device (glasses): This device has a built-in camera and image analysis capabilities, and automatically captures the environment and actions in front of the user's eyes. The glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that images are being recorded.

[0044] 2. Transmission and storage of image data

[0045] Device (glasses): Sends collected image data to the server at regular intervals. This data includes the time of shooting and location information.

[0046] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[0047] 3. Analysis of image data

[0048] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (food, exercise equipment, scenery, etc.) are identified and assigned classification tags.

[0049] 4. Lifestyle Analytics

[0050] Server: Based on the analysis results, the server evaluates the user's diet, exercise, and behavioral data, including calorie calculations, nutritional balance evaluations, and exercise frequency and intensity evaluations.

[0051] Server: Generates specific advice based on the evaluation results and stores it in a database.

[0052] 5. Advice Notification

[0053] Server: Sends the generated advice to the user's device (glasses, smartphone, etc.) in a format that is easy for the user to understand and follow.

[0054] Device (glasses or smartphone): Notifies the user of the received advice by displaying a visual message or making an audio notification.

[0055] Specific examples

[0056] Food records and advice

[0057] User: Eat lunch.

[0058] Device (glasses): Automatically takes photos of the user while they are eating and stores the image data in local storage.

[0059] Device (glasses): Sends image data of lunch to the server at the 12:00 PM interval.

[0060] Server: Analyzes the received image data and identifies the type and quantity of food.

[0061] Server: Calculates calories and nutrients and stores the evaluation results in a database.

[0062] Server: Generates advice such as "There weren't many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[0063] Device (smartphone): Displays a message to the user saying, "Today's lunch is high in calories, so you should add a salad to your dinner."

[0064] Travel Log

[0065] User: Walking around tourist spots.

[0066] Device (glasses): Automatically takes photos of important tourist spots and stores the image data in local storage.

[0067] Device (glasses): While walking around the tourist spot, the device periodically sends image data to the server.

[0068] Server: Analyzes the received image data and automatically identifies tourist spots. The image data also contains metadata (location, time, etc.).

[0069] Server: After the trip, the user's travel album is automatically generated based on the saved image data.

[0070] Server: Once the album is created, notify the user and provide an access link.

[0071] Device (smartphone): Display the travel album so that the user can check it.

[0072] This system allows users to improve their lifestyle habits and record their activities efficiently without any hassle. Furthermore, it has mechanisms for protecting privacy (such as notifications during recording), so users can use it with peace of mind.

[0073] The processing flow will be explained below.

[0074] Food record and advice processing steps

[0075] Step 1:

[0076] User: The user starts eating. The device (glasses) is turned on.

[0077] Step 2:

[0078] Device (glasses): The camera takes a picture of the food in front of the user's eyes.

[0079] Step 3:

[0080] Device (glasses): Captured image data is temporarily stored in local storage. When saved, the time of capture and location information are also added.

[0081] Step 4:

[0082] Device (glasses): At regular intervals (e.g., every hour), the image data stored in the local storage is sent to the server. The data includes the image file, the time of shooting, and location information.

[0083] Step 5:

[0084] Server: Stores the received image data in a database for analysis.

[0085] Step 6:

[0086] Server: Applying image recognition algorithms to analyze the received image data. This process involves identifying objects in the image (e.g., type and quantity of food).

[0087] Step 7:

[0088] Server: Based on the analysis results, calculates the calories of the food and evaluates the nutrients. Stores the results in a database.

[0089] Step 8:

[0090] Server: Generates specific advice based on the evaluation results. For example, "There weren't many vegetables today, so it would be good to add a salad to dinner."

[0091] Step 9:

[0092] Server: Sends the generated advice to the user's smartphone.

[0093] Step 10:

[0094] Device (smartphone): The received advice is displayed visually or notified by voice, allowing the user to check it and take action.

[0095] Travel Record Processing Steps

[0096] Step 1:

[0097] User: The user begins walking around the tourist spot. The device (glasses) is turned on.

[0098] Step 2:

[0099] Device (glasses): Automatically captures important tourist spots within the user's field of view.

[0100] Step 3:

[0101] Device (glasses): Captured image data is temporarily stored in local storage, along with the capture time and location information.

[0102] Step 4:

[0103] Device (glasses): Sends stored image data to the server at regular intervals (e.g., every hour).

[0104] Step 5:

[0105] Server: Stores the received image data in a database for analysis.

[0106] Step 6:

[0107] Server: Using an image recognition algorithm, the received image data is analyzed and tourist spots are automatically recognized. Metadata (location, time) is also associated with the images.

[0108] Step 7:

[0109] Server: Combines image data from multiple tourist spots to generate a travel album.

[0110] Step 8:

[0111] Server: Once the album is created, notify the user and provide an access link.

[0112] Step 9:

[0113] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album.

[0114] In this way, the process steps can be used to efficiently improve lifestyle habits and record travel. Privacy is also protected by a user notification function.

[0115] Example 1

[0116] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0117] With current technology, it takes a lot of time and effort to record a user's behavior in detail and provide appropriate advice. Furthermore, manual recording and analysis can lack accuracy and consistency. Furthermore, there is a risk that users will lose motivation to improve their lifestyle habits because there is a lack of a way for them to receive specific feedback on their improvements. For this reason, there is a need for a system that can automatically record a user's behavior, analyze it, and provide advice.

[0118] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0119] In this invention, the server includes means for analyzing image data to identify objects and assign classification tags, means for evaluating user behavior based on the analysis results, and means for generating specific advice using a generative AI model, thereby enabling automatic analysis of user behavior data and prompt provision of appropriate advice.

[0120] "Image recognition" refers to a device's ability to analyze images and identify objects.

[0121] "Device" refers to an electronic device that photographs and records user behavior and transmits the data to a server at regular intervals.

[0122] "Server" refers to a computer system that receives data sent from devices via a network, analyzes the data, and processes and stores the results.

[0123] "Image data" refers to image information captured by a device and recording the user's behavior and environment.

[0124] "Local storage" refers to a storage device installed within a device for temporarily storing data.

[0125] "Analysis results" refers to data that the server analyzes images to extract information about the user's behavior and environment, and assigns classification tags to.

[0126] A "generative AI model" is a pre-trained artificial intelligence model that uses algorithms to generate appropriate responses and advice based on input data.

[0127] "Advice" refers to a message that suggests specific guidelines for action or improvements to the user based on the analysis results.

[0128] "User terminal" refers to electronic devices that are directly used by users, such as devices and smartphones.

[0129] "Metadata" refers to supplementary information that accompanies image data, including the time of shooting, location information, and the like.

[0130] This invention proposes a system that automatically records user behavior and provides advice based on the analysis results. The system mainly includes a device with image recognition capabilities, a terminal that transmits and stores data, a server that analyzes and evaluates image data, and a means for generating specific advice using a generative AI model.

[0131] System configuration

[0132] 1. Devices with image recognition capabilities

[0133] Device (glasses):

[0134] The device has a built-in camera and image analysis capabilities, and captures real-time images of the user's behavior and environment. The captured image data is temporarily stored in local storage. The device also has a visual (e.g., LED) or audio interface to notify the user that images are being recorded.

[0135] 2. Data transmission and storage

[0136] Device (glasses):

[0137] The collected image data is sent to a server at regular intervals, including the time and location of the image.

[0138] server:

[0139] The received image data is stored in an analytical database, allowing for subsequent analysis.

[0140] 3. Analysis of image data

[0141] server:

[0142] The system applies image recognition algorithms (e.g., TensorFlow, OpenCV) to the received image data, identifies objects in the data, and assigns classification tags. For example, it identifies the type and quantity of food in a photo of a meal.

[0143] 4. Evaluation of behavioral data

[0144] server:

[0145] Based on the analysis results, the user's behavioral data is evaluated. Specific evaluation items include calorie calculation of meal contents, evaluation of nutritional balance, and evaluation of exercise frequency and intensity.

[0146] 5. Automatic Advice Generation

[0147] server:

[0148] Based on the evaluation data obtained from the analysis results, specific advice is generated using a generative AI model (e.g., GPT-3). The advice is in a format that is easy for users to understand and follow.

[0149] 6. Advice Notification

[0150] server:

[0151] The generated advice is sent to the user's device (glasses or smartphone).

[0152] Device (glasses or smartphone):

[0153] The user is notified of the received advice, either visually or by audio.

[0154] Specific examples

[0155] Food records and advice

[0156] User:

[0157] Eat lunch.

[0158] Device (glasses):

[0159] The system takes a photo of the user's meal and stores the image data in local storage. This data is sent to the server at 12:00 PM intervals.

[0160] server:

[0161] The received image data is analyzed to identify the type and quantity of food, then the calories and nutrients are calculated and the evaluation results are stored in a database.

[0162] server:

[0163] The system generates advice such as "You didn't have many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[0164] Device (smartphone):

[0165] The message "Today's lunch is high in calories, so you should add a salad to your dinner" is displayed to notify the user.

[0166] Prompt Sentence Examples

[0167] Analyze an image of a user eating lunch, identify the type and amount of food, calculate calories and nutrients, and generate appropriate recommendations.

[0168] Generate a trip record

[0169] User:

[0170] Walk around the tourist spots.

[0171] Device (glasses):

[0172] Photograph important tourist spots and store the image data in local storage.

[0173] Device (glasses):

[0174] During sightseeing, image data is periodically sent to the server.

[0175] server:

[0176] The received image data is analyzed and tourist spots are automatically recognized. The image data includes the time of shooting and location information.

[0177] server:

[0178] After the trip, a travel album is automatically generated based on the saved image data.

[0179] server:

[0180] Once the album has been created, the user will be notified and provided with an access link.

[0181] Device (smartphone):

[0182] Display the travel album so that the user can check it.

[0183] Prompt Sentence Examples

[0184] Analyze the received tourist attraction image data, identify tourist attractions based on the metadata, and generate a travel album.

[0185] This system allows users to efficiently improve their lifestyle habits and record their activities without any hassle. It also has privacy protection features (such as notifications during recording), so users can use it with peace of mind.

[0186] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0187] Step 1:

[0188] Device (glasses): The built-in camera captures the user's actions. The captured image data is temporarily stored in the device's local storage. At this time, the LED on the device lights up to notify the user that a photo is being taken.

[0189] Input: User behavior, environment

[0190] Output: Captured image data

[0191] Step 2:

[0192] Device (glasses): Image data, shooting time, and location information stored in local storage are sent to the server at regular intervals. While the data is being sent, the device's notification function is activated to notify the user that data is being sent.

[0193] Input: Image data from local storage, shooting time, location information

[0194] Output: Image data and metadata (photo time, location information) sent to the server

[0195] Step 3:

[0196] Server: Apply image recognition algorithms (e.g., TensorFlow, OpenCV) to the received image data to identify objects in the data and assign classification tags. As a result of the analysis process, noteworthy content in the image (e.g., food, tourist attractions) is identified.

[0197] Input: Image data, metadata (shooting time, location information)

[0198] Output: Analysis results (object classification tags)

[0199] Step 4:

[0200] Server: Evaluates the user's behavioral data based on the analysis results. Analyzes meal content, calculates calories, evaluates nutritional balance, and analyzes exercise to evaluate frequency and intensity.

[0201] Input: Analysis results (object classification tags)

[0202] Output: Evaluation results of behavioral data (calorie calculation, nutritional balance evaluation, exercise frequency and intensity)

[0203] Step 5:

[0204] Server: Based on the evaluation results, a generative AI model (e.g., GPT-3) is used to generate specific advice. The generated advice is easy for users to understand and follow.

[0205] Input: Evaluation results of behavioral data

[0206] Output: The generated advice

[0207] Step 6:

[0208] Server: Sends the generated advice to the user's device (glasses or smartphone). While sending the advice, the server notifies the user that a message has been received.

[0209] Input: Generated advice

[0210] Output: Advice sent to the user device (glasses or smartphone)

[0211] Step 7:

[0212] Device (smartphone or glasses): Notifies the user of the received advice, either visually (as a pop-up message) or by voice.

[0213] Input: Advice sent by the server

[0214] Output: Advice displayed to the user

[0215] As described above, this system automatically records the user's behavior and provides advice based on the analysis results, allowing for efficient lifestyle improvement and behavioral recording.It also has a notification function to protect privacy, so it can be used with peace of mind.

[0216] (Application example 1)

[0217] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0218] With traditional methods of providing customer service in brick-and-mortar stores, it is difficult to understand individual customer behavior and interests in real time, making it difficult to provide personalized service. Furthermore, staff often lack the information they need to make appropriate product recommendations to customers as needed, resulting in missed opportunities to maximize customer satisfaction. A system that can solve these issues and improve the quality of customer service is needed.

[0219] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0220] In this invention, the server includes a device with an image recognition function that captures images of customer behavior in a physical store and transmits the obtained image data to the server, a means for the server to analyze customer behavior patterns and interests and provide specific advice to store staff, and a means for identifying customer interests contained in the image data and providing advice recommending related products. This enables real-time analysis of customer behavior and interests in a physical store, enabling staff to provide appropriate advice and product recommendations.

[0221] A "device with image recognition functionality" is a device that has a built-in camera and image analysis functionality and is used to capture and analyze the surrounding environment and people's behavior.

[0222] A "server" is a computer system that receives, stores, and analyzes data over a network and provides necessary information.

[0223] "Image data" refers to data that includes visual information acquired by a photographic device such as a camera.

[0224] "Analysis" is the process of using algorithms to understand the content of acquired image data and recognize specific patterns or objects.

[0225] A "user" is a person using a device with image recognition capabilities or a person receiving advice from a server.

[0226] "Advice" is any instruction, advice, or suggestion provided to the user based on the results of the analysis.

[0227] A "brick and mortar store" is a retail establishment that exists in a physical location and offers goods and services.

[0228] A "customer" is a consumer who visits a physical store and purchases or uses goods or services.

[0229] A "behavioral pattern" is a series of actions and trends in interests that indicate how customers move around the store and what they are interested in.

[0230] "Recommendations" are the suggestion of products or services based on a customer's interests.

[0231] This invention provides a system for improving customer service by analyzing customer behavior in a physical store in real time using a device and server equipped with image recognition functionality.

[0232] System configuration

[0233] 1. Devices with image recognition capabilities

[0234] Terminal (smart glasses): This device has a built-in camera and image analysis capabilities, and automatically captures customer behavior in the store. The smart glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that an image is being recorded.

[0235] 2. Transmission and storage of image data

[0236] Terminal (smart glasses): Sends collected image data to the server at regular intervals, including the time of shooting and location information.

[0237] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[0238] 3. Analysis of image data

[0239] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (e.g., products, behavioral patterns, etc.) are identified and assigned classification tags.

[0240] 4. Customer Service Analytics

[0241] Server: Evaluates customer behavior patterns and interests based on the analysis results, including identifying customer areas of interest and suggesting related products.

[0242] Server: Generates specific advice based on the evaluation results and stores it in a database.

[0243] 5. Advice Notification

[0244] Server: Sends the generated advice to the store staff's devices (smart glasses, smartphones, etc.) in a format that is easy for staff to understand and implement.

[0245] Terminal (smart glasses or smartphone): Notifies staff of received advice, either by displaying a visual message or by making an audio notification.

[0246] Specific examples

[0247] Product Recommendations

[0248] Customer: Spends a long time in the store looking at new jackets.

[0249] Terminal (smart glasses): Automatically captures customer behavior and stores the image data in local storage.

[0250] Terminal (smart glasses): Sends image data to the server at regular intervals.

[0251] Server: Analyzes the received image data and identifies the customer's interests. For example, the analysis result may be, "This customer is interested in new jackets."

[0252] Server: Generates advice such as "Suggest sales information for related products to this customer" and sends it to the staff member's smartphone.

[0253] Device (smartphone): Displays the message "Please suggest related products for this customer's new jacket" and notifies the staff.

[0254] Improved customer service

[0255] Customer: Walks around the store, checking out multiple products.

[0256] Terminal (smart glasses): Automatically captures customer behavior and stores the image data in local storage.

[0257] Terminal (smart glasses): Sends image data to the server at regular intervals.

[0258] Server: Analyzes the received image data and identifies customer behavior patterns. For example, it recognizes a behavior pattern of "looking at many products in a short period of time."

[0259] Server: Generates advice such as, "This customer is interested in many products, so we will approach them individually and introduce them to special offers," and sends this advice to the staff member's smartphone.

[0260] Terminal (smartphone): Display the message "Inform customers who are viewing many products in a short time about special offers" and notify the staff.

[0261] Example of input prompt for generative AI model

[0262] Image data showing a user looking at a particular product for a long time in the store was sent to a server. The server used an image recognition algorithm to identify the customer's interests and generate a recommendation: "This customer is interested in a new jacket." Staff receive this notification and suggest sales information for the new jacket and related items to the customer. Other visual information used by the AI ​​model includes pop-up information about related products and promotional codes.

[0263] This system allows physical stores to understand customer behavior and interests in real time, enabling staff to provide appropriate advice and product recommendations to customers.

[0264] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0265] Step 1:

[0266] The device (smart glasses) acquires the following information: The input includes an image of the environment seen by the user. The output is to temporarily store the acquired image data in local storage. Specifically, the device's camera takes pictures at regular intervals and stores the data in local storage.

[0267] Step 2:

[0268] The device (smart glasses) transmits image data stored in local storage to a server at regular intervals. The input includes the stored image data and its metadata (time of capture, location information). As output, this data is transmitted to the server. Specifically, the communication module in the device transmits the data to the server and confirms that the transmission was successful.

[0269] Step 3:

[0270] The server stores the received image data in an analysis database. The input includes the image data sent from the terminal and the associated metadata. The output is stored in a specified format in the database. Specifically, the server's storage system stores the received data in the appropriate format.

[0271] Step 4:

[0272] The server analyzes image data stored in a database based on image recognition algorithms. The input includes the stored image data. The output is the identification of objects and behavioral patterns in the image and the assignment of classification tags. Specifically, the server's processing unit executes the image recognition model and analyzes the results.

[0273] Step 5:

[0274] The server evaluates the customer's behavioral patterns and interests based on the analysis results. The input includes the analyzed data. The output is the evaluation results stored in a database. Specifically, the evaluation algorithm extracts behavioral patterns based on the analysis results and generates evaluation results.

[0275] Step 6:

[0276] The server generates specific advice based on the evaluation results. The input includes the evaluation results stored in the database. As an output, the generated advice message is sent to the database and the user terminal. As a specific operation, the advice generation algorithm is executed and the result is formatted in a format suitable for the user interface.

[0277] Step 7:

[0278] The terminal (smart glasses or smartphone) notifies the staff of the advice sent from the server. The input includes the advice message sent from the server. The output is a visual message display or a voice notification. As a specific operation, the display module or voice output module of the terminal is activated and the advice is provided to the staff.

[0279] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0280] This invention proposes a system that uses a device with image recognition capabilities, a server, and an emotion engine to automatically record a user's behavior, identify the user's emotional state, and provide advice based on the analysis results. Hereinafter, an embodiment of the present invention will be described in detail.

[0281] System configuration

[0282] 1. Devices with image recognition capabilities

[0283] Device (glasses): This device has a built-in camera and image analysis capabilities, and automatically captures the environment and actions in front of the user's eyes. The glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that images are being recorded.

[0284] 2. Transmission and storage of image data

[0285] Device (glasses): Sends collected image data to the server at regular intervals. This data includes the time of shooting and location information.

[0286] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[0287] 3. Analysis of image data

[0288] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (food, exercise equipment, scenery, etc.) are identified and assigned classification tags.

[0289] 4. Emotional state analysis

[0290] Server: Using the emotion engine, analyzes the user's facial expressions from the received image data and identifies their emotional state. For example, it determines whether the user is happy or stressed.

[0291] 5. Lifestyle Analytics

[0292] Server: Based on the analysis results, the server evaluates the user's diet, exercise, and emotional state, including calorie calculations, nutritional balance assessments, exercise frequency and intensity, and emotional state assessments.

[0293] Server: Generates specific advice based on the evaluation results and stores it in a database. For example, if the user is feeling emotionally stressed, it will suggest relaxation techniques.

[0294] 6. Advice Notification

[0295] Server: Sends the generated advice to the user's device (glasses, smartphone, etc.) in a format that is easy for the user to understand and follow.

[0296] Device (glasses or smartphone): Notifies the user of the received advice by displaying a visual message or making an audio notification.

[0297] Specific examples

[0298] Food records and advice

[0299] User: Eat lunch.

[0300] Device (glasses): Automatically takes photos of the user while they are eating and stores the image data in local storage.

[0301] Device (glasses): Sends image data of lunch to the server at the 12:00 PM interval.

[0302] Server: Analyzes the received image data and identifies the type and quantity of food.

[0303] Server: Calculates calories and nutrients and stores the evaluation results in a database.

[0304] Server: Generates advice such as "There weren't many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[0305] Device (smartphone): Displays a message to the user saying, "Today's lunch is high in calories, so you should add a salad to your dinner."

[0306] Server: The emotion engine analyzes the user's facial expressions while they are eating and evaluates whether they are enjoying the meal. For example, if the user is not enjoying the meal, it can add advice such as "Next time, try incorporating your favorite dishes."

[0307] Travel Records and Sentiment Analysis

[0308] User: Walking around tourist spots.

[0309] Device (glasses): Automatically takes photos of important tourist spots and stores the image data in local storage.

[0310] Device (glasses): While walking around the tourist spot, the device periodically sends image data to the server.

[0311] Server: Analyzes the received image data and automatically identifies tourist spots. The image data also contains metadata (location, time).

[0312] Server: After the trip, the user's travel album is automatically generated based on the saved image data.

[0313] Server: The emotion engine analyzes the user's facial expressions while sightseeing to identify the places and moments they particularly enjoyed. Based on this information, it decides which points to highlight in the travel album.

[0314] Server: Once the album is created, notify the user and provide an access link.

[0315] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album.

[0316] This system not only allows users to efficiently improve their lifestyle habits and record their activities without any hassle, but also analyzes their emotional state and provides more personalized advice.It also has mechanisms for protecting privacy (such as notifications during recording), so users can use it with peace of mind.

[0317] The processing flow will be explained below.

[0318] Food Record and Emotion Analysis Processing Steps

[0319] Step 1:

[0320] User: The user starts eating. The device (glasses) is turned on.

[0321] Step 2:

[0322] Device (glasses): The camera takes a picture of the food in front of the user's eyes.

[0323] Step 3:

[0324] Device (glasses): Captured image data is temporarily stored in local storage. When saved, the time of capture and location information are also added.

[0325] Step 4:

[0326] Device (glasses): At regular intervals (e.g., every hour), the image data stored in the local storage is sent to the server. The data includes the image file, the time of shooting, and location information.

[0327] Step 5:

[0328] Server: Stores the received image data in a database for analysis.

[0329] Step 6:

[0330] Server: Applying image recognition algorithms to analyze the received image data. This process involves identifying objects in the image (e.g., type and quantity of food).

[0331] Step 7:

[0332] Server: Based on the analysis results, calculates the calories of the food and evaluates the nutrients. Stores the results in a database.

[0333] Step 8:

[0334] Server: Using the emotion engine, analyze the user's facial expressions from the image data and identify their emotional state. Evaluate whether the user is enjoying the meal or feeling stressed.

[0335] Step 9:

[0336] Server: Generates specific advice based on the evaluation results. For example, "You didn't eat many vegetables today, so you should add a salad to your dinner." It also generates emotion-based advice, such as "You seemed a little nervous during the meal, so next time try eating in a relaxed environment."

[0337] Step 10:

[0338] Server: Sends the generated advice to the user's smartphone.

[0339] Step 11:

[0340] Device (smartphone): The received advice is displayed visually or notified by voice, allowing the user to check it and take action.

[0341] Processing steps for travel records and sentiment analysis

[0342] Step 1:

[0343] User: The user begins walking around the tourist spot. The device (glasses) is turned on.

[0344] Step 2:

[0345] Device (glasses): Automatically captures important tourist spots within the user's field of view.

[0346] Step 3:

[0347] Device (glasses): Captured image data is temporarily stored in local storage, along with the capture time and location information.

[0348] Step 4:

[0349] Device (glasses): Sends stored image data to the server at regular intervals (e.g., every hour).

[0350] Step 5:

[0351] Server: Stores the received image data in a database for analysis.

[0352] Step 6:

[0353] Server: Using an image recognition algorithm, the received image data is analyzed and tourist spots are automatically recognized. Metadata (location, time) is also associated with the images.

[0354] Step 7:

[0355] Server: Using the emotion engine, analyze the user's facial expressions while sightseeing to identify their emotional state, and identify places and moments that they are particularly enjoying.

[0356] Step 8:

[0357] Server: Combines image data from multiple tourist spots to generate a travel album, placing the images in a way that highlights moments of positive emotional states.

[0358] Step 9:

[0359] Server: Once the album is created, notify the user and provide an access link.

[0360] Step 10:

[0361] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album, highlighting the moments they particularly enjoyed.

[0362] In this way, a system can be realized that records and analyzes the user's behavior and emotions in detail through specific processing steps and provides appropriate advice.

[0363] Example 2

[0364] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0365] Conventional lifestyle improvement systems require users to manually record their behavior, diet, and exercise, which is time-consuming for users. Furthermore, it is difficult to analyze emotional states and provide advice based on them, which means that more personalized advice cannot be provided to users. There is a need for a system that can solve these issues and enable users to efficiently improve their lifestyles and record their behavior.

[0366] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0367] In this invention, the server includes a terminal with an image recognition function, means for transmitting image data acquired by the terminal to the server, means for analyzing the image data in the server and identifying the user's behavior and emotional state, and means for providing advice to the user based on the analysis. This makes it possible to automatically record the user's behavior and emotional state and provide personalized, specific advice based on the analysis results.

[0368] A "terminal with image recognition functionality" is a device that has the ability to automatically capture images of the environment and behavior in front of the user's line of sight and acquire them as image data.

[0369] "Means for transmitting acquired image data to a server" refers to a function for transferring image data from a terminal to a server at regular intervals.

[0370] "Means for analyzing image data and identifying user behavior and emotional state" refers to the process of analyzing and identifying user behavior (e.g., eating, exercise) and emotional state (e.g., joy, anger) based on the image data received by the server using an image recognition algorithm and an emotion engine.

[0371] "Means for providing advice" refers to the function of generating specific guidelines for the user (e.g., adding a salad to dinner) based on the analysis results and notifying the user's device.

[0372] "Means for providing advice based on nutritional balance and calorie calculation" refers to the process of analyzing dietary content based on image data, calculating nutrient balance and calories, and making suggestions for dietary improvement based on that.

[0373] "Means for providing advice on the frequency and intensity of exercise" refers to the process of analyzing the exercise captured in the image data, evaluating the frequency and intensity of exercise, and making suggestions that will help improve lifestyle habits.

[0374] An "emotion engine" refers to an algorithm or software that analyzes a user's facial expressions contained in image data and identifies emotional states such as joy, anger, and sadness.

[0375] A "generative AI model" refers to a model that uses artificial intelligence to generate personalized advice based on the user's analysis results.

[0376] System Overview

[0377] This invention is a system that uses a device with image recognition capabilities (e.g., glasses with a built-in camera), a server, and an emotion engine to automatically record a user's behavior, identify the user's emotional state, and provide advice based on the analysis results. The main hardware and software configuration of this system is as follows:

[0378] Hardware used

[0379] 1. Device (glasses with built-in camera): Automatically captures the environment and actions in front of the user's line of sight and acquires image data.

[0380] 2. Server: Analyzes the acquired image data, generates advice based on the analysis results, and notifies the user.

[0381] 3. User device (smartphone or tablet): Notifies the user of the advice sent from the server.

[0382] Software and algorithms used

[0383] 1. Image recognition algorithms: Identify objects in images using OpenCV, TensorFlow, etc.

[0384] 2. Emotion engine: Analyzes the user's emotional state using Microsoft Azure's Face API and Emotion API.

[0385] 3. Generative AI model: An artificial intelligence model that generates personalized advice based on the user's analysis results.

[0386] Process Overview

[0387] Automatic image capture and data collection

[0388] The device (glasses with a built-in camera) automatically captures the environment and actions in front of the user's eyes and saves the image data in local storage. The glasses are equipped with a visual or audio interface to notify the user that a photo is being taken.

[0389] Sending image data

[0390] The device (glasses with a built-in camera) collects image data and sends it to the server at regular intervals. The image data includes the time of capture and location information.

[0391] Image data storage and analysis

[0392] The server stores the received image data in an analysis database and applies image recognition algorithms to analyze it, identifying objects in the image and assigning them classification tags.

[0393] Emotional state analysis

[0394] The server uses an emotion engine to analyze the user's facial expressions contained in the image data, thereby identifying the user's emotional state, for example, whether they are happy or stressed.

[0395] Advice Generation

[0396] The server evaluates the user's diet, exercise habits, and emotional state based on the analysis results and generates specific advice, which is personalized using a generative AI model.

[0397] Specific examples

[0398] Food records and advice

[0399] When a user eats lunch, the device (glasses with a built-in camera) automatically takes pictures of the meal and collects image data. At 12:00 PM intervals, the collected image data of the lunch is sent to the server. The server analyzes the received image data and identifies the type and amount of food. It then calculates calories and nutrients and generates advice such as "You didn't eat many vegetables today, so you should add a salad to your dinner." This advice is sent to the user's device (such as a smartphone) and notified to the user. The user's facial expressions while eating can also be analyzed to evaluate whether they are enjoying their meal. For example, if the user is not enjoying their meal, the system can add advice such as "Next time, try incorporating your favorite dishes."

[0400] Prompt Sentence Examples

[0401] "Please explain how you can analyze the user's emotional state and generate advice to suggest relaxation if they are feeling stressed."

[0402] In this way, the present invention not only allows users to efficiently improve their lifestyle habits and record their activities without any hassle, but also analyzes their emotional state and allows them to receive more personalized advice. Furthermore, the system is equipped with mechanisms for protecting privacy (such as notifications during recording), allowing users to use the system with peace of mind.

[0403] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0404] System processing steps

[0405] Step 1: Automatic image capture and data collection

[0406] The device (glasses with a built-in camera) automatically captures the environment and actions in front of the user's eyes in real time.

[0407] Input: The scene in front of the user's line of sight

[0408] Output: Acquired image data (photos of the environment and behavior)

[0409] Specifically, the glasses' built-in camera periodically takes a picture and temporarily stores the video data in local storage. The glasses provide visual and audio feedback to let the user know that they are recording.

[0410] Step 2: Sending image data

[0411] The terminal (glasses with a built-in camera) sends the collected image data to a server at a predetermined interval (for example, every 5 minutes).

[0412] Input: Image data stored in local storage

[0413] Output: Image data sent to the server (including shooting time and location information)

[0414] Specifically, the device periodically divides image data into packets and sends them to the server via a wireless network (Wi-Fi or mobile data).

[0415] Step 3: Save the image data

[0416] The server stores the received image data in an analysis database.

[0417] Input: Image data sent from the device (including shooting time and location information)

[0418] Output: Image data stored in a database for analysis

[0419] Specifically, the server checks the format and integrity of the received data, stores it correctly in a database for analysis, and implements security measures to ensure the data is managed in a protected environment.

[0420] Step 4: Image Recognition and Classification

[0421] The server applies an image recognition algorithm to the image data stored in the analysis database and performs analysis.

[0422] Input: Image data stored in a database

[0423] Output: Image data with identified objects and classification tags

[0424] Specifically, it uses image recognition libraries such as OpenCV and TensorFlow to identify objects in the image (e.g., food, exercise equipment, landscapes, etc.) The objects obtained as a result of the analysis are given classification tags and stored again in the database.

[0425] Step 5: Analyze emotional state

[0426] The server uses an emotion engine to analyze the user's facial expressions in the image data.

[0427] Input: Image data that has been recognized and classified

[0428] Output: Identification of the user's emotional state (e.g., happy, angry, sad)

[0429] Specifically, it uses Microsoft Azure's Face API and Emotion API to identify the user's emotional state from their facial expressions. The analysis results are stored in a database and used for subsequent processing.

[0430] Step 6: Data analysis and advice generation

[0431] The server evaluates the user's daily activities and emotional state based on the results of image recognition and emotion analysis.

[0432] Input: Emotional state and behavior analysis results

[0433] Output: personalized advice provided to the user

[0434] Specifically, it calculates the calories in meals, evaluates nutritional balance, and evaluates the frequency and intensity of exercise, and then uses a generative AI model to generate personalized advice, such as "Since you didn't eat many vegetables today, it would be good to add a salad to your dinner."

[0435] Step 7: Advice Notification

[0436] The server transmits the generated advice to a user terminal (such as a smartphone).

[0437] Input: Advice generation results

[0438] Output: Advice given to the user

[0439] Specifically, the server sends the generated advice to the user's device in an appropriate format (text, voice, etc.), and the user is notified. The user's device (e.g., a smartphone) then notifies the user of the received advice by displaying a visual message or by voice notification.

[0440] (Application example 2)

[0441] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0442] There are limitations to analyzing consumer behavior and providing customized advice in modern brick-and-mortar stores. Specifically, it is difficult for consumers to grasp product information in real time and receive personalized purchasing advice when selecting products in the store. Furthermore, advice is not provided that takes into account the consumer's emotional state. Under these circumstances, it is difficult to increase consumer satisfaction and maximize purchasing motivation.

[0443] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0444] In this invention, the server includes means for analyzing image data acquired by a device equipped with an image recognition function and identifying user behavior, means for providing advice to the user based on the analysis, means for analyzing the emotional state of the user and generating advice based on the emotional state, and means for analyzing product information in a physical store and providing purchasing advice based on the user's emotional state. This makes it possible to provide consumers in the physical store with product information and purchasing advice based on their emotional state in real time.

[0445] A "device with image recognition capabilities" is a device that incorporates a camera and image analysis capabilities to capture and identify visual data.

[0446] The "means for transmitting image data to a server" is a communication function for transferring image data to a server via a network.

[0447] The "means for identifying user behavior" is an algorithm or program for analyzing acquired image data and extracting and identifying specific user behavior from the data.

[0448] The "means for providing advice to the user" refers to an interface or notification function for providing the user with appropriate information or instructions based on the analysis results.

[0449] The "means for analyzing the user's emotional state and generating advice based on the emotional state" is a program that analyzes data such as the user's facial expressions and voice to identify emotions and create advice based on those emotions.

[0450] The "means for analyzing product information in a physical store and providing purchasing advice based on the user's emotional state" is a system for analyzing product information from image data acquired in a physical store and generating purchasing advice based on that information, taking into account the user's emotional state.

[0451] An embodiment of the present invention is a shopping assistant system for brick-and-mortar stores that uses a device with image recognition capabilities, a server, an emotion analysis engine, and a user interface. The system analyzes, in particular, the behavior and emotional state of a user and provides purchasing advice based on the analysis in real time.

[0452] Hardware and software used

[0453] A device with image recognition capabilities: Specifically, smart glasses with a built-in camera and image analysis capabilities that automatically capture images of products in front of the user's eyes and store them in local storage.

[0454] Server: A computer that receives and analyzes collected image data using image recognition algorithms and emotion analysis engines.

[0455] Emotion analysis engine: A program that analyzes data such as a user's facial expressions and voice to identify their emotional state.

[0456] User Interface: Visual and audio notifications to provide advice to the user, specifically via the smart glasses display and audio alerts.

[0457] Data processing and calculation

[0458] Acquisition and transmission of image data:

[0459] When a user looks at a product in a physical store, the smart glasses capture an image of the product, which is then stored in local storage and sent to a server at regular intervals.

[0460] Image data analysis:

[0461] The server analyzes the received image data and extracts product features (e.g., brand, price, ingredients, etc.) using a pre-trained image recognition algorithm.

[0462] Emotional State Analysis:

[0463] Using the camera and microphone installed in the smart glasses, the user's facial expressions and voice are analyzed by an emotion analysis engine to identify the user's emotional state (e.g., interest, curiosity, dissatisfaction, etc.).

[0464] Advice generation and notification:

[0465] Based on the analysis results, the server generates advice according to the user's behavior and emotional state. The advice is then sent to the smart glasses and presented to the user visually or audibly.

[0466] Specific examples

[0467] When a user looks at a particular product, the barcode of that product is scanned and sent to the server. The server analyzes the product information and generates a notification such as "This product is on sale" and displays it on the smart glasses' display. If the server determines that the user's facial expression indicates interest, it can also provide additional information such as "You can purchase this product at a lower price than other stores."

[0468] Example prompt sentence:

[0469] Analyze information about the product the user is looking at and generate recommendations based on the product's features, price, and the user's emotional state. For example, if the user is interested, notify them that "This product is 15% off."

[0470] In this way, the system can provide real-time advice based on the user's behavior and emotional state to enhance the shopping experience in a physical store.

[0471] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0472] Step 1:

[0473] Acquisition of image data

[0474] Input: User looks at product.

[0475] How it works: The device (smart glasses) uses its built-in camera to capture images of products in front of the user's line of sight.

[0476] Output: Captured image data.

[0477] Step 2:

[0478] Local storage of image data

[0479] Input: Captured image data.

[0480] Operation: The device (smart glasses) temporarily stores image data in local storage.

[0481] Output: Image data saved to local storage.

[0482] Step 3:

[0483] Sending image data

[0484] Input: Image data stored in local storage.

[0485] Operation: The device (smart glasses) sends image data from its local storage to the server at regular intervals.

[0486] Output: Image data sent to the server.

[0487] Step 4:

[0488] Image data analysis

[0489] Input: Image data sent to the server.

[0490] How it works: The server applies image recognition algorithms to identify product features (brand, price, ingredients, etc.) in the image.

[0491] Output: Product feature data as the analysis result.

[0492] Step 5:

[0493] Emotional state analysis

[0494] Input: User's facial expression data and voice data.

[0495] How it works: The device (smart glasses) sends the user's facial expressions and voice to an emotion analysis engine, which then analyzes them to determine the user's emotional state.

[0496] Output: User's emotional state data as the analysis result.

[0497] Step 6:

[0498] Generating Advice

[0499] Input: Product feature data and user emotional state data.

[0500] Operation: The server generates optimal advice based on the product features and the user's emotional state.

[0501] Output: The generated advice.

[0502] Step 7:

[0503] Advice Notification

[0504] Input: The generated advice.

[0505] Operation: The device (smart glasses) notifies the user of the generated advice visually or audibly.

[0506] Output: Advice given to the user.

[0507] In this way, when users choose products in a physical store, they can obtain real-time product information on the spot and receive personalized purchasing advice based on their emotional state. This system is expected to increase user satisfaction and maximize purchasing motivation.

[0508] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0509] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0510] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0511] [Second embodiment]

[0512] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0513] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0514] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0515] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0516] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0517] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0518] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0519] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0520] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0521] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0522] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0523] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0524] This invention proposes a system that uses a device and a server equipped with image recognition functionality to automatically record user behavior and provide advice based on the analysis results. Hereinafter, an embodiment of the invention will be described in detail.

[0525] System configuration

[0526] 1. Devices with image recognition capabilities

[0527] Device (glasses): This device has a built-in camera and image analysis capabilities, and automatically captures the environment and actions in front of the user's eyes. The glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that images are being recorded.

[0528] 2. Transmission and storage of image data

[0529] Device (glasses): Sends collected image data to the server at regular intervals. This data includes the time of shooting and location information.

[0530] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[0531] 3. Analysis of image data

[0532] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (food, exercise equipment, scenery, etc.) are identified and assigned classification tags.

[0533] 4. Lifestyle Analytics

[0534] Server: Based on the analysis results, the server evaluates the user's diet, exercise, and behavioral data, including calorie calculations, nutritional balance evaluations, and exercise frequency and intensity evaluations.

[0535] Server: Generates specific advice based on the evaluation results and stores it in a database.

[0536] 5. Advice Notification

[0537] Server: Sends the generated advice to the user's device (glasses, smartphone, etc.) in a format that is easy for the user to understand and follow.

[0538] Device (glasses or smartphone): Notifies the user of the received advice by displaying a visual message or making an audio notification.

[0539] Specific examples

[0540] Food records and advice

[0541] User: Eat lunch.

[0542] Device (glasses): Automatically takes photos of the user while they are eating and stores the image data in local storage.

[0543] Device (glasses): Sends image data of lunch to the server at the 12:00 PM interval.

[0544] Server: Analyzes the received image data and identifies the type and quantity of food.

[0545] Server: Calculates calories and nutrients and stores the evaluation results in a database.

[0546] Server: Generates advice such as "There weren't many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[0547] Device (smartphone): Displays a message to the user saying, "Today's lunch is high in calories, so you should add a salad to your dinner."

[0548] Travel Log

[0549] User: Walking around tourist spots.

[0550] Device (glasses): Automatically takes photos of important tourist spots and stores the image data in local storage.

[0551] Device (glasses): While walking around the tourist spot, the device periodically sends image data to the server.

[0552] Server: Analyzes the received image data and automatically identifies tourist spots. The image data also contains metadata (location, time, etc.).

[0553] Server: After the trip, the user's travel album is automatically generated based on the saved image data.

[0554] Server: Once the album is created, notify the user and provide an access link.

[0555] Device (smartphone): Display the travel album so that the user can check it.

[0556] This system allows users to improve their lifestyle habits and record their activities efficiently without any hassle. Furthermore, it has mechanisms for protecting privacy (such as notifications during recording), so users can use it with peace of mind.

[0557] The processing flow will be explained below.

[0558] Food record and advice processing steps

[0559] Step 1:

[0560] User: The user starts eating. The device (glasses) is turned on.

[0561] Step 2:

[0562] Device (glasses): The camera takes a picture of the food in front of the user's eyes.

[0563] Step 3:

[0564] Device (glasses): Captured image data is temporarily stored in local storage. When saved, the time of capture and location information are also added.

[0565] Step 4:

[0566] Device (glasses): At regular intervals (e.g., every hour), the image data stored in the local storage is sent to the server. The data includes the image file, the time of shooting, and location information.

[0567] Step 5:

[0568] Server: Stores the received image data in a database for analysis.

[0569] Step 6:

[0570] Server: Applying image recognition algorithms to analyze the received image data. This process involves identifying objects in the image (e.g., type and quantity of food).

[0571] Step 7:

[0572] Server: Based on the analysis results, calculates the calories of the food and evaluates the nutrients. Stores the results in a database.

[0573] Step 8:

[0574] Server: Generates specific advice based on the evaluation results. For example, "There weren't many vegetables today, so it would be good to add a salad to dinner."

[0575] Step 9:

[0576] Server: Sends the generated advice to the user's smartphone.

[0577] Step 10:

[0578] Device (smartphone): The received advice is displayed visually or notified by voice, allowing the user to check it and take action.

[0579] Travel Record Processing Steps

[0580] Step 1:

[0581] User: The user begins walking around the tourist spot. The device (glasses) is turned on.

[0582] Step 2:

[0583] Device (glasses): Automatically captures important tourist spots within the user's field of view.

[0584] Step 3:

[0585] Device (glasses): Captured image data is temporarily stored in local storage, along with the capture time and location information.

[0586] Step 4:

[0587] Device (glasses): Sends stored image data to the server at regular intervals (e.g., every hour).

[0588] Step 5:

[0589] Server: Stores the received image data in a database for analysis.

[0590] Step 6:

[0591] Server: Using an image recognition algorithm, the received image data is analyzed and tourist spots are automatically recognized. Metadata (location, time) is also associated with the images.

[0592] Step 7:

[0593] Server: Combines image data from multiple tourist spots to generate a travel album.

[0594] Step 8:

[0595] Server: Once the album is created, notify the user and provide an access link.

[0596] Step 9:

[0597] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album.

[0598] In this way, the process steps can be used to efficiently improve lifestyle habits and record travel. Privacy is also protected by a user notification function.

[0599] Example 1

[0600] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0601] With current technology, it takes a lot of time and effort to record a user's behavior in detail and provide appropriate advice. Furthermore, manual recording and analysis can lack accuracy and consistency. Furthermore, there is a risk that users will lose motivation to improve their lifestyle habits because there is a lack of a way for them to receive specific feedback on their improvements. For this reason, there is a need for a system that can automatically record a user's behavior, analyze it, and provide advice.

[0602] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0603] In this invention, the server includes means for analyzing image data to identify objects and assign classification tags, means for evaluating user behavior based on the analysis results, and means for generating specific advice using a generative AI model, thereby enabling automatic analysis of user behavior data and prompt provision of appropriate advice.

[0604] "Image recognition" refers to a device's ability to analyze images and identify objects.

[0605] "Device" refers to an electronic device that photographs and records user behavior and transmits the data to a server at regular intervals.

[0606] "Server" refers to a computer system that receives data sent from devices via a network, analyzes the data, and processes and stores the results.

[0607] "Image data" refers to image information captured by a device and recording the user's behavior and environment.

[0608] "Local storage" refers to a storage device installed within a device for temporarily storing data.

[0609] "Analysis results" refers to data that the server analyzes images to extract information about the user's behavior and environment, and assigns classification tags to.

[0610] A "generative AI model" is a pre-trained artificial intelligence model that uses algorithms to generate appropriate responses and advice based on input data.

[0611] "Advice" refers to a message that suggests specific guidelines for action or improvements to the user based on the analysis results.

[0612] "User terminal" refers to electronic devices that are directly used by users, such as devices and smartphones.

[0613] "Metadata" refers to supplementary information that accompanies image data, including the time of shooting, location information, and the like.

[0614] This invention proposes a system that automatically records user behavior and provides advice based on the analysis results. The system mainly includes a device with image recognition capabilities, a terminal that transmits and stores data, a server that analyzes and evaluates image data, and a means for generating specific advice using a generative AI model.

[0615] System configuration

[0616] 1. Devices with image recognition capabilities

[0617] Device (glasses):

[0618] The device has a built-in camera and image analysis capabilities, and captures real-time images of the user's behavior and environment. The captured image data is temporarily stored in local storage. The device also has a visual (e.g., LED) or audio interface to notify the user that images are being recorded.

[0619] 2. Data transmission and storage

[0620] Device (glasses):

[0621] The collected image data is sent to a server at regular intervals, including the time and location of the image.

[0622] server:

[0623] The received image data is stored in an analytical database, allowing for subsequent analysis.

[0624] 3. Analysis of image data

[0625] server:

[0626] The system applies image recognition algorithms (e.g., TensorFlow, OpenCV) to the received image data, identifies objects in the data, and assigns classification tags. For example, it identifies the type and quantity of food in a photo of a meal.

[0627] 4. Evaluation of behavioral data

[0628] server:

[0629] Based on the analysis results, the user's behavioral data is evaluated. Specific evaluation items include calorie calculation of meal contents, evaluation of nutritional balance, and evaluation of exercise frequency and intensity.

[0630] 5. Automatic Advice Generation

[0631] server:

[0632] Based on the evaluation data obtained from the analysis results, specific advice is generated using a generative AI model (e.g., GPT-3). The advice is in a format that is easy for users to understand and follow.

[0633] 6. Advice Notification

[0634] server:

[0635] The generated advice is sent to the user's device (glasses or smartphone).

[0636] Device (glasses or smartphone):

[0637] The user is notified of the received advice, either visually or by audio.

[0638] Specific examples

[0639] Food records and advice

[0640] User:

[0641] Eat lunch.

[0642] Device (glasses):

[0643] The system takes a photo of the user's meal and stores the image data in local storage. This data is sent to the server at 12:00 PM intervals.

[0644] server:

[0645] The received image data is analyzed to identify the type and quantity of food, then the calories and nutrients are calculated and the evaluation results are stored in a database.

[0646] server:

[0647] The system generates advice such as "You didn't have many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[0648] Device (smartphone):

[0649] The message "Today's lunch is high in calories, so you should add a salad to your dinner" is displayed to notify the user.

[0650] Prompt Sentence Examples

[0651] Analyze an image of a user eating lunch, identify the type and amount of food, calculate calories and nutrients, and generate appropriate recommendations.

[0652] Generate a trip record

[0653] User:

[0654] Walk around the tourist spots.

[0655] Device (glasses):

[0656] Photograph important tourist spots and store the image data in local storage.

[0657] Device (glasses):

[0658] During sightseeing, image data is periodically sent to the server.

[0659] server:

[0660] The received image data is analyzed and tourist spots are automatically recognized. The image data includes the time of shooting and location information.

[0661] server:

[0662] After the trip, a travel album is automatically generated based on the saved image data.

[0663] server:

[0664] Once the album has been created, the user will be notified and provided with an access link.

[0665] Device (smartphone):

[0666] Display the travel album so that the user can check it.

[0667] Prompt Sentence Examples

[0668] Analyze the received tourist attraction image data, identify tourist attractions based on the metadata, and generate a travel album.

[0669] This system allows users to efficiently improve their lifestyle habits and record their activities without any hassle. It also has privacy protection features (such as notifications during recording), so users can use it with peace of mind.

[0670] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0671] Step 1:

[0672] Device (glasses): The built-in camera captures the user's actions. The captured image data is temporarily stored in the device's local storage. At this time, the LED on the device lights up to notify the user that a photo is being taken.

[0673] Input: User behavior, environment

[0674] Output: Captured image data

[0675] Step 2:

[0676] Device (glasses): Image data, shooting time, and location information stored in local storage are sent to the server at regular intervals. While the data is being sent, the device's notification function is activated to notify the user that data is being sent.

[0677] Input: Image data from local storage, shooting time, location information

[0678] Output: Image data and metadata (photo time, location information) sent to the server

[0679] Step 3:

[0680] Server: Apply image recognition algorithms (e.g., TensorFlow, OpenCV) to the received image data to identify objects in the data and assign classification tags. As a result of the analysis process, noteworthy content in the image (e.g., food, tourist attractions) is identified.

[0681] Input: Image data, metadata (shooting time, location information)

[0682] Output: Analysis results (object classification tags)

[0683] Step 4:

[0684] Server: Evaluates the user's behavioral data based on the analysis results. Analyzes meal content, calculates calories, evaluates nutritional balance, and analyzes exercise to evaluate frequency and intensity.

[0685] Input: Analysis results (object classification tags)

[0686] Output: Evaluation results of behavioral data (calorie calculation, nutritional balance evaluation, exercise frequency and intensity)

[0687] Step 5:

[0688] Server: Based on the evaluation results, a generative AI model (e.g., GPT-3) is used to generate specific advice. The generated advice is easy for users to understand and follow.

[0689] Input: Evaluation results of behavioral data

[0690] Output: The generated advice

[0691] Step 6:

[0692] Server: Sends the generated advice to the user's device (glasses or smartphone). While sending the advice, the server notifies the user that a message has been received.

[0693] Input: Generated advice

[0694] Output: Advice sent to the user device (glasses or smartphone)

[0695] Step 7:

[0696] Device (smartphone or glasses): Notifies the user of the received advice, either visually (as a pop-up message) or by voice.

[0697] Input: Advice sent by the server

[0698] Output: Advice displayed to the user

[0699] As described above, this system automatically records the user's behavior and provides advice based on the analysis results, allowing for efficient lifestyle improvement and behavioral recording.It also has a notification function to protect privacy, so it can be used with peace of mind.

[0700] (Application example 1)

[0701] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0702] With traditional methods of providing customer service in brick-and-mortar stores, it is difficult to understand individual customer behavior and interests in real time, making it difficult to provide personalized service. Furthermore, staff often lack the information they need to make appropriate product recommendations to customers as needed, resulting in missed opportunities to maximize customer satisfaction. A system that can solve these issues and improve the quality of customer service is needed.

[0703] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0704] In this invention, the server includes a device with an image recognition function that captures images of customer behavior in a physical store and transmits the obtained image data to the server, a means for the server to analyze customer behavior patterns and interests and provide specific advice to store staff, and a means for identifying customer interests contained in the image data and providing advice recommending related products. This enables real-time analysis of customer behavior and interests in a physical store, enabling staff to provide appropriate advice and product recommendations.

[0705] A "device with image recognition functionality" is a device that has a built-in camera and image analysis functionality and is used to capture and analyze the surrounding environment and people's behavior.

[0706] A "server" is a computer system that receives, stores, and analyzes data over a network and provides necessary information.

[0707] "Image data" refers to data that includes visual information acquired by a photographic device such as a camera.

[0708] "Analysis" is the process of using algorithms to understand the content of acquired image data and recognize specific patterns or objects.

[0709] A "user" is a person using a device with image recognition capabilities or a person receiving advice from a server.

[0710] "Advice" is any instruction, advice, or suggestion provided to the user based on the results of the analysis.

[0711] A "brick and mortar store" is a retail establishment that exists in a physical location and offers goods and services.

[0712] A "customer" is a consumer who visits a physical store and purchases or uses goods or services.

[0713] A "behavioral pattern" is a series of actions and trends in interests that indicate how customers move around the store and what they are interested in.

[0714] "Recommendations" are the suggestion of products or services based on a customer's interests.

[0715] This invention provides a system for improving customer service by analyzing customer behavior in a physical store in real time using a device and server equipped with image recognition functionality.

[0716] System configuration

[0717] 1. Devices with image recognition capabilities

[0718] Terminal (smart glasses): This device has a built-in camera and image analysis capabilities, and automatically captures customer behavior in the store. The smart glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that an image is being recorded.

[0719] 2. Transmission and storage of image data

[0720] Terminal (smart glasses): Sends collected image data to the server at regular intervals, including the time of shooting and location information.

[0721] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[0722] 3. Analysis of image data

[0723] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (e.g., products, behavioral patterns, etc.) are identified and assigned classification tags.

[0724] 4. Customer Service Analytics

[0725] Server: Evaluates customer behavior patterns and interests based on the analysis results, including identifying customer areas of interest and suggesting related products.

[0726] Server: Generates specific advice based on the evaluation results and stores it in a database.

[0727] 5. Advice Notification

[0728] Server: Sends the generated advice to the store staff's devices (smart glasses, smartphones, etc.) in a format that is easy for staff to understand and implement.

[0729] Terminal (smart glasses or smartphone): Notifies staff of received advice, either by displaying a visual message or by making an audio notification.

[0730] Specific examples

[0731] Product Recommendations

[0732] Customer: Spends a long time in the store looking at new jackets.

[0733] Terminal (smart glasses): Automatically captures customer behavior and stores the image data in local storage.

[0734] Terminal (smart glasses): Sends image data to the server at regular intervals.

[0735] Server: Analyzes the received image data and identifies the customer's interests. For example, the analysis result may be, "This customer is interested in new jackets."

[0736] Server: Generates advice such as "Suggest sales information for related products to this customer" and sends it to the staff member's smartphone.

[0737] Device (smartphone): Displays the message "Please suggest related products for this customer's new jacket" and notifies the staff.

[0738] Improved customer service

[0739] Customer: Walks around the store, checking out multiple products.

[0740] Terminal (smart glasses): Automatically captures customer behavior and stores the image data in local storage.

[0741] Terminal (smart glasses): Sends image data to the server at regular intervals.

[0742] Server: Analyzes the received image data and identifies customer behavior patterns. For example, it recognizes a behavior pattern of "looking at many products in a short period of time."

[0743] Server: Generates advice such as, "This customer is interested in many products, so we will approach them individually and introduce them to special offers," and sends this advice to the staff member's smartphone.

[0744] Terminal (smartphone): Display the message "Inform customers who are viewing many products in a short time about special offers" and notify the staff.

[0745] Example of input prompt for generative AI model

[0746] Image data showing a user looking at a particular product for a long time in the store was sent to a server. The server used an image recognition algorithm to identify the customer's interests and generate a recommendation: "This customer is interested in a new jacket." Staff receive this notification and suggest sales information for the new jacket and related items to the customer. Other visual information used by the AI ​​model includes pop-up information about related products and promotional codes.

[0747] This system allows physical stores to understand customer behavior and interests in real time, enabling staff to provide appropriate advice and product recommendations to customers.

[0748] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0749] Step 1:

[0750] The device (smart glasses) acquires the following information: The input includes an image of the environment seen by the user. The output is to temporarily store the acquired image data in local storage. Specifically, the device's camera takes pictures at regular intervals and stores the data in local storage.

[0751] Step 2:

[0752] The device (smart glasses) transmits image data stored in local storage to a server at regular intervals. The input includes the stored image data and its metadata (time of capture, location information). As output, this data is transmitted to the server. Specifically, the communication module in the device transmits the data to the server and confirms that the transmission was successful.

[0753] Step 3:

[0754] The server stores the received image data in an analysis database. The input includes the image data sent from the terminal and the associated metadata. The output is stored in a specified format in the database. Specifically, the server's storage system stores the received data in the appropriate format.

[0755] Step 4:

[0756] The server analyzes image data stored in a database based on image recognition algorithms. The input includes the stored image data. The output is the identification of objects and behavioral patterns in the image and the assignment of classification tags. Specifically, the server's processing unit executes the image recognition model and analyzes the results.

[0757] Step 5:

[0758] The server evaluates the customer's behavioral patterns and interests based on the analysis results. The input includes the analyzed data. The output is the evaluation results stored in a database. Specifically, the evaluation algorithm extracts behavioral patterns based on the analysis results and generates evaluation results.

[0759] Step 6:

[0760] The server generates specific advice based on the evaluation results. The input includes the evaluation results stored in the database. As an output, the generated advice message is sent to the database and the user terminal. As a specific operation, the advice generation algorithm is executed and the result is formatted in a format suitable for the user interface.

[0761] Step 7:

[0762] The terminal (smart glasses or smartphone) notifies the staff of the advice sent from the server. The input includes the advice message sent from the server. The output is a visual message display or a voice notification. As a specific operation, the display module or voice output module of the terminal is activated and the advice is provided to the staff.

[0763] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0764] This invention proposes a system that uses a device with image recognition capabilities, a server, and an emotion engine to automatically record a user's behavior, identify the user's emotional state, and provide advice based on the analysis results. Hereinafter, an embodiment of the present invention will be described in detail.

[0765] System configuration

[0766] 1. Devices with image recognition capabilities

[0767] Device (glasses): This device has a built-in camera and image analysis capabilities, and automatically captures the environment and actions in front of the user's eyes. The glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that images are being recorded.

[0768] 2. Transmission and storage of image data

[0769] Device (glasses): Sends collected image data to the server at regular intervals. This data includes the time of shooting and location information.

[0770] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[0771] 3. Analysis of image data

[0772] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (food, exercise equipment, scenery, etc.) are identified and assigned classification tags.

[0773] 4. Emotional state analysis

[0774] Server: Using the emotion engine, analyzes the user's facial expressions from the received image data and identifies their emotional state. For example, it determines whether the user is happy or stressed.

[0775] 5. Lifestyle Analytics

[0776] Server: Based on the analysis results, the server evaluates the user's diet, exercise, and emotional state, including calorie calculations, nutritional balance assessments, exercise frequency and intensity, and emotional state assessments.

[0777] Server: Generates specific advice based on the evaluation results and stores it in a database. For example, if the user is feeling emotionally stressed, it will suggest relaxation techniques.

[0778] 6. Advice Notification

[0779] Server: Sends the generated advice to the user's device (glasses, smartphone, etc.) in a format that is easy for the user to understand and follow.

[0780] Device (glasses or smartphone): Notifies the user of the received advice by displaying a visual message or making an audio notification.

[0781] Specific examples

[0782] Food records and advice

[0783] User: Eat lunch.

[0784] Device (glasses): Automatically takes photos of the user while they are eating and stores the image data in local storage.

[0785] Device (glasses): Sends image data of lunch to the server at the 12:00 PM interval.

[0786] Server: Analyzes the received image data and identifies the type and quantity of food.

[0787] Server: Calculates calories and nutrients and stores the evaluation results in a database.

[0788] Server: Generates advice such as "There weren't many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[0789] Device (smartphone): Displays a message to the user saying, "Today's lunch is high in calories, so you should add a salad to your dinner."

[0790] Server: The emotion engine analyzes the user's facial expressions while they are eating and evaluates whether they are enjoying the meal. For example, if the user is not enjoying the meal, it can add advice such as "Next time, try incorporating your favorite dishes."

[0791] Travel Records and Sentiment Analysis

[0792] User: Walking around tourist spots.

[0793] Device (glasses): Automatically takes photos of important tourist spots and stores the image data in local storage.

[0794] Device (glasses): While walking around the tourist spot, the device periodically sends image data to the server.

[0795] Server: Analyzes the received image data and automatically identifies tourist spots. The image data also contains metadata (location, time).

[0796] Server: After the trip, the user's travel album is automatically generated based on the saved image data.

[0797] Server: The emotion engine analyzes the user's facial expressions while sightseeing to identify the places and moments they particularly enjoyed. Based on this information, it decides which points to highlight in the travel album.

[0798] Server: Once the album is created, notify the user and provide an access link.

[0799] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album.

[0800] This system not only allows users to efficiently improve their lifestyle habits and record their activities without any hassle, but also analyzes their emotional state and provides more personalized advice.It also has mechanisms for protecting privacy (such as notifications during recording), so users can use it with peace of mind.

[0801] The processing flow will be explained below.

[0802] Food Record and Emotion Analysis Processing Steps

[0803] Step 1:

[0804] User: The user starts eating. The device (glasses) is turned on.

[0805] Step 2:

[0806] Device (glasses): The camera takes a picture of the food in front of the user's eyes.

[0807] Step 3:

[0808] Device (glasses): Captured image data is temporarily stored in local storage. When saved, the time of capture and location information are also added.

[0809] Step 4:

[0810] Device (glasses): At regular intervals (e.g., every hour), the image data stored in the local storage is sent to the server. The data includes the image file, the time of shooting, and location information.

[0811] Step 5:

[0812] Server: Stores the received image data in a database for analysis.

[0813] Step 6:

[0814] Server: Applying image recognition algorithms to analyze the received image data. This process involves identifying objects in the image (e.g., type and quantity of food).

[0815] Step 7:

[0816] Server: Based on the analysis results, calculates the calories of the food and evaluates the nutrients. Stores the results in a database.

[0817] Step 8:

[0818] Server: Using the emotion engine, analyze the user's facial expressions from the image data and identify their emotional state. Evaluate whether the user is enjoying the meal or feeling stressed.

[0819] Step 9:

[0820] Server: Generates specific advice based on the evaluation results. For example, "You didn't eat many vegetables today, so you should add a salad to your dinner." It also generates emotion-based advice, such as "You seemed a little nervous during the meal, so next time try eating in a relaxed environment."

[0821] Step 10:

[0822] Server: Sends the generated advice to the user's smartphone.

[0823] Step 11:

[0824] Device (smartphone): The received advice is displayed visually or notified by voice, allowing the user to check it and take action.

[0825] Processing steps for travel records and sentiment analysis

[0826] Step 1:

[0827] User: The user begins walking around the tourist spot. The device (glasses) is turned on.

[0828] Step 2:

[0829] Device (glasses): Automatically captures important tourist spots within the user's field of view.

[0830] Step 3:

[0831] Device (glasses): Captured image data is temporarily stored in local storage, along with the capture time and location information.

[0832] Step 4:

[0833] Device (glasses): Sends stored image data to the server at regular intervals (e.g., every hour).

[0834] Step 5:

[0835] Server: Stores the received image data in a database for analysis.

[0836] Step 6:

[0837] Server: Using an image recognition algorithm, the received image data is analyzed and tourist spots are automatically recognized. Metadata (location, time) is also associated with the images.

[0838] Step 7:

[0839] Server: Using the emotion engine, analyze the user's facial expressions while sightseeing to identify their emotional state, and identify places and moments that they are particularly enjoying.

[0840] Step 8:

[0841] Server: Combines image data from multiple tourist spots to generate a travel album, placing the images in a way that highlights moments of positive emotional states.

[0842] Step 9:

[0843] Server: Once the album is created, notify the user and provide an access link.

[0844] Step 10:

[0845] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album, highlighting the moments they particularly enjoyed.

[0846] In this way, a system can be realized that records and analyzes the user's behavior and emotions in detail through specific processing steps and provides appropriate advice.

[0847] Example 2

[0848] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0849] Conventional lifestyle improvement systems require users to manually record their behavior, diet, and exercise, which is time-consuming for users. Furthermore, it is difficult to analyze emotional states and provide advice based on them, which means that more personalized advice cannot be provided to users. There is a need for a system that can solve these issues and enable users to efficiently improve their lifestyles and record their behavior.

[0850] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0851] In this invention, the server includes a terminal with an image recognition function, means for transmitting image data acquired by the terminal to the server, means for analyzing the image data in the server and identifying the user's behavior and emotional state, and means for providing advice to the user based on the analysis. This makes it possible to automatically record the user's behavior and emotional state and provide personalized, specific advice based on the analysis results.

[0852] A "terminal with image recognition functionality" is a device that has the ability to automatically capture images of the environment and behavior in front of the user's line of sight and acquire them as image data.

[0853] "Means for transmitting acquired image data to a server" refers to a function for transferring image data from a terminal to a server at regular intervals.

[0854] "Means for analyzing image data and identifying user behavior and emotional state" refers to the process of analyzing and identifying user behavior (e.g., eating, exercise) and emotional state (e.g., joy, anger) based on the image data received by the server using an image recognition algorithm and an emotion engine.

[0855] "Means for providing advice" refers to the function of generating specific guidelines for the user (e.g., adding a salad to dinner) based on the analysis results and notifying the user's device.

[0856] "Means for providing advice based on nutritional balance and calorie calculation" refers to the process of analyzing dietary content based on image data, calculating nutrient balance and calories, and making suggestions for dietary improvement based on that.

[0857] "Means for providing advice on the frequency and intensity of exercise" refers to the process of analyzing the exercise captured in the image data, evaluating the frequency and intensity of exercise, and making suggestions that will help improve lifestyle habits.

[0858] An "emotion engine" refers to an algorithm or software that analyzes a user's facial expressions contained in image data and identifies emotional states such as joy, anger, and sadness.

[0859] A "generative AI model" refers to a model that uses artificial intelligence to generate personalized advice based on the user's analysis results.

[0860] System Overview

[0861] This invention is a system that uses a device with image recognition capabilities (e.g., glasses with a built-in camera), a server, and an emotion engine to automatically record a user's behavior, identify the user's emotional state, and provide advice based on the analysis results. The main hardware and software configuration of this system is as follows:

[0862] Hardware used

[0863] 1. Device (glasses with built-in camera): Automatically captures the environment and actions in front of the user's line of sight and acquires image data.

[0864] 2. Server: Analyzes the acquired image data, generates advice based on the analysis results, and notifies the user.

[0865] 3. User device (smartphone or tablet): Notifies the user of the advice sent from the server.

[0866] Software and algorithms used

[0867] 1. Image recognition algorithms: Identify objects in images using OpenCV, TensorFlow, etc.

[0868] 2. Emotion engine: Analyzes the user's emotional state using Microsoft Azure's Face API and Emotion API.

[0869] 3. Generative AI model: An artificial intelligence model that generates personalized advice based on the user's analysis results.

[0870] Process Overview

[0871] Automatic image capture and data collection

[0872] The device (glasses with a built-in camera) automatically captures the environment and actions in front of the user's eyes and saves the image data in local storage. The glasses are equipped with a visual or audio interface to notify the user that a photo is being taken.

[0873] Sending image data

[0874] The device (glasses with a built-in camera) collects image data and sends it to the server at regular intervals. The image data includes the time of capture and location information.

[0875] Image data storage and analysis

[0876] The server stores the received image data in an analysis database and applies image recognition algorithms to analyze it, identifying objects in the image and assigning them classification tags.

[0877] Emotional state analysis

[0878] The server uses an emotion engine to analyze the user's facial expressions contained in the image data, thereby identifying the user's emotional state, for example, whether they are happy or stressed.

[0879] Advice Generation

[0880] The server evaluates the user's diet, exercise habits, and emotional state based on the analysis results and generates specific advice, which is personalized using a generative AI model.

[0881] Specific examples

[0882] Food records and advice

[0883] When a user eats lunch, the device (glasses with a built-in camera) automatically takes pictures of the meal and collects image data. At 12:00 PM intervals, the collected image data of the lunch is sent to the server. The server analyzes the received image data and identifies the type and amount of food. It then calculates calories and nutrients and generates advice such as "You didn't eat many vegetables today, so you should add a salad to your dinner." This advice is sent to the user's device (such as a smartphone) and notified to the user. The user's facial expressions while eating can also be analyzed to evaluate whether they are enjoying their meal. For example, if the user is not enjoying their meal, the system can add advice such as "Next time, try incorporating your favorite dishes."

[0884] Prompt Sentence Examples

[0885] "Please explain how you can analyze the user's emotional state and generate advice to suggest relaxation if they are feeling stressed."

[0886] In this way, the present invention not only allows users to efficiently improve their lifestyle habits and record their activities without any hassle, but also analyzes their emotional state and allows them to receive more personalized advice. Furthermore, the system is equipped with mechanisms for protecting privacy (such as notifications during recording), allowing users to use the system with peace of mind.

[0887] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0888] System processing steps

[0889] Step 1: Automatic image capture and data collection

[0890] The device (glasses with a built-in camera) automatically captures the environment and actions in front of the user's eyes in real time.

[0891] Input: The scene in front of the user's line of sight

[0892] Output: Acquired image data (photos of the environment and behavior)

[0893] Specifically, the glasses' built-in camera periodically takes a picture and temporarily stores the video data in local storage. The glasses provide visual and audio feedback to let the user know that they are recording.

[0894] Step 2: Sending image data

[0895] The terminal (glasses with a built-in camera) sends the collected image data to a server at a predetermined interval (for example, every 5 minutes).

[0896] Input: Image data stored in local storage

[0897] Output: Image data sent to the server (including shooting time and location information)

[0898] Specifically, the device periodically divides image data into packets and sends them to the server via a wireless network (Wi-Fi or mobile data).

[0899] Step 3: Save the image data

[0900] The server stores the received image data in an analysis database.

[0901] Input: Image data sent from the device (including shooting time and location information)

[0902] Output: Image data stored in a database for analysis

[0903] Specifically, the server checks the format and integrity of the received data, stores it correctly in a database for analysis, and implements security measures to ensure the data is managed in a protected environment.

[0904] Step 4: Image Recognition and Classification

[0905] The server applies an image recognition algorithm to the image data stored in the analysis database and performs analysis.

[0906] Input: Image data stored in a database

[0907] Output: Image data with identified objects and classification tags

[0908] Specifically, it uses image recognition libraries such as OpenCV and TensorFlow to identify objects in the image (e.g., food, exercise equipment, landscapes, etc.) The objects obtained as a result of the analysis are given classification tags and stored again in the database.

[0909] Step 5: Analyze emotional state

[0910] The server uses an emotion engine to analyze the user's facial expressions in the image data.

[0911] Input: Image data that has been recognized and classified

[0912] Output: Identification of the user's emotional state (e.g., happy, angry, sad)

[0913] Specifically, it uses Microsoft Azure's Face API and Emotion API to identify the user's emotional state from their facial expressions. The analysis results are stored in a database and used for subsequent processing.

[0914] Step 6: Data analysis and advice generation

[0915] The server evaluates the user's daily activities and emotional state based on the results of image recognition and emotion analysis.

[0916] Input: Emotional state and behavior analysis results

[0917] Output: personalized advice provided to the user

[0918] Specifically, it calculates the calories in meals, evaluates nutritional balance, and evaluates the frequency and intensity of exercise, and then uses a generative AI model to generate personalized advice, such as "Since you didn't eat many vegetables today, it would be good to add a salad to your dinner."

[0919] Step 7: Advice Notification

[0920] The server transmits the generated advice to a user terminal (such as a smartphone).

[0921] Input: Advice generation results

[0922] Output: Advice given to the user

[0923] Specifically, the server sends the generated advice to the user's device in an appropriate format (text, voice, etc.), and the user is notified. The user's device (e.g., a smartphone) then notifies the user of the received advice by displaying a visual message or by voice notification.

[0924] (Application example 2)

[0925] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0926] There are limitations to analyzing consumer behavior and providing customized advice in modern brick-and-mortar stores. Specifically, it is difficult for consumers to grasp product information in real time and receive personalized purchasing advice when selecting products in the store. Furthermore, advice is not provided that takes into account the consumer's emotional state. Under these circumstances, it is difficult to increase consumer satisfaction and maximize purchasing motivation.

[0927] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0928] In this invention, the server includes means for analyzing image data acquired by a device equipped with an image recognition function and identifying user behavior, means for providing advice to the user based on the analysis, means for analyzing the emotional state of the user and generating advice based on the emotional state, and means for analyzing product information in a physical store and providing purchasing advice based on the user's emotional state. This makes it possible to provide consumers in the physical store with product information and purchasing advice based on their emotional state in real time.

[0929] A "device with image recognition capabilities" is a device that incorporates a camera and image analysis capabilities to capture and identify visual data.

[0930] The "means for transmitting image data to a server" is a communication function for transferring image data to a server via a network.

[0931] The "means for identifying user behavior" is an algorithm or program for analyzing acquired image data and extracting and identifying specific user behavior from the data.

[0932] The "means for providing advice to the user" refers to an interface or notification function for providing the user with appropriate information or instructions based on the analysis results.

[0933] The "means for analyzing the user's emotional state and generating advice based on the emotional state" is a program that analyzes data such as the user's facial expressions and voice to identify emotions and create advice based on those emotions.

[0934] The "means for analyzing product information in a physical store and providing purchasing advice based on the user's emotional state" is a system for analyzing product information from image data acquired in a physical store and generating purchasing advice based on that information, taking into account the user's emotional state.

[0935] An embodiment of the present invention is a shopping assistant system for brick-and-mortar stores that uses a device with image recognition capabilities, a server, an emotion analysis engine, and a user interface. The system analyzes, in particular, the behavior and emotional state of a user and provides purchasing advice based on the analysis in real time.

[0936] Hardware and software used

[0937] A device with image recognition capabilities: Specifically, smart glasses with a built-in camera and image analysis capabilities that automatically capture images of products in front of the user's eyes and store them in local storage.

[0938] Server: A computer that receives and analyzes collected image data using image recognition algorithms and emotion analysis engines.

[0939] Emotion analysis engine: A program that analyzes data such as a user's facial expressions and voice to identify their emotional state.

[0940] User Interface: Visual and audio notifications to provide advice to the user, specifically via the smart glasses display and audio alerts.

[0941] Data processing and calculation

[0942] Acquisition and transmission of image data:

[0943] When a user looks at a product in a physical store, the smart glasses capture an image of the product, which is then stored in local storage and sent to a server at regular intervals.

[0944] Image data analysis:

[0945] The server analyzes the received image data and extracts product features (e.g., brand, price, ingredients, etc.) using a pre-trained image recognition algorithm.

[0946] Emotional State Analysis:

[0947] Using the camera and microphone installed in the smart glasses, the user's facial expressions and voice are analyzed by an emotion analysis engine to identify the user's emotional state (e.g., interest, curiosity, dissatisfaction, etc.).

[0948] Advice generation and notification:

[0949] Based on the analysis results, the server generates advice according to the user's behavior and emotional state. The advice is then sent to the smart glasses and presented to the user visually or audibly.

[0950] Specific examples

[0951] When a user looks at a particular product, the barcode of that product is scanned and sent to the server. The server analyzes the product information and generates a notification such as "This product is on sale" and displays it on the smart glasses' display. If the server determines that the user's facial expression indicates interest, it can also provide additional information such as "You can purchase this product at a lower price than other stores."

[0952] Example prompt sentence:

[0953] Analyze information about the product the user is looking at and generate recommendations based on the product's features, price, and the user's emotional state. For example, if the user is interested, notify them that "This product is 15% off."

[0954] In this way, the system can provide real-time advice based on the user's behavior and emotional state to enhance the shopping experience in a physical store.

[0955] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0956] Step 1:

[0957] Acquisition of image data

[0958] Input: User looks at product.

[0959] How it works: The device (smart glasses) uses its built-in camera to capture images of products in front of the user's line of sight.

[0960] Output: Captured image data.

[0961] Step 2:

[0962] Local storage of image data

[0963] Input: Captured image data.

[0964] Operation: The device (smart glasses) temporarily stores image data in local storage.

[0965] Output: Image data saved to local storage.

[0966] Step 3:

[0967] Sending image data

[0968] Input: Image data stored in local storage.

[0969] Operation: The device (smart glasses) sends image data from its local storage to the server at regular intervals.

[0970] Output: Image data sent to the server.

[0971] Step 4:

[0972] Image data analysis

[0973] Input: Image data sent to the server.

[0974] How it works: The server applies image recognition algorithms to identify product features (brand, price, ingredients, etc.) in the image.

[0975] Output: Product feature data as the analysis result.

[0976] Step 5:

[0977] Emotional state analysis

[0978] Input: User's facial expression data and voice data.

[0979] How it works: The device (smart glasses) sends the user's facial expressions and voice to an emotion analysis engine, which then analyzes them to determine the user's emotional state.

[0980] Output: User's emotional state data as the analysis result.

[0981] Step 6:

[0982] Generating Advice

[0983] Input: Product feature data and user emotional state data.

[0984] Operation: The server generates optimal advice based on the product features and the user's emotional state.

[0985] Output: The generated advice.

[0986] Step 7:

[0987] Advice Notification

[0988] Input: The generated advice.

[0989] Operation: The device (smart glasses) notifies the user of the generated advice visually or audibly.

[0990] Output: Advice given to the user.

[0991] In this way, when users choose products in a physical store, they can obtain real-time product information on the spot and receive personalized purchasing advice based on their emotional state. This system is expected to increase user satisfaction and maximize purchasing motivation.

[0992] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0993] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0994] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0995] [Third embodiment]

[0996] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0997] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0998] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0999] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1000] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1001] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1002] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1003] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1004] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1005] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1006] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1007] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1008] This invention proposes a system that uses a device and a server equipped with image recognition functionality to automatically record user behavior and provide advice based on the analysis results. Hereinafter, an embodiment of the invention will be described in detail.

[1009] System configuration

[1010] 1. Devices with image recognition capabilities

[1011] Device (glasses): This device has a built-in camera and image analysis capabilities, and automatically captures the environment and actions in front of the user's eyes. The glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that images are being recorded.

[1012] 2. Transmission and storage of image data

[1013] Device (glasses): Sends collected image data to the server at regular intervals. This data includes the time of shooting and location information.

[1014] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[1015] 3. Analysis of image data

[1016] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (food, exercise equipment, scenery, etc.) are identified and assigned classification tags.

[1017] 4. Lifestyle Analytics

[1018] Server: Based on the analysis results, the server evaluates the user's diet, exercise, and behavioral data, including calorie calculations, nutritional balance evaluations, and exercise frequency and intensity evaluations.

[1019] Server: Generates specific advice based on the evaluation results and stores it in a database.

[1020] 5. Advice Notification

[1021] Server: Sends the generated advice to the user's device (glasses, smartphone, etc.) in a format that is easy for the user to understand and follow.

[1022] Device (glasses or smartphone): Notifies the user of the received advice by displaying a visual message or making an audio notification.

[1023] Specific examples

[1024] Food records and advice

[1025] User: Eat lunch.

[1026] Device (glasses): Automatically takes photos of the user while they are eating and stores the image data in local storage.

[1027] Device (glasses): Sends image data of lunch to the server at the 12:00 PM interval.

[1028] Server: Analyzes the received image data and identifies the type and quantity of food.

[1029] Server: Calculates calories and nutrients and stores the evaluation results in a database.

[1030] Server: Generates advice such as "There weren't many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[1031] Device (smartphone): Displays a message to the user saying, "Today's lunch is high in calories, so you should add a salad to your dinner."

[1032] Travel Log

[1033] User: Walking around tourist spots.

[1034] Device (glasses): Automatically takes photos of important tourist spots and stores the image data in local storage.

[1035] Device (glasses): While walking around the tourist spot, the device periodically sends image data to the server.

[1036] Server: Analyzes the received image data and automatically identifies tourist spots. The image data also contains metadata (location, time, etc.).

[1037] Server: After the trip, the user's travel album is automatically generated based on the saved image data.

[1038] Server: Once the album is created, notify the user and provide an access link.

[1039] Device (smartphone): Display the travel album so that the user can check it.

[1040] This system allows users to improve their lifestyle habits and record their activities efficiently without any hassle. Furthermore, it has mechanisms for protecting privacy (such as notifications during recording), so users can use it with peace of mind.

[1041] The processing flow will be explained below.

[1042] Food record and advice processing steps

[1043] Step 1:

[1044] User: The user starts eating. The device (glasses) is turned on.

[1045] Step 2:

[1046] Device (glasses): The camera takes a picture of the food in front of the user's eyes.

[1047] Step 3:

[1048] Device (glasses): Captured image data is temporarily stored in local storage. When saved, the time of capture and location information are also added.

[1049] Step 4:

[1050] Device (glasses): At regular intervals (e.g., every hour), the image data stored in the local storage is sent to the server. The data includes the image file, the time of shooting, and location information.

[1051] Step 5:

[1052] Server: Stores the received image data in a database for analysis.

[1053] Step 6:

[1054] Server: Applying image recognition algorithms to analyze the received image data. This process involves identifying objects in the image (e.g., type and quantity of food).

[1055] Step 7:

[1056] Server: Based on the analysis results, calculates the calories of the food and evaluates the nutrients. Stores the results in a database.

[1057] Step 8:

[1058] Server: Generates specific advice based on the evaluation results. For example, "There weren't many vegetables today, so it would be good to add a salad to dinner."

[1059] Step 9:

[1060] Server: Sends the generated advice to the user's smartphone.

[1061] Step 10:

[1062] Device (smartphone): The received advice is displayed visually or notified by voice, allowing the user to check it and take action.

[1063] Travel Record Processing Steps

[1064] Step 1:

[1065] User: The user begins walking around the tourist spot. The device (glasses) is turned on.

[1066] Step 2:

[1067] Device (glasses): Automatically captures important tourist spots within the user's field of view.

[1068] Step 3:

[1069] Device (glasses): Captured image data is temporarily stored in local storage, along with the capture time and location information.

[1070] Step 4:

[1071] Device (glasses): Sends stored image data to the server at regular intervals (e.g., every hour).

[1072] Step 5:

[1073] Server: Stores the received image data in a database for analysis.

[1074] Step 6:

[1075] Server: Using an image recognition algorithm, the received image data is analyzed and tourist spots are automatically recognized. Metadata (location, time) is also associated with the images.

[1076] Step 7:

[1077] Server: Combines image data from multiple tourist spots to generate a travel album.

[1078] Step 8:

[1079] Server: Once the album is created, notify the user and provide an access link.

[1080] Step 9:

[1081] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album.

[1082] In this way, the process steps can be used to efficiently improve lifestyle habits and record travel. Privacy is also protected by a user notification function.

[1083] Example 1

[1084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1085] With current technology, it takes a lot of time and effort to record a user's behavior in detail and provide appropriate advice. Furthermore, manual recording and analysis can lack accuracy and consistency. Furthermore, there is a risk that users will lose motivation to improve their lifestyle habits because there is a lack of a way for them to receive specific feedback on their improvements. For this reason, there is a need for a system that can automatically record a user's behavior, analyze it, and provide advice.

[1086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1087] In this invention, the server includes means for analyzing image data to identify objects and assign classification tags, means for evaluating user behavior based on the analysis results, and means for generating specific advice using a generative AI model, thereby enabling automatic analysis of user behavior data and prompt provision of appropriate advice.

[1088] "Image recognition" refers to a device's ability to analyze images and identify objects.

[1089] "Device" refers to an electronic device that photographs and records user behavior and transmits the data to a server at regular intervals.

[1090] "Server" refers to a computer system that receives data sent from devices via a network, analyzes the data, and processes and stores the results.

[1091] "Image data" refers to image information captured by a device and recording the user's behavior and environment.

[1092] "Local storage" refers to a storage device installed within a device for temporarily storing data.

[1093] "Analysis results" refers to data that the server analyzes images to extract information about the user's behavior and environment, and assigns classification tags to.

[1094] A "generative AI model" is a pre-trained artificial intelligence model that uses algorithms to generate appropriate responses and advice based on input data.

[1095] "Advice" refers to a message that suggests specific guidelines for action or improvements to the user based on the analysis results.

[1096] "User terminal" refers to electronic devices that are directly used by users, such as devices and smartphones.

[1097] "Metadata" refers to supplementary information that accompanies image data, including the time of shooting, location information, and the like.

[1098] This invention proposes a system that automatically records user behavior and provides advice based on the analysis results. The system mainly includes a device with image recognition capabilities, a terminal that transmits and stores data, a server that analyzes and evaluates image data, and a means for generating specific advice using a generative AI model.

[1099] System configuration

[1100] 1. Devices with image recognition capabilities

[1101] Device (glasses):

[1102] The device has a built-in camera and image analysis capabilities, and captures real-time images of the user's behavior and environment. The captured image data is temporarily stored in local storage. The device also has a visual (e.g., LED) or audio interface to notify the user that images are being recorded.

[1103] 2. Data transmission and storage

[1104] Device (glasses):

[1105] The collected image data is sent to a server at regular intervals, including the time and location of the image.

[1106] server:

[1107] The received image data is stored in an analytical database, allowing for subsequent analysis.

[1108] 3. Analysis of image data

[1109] server:

[1110] The system applies image recognition algorithms (e.g., TensorFlow, OpenCV) to the received image data, identifies objects in the data, and assigns classification tags. For example, it identifies the type and quantity of food in a photo of a meal.

[1111] 4. Evaluation of behavioral data

[1112] server:

[1113] Based on the analysis results, the user's behavioral data is evaluated. Specific evaluation items include calorie calculation of meal contents, evaluation of nutritional balance, and evaluation of exercise frequency and intensity.

[1114] 5. Automatic Advice Generation

[1115] server:

[1116] Based on the evaluation data obtained from the analysis results, specific advice is generated using a generative AI model (e.g., GPT-3). The advice is in a format that is easy for users to understand and follow.

[1117] 6. Advice Notification

[1118] server:

[1119] The generated advice is sent to the user's device (glasses or smartphone).

[1120] Device (glasses or smartphone):

[1121] The user is notified of the received advice, either visually or by audio.

[1122] Specific examples

[1123] Food records and advice

[1124] User:

[1125] Eat lunch.

[1126] Device (glasses):

[1127] The system takes a photo of the user's meal and stores the image data in local storage. This data is sent to the server at 12:00 PM intervals.

[1128] server:

[1129] The received image data is analyzed to identify the type and quantity of food, then the calories and nutrients are calculated and the evaluation results are stored in a database.

[1130] server:

[1131] The system generates advice such as "You didn't have many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[1132] Device (smartphone):

[1133] The message "Today's lunch is high in calories, so you should add a salad to your dinner" is displayed to notify the user.

[1134] Prompt Sentence Examples

[1135] Analyze an image of a user eating lunch, identify the type and amount of food, calculate calories and nutrients, and generate appropriate recommendations.

[1136] Generate a trip record

[1137] User:

[1138] Walk around the tourist spots.

[1139] Device (glasses):

[1140] Photograph important tourist spots and store the image data in local storage.

[1141] Device (glasses):

[1142] During sightseeing, image data is periodically sent to the server.

[1143] server:

[1144] The received image data is analyzed and tourist spots are automatically recognized. The image data includes the time of shooting and location information.

[1145] server:

[1146] After the trip, a travel album is automatically generated based on the saved image data.

[1147] server:

[1148] Once the album has been created, the user will be notified and provided with an access link.

[1149] Device (smartphone):

[1150] Display the travel album so that the user can check it.

[1151] Prompt Sentence Examples

[1152] Analyze the received tourist attraction image data, identify tourist attractions based on the metadata, and generate a travel album.

[1153] This system allows users to efficiently improve their lifestyle habits and record their activities without any hassle. It also has privacy protection features (such as notifications during recording), so users can use it with peace of mind.

[1154] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1155] Step 1:

[1156] Device (glasses): The built-in camera captures the user's actions. The captured image data is temporarily stored in the device's local storage. At this time, the LED on the device lights up to notify the user that a photo is being taken.

[1157] Input: User behavior, environment

[1158] Output: Captured image data

[1159] Step 2:

[1160] Device (glasses): Image data, shooting time, and location information stored in local storage are sent to the server at regular intervals. While the data is being sent, the device's notification function is activated to notify the user that data is being sent.

[1161] Input: Image data from local storage, shooting time, location information

[1162] Output: Image data and metadata (photo time, location information) sent to the server

[1163] Step 3:

[1164] Server: Apply image recognition algorithms (e.g., TensorFlow, OpenCV) to the received image data to identify objects in the data and assign classification tags. As a result of the analysis process, noteworthy content in the image (e.g., food, tourist attractions) is identified.

[1165] Input: Image data, metadata (shooting time, location information)

[1166] Output: Analysis results (object classification tags)

[1167] Step 4:

[1168] Server: Evaluates the user's behavioral data based on the analysis results. Analyzes meal content, calculates calories, evaluates nutritional balance, and analyzes exercise to evaluate frequency and intensity.

[1169] Input: Analysis results (object classification tags)

[1170] Output: Evaluation results of behavioral data (calorie calculation, nutritional balance evaluation, exercise frequency and intensity)

[1171] Step 5:

[1172] Server: Based on the evaluation results, a generative AI model (e.g., GPT-3) is used to generate specific advice. The generated advice is easy for users to understand and follow.

[1173] Input: Evaluation results of behavioral data

[1174] Output: The generated advice

[1175] Step 6:

[1176] Server: Sends the generated advice to the user's device (glasses or smartphone). While sending the advice, the server notifies the user that a message has been received.

[1177] Input: Generated advice

[1178] Output: Advice sent to the user device (glasses or smartphone)

[1179] Step 7:

[1180] Device (smartphone or glasses): Notifies the user of the received advice, either visually (as a pop-up message) or by voice.

[1181] Input: Advice sent by the server

[1182] Output: Advice displayed to the user

[1183] As described above, this system automatically records the user's behavior and provides advice based on the analysis results, allowing for efficient lifestyle improvement and behavioral recording.It also has a notification function to protect privacy, so it can be used with peace of mind.

[1184] (Application example 1)

[1185] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1186] With traditional methods of providing customer service in brick-and-mortar stores, it is difficult to understand individual customer behavior and interests in real time, making it difficult to provide personalized service. Furthermore, staff often lack the information they need to make appropriate product recommendations to customers as needed, resulting in missed opportunities to maximize customer satisfaction. A system that can solve these issues and improve the quality of customer service is needed.

[1187] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1188] In this invention, the server includes a device with an image recognition function that captures images of customer behavior in a physical store and transmits the obtained image data to the server, a means for the server to analyze customer behavior patterns and interests and provide specific advice to store staff, and a means for identifying customer interests contained in the image data and providing advice recommending related products. This enables real-time analysis of customer behavior and interests in a physical store, enabling staff to provide appropriate advice and product recommendations.

[1189] A "device with image recognition functionality" is a device that has a built-in camera and image analysis functionality and is used to capture and analyze the surrounding environment and people's behavior.

[1190] A "server" is a computer system that receives, stores, and analyzes data over a network and provides necessary information.

[1191] "Image data" refers to data that includes visual information acquired by a photographic device such as a camera.

[1192] "Analysis" is the process of using algorithms to understand the content of acquired image data and recognize specific patterns or objects.

[1193] A "user" is a person using a device with image recognition capabilities or a person receiving advice from a server.

[1194] "Advice" is any instruction, advice, or suggestion provided to the user based on the results of the analysis.

[1195] A "brick and mortar store" is a retail establishment that exists in a physical location and offers goods and services.

[1196] A "customer" is a consumer who visits a physical store and purchases or uses goods or services.

[1197] A "behavioral pattern" is a series of actions and trends in interests that indicate how customers move around the store and what they are interested in.

[1198] "Recommendations" are the suggestion of products or services based on a customer's interests.

[1199] This invention provides a system for improving customer service by analyzing customer behavior in a physical store in real time using a device and server equipped with image recognition functionality.

[1200] System configuration

[1201] 1. Devices with image recognition capabilities

[1202] Terminal (smart glasses): This device has a built-in camera and image analysis capabilities, and automatically captures customer behavior in the store. The smart glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that an image is being recorded.

[1203] 2. Transmission and storage of image data

[1204] Terminal (smart glasses): Sends collected image data to the server at regular intervals, including the time of shooting and location information.

[1205] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[1206] 3. Analysis of image data

[1207] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (e.g., products, behavioral patterns, etc.) are identified and assigned classification tags.

[1208] 4. Customer Service Analytics

[1209] Server: Evaluates customer behavior patterns and interests based on the analysis results, including identifying customer areas of interest and suggesting related products.

[1210] Server: Generates specific advice based on the evaluation results and stores it in a database.

[1211] 5. Advice Notification

[1212] Server: Sends the generated advice to the store staff's devices (smart glasses, smartphones, etc.) in a format that is easy for staff to understand and implement.

[1213] Terminal (smart glasses or smartphone): Notifies staff of received advice, either by displaying a visual message or by making an audio notification.

[1214] Specific examples

[1215] Product Recommendations

[1216] Customer: Spends a long time in the store looking at new jackets.

[1217] Terminal (smart glasses): Automatically captures customer behavior and stores the image data in local storage.

[1218] Terminal (smart glasses): Sends image data to the server at regular intervals.

[1219] Server: Analyzes the received image data and identifies the customer's interests. For example, the analysis result may be, "This customer is interested in new jackets."

[1220] Server: Generates advice such as "Suggest sales information for related products to this customer" and sends it to the staff member's smartphone.

[1221] Device (smartphone): Displays the message "Please suggest related products for this customer's new jacket" and notifies the staff.

[1222] Improved customer service

[1223] Customer: Walks around the store, checking out multiple products.

[1224] Terminal (smart glasses): Automatically captures customer behavior and stores the image data in local storage.

[1225] Terminal (smart glasses): Sends image data to the server at regular intervals.

[1226] Server: Analyzes the received image data and identifies customer behavior patterns. For example, it recognizes a behavior pattern of "looking at many products in a short period of time."

[1227] Server: Generates advice such as, "This customer is interested in many products, so we will approach them individually and introduce them to special offers," and sends this advice to the staff member's smartphone.

[1228] Terminal (smartphone): Display the message "Inform customers who are viewing many products in a short time about special offers" and notify the staff.

[1229] Example of input prompt for generative AI model

[1230] Image data showing a user looking at a particular product for a long time in the store was sent to a server. The server used an image recognition algorithm to identify the customer's interests and generate a recommendation: "This customer is interested in a new jacket." Staff receive this notification and suggest sales information for the new jacket and related items to the customer. Other visual information used by the AI ​​model includes pop-up information about related products and promotional codes.

[1231] This system allows physical stores to understand customer behavior and interests in real time, enabling staff to provide appropriate advice and product recommendations to customers.

[1232] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1233] Step 1:

[1234] The device (smart glasses) acquires the following information: The input includes an image of the environment seen by the user. The output is to temporarily store the acquired image data in local storage. Specifically, the device's camera takes pictures at regular intervals and stores the data in local storage.

[1235] Step 2:

[1236] The device (smart glasses) transmits image data stored in local storage to a server at regular intervals. The input includes the stored image data and its metadata (time of capture, location information). As output, this data is transmitted to the server. Specifically, the communication module in the device transmits the data to the server and confirms that the transmission was successful.

[1237] Step 3:

[1238] The server stores the received image data in an analysis database. The input includes the image data sent from the terminal and the associated metadata. The output is stored in a specified format in the database. Specifically, the server's storage system stores the received data in the appropriate format.

[1239] Step 4:

[1240] The server analyzes image data stored in a database based on image recognition algorithms. The input includes the stored image data. The output is the identification of objects and behavioral patterns in the image and the assignment of classification tags. Specifically, the server's processing unit executes the image recognition model and analyzes the results.

[1241] Step 5:

[1242] The server evaluates the customer's behavioral patterns and interests based on the analysis results. The input includes the analyzed data. The output is the evaluation results stored in a database. Specifically, the evaluation algorithm extracts behavioral patterns based on the analysis results and generates evaluation results.

[1243] Step 6:

[1244] The server generates specific advice based on the evaluation results. The input includes the evaluation results stored in the database. As an output, the generated advice message is sent to the database and the user terminal. As a specific operation, the advice generation algorithm is executed and the result is formatted in a format suitable for the user interface.

[1245] Step 7:

[1246] The terminal (smart glasses or smartphone) notifies the staff of the advice sent from the server. The input includes the advice message sent from the server. The output is a visual message display or a voice notification. As a specific operation, the display module or voice output module of the terminal is activated and the advice is provided to the staff.

[1247] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1248] This invention proposes a system that uses a device with image recognition capabilities, a server, and an emotion engine to automatically record a user's behavior, identify the user's emotional state, and provide advice based on the analysis results. Hereinafter, an embodiment of the present invention will be described in detail.

[1249] System configuration

[1250] 1. Devices with image recognition capabilities

[1251] Device (glasses): This device has a built-in camera and image analysis capabilities, and automatically captures the environment and actions in front of the user's eyes. The glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that images are being recorded.

[1252] 2. Transmission and storage of image data

[1253] Device (glasses): Sends collected image data to the server at regular intervals. This data includes the time of shooting and location information.

[1254] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[1255] 3. Analysis of image data

[1256] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (food, exercise equipment, scenery, etc.) are identified and assigned classification tags.

[1257] 4. Emotional state analysis

[1258] Server: Using the emotion engine, analyzes the user's facial expressions from the received image data and identifies their emotional state. For example, it determines whether the user is happy or stressed.

[1259] 5. Lifestyle Analytics

[1260] Server: Based on the analysis results, the server evaluates the user's diet, exercise, and emotional state, including calorie calculations, nutritional balance assessments, exercise frequency and intensity, and emotional state assessments.

[1261] Server: Generates specific advice based on the evaluation results and stores it in a database. For example, if the user is feeling emotionally stressed, it will suggest relaxation techniques.

[1262] 6. Advice Notification

[1263] Server: Sends the generated advice to the user's device (glasses, smartphone, etc.) in a format that is easy for the user to understand and follow.

[1264] Device (glasses or smartphone): Notifies the user of the received advice by displaying a visual message or making an audio notification.

[1265] Specific examples

[1266] Food records and advice

[1267] User: Eat lunch.

[1268] Device (glasses): Automatically takes photos of the user while they are eating and stores the image data in local storage.

[1269] Device (glasses): Sends image data of lunch to the server at the 12:00 PM interval.

[1270] Server: Analyzes the received image data and identifies the type and quantity of food.

[1271] Server: Calculates calories and nutrients and stores the evaluation results in a database.

[1272] Server: Generates advice such as "There weren't many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[1273] Device (smartphone): Displays a message to the user saying, "Today's lunch is high in calories, so you should add a salad to your dinner."

[1274] Server: The emotion engine analyzes the user's facial expressions while they are eating and evaluates whether they are enjoying the meal. For example, if the user is not enjoying the meal, it can add advice such as "Next time, try incorporating your favorite dishes."

[1275] Travel Records and Sentiment Analysis

[1276] User: Walking around tourist spots.

[1277] Device (glasses): Automatically takes photos of important tourist spots and stores the image data in local storage.

[1278] Device (glasses): While walking around the tourist spot, the device periodically sends image data to the server.

[1279] Server: Analyzes the received image data and automatically identifies tourist spots. The image data also contains metadata (location, time).

[1280] Server: After the trip, the user's travel album is automatically generated based on the saved image data.

[1281] Server: The emotion engine analyzes the user's facial expressions while sightseeing to identify the places and moments they particularly enjoyed. Based on this information, it decides which points to highlight in the travel album.

[1282] Server: Once the album is created, notify the user and provide an access link.

[1283] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album.

[1284] This system not only allows users to efficiently improve their lifestyle habits and record their activities without any hassle, but also analyzes their emotional state and provides more personalized advice.It also has mechanisms for protecting privacy (such as notifications during recording), so users can use it with peace of mind.

[1285] The processing flow will be explained below.

[1286] Food Record and Emotion Analysis Processing Steps

[1287] Step 1:

[1288] User: The user starts eating. The device (glasses) is turned on.

[1289] Step 2:

[1290] Device (glasses): The camera takes a picture of the food in front of the user's eyes.

[1291] Step 3:

[1292] Device (glasses): Captured image data is temporarily stored in local storage. When saved, the time of capture and location information are also added.

[1293] Step 4:

[1294] Device (glasses): At regular intervals (e.g., every hour), the image data stored in the local storage is sent to the server. The data includes the image file, the time of shooting, and location information.

[1295] Step 5:

[1296] Server: Stores the received image data in a database for analysis.

[1297] Step 6:

[1298] Server: Applying image recognition algorithms to analyze the received image data. This process involves identifying objects in the image (e.g., type and quantity of food).

[1299] Step 7:

[1300] Server: Based on the analysis results, calculates the calories of the food and evaluates the nutrients. Stores the results in a database.

[1301] Step 8:

[1302] Server: Using the emotion engine, analyze the user's facial expressions from the image data and identify their emotional state. Evaluate whether the user is enjoying the meal or feeling stressed.

[1303] Step 9:

[1304] Server: Generates specific advice based on the evaluation results. For example, "You didn't eat many vegetables today, so you should add a salad to your dinner." It also generates emotion-based advice, such as "You seemed a little nervous during the meal, so next time try eating in a relaxed environment."

[1305] Step 10:

[1306] Server: Sends the generated advice to the user's smartphone.

[1307] Step 11:

[1308] Device (smartphone): The received advice is displayed visually or notified by voice, allowing the user to check it and take action.

[1309] Processing steps for travel records and sentiment analysis

[1310] Step 1:

[1311] User: The user begins walking around the tourist spot. The device (glasses) is turned on.

[1312] Step 2:

[1313] Device (glasses): Automatically captures important tourist spots within the user's field of view.

[1314] Step 3:

[1315] Device (glasses): Captured image data is temporarily stored in local storage, along with the capture time and location information.

[1316] Step 4:

[1317] Device (glasses): Sends stored image data to the server at regular intervals (e.g., every hour).

[1318] Step 5:

[1319] Server: Stores the received image data in a database for analysis.

[1320] Step 6:

[1321] Server: Using an image recognition algorithm, the received image data is analyzed and tourist spots are automatically recognized. Metadata (location, time) is also associated with the images.

[1322] Step 7:

[1323] Server: Using the emotion engine, analyze the user's facial expressions while sightseeing to identify their emotional state, and identify places and moments that they are particularly enjoying.

[1324] Step 8:

[1325] Server: Combines image data from multiple tourist spots to generate a travel album, placing the images in a way that highlights moments of positive emotional states.

[1326] Step 9:

[1327] Server: Once the album is created, notify the user and provide an access link.

[1328] Step 10:

[1329] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album, highlighting the moments they particularly enjoyed.

[1330] In this way, a system can be realized that records and analyzes the user's behavior and emotions in detail through specific processing steps and provides appropriate advice.

[1331] Example 2

[1332] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1333] Conventional lifestyle improvement systems require users to manually record their behavior, diet, and exercise, which is time-consuming for users. Furthermore, it is difficult to analyze emotional states and provide advice based on them, which means that more personalized advice cannot be provided to users. There is a need for a system that can solve these issues and enable users to efficiently improve their lifestyles and record their behavior.

[1334] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1335] In this invention, the server includes a terminal with an image recognition function, means for transmitting image data acquired by the terminal to the server, means for analyzing the image data in the server and identifying the user's behavior and emotional state, and means for providing advice to the user based on the analysis. This makes it possible to automatically record the user's behavior and emotional state and provide personalized, specific advice based on the analysis results.

[1336] A "terminal with image recognition functionality" is a device that has the ability to automatically capture images of the environment and behavior in front of the user's line of sight and acquire them as image data.

[1337] "Means for transmitting acquired image data to a server" refers to a function for transferring image data from a terminal to a server at regular intervals.

[1338] "Means for analyzing image data and identifying user behavior and emotional state" refers to the process of analyzing and identifying user behavior (e.g., eating, exercise) and emotional state (e.g., joy, anger) based on the image data received by the server using an image recognition algorithm and an emotion engine.

[1339] "Means for providing advice" refers to the function of generating specific guidelines for the user (e.g., adding a salad to dinner) based on the analysis results and notifying the user's device.

[1340] "Means for providing advice based on nutritional balance and calorie calculation" refers to the process of analyzing dietary content based on image data, calculating nutrient balance and calories, and making suggestions for dietary improvement based on that.

[1341] "Means for providing advice on the frequency and intensity of exercise" refers to the process of analyzing the exercise captured in the image data, evaluating the frequency and intensity of exercise, and making suggestions that will help improve lifestyle habits.

[1342] An "emotion engine" refers to an algorithm or software that analyzes a user's facial expressions contained in image data and identifies emotional states such as joy, anger, and sadness.

[1343] A "generative AI model" refers to a model that uses artificial intelligence to generate personalized advice based on the user's analysis results.

[1344] System Overview

[1345] This invention is a system that uses a device with image recognition capabilities (e.g., glasses with a built-in camera), a server, and an emotion engine to automatically record a user's behavior, identify the user's emotional state, and provide advice based on the analysis results. The main hardware and software configuration of this system is as follows:

[1346] Hardware used

[1347] 1. Device (glasses with built-in camera): Automatically captures the environment and actions in front of the user's line of sight and acquires image data.

[1348] 2. Server: Analyzes the acquired image data, generates advice based on the analysis results, and notifies the user.

[1349] 3. User device (smartphone or tablet): Notifies the user of the advice sent from the server.

[1350] Software and algorithms used

[1351] 1. Image recognition algorithms: Identify objects in images using OpenCV, TensorFlow, etc.

[1352] 2. Emotion engine: Analyzes the user's emotional state using Microsoft Azure's Face API and Emotion API.

[1353] 3. Generative AI model: An artificial intelligence model that generates personalized advice based on the user's analysis results.

[1354] Process Overview

[1355] Automatic image capture and data collection

[1356] The device (glasses with a built-in camera) automatically captures the environment and actions in front of the user's eyes and saves the image data in local storage. The glasses are equipped with a visual or audio interface to notify the user that a photo is being taken.

[1357] Sending image data

[1358] The device (glasses with a built-in camera) collects image data and sends it to the server at regular intervals. The image data includes the time of capture and location information.

[1359] Image data storage and analysis

[1360] The server stores the received image data in an analysis database and applies image recognition algorithms to analyze it, identifying objects in the image and assigning them classification tags.

[1361] Emotional state analysis

[1362] The server uses an emotion engine to analyze the user's facial expressions contained in the image data, thereby identifying the user's emotional state, for example, whether they are happy or stressed.

[1363] Advice Generation

[1364] The server evaluates the user's diet, exercise habits, and emotional state based on the analysis results and generates specific advice, which is personalized using a generative AI model.

[1365] Specific examples

[1366] Food records and advice

[1367] When a user eats lunch, the device (glasses with a built-in camera) automatically takes pictures of the meal and collects image data. At 12:00 PM intervals, the collected image data of the lunch is sent to the server. The server analyzes the received image data and identifies the type and amount of food. It then calculates calories and nutrients and generates advice such as "You didn't eat many vegetables today, so you should add a salad to your dinner." This advice is sent to the user's device (such as a smartphone) and notified to the user. The user's facial expressions while eating can also be analyzed to evaluate whether they are enjoying their meal. For example, if the user is not enjoying their meal, the system can add advice such as "Next time, try incorporating your favorite dishes."

[1368] Prompt Sentence Examples

[1369] "Please explain how you can analyze the user's emotional state and generate advice to suggest relaxation if they are feeling stressed."

[1370] In this way, the present invention not only allows users to efficiently improve their lifestyle habits and record their activities without any hassle, but also analyzes their emotional state and allows them to receive more personalized advice. Furthermore, the system is equipped with mechanisms for protecting privacy (such as notifications during recording), allowing users to use the system with peace of mind.

[1371] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1372] System processing steps

[1373] Step 1: Automatic image capture and data collection

[1374] The device (glasses with a built-in camera) automatically captures the environment and actions in front of the user's eyes in real time.

[1375] Input: The scene in front of the user's line of sight

[1376] Output: Acquired image data (photos of the environment and behavior)

[1377] Specifically, the glasses' built-in camera periodically takes a picture and temporarily stores the video data in local storage. The glasses provide visual and audio feedback to let the user know that they are recording.

[1378] Step 2: Sending image data

[1379] The terminal (glasses with a built-in camera) sends the collected image data to a server at a predetermined interval (for example, every 5 minutes).

[1380] Input: Image data stored in local storage

[1381] Output: Image data sent to the server (including shooting time and location information)

[1382] Specifically, the device periodically divides image data into packets and sends them to the server via a wireless network (Wi-Fi or mobile data).

[1383] Step 3: Save the image data

[1384] The server stores the received image data in an analysis database.

[1385] Input: Image data sent from the device (including shooting time and location information)

[1386] Output: Image data stored in a database for analysis

[1387] Specifically, the server checks the format and integrity of the received data, stores it correctly in a database for analysis, and implements security measures to ensure the data is managed in a protected environment.

[1388] Step 4: Image Recognition and Classification

[1389] The server applies an image recognition algorithm to the image data stored in the analysis database and performs analysis.

[1390] Input: Image data stored in a database

[1391] Output: Image data with identified objects and classification tags

[1392] Specifically, it uses image recognition libraries such as OpenCV and TensorFlow to identify objects in the image (e.g., food, exercise equipment, landscapes, etc.) The objects obtained as a result of the analysis are given classification tags and stored again in the database.

[1393] Step 5: Analyze emotional state

[1394] The server uses an emotion engine to analyze the user's facial expressions in the image data.

[1395] Input: Image data that has been recognized and classified

[1396] Output: Identification of the user's emotional state (e.g., happy, angry, sad)

[1397] Specifically, it uses Microsoft Azure's Face API and Emotion API to identify the user's emotional state from their facial expressions. The analysis results are stored in a database and used for subsequent processing.

[1398] Step 6: Data analysis and advice generation

[1399] The server evaluates the user's daily activities and emotional state based on the results of image recognition and emotion analysis.

[1400] Input: Emotional state and behavior analysis results

[1401] Output: personalized advice provided to the user

[1402] Specifically, it calculates the calories in meals, evaluates nutritional balance, and evaluates the frequency and intensity of exercise, and then uses a generative AI model to generate personalized advice, such as "Since you didn't eat many vegetables today, it would be good to add a salad to your dinner."

[1403] Step 7: Advice Notification

[1404] The server transmits the generated advice to a user terminal (such as a smartphone).

[1405] Input: Advice generation results

[1406] Output: Advice given to the user

[1407] Specifically, the server sends the generated advice to the user's device in an appropriate format (text, voice, etc.), and the user is notified. The user's device (e.g., a smartphone) then notifies the user of the received advice by displaying a visual message or by voice notification.

[1408] (Application example 2)

[1409] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1410] There are limitations to analyzing consumer behavior and providing customized advice in modern brick-and-mortar stores. Specifically, it is difficult for consumers to grasp product information in real time and receive personalized purchasing advice when selecting products in the store. Furthermore, advice is not provided that takes into account the consumer's emotional state. Under these circumstances, it is difficult to increase consumer satisfaction and maximize purchasing motivation.

[1411] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1412] In this invention, the server includes means for analyzing image data acquired by a device equipped with an image recognition function and identifying user behavior, means for providing advice to the user based on the analysis, means for analyzing the emotional state of the user and generating advice based on the emotional state, and means for analyzing product information in a physical store and providing purchasing advice based on the user's emotional state. This makes it possible to provide consumers in the physical store with product information and purchasing advice based on their emotional state in real time.

[1413] A "device with image recognition capabilities" is a device that incorporates a camera and image analysis capabilities to capture and identify visual data.

[1414] The "means for transmitting image data to a server" is a communication function for transferring image data to a server via a network.

[1415] The "means for identifying user behavior" is an algorithm or program for analyzing acquired image data and extracting and identifying specific user behavior from the data.

[1416] The "means for providing advice to the user" refers to an interface or notification function for providing the user with appropriate information or instructions based on the analysis results.

[1417] The "means for analyzing the user's emotional state and generating advice based on the emotional state" is a program that analyzes data such as the user's facial expressions and voice to identify emotions and create advice based on those emotions.

[1418] The "means for analyzing product information in a physical store and providing purchasing advice based on the user's emotional state" is a system for analyzing product information from image data acquired in a physical store and generating purchasing advice based on that information, taking into account the user's emotional state.

[1419] An embodiment of the present invention is a shopping assistant system for brick-and-mortar stores that uses a device with image recognition capabilities, a server, an emotion analysis engine, and a user interface. The system analyzes, in particular, the behavior and emotional state of a user and provides purchasing advice based on the analysis in real time.

[1420] Hardware and software used

[1421] A device with image recognition capabilities: Specifically, smart glasses with a built-in camera and image analysis capabilities that automatically capture images of products in front of the user's eyes and store them in local storage.

[1422] Server: A computer that receives and analyzes collected image data using image recognition algorithms and emotion analysis engines.

[1423] Emotion analysis engine: A program that analyzes data such as a user's facial expressions and voice to identify their emotional state.

[1424] User Interface: Visual and audio notifications to provide advice to the user, specifically via the smart glasses display and audio alerts.

[1425] Data processing and calculation

[1426] Acquisition and transmission of image data:

[1427] When a user looks at a product in a physical store, the smart glasses capture an image of the product, which is then stored in local storage and sent to a server at regular intervals.

[1428] Image data analysis:

[1429] The server analyzes the received image data and extracts product features (e.g., brand, price, ingredients, etc.) using a pre-trained image recognition algorithm.

[1430] Emotional State Analysis:

[1431] Using the camera and microphone installed in the smart glasses, the user's facial expressions and voice are analyzed by an emotion analysis engine to identify the user's emotional state (e.g., interest, curiosity, dissatisfaction, etc.).

[1432] Advice generation and notification:

[1433] Based on the analysis results, the server generates advice according to the user's behavior and emotional state. The advice is then sent to the smart glasses and presented to the user visually or audibly.

[1434] Specific examples

[1435] When a user looks at a particular product, the barcode of that product is scanned and sent to the server. The server analyzes the product information and generates a notification such as "This product is on sale" and displays it on the smart glasses' display. If the server determines that the user's facial expression indicates interest, it can also provide additional information such as "You can purchase this product at a lower price than other stores."

[1436] Example prompt sentence:

[1437] Analyze information about the product the user is looking at and generate recommendations based on the product's features, price, and the user's emotional state. For example, if the user is interested, notify them that "This product is 15% off."

[1438] In this way, the system can provide real-time advice based on the user's behavior and emotional state to enhance the shopping experience in a physical store.

[1439] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1440] Step 1:

[1441] Acquisition of image data

[1442] Input: User looks at product.

[1443] How it works: The device (smart glasses) uses its built-in camera to capture images of products in front of the user's line of sight.

[1444] Output: Captured image data.

[1445] Step 2:

[1446] Local storage of image data

[1447] Input: Captured image data.

[1448] Operation: The device (smart glasses) temporarily stores image data in local storage.

[1449] Output: Image data saved to local storage.

[1450] Step 3:

[1451] Sending image data

[1452] Input: Image data stored in local storage.

[1453] Operation: The device (smart glasses) sends image data from its local storage to the server at regular intervals.

[1454] Output: Image data sent to the server.

[1455] Step 4:

[1456] Image data analysis

[1457] Input: Image data sent to the server.

[1458] How it works: The server applies image recognition algorithms to identify product features (brand, price, ingredients, etc.) in the image.

[1459] Output: Product feature data as the analysis result.

[1460] Step 5:

[1461] Emotional state analysis

[1462] Input: User's facial expression data and voice data.

[1463] How it works: The device (smart glasses) sends the user's facial expressions and voice to an emotion analysis engine, which then analyzes them to determine the user's emotional state.

[1464] Output: User's emotional state data as the analysis result.

[1465] Step 6:

[1466] Generating Advice

[1467] Input: Product feature data and user emotional state data.

[1468] Operation: The server generates optimal advice based on the product features and the user's emotional state.

[1469] Output: The generated advice.

[1470] Step 7:

[1471] Advice Notification

[1472] Input: The generated advice.

[1473] Operation: The device (smart glasses) notifies the user of the generated advice visually or audibly.

[1474] Output: Advice given to the user.

[1475] In this way, when users choose products in a physical store, they can obtain real-time product information on the spot and receive personalized purchasing advice based on their emotional state. This system is expected to increase user satisfaction and maximize purchasing motivation.

[1476] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1477] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1478] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1479] [Fourth embodiment]

[1480] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1481] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1482] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1483] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1484] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1485] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1486] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1487] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1488] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1489] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1490] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1491] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1492] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1493] This invention proposes a system that uses a device and a server equipped with image recognition functionality to automatically record user behavior and provide advice based on the analysis results. Hereinafter, an embodiment of the invention will be described in detail.

[1494] System configuration

[1495] 1. Devices with image recognition capabilities

[1496] Device (glasses): This device has a built-in camera and image analysis capabilities, and automatically captures the environment and actions in front of the user's eyes. The glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that images are being recorded.

[1497] 2. Transmission and storage of image data

[1498] Device (glasses): Sends collected image data to the server at regular intervals. This data includes the time of shooting and location information.

[1499] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[1500] 3. Analysis of image data

[1501] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (food, exercise equipment, scenery, etc.) are identified and assigned classification tags.

[1502] 4. Lifestyle Analytics

[1503] Server: Based on the analysis results, the server evaluates the user's diet, exercise, and behavioral data, including calorie calculations, nutritional balance evaluations, and exercise frequency and intensity evaluations.

[1504] Server: Generates specific advice based on the evaluation results and stores it in a database.

[1505] 5. Advice Notification

[1506] Server: Sends the generated advice to the user's device (glasses, smartphone, etc.) in a format that is easy for the user to understand and follow.

[1507] Device (glasses or smartphone): Notifies the user of the received advice by displaying a visual message or making an audio notification.

[1508] Specific examples

[1509] Food records and advice

[1510] User: Eat lunch.

[1511] Device (glasses): Automatically takes photos of the user while they are eating and stores the image data in local storage.

[1512] Device (glasses): Sends image data of lunch to the server at the 12:00 PM interval.

[1513] Server: Analyzes the received image data and identifies the type and quantity of food.

[1514] Server: Calculates calories and nutrients and stores the evaluation results in a database.

[1515] Server: Generates advice such as "There weren't many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[1516] Device (smartphone): Displays a message to the user saying, "Today's lunch is high in calories, so you should add a salad to your dinner."

[1517] Travel Log

[1518] User: Walking around tourist spots.

[1519] Device (glasses): Automatically takes photos of important tourist spots and stores the image data in local storage.

[1520] Device (glasses): While walking around the tourist spot, the device periodically sends image data to the server.

[1521] Server: Analyzes the received image data and automatically identifies tourist spots. The image data also contains metadata (location, time, etc.).

[1522] Server: After the trip, the user's travel album is automatically generated based on the saved image data.

[1523] Server: Once the album is created, notify the user and provide an access link.

[1524] Device (smartphone): Display the travel album so that the user can check it.

[1525] This system allows users to improve their lifestyle habits and record their activities efficiently without any hassle. Furthermore, it has mechanisms for protecting privacy (such as notifications during recording), so users can use it with peace of mind.

[1526] The processing flow will be explained below.

[1527] Food record and advice processing steps

[1528] Step 1:

[1529] User: The user starts eating. The device (glasses) is turned on.

[1530] Step 2:

[1531] Device (glasses): The camera takes a picture of the food in front of the user's eyes.

[1532] Step 3:

[1533] Device (glasses): Captured image data is temporarily stored in local storage. When saved, the time of capture and location information are also added.

[1534] Step 4:

[1535] Device (glasses): At regular intervals (e.g., every hour), the image data stored in the local storage is sent to the server. The data includes the image file, the time of shooting, and location information.

[1536] Step 5:

[1537] Server: Stores the received image data in a database for analysis.

[1538] Step 6:

[1539] Server: Applying image recognition algorithms to analyze the received image data. This process involves identifying objects in the image (e.g., type and quantity of food).

[1540] Step 7:

[1541] Server: Based on the analysis results, calculates the calories of the food and evaluates the nutrients. Stores the results in a database.

[1542] Step 8:

[1543] Server: Generates specific advice based on the evaluation results. For example, "There weren't many vegetables today, so it would be good to add a salad to dinner."

[1544] Step 9:

[1545] Server: Sends the generated advice to the user's smartphone.

[1546] Step 10:

[1547] Device (smartphone): The received advice is displayed visually or notified by voice, allowing the user to check it and take action.

[1548] Travel Record Processing Steps

[1549] Step 1:

[1550] User: The user begins walking around the tourist spot. The device (glasses) is turned on.

[1551] Step 2:

[1552] Device (glasses): Automatically captures important tourist spots within the user's field of view.

[1553] Step 3:

[1554] Device (glasses): Captured image data is temporarily stored in local storage, along with the capture time and location information.

[1555] Step 4:

[1556] Device (glasses): Sends stored image data to the server at regular intervals (e.g., every hour).

[1557] Step 5:

[1558] Server: Stores the received image data in a database for analysis.

[1559] Step 6:

[1560] Server: Using an image recognition algorithm, the received image data is analyzed and tourist spots are automatically recognized. Metadata (location, time) is also associated with the images.

[1561] Step 7:

[1562] Server: Combines image data from multiple tourist spots to generate a travel album.

[1563] Step 8:

[1564] Server: Once the album is created, notify the user and provide an access link.

[1565] Step 9:

[1566] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album.

[1567] In this way, the process steps can be used to efficiently improve lifestyle habits and record travel. Privacy is also protected by a user notification function.

[1568] Example 1

[1569] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1570] With current technology, it takes a lot of time and effort to record a user's behavior in detail and provide appropriate advice. Furthermore, manual recording and analysis can lack accuracy and consistency. Furthermore, there is a risk that users will lose motivation to improve their lifestyle habits because there is a lack of a way for them to receive specific feedback on their improvements. For this reason, there is a need for a system that can automatically record a user's behavior, analyze it, and provide advice.

[1571] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1572] In this invention, the server includes means for analyzing image data to identify objects and assign classification tags, means for evaluating user behavior based on the analysis results, and means for generating specific advice using a generative AI model, thereby enabling automatic analysis of user behavior data and prompt provision of appropriate advice.

[1573] "Image recognition" refers to a device's ability to analyze images and identify objects.

[1574] "Device" refers to an electronic device that photographs and records user behavior and transmits the data to a server at regular intervals.

[1575] "Server" refers to a computer system that receives data sent from devices via a network, analyzes the data, and processes and stores the results.

[1576] "Image data" refers to image information captured by a device and recording the user's behavior and environment.

[1577] "Local storage" refers to a storage device installed within a device for temporarily storing data.

[1578] "Analysis results" refers to data that the server analyzes images to extract information about the user's behavior and environment, and assigns classification tags to.

[1579] A "generative AI model" is a pre-trained artificial intelligence model that uses algorithms to generate appropriate responses and advice based on input data.

[1580] "Advice" refers to a message that suggests specific guidelines for action or improvements to the user based on the analysis results.

[1581] "User terminal" refers to electronic devices that are directly used by users, such as devices and smartphones.

[1582] "Metadata" refers to supplementary information that accompanies image data, including the time of shooting, location information, and the like.

[1583] This invention proposes a system that automatically records user behavior and provides advice based on the analysis results. The system mainly includes a device with image recognition capabilities, a terminal that transmits and stores data, a server that analyzes and evaluates image data, and a means for generating specific advice using a generative AI model.

[1584] System configuration

[1585] 1. Devices with image recognition capabilities

[1586] Device (glasses):

[1587] The device has a built-in camera and image analysis capabilities, and captures real-time images of the user's behavior and environment. The captured image data is temporarily stored in local storage. The device also has a visual (e.g., LED) or audio interface to notify the user that images are being recorded.

[1588] 2. Data transmission and storage

[1589] Device (glasses):

[1590] The collected image data is sent to a server at regular intervals, including the time and location of the image.

[1591] server:

[1592] The received image data is stored in an analytical database, allowing for subsequent analysis.

[1593] 3. Analysis of image data

[1594] server:

[1595] The system applies image recognition algorithms (e.g., TensorFlow, OpenCV) to the received image data, identifies objects in the data, and assigns classification tags. For example, it identifies the type and quantity of food in a photo of a meal.

[1596] 4. Evaluation of behavioral data

[1597] server:

[1598] Based on the analysis results, the user's behavioral data is evaluated. Specific evaluation items include calorie calculation of meal contents, evaluation of nutritional balance, and evaluation of exercise frequency and intensity.

[1599] 5. Automatic Advice Generation

[1600] server:

[1601] Based on the evaluation data obtained from the analysis results, specific advice is generated using a generative AI model (e.g., GPT-3). The advice is in a format that is easy for users to understand and follow.

[1602] 6. Advice Notification

[1603] server:

[1604] The generated advice is sent to the user's device (glasses or smartphone).

[1605] Device (glasses or smartphone):

[1606] The user is notified of the received advice, either visually or by audio.

[1607] Specific examples

[1608] Food records and advice

[1609] User:

[1610] Eat lunch.

[1611] Device (glasses):

[1612] The system takes a photo of the user's meal and stores the image data in local storage. This data is sent to the server at 12:00 PM intervals.

[1613] server:

[1614] The received image data is analyzed to identify the type and quantity of food, then the calories and nutrients are calculated and the evaluation results are stored in a database.

[1615] server:

[1616] The system generates advice such as "You didn't have many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[1617] Device (smartphone):

[1618] The message "Today's lunch is high in calories, so you should add a salad to your dinner" is displayed to notify the user.

[1619] Prompt Sentence Examples

[1620] Analyze an image of a user eating lunch, identify the type and amount of food, calculate calories and nutrients, and generate appropriate recommendations.

[1621] Generate a trip record

[1622] User:

[1623] Walk around the tourist spots.

[1624] Device (glasses):

[1625] Photograph important tourist spots and store the image data in local storage.

[1626] Device (glasses):

[1627] During sightseeing, image data is periodically sent to the server.

[1628] server:

[1629] The received image data is analyzed and tourist spots are automatically recognized. The image data includes the time of shooting and location information.

[1630] server:

[1631] After the trip, a travel album is automatically generated based on the saved image data.

[1632] server:

[1633] Once the album has been created, the user will be notified and provided with an access link.

[1634] Device (smartphone):

[1635] Display the travel album so that the user can check it.

[1636] Prompt Sentence Examples

[1637] Analyze the received tourist attraction image data, identify tourist attractions based on the metadata, and generate a travel album.

[1638] This system allows users to efficiently improve their lifestyle habits and record their activities without any hassle. It also has privacy protection features (such as notifications during recording), so users can use it with peace of mind.

[1639] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1640] Step 1:

[1641] Device (glasses): The built-in camera captures the user's actions. The captured image data is temporarily stored in the device's local storage. At this time, the LED on the device lights up to notify the user that a photo is being taken.

[1642] Input: User behavior, environment

[1643] Output: Captured image data

[1644] Step 2:

[1645] Device (glasses): Image data, shooting time, and location information stored in local storage are sent to the server at regular intervals. While the data is being sent, the device's notification function is activated to notify the user that data is being sent.

[1646] Input: Image data from local storage, shooting time, location information

[1647] Output: Image data and metadata (photo time, location information) sent to the server

[1648] Step 3:

[1649] Server: Apply image recognition algorithms (e.g., TensorFlow, OpenCV) to the received image data to identify objects in the data and assign classification tags. As a result of the analysis process, noteworthy content in the image (e.g., food, tourist attractions) is identified.

[1650] Input: Image data, metadata (shooting time, location information)

[1651] Output: Analysis results (object classification tags)

[1652] Step 4:

[1653] Server: Evaluates the user's behavioral data based on the analysis results. Analyzes meal content, calculates calories, evaluates nutritional balance, and analyzes exercise to evaluate frequency and intensity.

[1654] Input: Analysis results (object classification tags)

[1655] Output: Evaluation results of behavioral data (calorie calculation, nutritional balance evaluation, exercise frequency and intensity)

[1656] Step 5:

[1657] Server: Based on the evaluation results, a generative AI model (e.g., GPT-3) is used to generate specific advice. The generated advice is easy for users to understand and follow.

[1658] Input: Evaluation results of behavioral data

[1659] Output: The generated advice

[1660] Step 6:

[1661] Server: Sends the generated advice to the user's device (glasses or smartphone). While sending the advice, the server notifies the user that a message has been received.

[1662] Input: Generated advice

[1663] Output: Advice sent to the user device (glasses or smartphone)

[1664] Step 7:

[1665] Device (smartphone or glasses): Notifies the user of the received advice, either visually (as a pop-up message) or by voice.

[1666] Input: Advice sent by the server

[1667] Output: Advice displayed to the user

[1668] As described above, this system automatically records the user's behavior and provides advice based on the analysis results, allowing for efficient lifestyle improvement and behavioral recording.It also has a notification function to protect privacy, so it can be used with peace of mind.

[1669] (Application example 1)

[1670] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1671] With traditional methods of providing customer service in brick-and-mortar stores, it is difficult to understand individual customer behavior and interests in real time, making it difficult to provide personalized service. Furthermore, staff often lack the information they need to make appropriate product recommendations to customers as needed, resulting in missed opportunities to maximize customer satisfaction. A system that can solve these issues and improve the quality of customer service is needed.

[1672] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1673] In this invention, the server includes a device with an image recognition function that captures images of customer behavior in a physical store and transmits the obtained image data to the server, a means for the server to analyze customer behavior patterns and interests and provide specific advice to store staff, and a means for identifying customer interests contained in the image data and providing advice recommending related products. This enables real-time analysis of customer behavior and interests in a physical store, enabling staff to provide appropriate advice and product recommendations.

[1674] A "device with image recognition functionality" is a device that has a built-in camera and image analysis functionality and is used to capture and analyze the surrounding environment and people's behavior.

[1675] A "server" is a computer system that receives, stores, and analyzes data over a network and provides necessary information.

[1676] "Image data" refers to data that includes visual information acquired by a photographic device such as a camera.

[1677] "Analysis" is the process of using algorithms to understand the content of acquired image data and recognize specific patterns or objects.

[1678] A "user" is a person using a device with image recognition capabilities or a person receiving advice from a server.

[1679] "Advice" is any instruction, advice, or suggestion provided to the user based on the results of the analysis.

[1680] A "brick and mortar store" is a retail establishment that exists in a physical location and offers goods and services.

[1681] A "customer" is a consumer who visits a physical store and purchases or uses goods or services.

[1682] A "behavioral pattern" is a series of actions and trends in interests that indicate how customers move around the store and what they are interested in.

[1683] "Recommendations" are the suggestion of products or services based on a customer's interests.

[1684] This invention provides a system for improving customer service by analyzing customer behavior in a physical store in real time using a device and server equipped with image recognition functionality.

[1685] System configuration

[1686] 1. Devices with image recognition capabilities

[1687] Terminal (smart glasses): This device has a built-in camera and image analysis capabilities, and automatically captures customer behavior in the store. The smart glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that an image is being recorded.

[1688] 2. Transmission and storage of image data

[1689] Terminal (smart glasses): Sends collected image data to the server at regular intervals, including the time of shooting and location information.

[1690] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[1691] 3. Analysis of image data

[1692] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (e.g., products, behavioral patterns, etc.) are identified and assigned classification tags.

[1693] 4. Customer Service Analytics

[1694] Server: Evaluates customer behavior patterns and interests based on the analysis results, including identifying customer areas of interest and suggesting related products.

[1695] Server: Generates specific advice based on the evaluation results and stores it in a database.

[1696] 5. Advice Notification

[1697] Server: Sends the generated advice to the store staff's devices (smart glasses, smartphones, etc.) in a format that is easy for staff to understand and implement.

[1698] Terminal (smart glasses or smartphone): Notifies staff of received advice, either by displaying a visual message or by making an audio notification.

[1699] Specific examples

[1700] Product Recommendations

[1701] Customer: Spends a long time in the store looking at new jackets.

[1702] Terminal (smart glasses): Automatically captures customer behavior and stores the image data in local storage.

[1703] Terminal (smart glasses): Sends image data to the server at regular intervals.

[1704] Server: Analyzes the received image data and identifies the customer's interests. For example, the analysis result may be, "This customer is interested in new jackets."

[1705] Server: Generates advice such as "Suggest sales information for related products to this customer" and sends it to the staff member's smartphone.

[1706] Device (smartphone): Displays the message "Please suggest related products for this customer's new jacket" and notifies the staff.

[1707] Improved customer service

[1708] Customer: Walks around the store, checking out multiple products.

[1709] Terminal (smart glasses): Automatically captures customer behavior and stores the image data in local storage.

[1710] Terminal (smart glasses): Sends image data to the server at regular intervals.

[1711] Server: Analyzes the received image data and identifies customer behavior patterns. For example, it recognizes a behavior pattern of "looking at many products in a short period of time."

[1712] Server: Generates advice such as, "This customer is interested in many products, so we will approach them individually and introduce them to special offers," and sends this advice to the staff member's smartphone.

[1713] Terminal (smartphone): Display the message "Inform customers who are viewing many products in a short time about special offers" and notify the staff.

[1714] Example of input prompt for generative AI model

[1715] Image data showing a user looking at a particular product for a long time in the store was sent to a server. The server used an image recognition algorithm to identify the customer's interests and generate a recommendation: "This customer is interested in a new jacket." Staff receive this notification and suggest sales information for the new jacket and related items to the customer. Other visual information used by the AI ​​model includes pop-up information about related products and promotional codes.

[1716] This system allows physical stores to understand customer behavior and interests in real time, enabling staff to provide appropriate advice and product recommendations to customers.

[1717] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1718] Step 1:

[1719] The device (smart glasses) acquires the following information: The input includes an image of the environment seen by the user. The output is to temporarily store the acquired image data in local storage. Specifically, the device's camera takes pictures at regular intervals and stores the data in local storage.

[1720] Step 2:

[1721] The device (smart glasses) transmits image data stored in local storage to a server at regular intervals. The input includes the stored image data and its metadata (time of capture, location information). As output, this data is transmitted to the server. Specifically, the communication module in the device transmits the data to the server and confirms that the transmission was successful.

[1722] Step 3:

[1723] The server stores the received image data in an analysis database. The input includes the image data sent from the terminal and the associated metadata. The output is stored in a specified format in the database. Specifically, the server's storage system stores the received data in the appropriate format.

[1724] Step 4:

[1725] The server analyzes image data stored in a database based on image recognition algorithms. The input includes the stored image data. The output is the identification of objects and behavioral patterns in the image and the assignment of classification tags. Specifically, the server's processing unit executes the image recognition model and analyzes the results.

[1726] Step 5:

[1727] The server evaluates the customer's behavioral patterns and interests based on the analysis results. The input includes the analyzed data. The output is the evaluation results stored in a database. Specifically, the evaluation algorithm extracts behavioral patterns based on the analysis results and generates evaluation results.

[1728] Step 6:

[1729] The server generates specific advice based on the evaluation results. The input includes the evaluation results stored in the database. As an output, the generated advice message is sent to the database and the user terminal. As a specific operation, the advice generation algorithm is executed and the result is formatted in a format suitable for the user interface.

[1730] Step 7:

[1731] The terminal (smart glasses or smartphone) notifies the staff of the advice sent from the server. The input includes the advice message sent from the server. The output is a visual message display or a voice notification. As a specific operation, the display module or voice output module of the terminal is activated and the advice is provided to the staff.

[1732] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1733] This invention proposes a system that uses a device with image recognition capabilities, a server, and an emotion engine to automatically record a user's behavior, identify the user's emotional state, and provide advice based on the analysis results. Hereinafter, an embodiment of the present invention will be described in detail.

[1734] System configuration

[1735] 1. Devices with image recognition capabilities

[1736] Device (glasses): This device has a built-in camera and image analysis capabilities, and automatically captures the environment and actions in front of the user's eyes. The glasses collect image data in real time and temporarily store it in local storage. It also has a visual or audio interface to notify the user that images are being recorded.

[1737] 2. Transmission and storage of image data

[1738] Device (glasses): Sends collected image data to the server at regular intervals. This data includes the time of shooting and location information.

[1739] Server: Stores the received image data in an analysis database, allowing subsequent analysis and processing.

[1740] 3. Analysis of image data

[1741] Server: Analyzes the received image data using image recognition algorithms. As a result of the analysis, objects in the image (food, exercise equipment, scenery, etc.) are identified and assigned classification tags.

[1742] 4. Emotional state analysis

[1743] Server: Using the emotion engine, analyzes the user's facial expressions from the received image data and identifies their emotional state. For example, it determines whether the user is happy or stressed.

[1744] 5. Lifestyle Analytics

[1745] Server: Based on the analysis results, the server evaluates the user's diet, exercise, and emotional state, including calorie calculations, nutritional balance assessments, exercise frequency and intensity, and emotional state assessments.

[1746] Server: Generates specific advice based on the evaluation results and stores it in a database. For example, if the user is feeling emotionally stressed, it will suggest relaxation techniques.

[1747] 6. Advice Notification

[1748] Server: Sends the generated advice to the user's device (glasses, smartphone, etc.) in a format that is easy for the user to understand and follow.

[1749] Device (glasses or smartphone): Notifies the user of the received advice by displaying a visual message or making an audio notification.

[1750] Specific examples

[1751] Food records and advice

[1752] User: Eat lunch.

[1753] Device (glasses): Automatically takes photos of the user while they are eating and stores the image data in local storage.

[1754] Device (glasses): Sends image data of lunch to the server at the 12:00 PM interval.

[1755] Server: Analyzes the received image data and identifies the type and quantity of food.

[1756] Server: Calculates calories and nutrients and stores the evaluation results in a database.

[1757] Server: Generates advice such as "There weren't many vegetables today, so you should add a salad to your dinner," and sends it to the user's smartphone.

[1758] Device (smartphone): Displays a message to the user saying, "Today's lunch is high in calories, so you should add a salad to your dinner."

[1759] Server: The emotion engine analyzes the user's facial expressions while they are eating and evaluates whether they are enjoying the meal. For example, if the user is not enjoying the meal, it can add advice such as "Next time, try incorporating your favorite dishes."

[1760] Travel Records and Sentiment Analysis

[1761] User: Walking around tourist spots.

[1762] Device (glasses): Automatically takes photos of important tourist spots and stores the image data in local storage.

[1763] Device (glasses): While walking around the tourist spot, the device periodically sends image data to the server.

[1764] Server: Analyzes the received image data and automatically identifies tourist spots. The image data also contains metadata (location, time).

[1765] Server: After the trip, the user's travel album is automatically generated based on the saved image data.

[1766] Server: The emotion engine analyzes the user's facial expressions while sightseeing to identify the places and moments they particularly enjoyed. Based on this information, it decides which points to highlight in the travel album.

[1767] Server: Once the album is created, notify the user and provide an access link.

[1768] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album.

[1769] This system not only allows users to efficiently improve their lifestyle habits and record their activities without any hassle, but also analyzes their emotional state and provides more personalized advice.It also has mechanisms for protecting privacy (such as notifications during recording), so users can use it with peace of mind.

[1770] The processing flow will be explained below.

[1771] Food Record and Emotion Analysis Processing Steps

[1772] Step 1:

[1773] User: The user starts eating. The device (glasses) is turned on.

[1774] Step 2:

[1775] Device (glasses): The camera takes a picture of the food in front of the user's eyes.

[1776] Step 3:

[1777] Device (glasses): Captured image data is temporarily stored in local storage. When saved, the time of capture and location information are also added.

[1778] Step 4:

[1779] Device (glasses): At regular intervals (e.g., every hour), the image data stored in the local storage is sent to the server. The data includes the image file, the time of shooting, and location information.

[1780] Step 5:

[1781] Server: Stores the received image data in a database for analysis.

[1782] Step 6:

[1783] Server: Applying image recognition algorithms to analyze the received image data. This process involves identifying objects in the image (e.g., type and quantity of food).

[1784] Step 7:

[1785] Server: Based on the analysis results, calculates the calories of the food and evaluates the nutrients. Stores the results in a database.

[1786] Step 8:

[1787] Server: Using the emotion engine, analyze the user's facial expressions from the image data and identify their emotional state. Evaluate whether the user is enjoying the meal or feeling stressed.

[1788] Step 9:

[1789] Server: Generates specific advice based on the evaluation results. For example, "You didn't eat many vegetables today, so you should add a salad to your dinner." It also generates emotion-based advice, such as "You seemed a little nervous during the meal, so next time try eating in a relaxed environment."

[1790] Step 10:

[1791] Server: Sends the generated advice to the user's smartphone.

[1792] Step 11:

[1793] Device (smartphone): The received advice is displayed visually or notified by voice, allowing the user to check it and take action.

[1794] Processing steps for travel records and sentiment analysis

[1795] Step 1:

[1796] User: The user begins walking around the tourist spot. The device (glasses) is turned on.

[1797] Step 2:

[1798] Device (glasses): Automatically captures important tourist spots within the user's field of view.

[1799] Step 3:

[1800] Device (glasses): Captured image data is temporarily stored in local storage, along with the capture time and location information.

[1801] Step 4:

[1802] Device (glasses): Sends stored image data to the server at regular intervals (e.g., every hour).

[1803] Step 5:

[1804] Server: Stores the received image data in a database for analysis.

[1805] Step 6:

[1806] Server: Using an image recognition algorithm, the received image data is analyzed and tourist spots are automatically recognized. Metadata (location, time) is also associated with the images.

[1807] Step 7:

[1808] Server: Using the emotion engine, analyze the user's facial expressions while sightseeing to identify their emotional state, and identify places and moments that they are particularly enjoying.

[1809] Step 8:

[1810] Server: Combines image data from multiple tourist spots to generate a travel album, placing the images in a way that highlights moments of positive emotional states.

[1811] Step 9:

[1812] Server: Once the album is created, notify the user and provide an access link.

[1813] Step 10:

[1814] Device (smartphone): The user checks the album link and accesses it. They can look back on their travel memories through the album, highlighting the moments they particularly enjoyed.

[1815] In this way, a system can be realized that records and analyzes the user's behavior and emotions in detail through specific processing steps and provides appropriate advice.

[1816] Example 2

[1817] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1818] Conventional lifestyle improvement systems require users to manually record their behavior, diet, and exercise, which is time-consuming for users. Furthermore, it is difficult to analyze emotional states and provide advice based on them, which means that more personalized advice cannot be provided to users. There is a need for a system that can solve these issues and enable users to efficiently improve their lifestyles and record their behavior.

[1819] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1820] In this invention, the server includes a terminal with an image recognition function, means for transmitting image data acquired by the terminal to the server, means for analyzing the image data in the server and identifying the user's behavior and emotional state, and means for providing advice to the user based on the analysis. This makes it possible to automatically record the user's behavior and emotional state and provide personalized, specific advice based on the analysis results.

[1821] A "terminal with image recognition functionality" is a device that has the ability to automatically capture images of the environment and behavior in front of the user's line of sight and acquire them as image data.

[1822] "Means for transmitting acquired image data to a server" refers to a function for transferring image data from a terminal to a server at regular intervals.

[1823] "Means for analyzing image data and identifying user behavior and emotional state" refers to the process of analyzing and identifying user behavior (e.g., eating, exercise) and emotional state (e.g., joy, anger) based on the image data received by the server using an image recognition algorithm and an emotion engine.

[1824] "Means for providing advice" refers to the function of generating specific guidelines for the user (e.g., adding a salad to dinner) based on the analysis results and notifying the user's device.

[1825] "Means for providing advice based on nutritional balance and calorie calculation" refers to the process of analyzing dietary content based on image data, calculating nutrient balance and calories, and making suggestions for dietary improvement based on that.

[1826] "Means for providing advice on the frequency and intensity of exercise" refers to the process of analyzing the exercise captured in the image data, evaluating the frequency and intensity of exercise, and making suggestions that will help improve lifestyle habits.

[1827] An "emotion engine" refers to an algorithm or software that analyzes a user's facial expressions contained in image data and identifies emotional states such as joy, anger, and sadness.

[1828] A "generative AI model" refers to a model that uses artificial intelligence to generate personalized advice based on the user's analysis results.

[1829] System Overview

[1830] This invention is a system that uses a device with image recognition capabilities (e.g., glasses with a built-in camera), a server, and an emotion engine to automatically record a user's behavior, identify the user's emotional state, and provide advice based on the analysis results. The main hardware and software configuration of this system is as follows:

[1831] Hardware used

[1832] 1. Device (glasses with built-in camera): Automatically captures the environment and actions in front of the user's line of sight and acquires image data.

[1833] 2. Server: Analyzes the acquired image data, generates advice based on the analysis results, and notifies the user.

[1834] 3. User device (smartphone or tablet): Notifies the user of the advice sent from the server.

[1835] Software and algorithms used

[1836] 1. Image recognition algorithms: Identify objects in images using OpenCV, TensorFlow, etc.

[1837] 2. Emotion engine: Analyzes the user's emotional state using Microsoft Azure's Face API and Emotion API.

[1838] 3. Generative AI model: An artificial intelligence model that generates personalized advice based on the user's analysis results.

[1839] Process Overview

[1840] Automatic image capture and data collection

[1841] The device (glasses with a built-in camera) automatically captures the environment and actions in front of the user's eyes and saves the image data in local storage. The glasses are equipped with a visual or audio interface to notify the user that a photo is being taken.

[1842] Sending image data

[1843] The device (glasses with a built-in camera) collects image data and sends it to the server at regular intervals. The image data includes the time of capture and location information.

[1844] Image data storage and analysis

[1845] The server stores the received image data in an analysis database and applies image recognition algorithms to analyze it, identifying objects in the image and assigning them classification tags.

[1846] Emotional state analysis

[1847] The server uses an emotion engine to analyze the user's facial expressions contained in the image data, thereby identifying the user's emotional state, for example, whether they are happy or stressed.

[1848] Advice Generation

[1849] The server evaluates the user's diet, exercise habits, and emotional state based on the analysis results and generates specific advice, which is personalized using a generative AI model.

[1850] Specific examples

[1851] Food records and advice

[1852] When a user eats lunch, the device (glasses with a built-in camera) automatically takes pictures of the meal and collects image data. At 12:00 PM intervals, the collected image data of the lunch is sent to the server. The server analyzes the received image data and identifies the type and amount of food. It then calculates calories and nutrients and generates advice such as "You didn't eat many vegetables today, so you should add a salad to your dinner." This advice is sent to the user's device (such as a smartphone) and notified to the user. The user's facial expressions while eating can also be analyzed to evaluate whether they are enjoying their meal. For example, if the user is not enjoying their meal, the system can add advice such as "Next time, try incorporating your favorite dishes."

[1853] Prompt Sentence Examples

[1854] "Please explain how you can analyze the user's emotional state and generate advice to suggest relaxation if they are feeling stressed."

[1855] In this way, the present invention not only allows users to efficiently improve their lifestyle habits and record their activities without any hassle, but also analyzes their emotional state and allows them to receive more personalized advice. Furthermore, the system is equipped with mechanisms for protecting privacy (such as notifications during recording), allowing users to use the system with peace of mind.

[1856] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1857] System processing steps

[1858] Step 1: Automatic image capture and data collection

[1859] The device (glasses with a built-in camera) automatically captures the environment and actions in front of the user's eyes in real time.

[1860] Input: The scene in front of the user's line of sight

[1861] Output: Acquired image data (photos of the environment and behavior)

[1862] Specifically, the glasses' built-in camera periodically takes a picture and temporarily stores the video data in local storage. The glasses provide visual and audio feedback to let the user know that they are recording.

[1863] Step 2: Sending image data

[1864] The terminal (glasses with a built-in camera) sends the collected image data to a server at a predetermined interval (for example, every 5 minutes).

[1865] Input: Image data stored in local storage

[1866] Output: Image data sent to the server (including shooting time and location information)

[1867] Specifically, the device periodically divides image data into packets and sends them to the server via a wireless network (Wi-Fi or mobile data).

[1868] Step 3: Save the image data

[1869] The server stores the received image data in an analysis database.

[1870] Input: Image data sent from the device (including shooting time and location information)

[1871] Output: Image data stored in a database for analysis

[1872] Specifically, the server checks the format and integrity of the received data, stores it correctly in a database for analysis, and implements security measures to ensure the data is managed in a protected environment.

[1873] Step 4: Image Recognition and Classification

[1874] The server applies an image recognition algorithm to the image data stored in the analysis database and performs analysis.

[1875] Input: Image data stored in a database

[1876] Output: Image data with identified objects and classification tags

[1877] Specifically, it uses image recognition libraries such as OpenCV and TensorFlow to identify objects in the image (e.g., food, exercise equipment, landscapes, etc.) The objects obtained as a result of the analysis are given classification tags and stored again in the database.

[1878] Step 5: Analyze emotional state

[1879] The server uses an emotion engine to analyze the user's facial expressions in the image data.

[1880] Input: Image data that has been recognized and classified

[1881] Output: Identification of the user's emotional state (e.g., happy, angry, sad)

[1882] Specifically, it uses Microsoft Azure's Face API and Emotion API to identify the user's emotional state from their facial expressions. The analysis results are stored in a database and used for subsequent processing.

[1883] Step 6: Data analysis and advice generation

[1884] The server evaluates the user's daily activities and emotional state based on the results of image recognition and emotion analysis.

[1885] Input: Emotional state and behavior analysis results

[1886] Output: personalized advice provided to the user

[1887] Specifically, it calculates the calories in meals, evaluates nutritional balance, and evaluates the frequency and intensity of exercise, and then uses a generative AI model to generate personalized advice, such as "Since you didn't eat many vegetables today, it would be good to add a salad to your dinner."

[1888] Step 7: Advice Notification

[1889] The server transmits the generated advice to a user terminal (such as a smartphone).

[1890] Input: Advice generation results

[1891] Output: Advice given to the user

[1892] Specifically, the server sends the generated advice to the user's device in an appropriate format (text, voice, etc.), and the user is notified. The user's device (e.g., a smartphone) then notifies the user of the received advice by displaying a visual message or by voice notification.

[1893] (Application example 2)

[1894] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1895] There are limitations to analyzing consumer behavior and providing customized advice in modern brick-and-mortar stores. Specifically, it is difficult for consumers to grasp product information in real time and receive personalized purchasing advice when selecting products in the store. Furthermore, advice is not provided that takes into account the consumer's emotional state. Under these circumstances, it is difficult to increase consumer satisfaction and maximize purchasing motivation.

[1896] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1897] In this invention, the server includes means for analyzing image data acquired by a device equipped with an image recognition function and identifying user behavior, means for providing advice to the user based on the analysis, means for analyzing the emotional state of the user and generating advice based on the emotional state, and means for analyzing product information in a physical store and providing purchasing advice based on the user's emotional state. This makes it possible to provide consumers in the physical store with product information and purchasing advice based on their emotional state in real time.

[1898] A "device with image recognition capabilities" is a device that incorporates a camera and image analysis capabilities to capture and identify visual data.

[1899] The "means for transmitting image data to a server" is a communication function for transferring image data to a server via a network.

[1900] The "means for identifying user behavior" is an algorithm or program for analyzing acquired image data and extracting and identifying specific user behavior from the data.

[1901] The "means for providing advice to the user" refers to an interface or notification function for providing the user with appropriate information or instructions based on the analysis results.

[1902] The "means for analyzing the user's emotional state and generating advice based on the emotional state" is a program that analyzes data such as the user's facial expressions and voice to identify emotions and create advice based on those emotions.

[1903] The "means for analyzing product information in a physical store and providing purchasing advice based on the user's emotional state" is a system for analyzing product information from image data acquired in a physical store and generating purchasing advice based on that information, taking into account the user's emotional state.

[1904] An embodiment of the present invention is a shopping assistant system for brick-and-mortar stores that uses a device with image recognition capabilities, a server, an emotion analysis engine, and a user interface. The system analyzes, in particular, the behavior and emotional state of a user and provides purchasing advice based on the analysis in real time.

[1905] Hardware and software used

[1906] A device with image recognition capabilities: Specifically, smart glasses with a built-in camera and image analysis capabilities that automatically capture images of products in front of the user's eyes and store them in local storage.

[1907] Server: A computer that receives and analyzes collected image data using image recognition algorithms and emotion analysis engines.

[1908] Emotion analysis engine: A program that analyzes data such as a user's facial expressions and voice to identify their emotional state.

[1909] User Interface: Visual and audio notifications to provide advice to the user, specifically via the smart glasses display and audio alerts.

[1910] Data processing and calculation

[1911] Acquisition and transmission of image data:

[1912] When a user looks at a product in a physical store, the smart glasses capture an image of the product, which is then stored in local storage and sent to a server at regular intervals.

[1913] Image data analysis:

[1914] The server analyzes the received image data and extracts product features (e.g., brand, price, ingredients, etc.) using a pre-trained image recognition algorithm.

[1915] Emotional State Analysis:

[1916] Using the camera and microphone installed in the smart glasses, the user's facial expressions and voice are analyzed by an emotion analysis engine to identify the user's emotional state (e.g., interest, curiosity, dissatisfaction, etc.).

[1917] Advice generation and notification:

[1918] Based on the analysis results, the server generates advice according to the user's behavior and emotional state. The advice is then sent to the smart glasses and presented to the user visually or audibly.

[1919] Specific examples

[1920] When a user looks at a particular product, the barcode of that product is scanned and sent to the server. The server analyzes the product information and generates a notification such as "This product is on sale" and displays it on the smart glasses' display. If the server determines that the user's facial expression indicates interest, it can also provide additional information such as "You can purchase this product at a lower price than other stores."

[1921] Example prompt sentence:

[1922] Analyze information about the product the user is looking at and generate recommendations based on the product's features, price, and the user's emotional state. For example, if the user is interested, notify them that "This product is 15% off."

[1923] In this way, the system can provide real-time advice based on the user's behavior and emotional state to enhance the shopping experience in a physical store.

[1924] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1925] Step 1:

[1926] Acquisition of image data

[1927] Input: User looks at product.

[1928] How it works: The device (smart glasses) uses its built-in camera to capture images of products in front of the user's line of sight.

[1929] Output: Captured image data.

[1930] Step 2:

[1931] Local storage of image data

[1932] Input: Captured image data.

[1933] Operation: The device (smart glasses) temporarily stores image data in local storage.

[1934] Output: Image data saved to local storage.

[1935] Step 3:

[1936] Sending image data

[1937] Input: Image data stored in local storage.

[1938] Operation: The device (smart glasses) sends image data from its local storage to the server at regular intervals.

[1939] Output: Image data sent to the server.

[1940] Step 4:

[1941] Image data analysis

[1942] Input: Image data sent to the server.

[1943] How it works: The server applies image recognition algorithms to identify product features (brand, price, ingredients, etc.) in the image.

[1944] Output: Product feature data as the analysis result.

[1945] Step 5:

[1946] Emotional state analysis

[1947] Input: User's facial expression data and voice data.

[1948] How it works: The device (smart glasses) sends the user's facial expressions and voice to an emotion analysis engine, which then analyzes them to determine the user's emotional state.

[1949] Output: User's emotional state data as the analysis result.

[1950] Step 6:

[1951] Generating Advice

[1952] Input: Product feature data and user emotional state data.

[1953] Operation: The server generates optimal advice based on the product features and the user's emotional state.

[1954] Output: The generated advice.

[1955] Step 7:

[1956] Advice Notification

[1957] Input: The generated advice.

[1958] Operation: The device (smart glasses) notifies the user of the generated advice visually or audibly.

[1959] Output: Advice given to the user.

[1960] In this way, when users choose products in a physical store, they can obtain real-time product information on the spot and receive personalized purchasing advice based on their emotional state. This system is expected to increase user satisfaction and maximize purchasing motivation.

[1961] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1962] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1963] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1964] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1965] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1966] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1967] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1968] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1969] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1970] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1971] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1972] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1973] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1974] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1975] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1976] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1977] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1978] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1979] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1980] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1981] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1982] The following is further disclosed regarding the above embodiment.

[1983] (Claim 1)

[1984] A device with image recognition capabilities,

[1985] means for transmitting image data acquired by the device to a server;

[1986] means for analyzing the image data in the server and identifying user behavior;

[1987] means for providing advice to a user based on said analysis;

[1988] A system including:

[1989] (Claim 2)

[1990] The image data further includes a means for analyzing dietary content included in the image data and providing advice on nutritional balance.

[1991] 10. The system of claim 1.

[1992] (Claim 3)

[1993] The system further includes a means for analyzing the state of exercise included in the image data and providing advice on improving lifestyle habits.

[1994] 10. The system of claim 1.

[1995] (Claim 4)

[1996] The image data may further include a means for identifying and automatically recording scenery during travel.

[1997] 10. The system of claim 1.

[1998] (Claim 5)

[1999] further comprising means for notifying a user that the device is recording an image.

[2000] 10. The system of claim 1.

[2001] (Claim 6)

[2002] The method further includes means for classifying user behavior data into categories using image data stored in the server.

[2003] 10. The system of claim 1.

[2004] "Example 1"

[2005] (Claim 1)

[2006] A device with an image recognition function that captures user behavior;

[2007] means for temporarily storing image data acquired by the device and transmitting the image data to a server;

[2008] means for analyzing the image data in the server, identifying objects, and assigning classification tags;

[2009] A means for evaluating the user's behavior based on the analysis results and generating specific advice using a generative AI model;

[2010] means for notifying a user terminal of the advice;

[2011] A system including:

[2012] (Claim 2)

[2013] The image data further includes a means for analyzing the dietary content included in the image data, calculating calories and evaluating nutritional balance, and providing advice regarding nutritional balance.

[2014] 10. The system of claim 1.

[2015] (Claim 3)

[2016] The system further includes a means for analyzing the state of exercise contained in the image data, evaluating the frequency and intensity of the exercise, and providing advice on improving lifestyle habits.

[2017] 10. The system of claim 1.

[2018] "Application Example 1"

[2019] (Claim 1)

[2020] A device with image recognition capabilities,

[2021] means for transmitting image data acquired by the device to a server;

[2022] means for analyzing the image data in the server and identifying user behavior;

[2023] means for providing advice to a user based on said analysis;

[2024] A means for the device to capture images of customer behavior in a physical store and transmit the captured image data to a server;

[2025] A means for the server to analyze the behavioral patterns and interests of customers and provide specific advice to store staff;

[2026] A system including:

[2027] (Claim 2)

[2028] The system of claim 1 , further comprising means for identifying customer interests contained in the image data and providing advice recommending related products.

[2029] (Claim 3)

[2030] 10. The system of claim 1, further comprising means for analyzing customer behavior patterns contained in the image data and providing advice regarding improvement of customer service.

[2031] "Example 2: Combining Emotion Engines"

[2032] (Claim 1)

[2033] A device with image recognition capabilities,

[2034] means for transmitting image data acquired by the terminal to a server;

[2035] means at said server for analyzing said image data to identify a user's behavior and emotional state;

[2036] means for providing advice to a user based on said analysis;

[2037] A system including:

[2038] (Claim 2)

[2039] The device further includes a means for analyzing the dietary content included in the image data and providing advice based on nutritional balance and calorie calculation.

[2040] 10. The system of claim 1.

[2041] (Claim 3)

[2042] The system further includes a means for analyzing the state of exercise contained in the image data and providing advice regarding the frequency and intensity of exercise.

[2043] 10. The system of claim 1.

[2044] (Claim 4)

[2045] and means for analyzing the user's emotional state from the image data using the emotion engine and providing advice based on the emotional state.

[2046] 10. The system of claim 1.

[2047] (Claim 5)

[2048] and means for performing the assessment by a generative AI model and providing personalized advice.

[2049] 10. The system of claim 1.

[2050] "Application example 2 when combining emotion engines"

[2051] (Claim 1)

[2052] A device with image recognition capabilities,

[2053] means for transmitting image data acquired by the device to a server;

[2054] means for analyzing the image data in the server and identifying user behavior;

[2055] means for providing advice to a user based on said analysis;

[2056] means for analyzing the emotional state of the user and generating advice based on the emotional state;

[2057] A means for analyzing product information in a physical store and providing purchasing advice according to the user's emotional state;

[2058] A system including:

[2059] (Claim 2)

[2060] The image data further includes a means for analyzing dietary content included in the image data and providing advice on nutritional balance.

[2061] 10. The system of claim 1.

[2062] (Claim 3)

[2063] The system further includes a means for analyzing the state of exercise included in the image data and providing advice on improving lifestyle habits.

[2064] 10. The system of claim 1. [Explanation of symbols]

[2065] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A device with image recognition capabilities, means for transmitting image data acquired by the device to a server; means for analyzing the image data in the server and identifying user behavior; means for providing advice to a user based on said analysis; A system including:

2. The image data further includes a means for analyzing dietary content included in the image data and providing advice on nutritional balance. The system of claim 1 .

3. The system further includes a means for analyzing the state of exercise included in the image data and providing advice on improving lifestyle habits. The system of claim 1 .

4. The image data may further include a means for identifying and automatically recording scenery during travel. The system of claim 1 .

5. further comprising means for notifying a user that the device is recording an image. The system of claim 1 .

6. The method further includes means for classifying user behavior data into categories using image data stored in the server. The system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A