System

The system addresses the challenge of real-time user behavior analysis by using a glasses-type device to capture gaze and audio data, encrypting it, and transmitting it to a server for personalized recommendations, thereby improving user experience through continuous learning and feedback.

JP2026033977APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137098
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing portable AI devices fail to accurately learn and analyze user behavior and interests in real time, leading to insufficiently personalized information and services due to the lack of comprehensive utilization of gaze and voice data.

Method used

A system utilizing a glasses-type device that captures gaze and audio data in real time, encrypts it, and transmits it to a server for analysis, which generates personalized recommendations based on user behavior and interests, with feedback loops for continuous improvement.

Benefits of technology

The system provides timely and accurate personalized information and services by continuously learning user behavior and interests, enhancing user experience through real-time data analysis and feedback integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026033977000001_ABST
    Figure 2026033977000001_ABST
Patent Text Reader

Abstract

To provide a system for improving user experience by quickly providing information corresponding to the action or interest of a user.SOLUTION: The eyeglass-type device includes a line-of-sight detection unit configured to acquire line-of-sight information of a user, a unit configured to capture image data of an object viewed by the user, a voice capture unit configured to record a speech of the user, a unit configured to transmit the captured data to a server, a unit configured to receive and decode the data by the server, an image analysis unit configured to analyze the image data and specify an object in which the user is interested, a voice analysis unit configured to analyze the voice data and specify a content heard by the user, a unit configured to generate recommendation information on the basis of the line-of-sight data and an analysis result, and a unit configured to encrypt the recommendation information and transmit the encrypted recommendation information to the eyeglass-type device. The eyeglass-type device includes means for decoding the received recommendation information and displaying the decoded recommendation information on the display.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] While there are many portable AI devices available today, these devices are unable to constantly learn and analyze user behavior and interests in real time, making it difficult to provide personalized information. Furthermore, few devices comprehensively utilize a variety of data, such as the user's gaze and voice data, making the information and services provided to users often insufficiently accurate. The present invention aims to solve these problems by learning and analyzing users' daily behavior with high accuracy and providing more effective personalized information and services. [Means for solving the problem]

[0005] The present invention provides a system that uses a glasses-type device to acquire a user's gaze information, image data, and audio data in real time, encrypts the data, and transmits it to a server. The server decrypts the received data, performs image and audio analysis, identifies the user's interests, and generates recommendation information. This recommendation information is encrypted and transmitted to the glasses-type device, where it is displayed on the display, providing the user with appropriate information. The system also has a function for acquiring feedback from the user and reflecting this feedback in improving the accuracy of the next recommendation. This allows for the rapid provision of information tailored to the user's behavior and interests, improving the user experience.

[0006] A "glasses-type device" is a device worn by a user that has the function of acquiring, displaying, and transmitting information related to the user's vision and hearing.

[0007] "Gaze detection means" refers to a device or technology for detecting the user's gaze and pupil movement and acquiring that data.

[0008] An "image capture means" is a device for acquiring image data of a scene or object that a user is viewing.

[0009] The "audio capture means" is a device for recording the ambient sounds around the user and the user's own speech.

[0010] "Data transmission means" refers to a device or technology for encrypting acquired data and transmitting it to a server via a communication network.

[0011] The "server" is a computer system that receives and analyzes data sent from the glasses-type device and sends the information back to the glasses-type device.

[0012] "Data receiving and decoding means" refers to the devices and technologies used to decode data received by the server and make it usable.

[0013] "Image analysis means" refers to a device or technology that analyzes received image data and identifies the object or scene the user is viewing.

[0014] "Audio analysis means" refers to a device or technology that analyzes received audio data and identifies what the user is listening to and what is being said.

[0015] A "recommendation generation means" is a device or technology that generates appropriate information or services to provide to a user based on the analysis results and the user's past behavioral history.

[0016] The "information display means" refers to a device or technology for displaying received recommendation information on the display of the glasses-type device.

[0017] The "feedback transmission means" refers to a device or technology for obtaining feedback such as voice instructions or touch operations from the user and transmitting it to the server. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] System overview and equipment configuration

[0040] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[0041] 1. Glasses-type device

[0042] The glasses-type device is equipped with the following functions:

[0043] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[0044] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[0045] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[0046] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[0047] Information display means: Received recommendation information is displayed on a waveguide type display.

[0048] Feedback sending means: Obtains feedback from users and sends it to the server.

[0049] 2. Server

[0050] The server has the following functions:

[0051] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[0052] Image analysis means: Analyzes the received image data and identifies the object or scene the user is looking at.

[0053] Audio analysis means: Analyzes received audio data to determine what the user is hearing and saying.

[0054] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[0055] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[0056] Program processing

[0057] Information collection and transmission

[0058] Terminal

[0059] While the user is wearing the glasses-type device and moving around, the gaze detection module, camera, and microphone capture the user's gaze, image data, and audio data in real time.

[0060] The acquired data is encrypted and sent to the server at regular intervals.

[0061] Analysis and recommendation generation

[0062] server

[0063] The received data is decoded and the image data is analyzed using an image analysis model to determine what the user is looking at.

[0064] The voice data is analyzed using a voice recognition model, and what the user is hearing and saying is converted into text data.

[0065] Identify the user's interests based on gaze data, image analysis results, and audio analysis results.

[0066] This is compared with the user's past behavioral history to generate appropriate recommendation information.

[0067] The generated recommendation information is encrypted and sent to the glasses-type device.

[0068] Recommendation and feedback

[0069] Terminal

[0070] The recommendation information sent from the server is received, decrypted, and displayed on the screen.

[0071] Get feedback from the user and send this feedback information back to the server.

[0072] Specific examples

[0073] Example 1: Shopping recommendations

[0074] User

[0075] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[0076] Terminal

[0077] The glasses-type device captures the gaze and images and sends them to a server.

[0078] server

[0079] The received data is analyzed to identify the clothes the user is looking at and recommend accessories and shoes that go well with them.

[0080] Terminal

[0081] Recommendation information is displayed on the screen and presented to the user.

[0082] Example 2: Support for calorie counting

[0083] User

[0084] When a user is looking at a menu at a restaurant, their eyes are drawn to the menu.

[0085] Terminal

[0086] The glasses-type device captures the gaze and images and sends them to a server.

[0087] server

[0088] It analyzes images of menu items, identifies the names of dishes, and retrieves calorie information for the dishes from a database.

[0089] Terminal

[0090] Calorie information is displayed on the display and presented to the user.

[0091] In this way, the present invention is a system that improves the user experience by learning user behavior and interests in real time and providing appropriate recommendation information.

[0092] The processing flow will be explained below.

[0093] Step 1:

[0094] Terminal

[0095] When the user puts on the glasses, the device automatically activates its gaze detection module, camera, and microphone. It detects and collects data on the user's gaze and pupil movements in real time. It also captures image data of the scenery and objects the user is looking at at a frequency of several frames per second, and records surrounding sounds and the user's speech.

[0096] Step 2:

[0097] Terminal

[0098] The acquired gaze data, image data, and audio data are temporarily stored in local memory, and this data is encrypted for later transmission to the server.

[0099] Step 3:

[0100] Terminal

[0101] The stored data is encrypted using an encryption algorithm such as AES-256, and the encrypted data is sent to the server via a secure communication network.

[0102] Step 4:

[0103] server

[0104] The server receives the encrypted data sent over the Internet and decrypts it using a decryption algorithm such as AES-256.

[0105] Step 5:

[0106] server

[0107] The decoded image data is analyzed using an image recognition model (e.g., YOLO, ResNet), which identifies the object or scene the user is looking at.

[0108] Step 6:

[0109] server

[0110] The decoded audio data is analyzed using a speech recognition model (e.g., DeepSpeech, Wav2Vec), which converts the audio data into text data and identifies what the user is hearing and saying.

[0111] Step 7:

[0112] server

[0113] By combining gaze data, image analysis results, and audio analysis results, the system identifies the user's interests. Based on this data, it compares it with the user's behavioral history and generates appropriate recommendation information.

[0114] Step 8:

[0115] server

[0116] The generated recommendation information is encrypted and retransmitted to the glasses-type device via the data transmission means.

[0117] Step 9:

[0118] Terminal

[0119] The glasses-type device decodes the received recommendation information, and the decoded information is displayed on a waveguide display and presented to the user.

[0120] Step 10:

[0121] Terminal

[0122] A function to obtain feedback from the user is activated. For example, feedback is collected in the form of voice commands or touch operations, and this feedback is sent back to the server. The feedback is used to improve the accuracy of future recommendations.

[0123] Example 1

[0124] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0125] Conventional personalized information provision systems have difficulty accurately understanding user behavior and interests in real time. They also have difficulty providing appropriate recommendations in a timely manner. Furthermore, they have been unable to effectively utilize user feedback and reflect it in system improvements.

[0126] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0127] In this invention, the server includes a data receiving / decoding means for receiving and decoding data, an image analysis means for analyzing image data to identify objects that the user is interested in, and an audio analysis means for analyzing audio data to identify what the user is listening to. This makes it possible to identify the user's interests based on the user's gaze information, image analysis results, and audio analysis results, and to dynamically update recommendation information.

[0128] The "gaze detection means" is a device or tool for acquiring information about the user's gaze.

[0129] An "image capture device" is a device or tool for capturing image data of an object that a user is looking at.

[0130] An "audio capture device" is a device or tool used to record a user's speech.

[0131] The "data transmission means" is a device or tool for encrypting the captured data and transmitting it to the server via a communication network.

[0132] "Data receiving and decrypting means" refers to a device or tool for receiving and decrypting encrypted data sent from the glasses-type device.

[0133] "Image analysis means" refers to a device or tool that analyzes received image data and identifies objects in which the user is interested.

[0134] "Audio analysis means" refers to a device or tool that analyzes received audio data and identifies what the user is listening to.

[0135] The "recommendation generating means" is a device or tool for generating recommendation information based on gaze data and analysis results.

[0136] The "information display means" is a device or tool for decoding the received recommendation information and displaying it on a display.

[0137] The "feedback sending means" is a device or tool for obtaining feedback from a user and sending it to the server.

[0138] The "analysis means" is a device or tool for identifying the user's subject of interest based on gaze information, image analysis results, and audio analysis results, and is used to dynamically update recommendation information.

[0139] System overview and equipment configuration

[0140] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[0141] 1. Glasses-type device

[0142] The glasses-type device is equipped with the following functions:

[0143] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[0144] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[0145] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[0146] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[0147] Information display means: Received recommendation information is displayed on a waveguide type display.

[0148] Feedback sending means: Obtains feedback from users and sends it to the server.

[0149] 2. Server

[0150] The server has the following functions:

[0151] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[0152] Image analysis means: Analyzes the received image data and identifies the objects and scenery the user is looking at. Specifically, an image analysis model (e.g., YOLO or ResNet) is used.

[0153] Speech analysis means: Analyzes the received voice data and converts what the user is listening to and saying into text data. A speech recognition model (e.g., Google® Speech-to-Text) is used.

[0154] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[0155] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[0156] Program processing

[0157] Information collection and transmission

[0158] User

[0159] The user wears a glasses-type device and acts accordingly.

[0160] Terminal

[0161] The glasses-type device uses a gaze detection module, camera, and microphone to detect the user's gaze and acquires gaze, image data, and audio data in real time.

[0162] The acquired data is encrypted and sent to the server via the communication network. Specifically, encryption technology (e.g., AES-256) is used.

[0163] Analysis and recommendation generation

[0164] server

[0165] Decrypt the received data.

[0166] Image data is analyzed using an image analysis model (e.g., YOLO or ResNet) to identify objects of interest to the user.

[0167] The voice data is analyzed using a voice recognition model (e.g., Google Speech-to-Text) and converted into text data.

[0168] Based on this data, the user's interests are identified and compared with past behavioral history to generate appropriate recommendation information.

[0169] The generated recommendation information is encrypted and sent to the glasses-type device.

[0170] Recommendation and feedback

[0171] Terminal

[0172] The recommendation information received from the server is decoded and displayed on a waveguide display.

[0173] The feedback from the user is taken, re-encrypted and sent to the server.

[0174] Specific examples

[0175] Example 1: Shopping recommendations

[0176] User

[0177] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[0178] Terminal

[0179] The glasses-type device captures the gaze and images and sends them to a server.

[0180] server

[0181] The received data is analyzed to identify the clothes the user is looking at and recommend accessories and shoes that go well with them.

[0182] Terminal

[0183] Recommendation information is displayed on the screen and presented to the user.

[0184] Example 2: Support for calorie counting

[0185] User

[0186] When a user is looking at a menu at a restaurant, their eyes are drawn to the menu.

[0187] Terminal

[0188] The glasses-type device captures the gaze and images and sends them to a server.

[0189] server

[0190] It analyzes images of menu items, identifies the names of dishes, and retrieves calorie information for the dishes from a database.

[0191] Terminal

[0192] Calorie information is displayed on the display and presented to the user.

[0193] In this way, the present invention is a system that improves the user experience by learning user behavior and interests in real time and providing appropriate recommendation information.

[0194] Prompt Sentence Examples

[0195] The prompt to be input to the generative AI model is as follows:

[0196] "Please explain what kind of recommendations will be provided when a user shows interest in a particular product in the shopping mall."

[0197] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0198] Processing Steps

[0199] Step 1: Collect data

[0200] User

[0201] Users wear the glasses-type device and go about their daily lives.

[0202] Input: User gaze, surrounding images, ambient sounds, and speech.

[0203] Output: Raw gaze data, image data, and audio data.

[0204] Terminal

[0205] The glasses-type device tracks the user's gaze using an eye-gaze detection module.

[0206] An image capture means captures the scenery or object the user is looking at in real time.

[0207] An audio capture method records the ambient sounds around the user and the user's own speech.

[0208] Specifically, when a user is looking at a menu in a cafe, the gaze detection means detects that the user is looking at the menu, and the image capture means takes an image of the menu, while the audio capture means records the sounds of the cafe.

[0209] Step 2: Encrypt and send data

[0210] Terminal

[0211] The collected gaze data, image data, and audio data will be encrypted using encryption technology (e.g., AES-256).

[0212] Input: Raw gaze data, image data, audio data.

[0213] Output: Encrypted gaze data, image data, and audio data.

[0214] The encrypted data is sent to the server at regular intervals (e.g., once per second).

[0215] Specifically, the terminal encrypts the gaze data, menu images, and ambient sounds of the cafe, and then transmits the encrypted data to a server via the Internet.

[0216] Step 3: Receiving and Decrypting Data

[0217] server

[0218] The server receives the encrypted data sent from the eyeglass-type device.

[0219] The received data includes line-of-sight data, image data, and audio data.

[0220] The encrypted data is decrypted using a data receiving and decrypting means.

[0221] Input: Encrypted gaze data, image data, and audio data.

[0222] Output: Decoded gaze data, image data, and audio data.

[0223] Specifically, the server receives the encrypted data sent from the terminal and uses encryption / decryption technology to restore the original gaze data, image data, and audio data.

[0224] Step 4: Data analysis

[0225] server

[0226] The received image data is analyzed using an image analysis model (e.g., YOLO or ResNet) to determine what the user is looking at.

[0227] The voice data is analyzed using a speech recognition model (e.g., Google Speech-to-Text) and the user's speech is converted into text data.

[0228] Input: Decoded image and audio data.

[0229] Output: Image analysis results, audio analysis results.

[0230] Specifically, the server analyzes the menu image to determine that the user is looking at a pizza menu, and analyzes the voice data to determine that the user is saying, "How many calories?"

[0231] Step 5: Generate recommendations

[0232] server

[0233] Identify the user's interests based on gaze data, image analysis results, and audio analysis results.

[0234] This is compared with the user's past behavioral history to generate appropriate recommendation information.

[0235] Input: Gaze data, image analysis results, audio analysis results, past behavioral history.

[0236] Output: Recommendation information.

[0237] Specifically, the server retrieves pizza calorie information from a database based on the user's gaze data and analysis results, and generates recommendation information.

[0238] Step 6: Encrypt and send recommendation information

[0239] server

[0240] The generated recommendation information is encrypted.

[0241] The encrypted recommendation information is sent to the glasses-type device.

[0242] Input: Recommendation information.

[0243] Output: Encrypted recommendation information.

[0244] Specifically, the server encrypts the calorie information of the pizza and sends it to the terminal.

[0245] Step 7: Present recommendations

[0246] Terminal

[0247] The glasses-type device receives the recommendation information transmitted from the server.

[0248] The received recommendation information is decoded and displayed on a waveguide display.

[0249] Input: Encrypted recommendation information.

[0250] Output: Display of decoded recommendation information.

[0251] Specifically, the terminal displays the decrypted calorie information of the pizza on the display.

[0252] Step 8: Getting and sending feedback

[0253] User

[0254] Providing feedback from the user, for example, through voice input or gaze input.

[0255] Terminal

[0256] The glasses-type device collects feedback from the user.

[0257] The obtained feedback is encrypted and sent to the server.

[0258] Input: User feedback.

[0259] Output: Encrypted feedback data.

[0260] Specifically, the user says, "This information was helpful," and the feedback is acquired by the device, encrypted, and sent to the server.

[0261] server

[0262] The server decodes and analyzes the received feedback.

[0263] The results of this analysis will be used to generate the next recommendation information.

[0264] Input: Encrypted feedback data.

[0265] Output: Feedback analysis results.

[0266] Specifically, the server decrypts the encrypted feedback and stores the analysis results for future recommendation generation.

[0267] Prompt Sentence Examples

[0268] The prompt to be input to the generative AI model is as follows:

[0269] "Please explain what kind of recommendations will be provided when a user shows interest in a particular product in the shopping mall."

[0270] (Application example 1)

[0271] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0272] Traditional shopping experiences lack personalized product recommendations based on a user's specific interests and behavioral history, resulting in the significant time and effort required for users to select the right products. Even in brick-and-mortar stores, product recommendations often remain general and do not adequately address individual user needs. Furthermore, there is a lack of a system for instantly incorporating user feedback and improving the next recommendation. This prevents an improved user experience and efficient shopping.

[0273] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0274] In this invention, the server includes a data receiving / decoding means, an image analysis means, a voice analysis means, a product recommendation means, a recommendation generation means, and a feedback transmission means. This makes it possible to make personalized product recommendations based on detailed analysis results of the user's gaze, voice, and behavioral history, and to send feedback from the user to the server and use the generation AI model to improve the accuracy of the next recommendation.

[0275] A "glasses-type device" is a device worn by the user that can collect gaze information, image data, audio data, etc. in real time and send it to a server for analysis.

[0276] The "gaze detection means" is a function installed in the glasses-type device that detects the user's gaze and pupil movement to collect gaze information.

[0277] The "image capture means" is a function that captures image data of the scenery or object that the user is viewing in real time.

[0278] "Audio capture means" is a function that records the user's speech and surrounding environmental sounds.

[0279] The "data transmission means" is a function that encrypts collected data and transmits it to a server via a communication network.

[0280] The "data receiving and decrypting means" is a function that receives and decrypts encrypted data sent from the glasses-type device.

[0281] The "image analysis means" is a function that analyzes the received image data and identifies the object or scene that the user is looking at.

[0282] The "voice analysis means" is a function that analyzes received voice data and identifies the content and statements that the user is listening to.

[0283] The "recommendation generation means" is a function that generates appropriate recommendation information based on the analysis results and the user's past behavioral history.

[0284] The "information display means" is a function that decodes the received recommendation information and displays it on a display.

[0285] The "product recommendation means" is a function that provides personalized product recommendations based on the user's past behavioral history when selecting products in a physical store.

[0286] The "feedback sending means" is a function that obtains feedback information from users and sends prompt sentences to the server using a generative AI model in order to improve the accuracy of the next recommendation based on the feedback information.

[0287] A "generative AI model" is an artificial intelligence model that uses machine learning to learn patterns from data and generate output based on input data.

[0288] A "prompt" is a text sentence entered into a generative AI model to instruct it on a specific task.

[0289] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[0290] System overview and equipment configuration

[0291] 1. Glasses-type device

[0292] The glasses-type device is equipped with the following functions:

[0293] Gaze detection means: A means of detecting the user's gaze and pupil movements and collecting that data.

[0294] Image capture means: A means for capturing image data of the scenery or object the user is viewing in real time.

[0295] Audio capture means: A means of recording the user's surroundings and their own speech.

[0296] Data transmission means: A means for encrypting the captured data and transmitting it to a server via a communication network.

[0297] Information display means: A means for decrypting the received recommendation information and displaying it on a display.

[0298] Feedback sending means: A means for obtaining feedback from users and sending it to the server.

[0299] 2. Server

[0300] The server has the following functions:

[0301] Data reception and decryption means: A means for receiving and decrypting encrypted data sent from the glasses-type device.

[0302] Image analysis means: A means for analyzing the received image data and identifying the object or scene the user is looking at. In this case, image analysis software such as TENSORFLOW (registered trademark) or OpenCV is used.

[0303] Voice analysis means: A means for analyzing received voice data and identifying what the user is listening to and saying. In this case, voice recognition software such as Google Cloud Speech-to-Text API or IBM Watson (registered trademark) is used.

[0304] Recommendation generation method: A method for generating appropriate recommendation information based on analysis results and the user's past behavioral history. In this case, a generative AI model is used.

[0305] Data transmission means: A means for encrypting the generated recommendation information and transmitting it to the glasses-type device.

[0306] Program processing and specific examples

[0307] Information collection and transmission

[0308] The glasses-type device uses a gaze detection module, camera, and microphone to capture the user's gaze, image data, and voice data in real time while the user is walking around the physical store. For example, when a user looks at a pair of sneakers, that information is collected. The collected data is encrypted and sent to a server at regular intervals.

[0309] Analysis and recommendation generation

[0310] The server decrypts the received data and analyzes the image data using an image analysis model to identify the object the user is looking at (e.g., sneakers). It also analyzes the audio data using a voice recognition model to identify what the user is listening to (e.g., product description) and what they are saying (e.g., "Do these sneakers come in other colors?"). Based on the gaze data, image analysis results, and audio analysis results, it identifies the user's interests and compares them with their past behavioral history to generate recommendation information. The recommendation information is encrypted and sent to the glasses-type device.

[0311] Recommendation and feedback

[0312] The glasses-type device receives recommendation information sent from the server, decodes it, and displays it on the screen. If the user is looking at sneakers, it will recommend matching socks and other color variations. Feedback from the user is obtained, and this feedback information is sent back to the server and input into the generative AI model as prompt sentences to improve the accuracy of the next recommendation.

[0313] Examples of prompt statements

[0314] "If a user expresses interest in a blue shirt, offer relevant recommendations, such as matching pants and accessories, or other color variations of the shirt. Take into account the user's previous purchases and style preferences."

[0315] "When a user is listening to information about sneakers, recommend related information, such as socks and other workout apparel to go with the sneakers, or information about current promotions."

[0316] As described above, the present invention is a system that improves the shopping experience in physical stores by learning user behavior and interests in real time and providing appropriate recommendation information.

[0317] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0318] Step 1:

[0319] Collection of information

[0320] The terminal (glasses-type device) acquires the user's gaze information using a gaze detection means, captures image data of the scenery or object the user is looking at in real time using an image capture means, and also collects the user's remarks and surrounding environmental sounds using an audio capture means.

[0321] Input: User gaze data, image data, and audio data.

[0322] Output: Encrypted gaze data, image data, and audio data.

[0323] Step 2:

[0324] Sending data

[0325] The terminal (glasses-type device) encrypts the collected gaze data, image data, and voice data, and transmits them to the server via a communication network using a data transmission means.

[0326] Input: Encrypted gaze data, image data, and audio data.

[0327] Output: The encrypted data sent to the server.

[0328] Step 3:

[0329] Receiving and Decrypting Data

[0330] The server receives the encrypted data sent from the glasses-type device using a data receiving / decrypting means and decrypts it.

[0331] Input: Encrypted gaze data, image data, and audio data.

[0332] Output: Decoded gaze data, image data, and audio data.

[0333] Step 4:

[0334] Image analysis

[0335] The server analyzes the decoded image data using image analysis tools (e.g., TensorFlow or OpenCV) to identify the objects and scenes the user is viewing.

[0336] Input: Decoded image data.

[0337] Output: Analysis results (identification of the object or scene the user is looking at).

[0338] Step 5:

[0339] Audio analysis

[0340] The server then analyzes the decoded audio data using a speech analysis tool (such as Google Cloud Speech-to-Text API or IBM Watson) to determine what the user is hearing and saying.

[0341] Input: Decoded audio data.

[0342] Output: Analysis results (identification of what the user is hearing and saying).

[0343] Step 6:

[0344] Generating recommendation information

[0345] The server identifies the user's interests based on gaze data, image analysis results, and audio analysis results, and generates recommendation information by comparing it with the user's past behavioral history. In this case, a generative AI model is used.

[0346] Input: Gaze data, image analysis results, audio analysis results, past behavioral history.

[0347] Output: Recommendation information (e.g. related products and campaign information).

[0348] Step 7:

[0349] Sending recommendation information

[0350] The server encrypts the generated recommendation information using a data transmission means and transmits it to the glasses-type device.

[0351] Input: Recommendation information.

[0352] Output: Encrypted recommendation information.

[0353] Step 8:

[0354] Displaying Information

[0355] The terminal (glasses-type device) receives the encrypted recommendation information transmitted from the server using the data receiving means and the decrypting means, decrypts it, and displays it to the user using the information displaying means.

[0356] Input: Encrypted recommendation information.

[0357] Output: Recommendation information displayed on the screen.

[0358] Step 9:

[0359] Get feedback

[0360] The terminal (glasses-type device) acquires feedback from the user and transmits the feedback information to the server using the feedback transmission means.

[0361] Input: User feedback information.

[0362] Output: Feedback information sent to the server.

[0363] Step 10:

[0364] Input to generative AI models

[0365] The server creates a prompt sentence for the generative AI model based on the feedback information, which is used to generate the next recommendation information.

[0366] Input: Feedback information.

[0367] Output: The prompt and the generative AI model after retraining.

[0368] Examples of prompt statements

[0369] "If a user expresses interest in a blue shirt, offer relevant recommendations, such as matching pants and accessories, or other color variations of the shirt. Take into account the user's previous purchases and style preferences."

[0370] "When a user is listening to information about sneakers, recommend related information, such as socks and other workout apparel to go with the sneakers, or information about current promotions."

[0371] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0372] System overview and equipment configuration

[0373] This invention is a system that uses a glasses-type device to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that recognizes the user's emotions.

[0374] 1. Glasses-type device

[0375] The glasses-type device is equipped with the following functions:

[0376] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[0377] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[0378] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[0379] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[0380] Information display means: Received recommendation information is displayed on a waveguide type display.

[0381] Feedback sending means: Obtains feedback from users and sends it to the server.

[0382] 2. Server

[0383] The server has the following functions:

[0384] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[0385] Image analysis means: Analyzes the received image data and identifies the object or scene the user is looking at.

[0386] Audio analysis means: Analyzes received audio data to determine what the user is hearing and saying.

[0387] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[0388] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[0389] 3. Emotion Engine

[0390] The emotion engine has the following functions:

[0391] Facial expression analysis means: Analyzes received image data and identifies the user's emotions.

[0392] Voice tone analysis means: Analyzes the tone and rhythm of received voice data to identify the user's emotions.

[0393] Emotion data transmission means: Transmits the identified emotion data to the server.

[0394] Program processing

[0395] Information collection and transmission

[0396] Terminal

[0397] While the user is wearing the glasses-type device and moving around, the gaze detection module, camera, and microphone capture the user's gaze, image data, and audio data in real time.

[0398] As you use it, the emotion engine analyzes your facial expressions and voice tone in real time to generate emotion data, which is also temporarily stored in local memory.

[0399] The acquired data and emotion data are encrypted and sent to the server at regular intervals.

[0400] Analysis and recommendation generation

[0401] server

[0402] The received data is decoded and the image data is analyzed using an image analysis model to determine what the user is looking at.

[0403] The voice data is analyzed using a voice recognition model, and what the user is hearing and saying is converted into text data.

[0404] The emotional data from the emotion engine is analyzed to determine the user's current emotional state.

[0405] Gaze data, image analysis results, audio analysis results, and emotional data are combined to identify the user's interests and emotions.

[0406] By comparing the user's past behavioral history, the system generates appropriate recommendations, especially by prioritizing information that reflects the user's emotional state.

[0407] The generated recommendation information is encrypted and sent to the glasses-type device.

[0408] Recommendation and feedback

[0409] Terminal

[0410] The recommendation information sent from the server is received, decrypted, and displayed on the screen.

[0411] A function to obtain feedback from the user is activated. For example, feedback is collected in the form of voice commands or touch operations, and this information is sent back to the server. The feedback information is used to improve the accuracy of the next recommendation.

[0412] Specific examples

[0413] Example 1: Shopping recommendations

[0414] User

[0415] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[0416] Terminal

[0417] The glasses-type device captures gaze and images, and if the user smiles at the clothes, it sends the data, including their emotions, to the server.

[0418] server

[0419] The system analyzes the received data to identify the clothes the user is looking at, determines that the user has a positive feeling toward the clothes, and recommends accessories and shoes that go well with them.

[0420] Terminal

[0421] Recommendation information is displayed on the screen and presented to the user.

[0422] Example 2: Support for calorie counting

[0423] User

[0424] While looking at a menu at a restaurant, the user glances at the menu and looks a little surprised.

[0425] Terminal

[0426] The glasses-type device captures gaze and images, and sends the user's surprised facial expression as emotional data to the server.

[0427] server

[0428] It analyzes menu images, identifies the names of dishes, and retrieves calorie information from a database, while simultaneously analyzing surprise emotions and generating advice to avoid high-calorie choices.

[0429] Terminal

[0430] Calorie information and advice is displayed on the screen and presented to the user.

[0431] In this way, the present invention is a system that can learn not only a user's behavior and interests but also their emotions in real time, and provide more appropriate and personalized information.

[0432] The processing flow will be explained below.

[0433] Step 1:

[0434] Terminal

[0435] When a user puts on the glasses-type device, the gaze detection module, camera, microphone, and emotion engine are automatically activated. The gaze detection module detects the user's gaze and pupil movements in real time and collects that data. The camera also captures image data of the scenery and objects the user is looking at at a frequency of several frames per second. The microphone records surrounding sounds and what the user is saying.

[0436] Step 2:

[0437] Terminal

[0438] The emotion engine operates, analyzing the user's facial expressions and voice tone in real time to generate emotion data, which is temporarily stored in local memory.

[0439] Step 3:

[0440] Terminal

[0441] The gaze data, image data, voice data, and emotion data are encrypted at regular intervals, and then transmitted to a server via the internet through a data transmission means.

[0442] Step 4:

[0443] server

[0444] The server receives the encrypted data via the Internet, and the received data is decrypted by the data receiving and decrypting means.

[0445] Step 5:

[0446] server

[0447] The decoded image data is analyzed using image analysis means to identify the object or scene the user is looking at.

[0448] Step 6:

[0449] server

[0450] The decoded voice data is analyzed using a voice analysis means, and what the user is hearing or saying is converted into text data.

[0451] Step 7:

[0452] server

[0453] The received emotional data is analyzed by an emotional analysis means to identify the emotional state of the user.

[0454] Step 8:

[0455] server

[0456] Gaze data, image analysis results, audio analysis results, and emotional data are aggregated to identify the user's interests and current emotions.

[0457] Step 9:

[0458] server

[0459] The recommendation generating means generates appropriate recommendation information based on the user's past behavior history and current emotional state, and this recommendation information includes more personalized content.

[0460] Step 10:

[0461] server

[0462] The generated recommendation information is encrypted and transmitted to the glasses-type device via the data transmission means.

[0463] Step 11:

[0464] Terminal

[0465] The glasses-type device decodes the received recommendation information, and the decoded information is displayed on a waveguide display and presented to the user.

[0466] Step 12:

[0467] Terminal

[0468] The user provides some kind of feedback (voice instruction or touch operation) to the presented information. The feedback data is encrypted and sent to the server via the data transmission means.

[0469] Step 13:

[0470] server

[0471] The server receives and analyzes the feedback data, and the results are used to generate the next set of recommendations, helping to improve the accuracy of the system.

[0472] Example 2

[0473] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0474] Conventional information provision systems using glasses-type devices were able to collect data such as the user's gaze, ambient sounds, and speech, but were limited in their ability to provide recommendation information that took the user's emotional state into account. This made it difficult to provide more personalized information based on the user's current emotional state. There was also a need for improvements in the accuracy and real-time nature of the analysis of collected data.

[0475] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0476] In this invention, the server includes a data receiving / decoding means for receiving and decoding the data, an image analysis means for analyzing the image data to identify objects in which the user is interested, a voice analysis means for analyzing the voice data to identify what the user is listening to, an emotion analysis means for analyzing the emotion data to identify the user's emotional state, and a recommendation generation means for integrating the gaze data, image analysis results, voice analysis results, and emotion data to generate recommendation information. This makes it possible to generate appropriate recommendation information that takes into account not only the user's gaze and voice information but also their emotional state.

[0477] A "glasses-type device" is a device that can acquire gaze information, image data, voice data, and emotional data when worn by a user and transmit the data to a server.

[0478] The "gaze detection means" is a mechanism for acquiring the user's gaze information in real time and storing the data internally or transmitting it externally.

[0479] The "image capture means" is a mechanism for capturing and recording image data of the scenery or object that the user is viewing in real time.

[0480] The "audio capture means" is a mechanism for collecting and recording the user's surrounding environmental sounds and the user's speech in real time.

[0481] The "emotion engine" is a mechanism that analyzes the user's facial expressions and tone of voice to identify the user's emotional state in real time.

[0482] The "data transmission means" is a mechanism for encrypting the collected gaze data, image data, voice data, and emotion data and transmitting them to the server.

[0483] The "server" is a central processing unit that receives, decodes, and analyzes data sent from the glasses-type device to generate recommendation information.

[0484] The "data receiving and decrypting means" is a mechanism for receiving and decrypting encrypted data sent from the glasses-type device.

[0485] The "image analysis means" is a mechanism for analyzing received image data and identifying the object or scenery the user is looking at.

[0486] The "voice analysis means" is a mechanism for converting received voice data into text and identifying what the user has said and what they are listening to.

[0487] The "emotion analysis means" is a mechanism for analyzing received emotion data and identifying the user's current emotional state.

[0488] The "recommendation generation means" is a mechanism that generates recommendation information to provide appropriate information to users based on gaze data, image analysis results, audio analysis results, and emotional data.

[0489] The "information display means" is a display device that is installed in the glasses-type device and visually presents the received recommendation information to the user.

[0490] The "feedback transmission means" is a mechanism for obtaining feedback from the user, encrypting it, and transmitting it to the server.

[0491] MODE FOR CARRYING OUT THE INVENTION

[0492] The present invention is a personalized information providing system that uses a glasses-type device and a server. The system configuration and operating principles of the present invention will be described in detail below.

[0493] System Configuration

[0494] This system consists of a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that identifies the user's emotional state.

[0495] Glasses-type device

[0496] The glasses-type device includes the following hardware and software:

[0497] 1. Gaze detection method: Obtaining user gaze information in real time. Specifically, using an infrared sensor or camera to detect gaze and generate gaze tracking data.

[0498] 2. Image capture method: Obtain image data of the scenery or object the user is looking at. Use a high-resolution camera to capture image data in real time.

[0499] 3. Audio capture: Collects the user's surrounding sounds and conversations. Audio data is acquired using the built-in microphone.

[0500] 4. Emotion Engine: Analyzes the user's facial expressions and tone of voice in real time to generate emotion data, for example, using facial expression recognition algorithms and voice analysis algorithms.

[0501] 5. Data transmission method: Collected data and emotion data are encrypted and transmitted to a server via a communication network using AES (Advanced Encryption Standard) via Wi-Fi or mobile networks.

[0502] 6. Information display means: The received recommendation information is displayed on a waveguide type display.

[0503] 7. Feedback sending method: Obtains feedback from the user and sends it to the server. Feedback is collected through voice commands and touch operations.

[0504] server

[0505] The server includes the following software:

[0506] 1. Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[0507] 2. Image analysis: Analyzes the received image data and identifies the object or scene the user is looking at. For image analysis, a deep learning model (such as YOLO or ResNet) is used.

[0508] 3. Voice analysis: Converts received voice data into text and identifies what the user is saying and the surrounding sounds. Google Speech-to-Text API is used for voice recognition.

[0509] 4. Emotion analysis means: Analyzes the received emotion data and identifies the user's current emotional state.

[0510] 5. Recommendation generation method: Integrates gaze data, image analysis results, audio analysis results, and emotion data to generate appropriate recommendations for users, taking into account the user's past behavioral history.

[0511] Specific examples

[0512] For example:

[0513] Example 1: Shopping recommendations

[0514] User

[0515] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[0516] Terminal

[0517] The gaze detection means and image capture means capture images of the user's gaze and clothes, and the emotion engine detects smiling facial expressions. These data are sent to the server.

[0518] server

[0519] The system analyzes the received data, identifies the clothes the user is looking at, determines whether they have a positive emotion, and recommends accessories and shoes that go well with the clothes.

[0520] Terminal

[0521] Recommendation information is displayed on the screen and presented to the user.

[0522] Example 2: Support for calorie counting

[0523] User

[0524] While looking at a menu at a restaurant, the user glances at the menu and looks a little surprised.

[0525] Terminal

[0526] The gaze detection means and image capture means capture the menu image, and the emotion engine detects surprised facial expressions. These data are sent to the server.

[0527] server

[0528] It analyzes menu images, identifies the names of dishes, retrieves calorie information from a database, analyzes surprise emotions, and generates advice to avoid high-calorie choices.

[0529] Terminal

[0530] Calorie information and advice is displayed on the screen and presented to the user.

[0531] Prompt Sentence Examples

[0532] An example of a prompt to be input to a generative AI model is, "Please consider the user's emotional state and generate recommendation information for the object they are looking at."

[0533] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0534] *The following explanation will detail the inputs, outputs, and specific operations at each step of the procedure.

[0535] Step 1: Gather information

[0536] Terminal

[0537] Input: The glasses detect the user's gaze, images of the surrounding scenery and objects, ambient sounds and speech, facial expressions, and tone of voice.

[0538] How it works: The gaze detection means captures the user's gaze data in real time (using infrared sensors and cameras). The image capture means uses a high-resolution camera to capture image data of scenery and objects. The audio capture means uses a built-in microphone to record the surrounding environmental sounds and the user's speech. The emotion engine analyzes facial expressions and voice tone to generate emotion data (using facial expression recognition algorithms and audio analysis algorithms).

[0539] Output: Gaze data, image data, audio data, emotion data.

[0540] Step 2: Sending data

[0541] Terminal

[0542] Input: Gaze data, image data, audio data, and emotion data collected in Step 1.

[0543] How it works: Collected data is encrypted using the AES encryption algorithm and sent to a server via Wi-Fi or mobile network.

[0544] Output: Encrypted gaze data, image data, audio data, and emotion data.

[0545] Step 3: Receiving and Decrypting Data

[0546] server

[0547] Input: Encrypted gaze data, image data, audio data, and emotion data.

[0548] How it works: The server receives encrypted data sent from the device via the HTTP communication protocol and decrypts the data using the same AES encryption algorithm.

[0549] Output: Decoded gaze data, image data, audio data, and emotion data.

[0550] Step 4: Analyze the data

[0551] server

[0552] Input: Decoded gaze data, image data, audio data, and emotion data.

[0553] Operation: Analyzes gaze data to identify the direction and object the user is looking at. The image analysis means analyzes the received image data using deep learning models (YOLO or ResNet) to identify the object the user is looking at. The audio analysis means converts audio data into text using the Google Speech-to-Text API to identify what is being said. The emotion analysis means analyzes emotion data to identify the user's emotional state.

[0554] Output: Gaze analysis results, image analysis results, audio analysis results, emotional state.

[0555] Step 5: Generate recommendations

[0556] server

[0557] Input: Gaze analysis results, image analysis results, audio analysis results, emotional state, and user's past behavior history.

[0558] How it works: Integrates gaze data, image analysis results, audio analysis results, and emotional state to identify the user's interests. References past behavioral history and generates appropriate recommendations based on the user's current interests and emotional state. Creates personalized recommendations using a generative AI model.

[0559] Output: Recommendation information.

[0560] Step 6: Sending recommendations

[0561] server

[0562] Input: Recommendation information.

[0563] How it works: The generated recommendation information is encrypted using the AES encryption algorithm and sent to the glasses-type device via Wi-Fi or mobile network.

[0564] Output: Encrypted recommendation information.

[0565] Step 7: Displaying Recommendations

[0566] Terminal

[0567] Input: Encrypted recommendation information.

[0568] How it works: Decodes received recommendation information and visually displays the information on a waveguide display.

[0569] Output: Recommendation information visually presented to the user.

[0570] Step 8: Collect and send feedback

[0571] Terminal

[0572] Input: User actions and feedback on recommendations.

[0573] How it works: The feedback sending means collects feedback from the user through voice commands and touch operations, encrypts it using the AES encryption algorithm, and sends it to the server.

[0574] Output: Encrypted feedback information.

[0575] As a specific example of the operation of this system, the prompt sentence is "Please consider the user's emotional state and generate recommendation information for the object they are looking at."

[0576] (Application example 2)

[0577] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0578] Today's brick-and-mortar stores are required to respond to diversifying customer needs and quickly and accurately provide each customer with the most appropriate product information. However, conventional methods have difficulty grasping customers' interests and emotional state in real time, limiting the accuracy of personalized recommendation information. Furthermore, customers must consciously search for and obtain product information, which makes it difficult to provide information at the right time to stimulate purchasing motivation.

[0579] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a data receiving / decoding means for receiving and decrypting the data, an image analysis means for analyzing the image data to identify an object in which the user is interested, an audio analysis means for analyzing the audio data to identify what the user is listening to, a recommendation generation means for generating recommendation information based on the gaze data and the analysis results, an emotion analysis means for generating emotion data in the glasses-type device and transmitting it to the server, and a data transmission means for encrypting the recommendation information and transmitting it to the glasses-type device. This makes it possible to provide personalized recommendation information in a timely manner by acquiring and analyzing the user's gaze information and emotion data in real time.

[0580] A "glasses-type device" is an electronic device that can be worn by the user to collect and transmit data such as gaze, images, and audio in real time, and has display functions.

[0581] The term "gaze detection means" refers to a device or method for acquiring information about a user's gaze.

[0582] "Image capture means" means a device or method for capturing image data of an object or scene viewed by a user.

[0583] "Audio capture device" means a device or method for recording a user's speech or surrounding sounds.

[0584] "Data transmission means" refers to a device or method for encrypting the captured data and transmitting it to a server via a communication network.

[0585] "Data receiving and decrypting means" refers to a device or method for receiving data transmitted from the glasses-type device and decrypting encrypted data.

[0586] "Image analysis means" refers to a device or method for analyzing received image data to identify objects of interest to the user.

[0587] "Audio analysis means" means a device or method for analyzing received audio data to determine what the user is hearing.

[0588] "Emotion analysis means" refers to a device or method for analyzing a user's facial expressions and tone of voice to identify the user's emotions.

[0589] The "recommendation generation means" refers to a device or method for generating recommendation information to be provided to a user based on gaze data and analysis results.

[0590] The "information display means" refers to a device or method for decoding the received recommendation information and displaying it on a display.

[0591] "Feedback sending means" refers to a device or method for obtaining feedback from a user and sending it to a server.

[0592] System Overview and Configuration

[0593] The present invention is a system that uses a glasses-type device to learn a user's daily behavior and emotions in real time and provide personalized information. This system includes a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that recognizes the user's emotions.

[0594] 1. Glasses-type device

[0595] Device configuration:

[0596] Gaze detection means: Includes devices that detect the user's gaze and pupil movement and collect that data.

[0597] Image capture means: Includes a camera that captures image data of the scene or object the user is viewing in real time.

[0598] Audio capture means: Includes a microphone that records the ambient sounds around the user and the user's own speech.

[0599] Data transmission means: Includes a communication module for encrypting collected data and transmitting it to a server via a communication network.

[0600] Information display means: includes a display for displaying the received recommendation information on a waveguide type display.

[0601] Feedback sending means: Includes an interface for obtaining feedback from the user and sending it to the server.

[0602] 2. Server

[0603] Server configuration:

[0604] Data receiving and decoding means: Includes a module for receiving and decoding data sent from the glasses-type device.

[0605] Image analysis means: Includes an image analysis model (e.g., TensorFlow, Keras) to analyze the received image data and identify the object the user is looking at.

[0606] Speech analysis means, including speech recognition models (e.g., Google Cloud Speech-to-Text) to analyze received audio data and determine what the user is hearing.

[0607] Emotion analysis tools include emotion recognition models (e.g., Emotion API) that analyze facial expressions and vocal tone to identify a user's emotions.

[0608] Recommendation generation means: Includes a generative AI model that combines gaze data, image analysis results, audio analysis results, and emotion data to generate recommendation information to be provided to users.

[0609] Data transmission means: Includes a communication module for encrypting the generated recommendation information and transmitting it to the glasses-type device.

[0610] 3. Emotion Engine

[0611] Emotion engine configuration:

[0612] Facial expression analysis means: includes a module that analyzes the received image data and identifies the user's emotions.

[0613] Voice tone analysis means: Includes a module that analyzes the tone and rhythm of received voice data to identify the user's emotions.

[0614] Emotion data transmission means: includes a module for transmitting the identified emotion data to the server.

[0615] Component Processing

[0616] Glasses-type device

[0617] While the user is wearing the glasses, the gaze detection module captures the user's gaze, image data, and voice data in real time. The emotion engine analyzes facial expressions and voice tone to generate emotion data. This data is temporarily stored in local memory and then encrypted and sent to the server.

[0618] server

[0619] The server decrypts the received data, uses an image analysis model to identify the object the user is looking at, and converts what the user is hearing into text using a voice recognition model. It then analyzes the user's emotions using an emotion recognition model and generates recommendation information by integrating gaze data, image analysis results, voice analysis results, and emotion data. The generated recommendation information is then encrypted and sent to the glasses-type device.

[0620] Recommendation and feedback

[0621] The glasses-type device decodes the received recommendation information, displays it on the display, and presents it to the user. It also collects feedback from the user and sends it back to the server to use in improving the accuracy of the next recommendation.

[0622] Specific examples

[0623] Use in shopping guide apps

[0624] When a user looks at a particular product in a shopping mall, the camera in the glasses-type device captures image data of the product and sends it along with gaze data to a server. The server identifies the user's interests based on image analysis and gaze data, and also analyzes the user's emotions using an emotion recognition model. Based on the results, it generates optimal recommendation information for the user and sends it to the glasses-type device, providing information to increase purchasing motivation in real time.

[0625] Prompt Sentence Examples

[0626] “Every time a user sees a new piece of clothing, analyze their facial expression and gaze data to see if they tend to be interested in that piece of clothing. Based on the analysis, generate a list of similar products and order them by proximity to past purchase history, rather than random sampling.”

[0627] In this way, the present invention improves the user experience in physical stores by acquiring and analyzing user gaze information and emotional data in real time and providing more personalized recommendation information.

[0628] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0629] Step 1:

[0630] A user puts on a glasses-type device and begins to move around in a physical store such as a shopping mall. The glasses-type device uses gaze detection, image capture, and audio capture means to acquire the user's gaze information, image data of products or objects in front of the user's gaze, surrounding audio, and the user's speech in real time. These data are temporarily stored in local memory. The inputs are the user's gaze, images, and audio, which are acquired by the device. The output is raw data stored in local memory.

[0631] Step 2:

[0632] The terminal generates emotional data. An emotion analysis means installed in the glasses-type device analyzes the acquired image data and audio data and identifies the emotion from the user's facial expression and vocal tone. This emotional data is also temporarily stored in local memory. The input is facial expression and vocal tone, and the output is the generated emotional data.

[0633] Step 3:

[0634] The data collected by the terminal (gaze information, image data, voice data, emotion data) is encrypted and sent to the server via a communication network using a data transmission means. The input is data stored in the local memory, and the output is encrypted data.

[0635] Step 4:

[0636] The server receives the encrypted data sent from the glasses-type device using the data receiving and decrypting means and decrypts it. The input is the encrypted data, and the output is the decrypted raw data.

[0637] Step 5:

[0638] The server uses image analysis means to analyze the decoded image data and identify the object or product the user is looking at. The input is the decoded image data, and the output is the identified object or product information. Specifically, image analysis models such as TensorFlow and Keras are used to identify products and objects.

[0639] Step 6:

[0640] The server uses a speech analysis tool to analyze the decoded audio data and convert what the user is listening to or saying into text data. The input is the decoded audio data, and the output is the text-translated audio information. Specifically, Google Cloud Speech-to-Text is used to convert the audio to text.

[0641] Step 7:

[0642] The server uses emotion analysis means to analyze the received emotion data and identify the user's current emotional state. The input is emotion data, and the output is the identified emotional state of the user. Specifically, the emotion data is analyzed using an Emotion API or similar.

[0643] Step 8:

[0644] The server integrates the gaze data and analysis results (objects, voice, emotions) and uses a recommendation generation means to generate recommendation information to be provided to the user. The input is gaze data, object information, voice information, and emotional state, and the output is recommendation information. Specifically, a generative AI model is used to generate personalized recommendation information while matching it with the user's behavioral history.

[0645] Step 9:

[0646] The server encrypts the generated recommendation information and transmits it to the glasses-type device using a data transmission means. The input is the generated recommendation information, and the output is the encrypted recommendation information.

[0647] Step 10:

[0648] The terminal decrypts the received recommendation information and displays the information on the display of the glasses-type device using the information display means. The input is the encrypted recommendation information, and the output is the information displayed on the display.

[0649] Step 11:

[0650] The terminal obtains feedback from the user and transmits it to the server via the feedback transmission means. The input is the user feedback, and the output is the feedback information transmitted to the server. The feedback is used when generating the next recommendation.

[0651] In this way, a system can be constructed that analyzes user behavior and emotions in real time through each step and provides optimal recommendation information.

[0652] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0653] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0654] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0655] [Second embodiment]

[0656] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0657] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0658] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0659] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0660] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0661] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0662] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0663] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0664] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0665] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0666] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0667] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0668] System overview and equipment configuration

[0669] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[0670] 1. Glasses-type device

[0671] The glasses-type device is equipped with the following functions:

[0672] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[0673] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[0674] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[0675] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[0676] Information display means: Received recommendation information is displayed on a waveguide type display.

[0677] Feedback sending means: Obtains feedback from users and sends it to the server.

[0678] 2. Server

[0679] The server has the following functions:

[0680] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[0681] Image analysis means: Analyzes the received image data and identifies the object or scene the user is looking at.

[0682] Audio analysis means: Analyzes received audio data to determine what the user is hearing and saying.

[0683] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[0684] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[0685] Program processing

[0686] Information collection and transmission

[0687] Terminal

[0688] While the user is wearing the glasses-type device and moving around, the gaze detection module, camera, and microphone capture the user's gaze, image data, and audio data in real time.

[0689] The acquired data is encrypted and sent to the server at regular intervals.

[0690] Analysis and recommendation generation

[0691] server

[0692] The received data is decoded and the image data is analyzed using an image analysis model to determine what the user is looking at.

[0693] The voice data is analyzed using a voice recognition model, and what the user is hearing and saying is converted into text data.

[0694] Identify the user's interests based on gaze data, image analysis results, and audio analysis results.

[0695] This is compared with the user's past behavioral history to generate appropriate recommendation information.

[0696] The generated recommendation information is encrypted and sent to the glasses-type device.

[0697] Recommendation and feedback

[0698] Terminal

[0699] The recommendation information sent from the server is received, decrypted, and displayed on the screen.

[0700] Get feedback from the user and send this feedback information back to the server.

[0701] Specific examples

[0702] Example 1: Shopping recommendations

[0703] User

[0704] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[0705] Terminal

[0706] The glasses-type device captures the gaze and images and sends them to a server.

[0707] server

[0708] The received data is analyzed to identify the clothes the user is looking at and recommend accessories and shoes that go well with them.

[0709] Terminal

[0710] Recommendation information is displayed on the screen and presented to the user.

[0711] Example 2: Support for calorie counting

[0712] User

[0713] When a user is looking at a menu at a restaurant, their eyes are drawn to the menu.

[0714] Terminal

[0715] The glasses-type device captures the gaze and images and sends them to a server.

[0716] server

[0717] It analyzes images of menu items, identifies the names of dishes, and retrieves calorie information for the dishes from a database.

[0718] Terminal

[0719] Calorie information is displayed on the display and presented to the user.

[0720] In this way, the present invention is a system that improves the user experience by learning user behavior and interests in real time and providing appropriate recommendation information.

[0721] The processing flow will be explained below.

[0722] Step 1:

[0723] Terminal

[0724] When the user puts on the glasses, the device automatically activates its gaze detection module, camera, and microphone. It detects and collects data on the user's gaze and pupil movements in real time. It also captures image data of the scenery and objects the user is looking at at a frequency of several frames per second, and records surrounding sounds and the user's speech.

[0725] Step 2:

[0726] Terminal

[0727] The acquired gaze data, image data, and audio data are temporarily stored in local memory, and this data is encrypted for later transmission to the server.

[0728] Step 3:

[0729] Terminal

[0730] The stored data is encrypted using an encryption algorithm such as AES-256, and the encrypted data is sent to the server via a secure communication network.

[0731] Step 4:

[0732] server

[0733] The server receives the encrypted data sent over the Internet and decrypts it using a decryption algorithm such as AES-256.

[0734] Step 5:

[0735] server

[0736] The decoded image data is analyzed using an image recognition model (e.g., YOLO, ResNet), which identifies the object or scene the user is looking at.

[0737] Step 6:

[0738] server

[0739] The decoded audio data is analyzed using a speech recognition model (e.g., DeepSpeech, Wav2Vec), which converts the audio data into text data and identifies what the user is hearing and saying.

[0740] Step 7:

[0741] server

[0742] By combining gaze data, image analysis results, and audio analysis results, the system identifies the user's interests. Based on this data, it compares it with the user's behavioral history and generates appropriate recommendation information.

[0743] Step 8:

[0744] server

[0745] The generated recommendation information is encrypted and retransmitted to the glasses-type device via the data transmission means.

[0746] Step 9:

[0747] Terminal

[0748] The glasses-type device decodes the received recommendation information, and the decoded information is displayed on a waveguide display and presented to the user.

[0749] Step 10:

[0750] Terminal

[0751] A function to obtain feedback from the user is activated. For example, feedback is collected in the form of voice commands or touch operations, and this feedback is sent back to the server. The feedback is used to improve the accuracy of future recommendations.

[0752] Example 1

[0753] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0754] Conventional personalized information provision systems have difficulty accurately understanding user behavior and interests in real time. They also have difficulty providing appropriate recommendations in a timely manner. Furthermore, they have been unable to effectively utilize user feedback and reflect it in system improvements.

[0755] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0756] In this invention, the server includes a data receiving / decoding means for receiving and decoding data, an image analysis means for analyzing image data to identify objects that the user is interested in, and an audio analysis means for analyzing audio data to identify what the user is listening to. This makes it possible to identify the user's interests based on the user's gaze information, image analysis results, and audio analysis results, and to dynamically update recommendation information.

[0757] The "gaze detection means" is a device or tool for acquiring information about the user's gaze.

[0758] An "image capture device" is a device or tool for capturing image data of an object that a user is looking at.

[0759] An "audio capture device" is a device or tool used to record a user's speech.

[0760] The "data transmission means" is a device or tool for encrypting the captured data and transmitting it to the server via a communication network.

[0761] "Data receiving and decrypting means" refers to a device or tool for receiving and decrypting encrypted data sent from the glasses-type device.

[0762] "Image analysis means" refers to a device or tool that analyzes received image data and identifies objects in which the user is interested.

[0763] "Audio analysis means" refers to a device or tool that analyzes received audio data and identifies what the user is listening to.

[0764] The "recommendation generating means" is a device or tool for generating recommendation information based on gaze data and analysis results.

[0765] The "information display means" is a device or tool for decoding the received recommendation information and displaying it on a display.

[0766] The "feedback sending means" is a device or tool for obtaining feedback from a user and sending it to the server.

[0767] The "analysis means" is a device or tool for identifying the user's subject of interest based on gaze information, image analysis results, and audio analysis results, and is used to dynamically update recommendation information.

[0768] System overview and equipment configuration

[0769] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[0770] 1. Glasses-type device

[0771] The glasses-type device is equipped with the following functions:

[0772] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[0773] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[0774] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[0775] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[0776] Information display means: Received recommendation information is displayed on a waveguide type display.

[0777] Feedback sending means: Obtains feedback from users and sends it to the server.

[0778] 2. Server

[0779] The server has the following functions:

[0780] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[0781] Image analysis means: Analyzes the received image data and identifies the objects and scenery the user is looking at. Specifically, an image analysis model (e.g., YOLO or ResNet) is used.

[0782] Speech analysis means: Analyzes the received audio data and converts what the user is hearing and saying into text data. A speech recognition model (e.g., Google Speech-to-Text) is used.

[0783] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[0784] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[0785] Program processing

[0786] Information collection and transmission

[0787] User

[0788] The user wears a glasses-type device and acts accordingly.

[0789] Terminal

[0790] The glasses-type device uses a gaze detection module, camera, and microphone to detect the user's gaze and acquires gaze, image data, and audio data in real time.

[0791] The acquired data is encrypted and sent to the server via the communication network. Specifically, encryption technology (e.g., AES-256) is used.

[0792] Analysis and recommendation generation

[0793] server

[0794] Decrypt the received data.

[0795] Image data is analyzed using an image analysis model (e.g., YOLO or ResNet) to identify objects of interest to the user.

[0796] The voice data is analyzed using a voice recognition model (e.g., Google Speech-to-Text) and converted into text data.

[0797] Based on this data, the user's interests are identified and compared with past behavioral history to generate appropriate recommendation information.

[0798] The generated recommendation information is encrypted and sent to the glasses-type device.

[0799] Recommendation and feedback

[0800] Terminal

[0801] The recommendation information received from the server is decoded and displayed on a waveguide display.

[0802] The feedback from the user is taken, re-encrypted and sent to the server.

[0803] Specific examples

[0804] Example 1: Shopping recommendations

[0805] User

[0806] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[0807] Terminal

[0808] The glasses-type device captures the gaze and images and sends them to a server.

[0809] server

[0810] The received data is analyzed to identify the clothes the user is looking at and recommend accessories and shoes that go well with them.

[0811] Terminal

[0812] Recommendation information is displayed on the screen and presented to the user.

[0813] Example 2: Support for calorie counting

[0814] User

[0815] When a user is looking at a menu at a restaurant, their eyes are drawn to the menu.

[0816] Terminal

[0817] The glasses-type device captures the gaze and images and sends them to a server.

[0818] server

[0819] It analyzes images of menu items, identifies the names of dishes, and retrieves calorie information for the dishes from a database.

[0820] Terminal

[0821] Calorie information is displayed on the display and presented to the user.

[0822] In this way, the present invention is a system that improves the user experience by learning user behavior and interests in real time and providing appropriate recommendation information.

[0823] Prompt Sentence Examples

[0824] The prompt to be input to the generative AI model is as follows:

[0825] "Please explain what kind of recommendations will be provided when a user shows interest in a particular product in the shopping mall."

[0826] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0827] Processing Steps

[0828] Step 1: Collect data

[0829] User

[0830] Users wear the glasses-type device and go about their daily lives.

[0831] Input: User gaze, surrounding images, ambient sounds, and speech.

[0832] Output: Raw gaze data, image data, and audio data.

[0833] Terminal

[0834] The glasses-type device tracks the user's gaze using an eye-gaze detection module.

[0835] An image capture means captures the scenery or object the user is looking at in real time.

[0836] An audio capture method records the ambient sounds around the user and the user's own speech.

[0837] Specifically, when a user is looking at a menu in a cafe, the gaze detection means detects that the user is looking at the menu, and the image capture means takes an image of the menu, while the audio capture means records the sounds of the cafe.

[0838] Step 2: Encrypt and send data

[0839] Terminal

[0840] The collected gaze data, image data, and audio data will be encrypted using encryption technology (e.g., AES-256).

[0841] Input: Raw gaze data, image data, audio data.

[0842] Output: Encrypted gaze data, image data, and audio data.

[0843] The encrypted data is sent to the server at regular intervals (e.g., once per second).

[0844] Specifically, the terminal encrypts the gaze data, menu images, and ambient sounds of the cafe, and then transmits the encrypted data to a server via the Internet.

[0845] Step 3: Receiving and Decrypting Data

[0846] server

[0847] The server receives the encrypted data sent from the eyeglass-type device.

[0848] The received data includes line-of-sight data, image data, and audio data.

[0849] The encrypted data is decrypted using a data receiving and decrypting means.

[0850] Input: Encrypted gaze data, image data, and audio data.

[0851] Output: Decoded gaze data, image data, and audio data.

[0852] Specifically, the server receives the encrypted data sent from the terminal and uses encryption / decryption technology to restore the original gaze data, image data, and audio data.

[0853] Step 4: Data analysis

[0854] server

[0855] The received image data is analyzed using an image analysis model (e.g., YOLO or ResNet) to determine what the user is looking at.

[0856] The voice data is analyzed using a speech recognition model (e.g., Google Speech-to-Text) and the user's speech is converted into text data.

[0857] Input: Decoded image and audio data.

[0858] Output: Image analysis results, audio analysis results.

[0859] Specifically, the server analyzes the menu image to determine that the user is looking at a pizza menu, and analyzes the voice data to determine that the user is saying, "How many calories?"

[0860] Step 5: Generate recommendations

[0861] server

[0862] Identify the user's interests based on gaze data, image analysis results, and audio analysis results.

[0863] This is compared with the user's past behavioral history to generate appropriate recommendation information.

[0864] Input: Gaze data, image analysis results, audio analysis results, past behavioral history.

[0865] Output: Recommendation information.

[0866] Specifically, the server retrieves pizza calorie information from a database based on the user's gaze data and analysis results, and generates recommendation information.

[0867] Step 6: Encrypt and send recommendation information

[0868] server

[0869] The generated recommendation information is encrypted.

[0870] The encrypted recommendation information is sent to the glasses-type device.

[0871] Input: Recommendation information.

[0872] Output: Encrypted recommendation information.

[0873] Specifically, the server encrypts the calorie information of the pizza and sends it to the terminal.

[0874] Step 7: Present recommendations

[0875] Terminal

[0876] The glasses-type device receives the recommendation information transmitted from the server.

[0877] The received recommendation information is decoded and displayed on a waveguide display.

[0878] Input: Encrypted recommendation information.

[0879] Output: Display of decoded recommendation information.

[0880] Specifically, the terminal displays the decrypted calorie information of the pizza on the display.

[0881] Step 8: Getting and sending feedback

[0882] User

[0883] Providing feedback from the user, for example, through voice input or gaze input.

[0884] Terminal

[0885] The glasses-type device collects feedback from the user.

[0886] The obtained feedback is encrypted and sent to the server.

[0887] Input: User feedback.

[0888] Output: Encrypted feedback data.

[0889] Specifically, the user says, "This information was helpful," and the feedback is acquired by the device, encrypted, and sent to the server.

[0890] server

[0891] The server decodes and analyzes the received feedback.

[0892] The results of this analysis will be used to generate the next recommendation information.

[0893] Input: Encrypted feedback data.

[0894] Output: Feedback analysis results.

[0895] Specifically, the server decrypts the encrypted feedback and stores the analysis results for future recommendation generation.

[0896] Prompt Sentence Examples

[0897] The prompt to be input to the generative AI model is as follows:

[0898] "Please explain what kind of recommendations will be provided when a user shows interest in a particular product in the shopping mall."

[0899] (Application example 1)

[0900] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0901] Traditional shopping experiences lack personalized product recommendations based on a user's specific interests and behavioral history, resulting in the significant time and effort required for users to select the right products. Even in brick-and-mortar stores, product recommendations often remain general and do not adequately address individual user needs. Furthermore, there is a lack of a system for instantly incorporating user feedback and improving the next recommendation. This prevents an improved user experience and efficient shopping.

[0902] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0903] In this invention, the server includes a data receiving / decoding means, an image analysis means, a voice analysis means, a product recommendation means, a recommendation generation means, and a feedback transmission means. This makes it possible to make personalized product recommendations based on detailed analysis results of the user's gaze, voice, and behavioral history, and to send feedback from the user to the server and use the generation AI model to improve the accuracy of the next recommendation.

[0904] A "glasses-type device" is a device worn by the user that can collect gaze information, image data, audio data, etc. in real time and send it to a server for analysis.

[0905] The "gaze detection means" is a function installed in the glasses-type device that detects the user's gaze and pupil movement to collect gaze information.

[0906] The "image capture means" is a function that captures image data of the scenery or object that the user is viewing in real time.

[0907] "Audio capture means" is a function that records the user's speech and surrounding environmental sounds.

[0908] The "data transmission means" is a function that encrypts collected data and transmits it to a server via a communication network.

[0909] The "data receiving and decrypting means" is a function that receives and decrypts encrypted data sent from the glasses-type device.

[0910] The "image analysis means" is a function that analyzes the received image data and identifies the object or scene that the user is looking at.

[0911] The "voice analysis means" is a function that analyzes received voice data and identifies the content and statements that the user is listening to.

[0912] The "recommendation generation means" is a function that generates appropriate recommendation information based on the analysis results and the user's past behavioral history.

[0913] The "information display means" is a function that decodes the received recommendation information and displays it on a display.

[0914] The "product recommendation means" is a function that provides personalized product recommendations based on the user's past behavioral history when selecting products in a physical store.

[0915] The "feedback sending means" is a function that obtains feedback information from users and sends prompt sentences to the server using a generative AI model in order to improve the accuracy of the next recommendation based on the feedback information.

[0916] A "generative AI model" is an artificial intelligence model that uses machine learning to learn patterns from data and generate output based on input data.

[0917] A "prompt" is a text sentence entered into a generative AI model to instruct it on a specific task.

[0918] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[0919] System overview and equipment configuration

[0920] 1. Glasses-type device

[0921] The glasses-type device is equipped with the following functions:

[0922] Gaze detection means: A means of detecting the user's gaze and pupil movements and collecting that data.

[0923] Image capture means: A means for capturing image data of the scenery or object the user is viewing in real time.

[0924] Audio capture means: A means of recording the user's surroundings and their own speech.

[0925] Data transmission means: A means for encrypting the captured data and transmitting it to a server via a communication network.

[0926] Information display means: A means for decrypting the received recommendation information and displaying it on a display.

[0927] Feedback sending means: A means for obtaining feedback from users and sending it to the server.

[0928] 2. Server

[0929] The server has the following functions:

[0930] Data reception and decryption means: A means for receiving and decrypting encrypted data sent from the glasses-type device.

[0931] Image analysis means: A means for analyzing the received image data and identifying the object or scene the user is looking at. In this case, image analysis software such as TensorFlow or OpenCV is used.

[0932] Speech analysis means: A means of analyzing the received audio data to identify what the user is listening to and saying, such as using voice recognition software like Google Cloud Speech-to-Text API or IBM Watson.

[0933] Recommendation generation method: A method for generating appropriate recommendation information based on analysis results and the user's past behavioral history. In this case, a generative AI model is used.

[0934] Data transmission means: A means for encrypting the generated recommendation information and transmitting it to the glasses-type device.

[0935] Program processing and specific examples

[0936] Information collection and transmission

[0937] The glasses-type device uses a gaze detection module, camera, and microphone to capture the user's gaze, image data, and voice data in real time while the user is walking around the physical store. For example, when a user looks at a pair of sneakers, that information is collected. The collected data is encrypted and sent to a server at regular intervals.

[0938] Analysis and recommendation generation

[0939] The server decrypts the received data and analyzes the image data using an image analysis model to identify the object the user is looking at (e.g., sneakers). It also analyzes the audio data using a voice recognition model to identify what the user is listening to (e.g., product description) and what they are saying (e.g., "Do these sneakers come in other colors?"). Based on the gaze data, image analysis results, and audio analysis results, it identifies the user's interests and compares them with their past behavioral history to generate recommendation information. The recommendation information is encrypted and sent to the glasses-type device.

[0940] Recommendation and feedback

[0941] The glasses-type device receives recommendation information sent from the server, decodes it, and displays it on the screen. If the user is looking at sneakers, it will recommend matching socks and other color variations. Feedback from the user is obtained, and this feedback information is sent back to the server and input into the generative AI model as prompt sentences to improve the accuracy of the next recommendation.

[0942] Examples of prompt statements

[0943] "If a user expresses interest in a blue shirt, offer relevant recommendations, such as matching pants and accessories, or other color variations of the shirt. Take into account the user's previous purchases and style preferences."

[0944] "When a user is listening to information about sneakers, recommend related information, such as socks and other workout apparel to go with the sneakers, or information about current promotions."

[0945] As described above, the present invention is a system that improves the shopping experience in physical stores by learning user behavior and interests in real time and providing appropriate recommendation information.

[0946] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0947] Step 1:

[0948] Collection of information

[0949] The terminal (glasses-type device) acquires the user's gaze information using a gaze detection means, captures image data of the scenery or object the user is looking at in real time using an image capture means, and also collects the user's remarks and surrounding environmental sounds using an audio capture means.

[0950] Input: User gaze data, image data, and audio data.

[0951] Output: Encrypted gaze data, image data, and audio data.

[0952] Step 2:

[0953] Sending data

[0954] The terminal (glasses-type device) encrypts the collected gaze data, image data, and voice data, and transmits them to the server via a communication network using a data transmission means.

[0955] Input: Encrypted gaze data, image data, and audio data.

[0956] Output: The encrypted data sent to the server.

[0957] Step 3:

[0958] Receiving and Decrypting Data

[0959] The server receives the encrypted data sent from the glasses-type device using a data receiving / decrypting means and decrypts it.

[0960] Input: Encrypted gaze data, image data, and audio data.

[0961] Output: Decoded gaze data, image data, and audio data.

[0962] Step 4:

[0963] Image analysis

[0964] The server analyzes the decoded image data using image analysis tools (e.g., TensorFlow or OpenCV) to identify the objects and scenes the user is viewing.

[0965] Input: Decoded image data.

[0966] Output: Analysis results (identification of the object or scene the user is looking at).

[0967] Step 5:

[0968] Audio analysis

[0969] The server then analyzes the decoded audio data using a speech analysis tool (such as Google Cloud Speech-to-Text API or IBM Watson) to determine what the user is hearing and saying.

[0970] Input: Decoded audio data.

[0971] Output: Analysis results (identification of what the user is hearing and saying).

[0972] Step 6:

[0973] Generating recommendation information

[0974] The server identifies the user's interests based on gaze data, image analysis results, and audio analysis results, and generates recommendation information by comparing it with the user's past behavioral history. In this case, a generative AI model is used.

[0975] Input: Gaze data, image analysis results, audio analysis results, past behavioral history.

[0976] Output: Recommendation information (e.g. related products and campaign information).

[0977] Step 7:

[0978] Sending recommendation information

[0979] The server encrypts the generated recommendation information using a data transmission means and transmits it to the glasses-type device.

[0980] Input: Recommendation information.

[0981] Output: Encrypted recommendation information.

[0982] Step 8:

[0983] Displaying Information

[0984] The terminal (glasses-type device) receives the encrypted recommendation information transmitted from the server using the data receiving means and the decrypting means, decrypts it, and displays it to the user using the information displaying means.

[0985] Input: Encrypted recommendation information.

[0986] Output: Recommendation information displayed on the screen.

[0987] Step 9:

[0988] Get feedback

[0989] The terminal (glasses-type device) acquires feedback from the user and transmits the feedback information to the server using the feedback transmission means.

[0990] Input: User feedback information.

[0991] Output: Feedback information sent to the server.

[0992] Step 10:

[0993] Input to generative AI models

[0994] The server creates a prompt sentence for the generative AI model based on the feedback information, which is used to generate the next recommendation information.

[0995] Input: Feedback information.

[0996] Output: The prompt and the generative AI model after retraining.

[0997] Examples of prompt statements

[0998] "If a user expresses interest in a blue shirt, offer relevant recommendations, such as matching pants and accessories, or other color variations of the shirt. Take into account the user's previous purchases and style preferences."

[0999] "When a user is listening to information about sneakers, recommend related information, such as socks and other workout apparel to go with the sneakers, or information about current promotions."

[1000] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1001] System overview and equipment configuration

[1002] This invention is a system that uses a glasses-type device to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that recognizes the user's emotions.

[1003] 1. Glasses-type device

[1004] The glasses-type device is equipped with the following functions:

[1005] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[1006] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[1007] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[1008] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[1009] Information display means: Received recommendation information is displayed on a waveguide type display.

[1010] Feedback sending means: Obtains feedback from users and sends it to the server.

[1011] 2. Server

[1012] The server has the following functions:

[1013] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[1014] Image analysis means: Analyzes the received image data and identifies the object or scene the user is looking at.

[1015] Audio analysis means: Analyzes received audio data to determine what the user is hearing and saying.

[1016] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[1017] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[1018] 3. Emotion Engine

[1019] The emotion engine has the following functions:

[1020] Facial expression analysis means: Analyzes received image data and identifies the user's emotions.

[1021] Voice tone analysis means: Analyzes the tone and rhythm of received voice data to identify the user's emotions.

[1022] Emotion data transmission means: Transmits the identified emotion data to the server.

[1023] Program processing

[1024] Information collection and transmission

[1025] Terminal

[1026] While the user is wearing the glasses-type device and moving around, the gaze detection module, camera, and microphone capture the user's gaze, image data, and audio data in real time.

[1027] As you use it, the emotion engine analyzes your facial expressions and voice tone in real time to generate emotion data, which is also temporarily stored in local memory.

[1028] The acquired data and emotion data are encrypted and sent to the server at regular intervals.

[1029] Analysis and recommendation generation

[1030] server

[1031] The received data is decoded and the image data is analyzed using an image analysis model to determine what the user is looking at.

[1032] The voice data is analyzed using a voice recognition model, and what the user is hearing and saying is converted into text data.

[1033] The emotional data from the emotion engine is analyzed to determine the user's current emotional state.

[1034] Gaze data, image analysis results, audio analysis results, and emotional data are combined to identify the user's interests and emotions.

[1035] By comparing the user's past behavioral history, the system generates appropriate recommendations, especially by prioritizing information that reflects the user's emotional state.

[1036] The generated recommendation information is encrypted and sent to the glasses-type device.

[1037] Recommendation and feedback

[1038] Terminal

[1039] The recommendation information sent from the server is received, decrypted, and displayed on the screen.

[1040] A function to obtain feedback from the user is activated. For example, feedback is collected in the form of voice commands or touch operations, and this information is sent back to the server. The feedback information is used to improve the accuracy of the next recommendation.

[1041] Specific examples

[1042] Example 1: Shopping recommendations

[1043] User

[1044] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[1045] Terminal

[1046] The glasses-type device captures gaze and images, and if the user smiles at the clothes, it sends the data, including their emotions, to the server.

[1047] server

[1048] The system analyzes the received data to identify the clothes the user is looking at, determines that the user has a positive feeling toward the clothes, and recommends accessories and shoes that go well with them.

[1049] Terminal

[1050] Recommendation information is displayed on the screen and presented to the user.

[1051] Example 2: Support for calorie counting

[1052] User

[1053] While looking at a menu at a restaurant, the user glances at the menu and looks a little surprised.

[1054] Terminal

[1055] The glasses-type device captures gaze and images, and sends the user's surprised facial expression as emotional data to the server.

[1056] server

[1057] It analyzes menu images, identifies the names of dishes, and retrieves calorie information from a database, while simultaneously analyzing surprise emotions and generating advice to avoid high-calorie choices.

[1058] Terminal

[1059] Calorie information and advice is displayed on the screen and presented to the user.

[1060] In this way, the present invention is a system that can learn not only a user's behavior and interests but also their emotions in real time, and provide more appropriate and personalized information.

[1061] The processing flow will be explained below.

[1062] Step 1:

[1063] Terminal

[1064] When a user puts on the glasses-type device, the gaze detection module, camera, microphone, and emotion engine are automatically activated. The gaze detection module detects the user's gaze and pupil movements in real time and collects that data. The camera also captures image data of the scenery and objects the user is looking at at a frequency of several frames per second. The microphone records surrounding sounds and what the user is saying.

[1065] Step 2:

[1066] Terminal

[1067] The emotion engine operates, analyzing the user's facial expressions and voice tone in real time to generate emotion data, which is temporarily stored in local memory.

[1068] Step 3:

[1069] Terminal

[1070] The gaze data, image data, voice data, and emotion data are encrypted at regular intervals, and then transmitted to a server via the internet through a data transmission means.

[1071] Step 4:

[1072] server

[1073] The server receives the encrypted data via the Internet, and the received data is decrypted by the data receiving and decrypting means.

[1074] Step 5:

[1075] server

[1076] The decoded image data is analyzed using image analysis means to identify the object or scene the user is looking at.

[1077] Step 6:

[1078] server

[1079] The decoded voice data is analyzed using a voice analysis means, and what the user is hearing or saying is converted into text data.

[1080] Step 7:

[1081] server

[1082] The received emotional data is analyzed by an emotional analysis means to identify the emotional state of the user.

[1083] Step 8:

[1084] server

[1085] Gaze data, image analysis results, audio analysis results, and emotional data are aggregated to identify the user's interests and current emotions.

[1086] Step 9:

[1087] server

[1088] The recommendation generating means generates appropriate recommendation information based on the user's past behavior history and current emotional state, and this recommendation information includes more personalized content.

[1089] Step 10:

[1090] server

[1091] The generated recommendation information is encrypted and transmitted to the glasses-type device via the data transmission means.

[1092] Step 11:

[1093] Terminal

[1094] The glasses-type device decodes the received recommendation information, and the decoded information is displayed on a waveguide display and presented to the user.

[1095] Step 12:

[1096] Terminal

[1097] The user provides some kind of feedback (voice instruction or touch operation) to the presented information. The feedback data is encrypted and sent to the server via the data transmission means.

[1098] Step 13:

[1099] server

[1100] The server receives and analyzes the feedback data, and the results are used to generate the next set of recommendations, helping to improve the accuracy of the system.

[1101] Example 2

[1102] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1103] Conventional information provision systems using glasses-type devices were able to collect data such as the user's gaze, ambient sounds, and speech, but were limited in their ability to provide recommendation information that took the user's emotional state into account. This made it difficult to provide more personalized information based on the user's current emotional state. There was also a need for improvements in the accuracy and real-time nature of the analysis of collected data.

[1104] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1105] In this invention, the server includes a data receiving / decoding means for receiving and decoding the data, an image analysis means for analyzing the image data to identify objects in which the user is interested, a voice analysis means for analyzing the voice data to identify what the user is listening to, an emotion analysis means for analyzing the emotion data to identify the user's emotional state, and a recommendation generation means for integrating the gaze data, image analysis results, voice analysis results, and emotion data to generate recommendation information. This makes it possible to generate appropriate recommendation information that takes into account not only the user's gaze and voice information but also their emotional state.

[1106] A "glasses-type device" is a device that can acquire gaze information, image data, voice data, and emotional data when worn by a user and transmit the data to a server.

[1107] The "gaze detection means" is a mechanism for acquiring the user's gaze information in real time and storing the data internally or transmitting it externally.

[1108] The "image capture means" is a mechanism for capturing and recording image data of the scenery or object that the user is viewing in real time.

[1109] The "audio capture means" is a mechanism for collecting and recording the user's surrounding environmental sounds and the user's speech in real time.

[1110] The "emotion engine" is a mechanism that analyzes the user's facial expressions and tone of voice to identify the user's emotional state in real time.

[1111] The "data transmission means" is a mechanism for encrypting the collected gaze data, image data, voice data, and emotion data and transmitting them to the server.

[1112] The "server" is a central processing unit that receives, decodes, and analyzes data sent from the glasses-type device to generate recommendation information.

[1113] The "data receiving and decrypting means" is a mechanism for receiving and decrypting encrypted data sent from the glasses-type device.

[1114] The "image analysis means" is a mechanism for analyzing received image data and identifying the object or scenery the user is looking at.

[1115] The "voice analysis means" is a mechanism for converting received voice data into text and identifying what the user has said and what they are listening to.

[1116] The "emotion analysis means" is a mechanism for analyzing received emotion data and identifying the user's current emotional state.

[1117] The "recommendation generation means" is a mechanism that generates recommendation information to provide appropriate information to users based on gaze data, image analysis results, audio analysis results, and emotional data.

[1118] The "information display means" is a display device that is installed in the glasses-type device and visually presents the received recommendation information to the user.

[1119] The "feedback transmission means" is a mechanism for obtaining feedback from the user, encrypting it, and transmitting it to the server.

[1120] MODE FOR CARRYING OUT THE INVENTION

[1121] The present invention is a personalized information providing system that uses a glasses-type device and a server. The system configuration and operating principles of the present invention will be described in detail below.

[1122] System Configuration

[1123] This system consists of a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that identifies the user's emotional state.

[1124] Glasses-type device

[1125] The glasses-type device includes the following hardware and software:

[1126] 1. Gaze detection method: Obtaining user gaze information in real time. Specifically, using an infrared sensor or camera to detect gaze and generate gaze tracking data.

[1127] 2. Image capture method: Obtain image data of the scenery or object the user is looking at. Use a high-resolution camera to capture image data in real time.

[1128] 3. Audio capture: Collects the user's surrounding sounds and conversations. Audio data is acquired using the built-in microphone.

[1129] 4. Emotion Engine: Analyzes the user's facial expressions and tone of voice in real time to generate emotion data, for example, using facial expression recognition algorithms and voice analysis algorithms.

[1130] 5. Data transmission method: Collected data and emotion data are encrypted and transmitted to a server via a communication network using AES (Advanced Encryption Standard) via Wi-Fi or mobile networks.

[1131] 6. Information display means: The received recommendation information is displayed on a waveguide type display.

[1132] 7. Feedback sending method: Obtains feedback from the user and sends it to the server. Feedback is collected through voice commands and touch operations.

[1133] server

[1134] The server includes the following software:

[1135] 1. Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[1136] 2. Image analysis: Analyzes the received image data and identifies the object or scene the user is looking at. For image analysis, a deep learning model (such as YOLO or ResNet) is used.

[1137] 3. Voice analysis: Converts received voice data into text and identifies what the user is saying and the surrounding sounds. Google Speech-to-Text API is used for voice recognition.

[1138] 4. Emotion analysis means: Analyzes the received emotion data and identifies the user's current emotional state.

[1139] 5. Recommendation generation method: Integrates gaze data, image analysis results, audio analysis results, and emotion data to generate appropriate recommendations for users, taking into account the user's past behavioral history.

[1140] Specific examples

[1141] For example:

[1142] Example 1: Shopping recommendations

[1143] User

[1144] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[1145] Terminal

[1146] The gaze detection means and image capture means capture images of the user's gaze and clothes, and the emotion engine detects smiling facial expressions. These data are sent to the server.

[1147] server

[1148] The system analyzes the received data, identifies the clothes the user is looking at, determines whether they have a positive emotion, and recommends accessories and shoes that go well with the clothes.

[1149] Terminal

[1150] Recommendation information is displayed on the screen and presented to the user.

[1151] Example 2: Support for calorie counting

[1152] User

[1153] While looking at a menu at a restaurant, the user glances at the menu and looks a little surprised.

[1154] Terminal

[1155] The gaze detection means and image capture means capture the menu image, and the emotion engine detects surprised facial expressions. These data are sent to the server.

[1156] server

[1157] It analyzes menu images, identifies the names of dishes, retrieves calorie information from a database, analyzes surprise emotions, and generates advice to avoid high-calorie choices.

[1158] Terminal

[1159] Calorie information and advice is displayed on the screen and presented to the user.

[1160] Prompt Sentence Examples

[1161] An example of a prompt to be input to a generative AI model is, "Please consider the user's emotional state and generate recommendation information for the object they are looking at."

[1162] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1163] *The following explanation will detail the inputs, outputs, and specific operations at each step of the procedure.

[1164] Step 1: Gather information

[1165] Terminal

[1166] Input: The glasses detect the user's gaze, images of the surrounding scenery and objects, ambient sounds and speech, facial expressions, and tone of voice.

[1167] How it works: The gaze detection means captures the user's gaze data in real time (using infrared sensors and cameras). The image capture means uses a high-resolution camera to capture image data of scenery and objects. The audio capture means uses a built-in microphone to record the surrounding environmental sounds and the user's speech. The emotion engine analyzes facial expressions and voice tone to generate emotion data (using facial expression recognition algorithms and audio analysis algorithms).

[1168] Output: Gaze data, image data, audio data, emotion data.

[1169] Step 2: Sending data

[1170] Terminal

[1171] Input: Gaze data, image data, audio data, and emotion data collected in Step 1.

[1172] How it works: Collected data is encrypted using the AES encryption algorithm and sent to a server via Wi-Fi or mobile network.

[1173] Output: Encrypted gaze data, image data, audio data, and emotion data.

[1174] Step 3: Receiving and Decrypting Data

[1175] server

[1176] Input: Encrypted gaze data, image data, audio data, and emotion data.

[1177] How it works: The server receives encrypted data sent from the device via the HTTP communication protocol and decrypts the data using the same AES encryption algorithm.

[1178] Output: Decoded gaze data, image data, audio data, and emotion data.

[1179] Step 4: Analyze the data

[1180] server

[1181] Input: Decoded gaze data, image data, audio data, and emotion data.

[1182] Operation: Analyzes gaze data to identify the direction and object the user is looking at. The image analysis means analyzes the received image data using deep learning models (YOLO or ResNet) to identify the object the user is looking at. The audio analysis means converts audio data into text using the Google Speech-to-Text API to identify what is being said. The emotion analysis means analyzes emotion data to identify the user's emotional state.

[1183] Output: Gaze analysis results, image analysis results, audio analysis results, emotional state.

[1184] Step 5: Generate recommendations

[1185] server

[1186] Input: Gaze analysis results, image analysis results, audio analysis results, emotional state, and user's past behavior history.

[1187] How it works: Integrates gaze data, image analysis results, audio analysis results, and emotional state to identify the user's interests. References past behavioral history and generates appropriate recommendations based on the user's current interests and emotional state. Creates personalized recommendations using a generative AI model.

[1188] Output: Recommendation information.

[1189] Step 6: Sending recommendations

[1190] server

[1191] Input: Recommendation information.

[1192] How it works: The generated recommendation information is encrypted using the AES encryption algorithm and sent to the glasses-type device via Wi-Fi or mobile network.

[1193] Output: Encrypted recommendation information.

[1194] Step 7: Displaying Recommendations

[1195] Terminal

[1196] Input: Encrypted recommendation information.

[1197] How it works: Decodes received recommendation information and visually displays the information on a waveguide display.

[1198] Output: Recommendation information visually presented to the user.

[1199] Step 8: Collect and send feedback

[1200] Terminal

[1201] Input: User actions and feedback on recommendations.

[1202] How it works: The feedback sending means collects feedback from the user through voice commands and touch operations, encrypts it using the AES encryption algorithm, and sends it to the server.

[1203] Output: Encrypted feedback information.

[1204] As a specific example of the operation of this system, the prompt sentence is "Please consider the user's emotional state and generate recommendation information for the object they are looking at."

[1205] (Application example 2)

[1206] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1207] Today's brick-and-mortar stores are required to respond to diversifying customer needs and quickly and accurately provide each customer with the most appropriate product information. However, conventional methods have difficulty grasping customers' interests and emotional state in real time, limiting the accuracy of personalized recommendation information. Furthermore, customers must consciously search for and obtain product information, which makes it difficult to provide information at the right time to stimulate purchasing motivation.

[1208] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a data receiving / decoding means for receiving and decrypting the data, an image analysis means for analyzing the image data to identify an object in which the user is interested, an audio analysis means for analyzing the audio data to identify what the user is listening to, a recommendation generation means for generating recommendation information based on the gaze data and the analysis results, an emotion analysis means for generating emotion data in the glasses-type device and transmitting it to the server, and a data transmission means for encrypting the recommendation information and transmitting it to the glasses-type device. This makes it possible to provide personalized recommendation information in a timely manner by acquiring and analyzing the user's gaze information and emotion data in real time.

[1209] A "glasses-type device" is an electronic device that can be worn by the user to collect and transmit data such as gaze, images, and audio in real time, and has display functions.

[1210] The term "gaze detection means" refers to a device or method for acquiring information about a user's gaze.

[1211] "Image capture means" means a device or method for capturing image data of an object or scene viewed by a user.

[1212] "Audio capture device" means a device or method for recording a user's speech or surrounding sounds.

[1213] "Data transmission means" refers to a device or method for encrypting the captured data and transmitting it to a server via a communication network.

[1214] "Data receiving and decrypting means" refers to a device or method for receiving data transmitted from the glasses-type device and decrypting encrypted data.

[1215] "Image analysis means" refers to a device or method for analyzing received image data to identify objects of interest to the user.

[1216] "Audio analysis means" means a device or method for analyzing received audio data to determine what the user is hearing.

[1217] "Emotion analysis means" refers to a device or method for analyzing a user's facial expressions and tone of voice to identify the user's emotions.

[1218] The "recommendation generation means" refers to a device or method for generating recommendation information to be provided to a user based on gaze data and analysis results.

[1219] The "information display means" refers to a device or method for decoding the received recommendation information and displaying it on a display.

[1220] "Feedback sending means" refers to a device or method for obtaining feedback from a user and sending it to a server.

[1221] System Overview and Configuration

[1222] The present invention is a system that uses a glasses-type device to learn a user's daily behavior and emotions in real time and provide personalized information. This system includes a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that recognizes the user's emotions.

[1223] 1. Glasses-type device

[1224] Device configuration:

[1225] Gaze detection means: Includes devices that detect the user's gaze and pupil movement and collect that data.

[1226] Image capture means: Includes a camera that captures image data of the scene or object the user is viewing in real time.

[1227] Audio capture means: Includes a microphone that records the ambient sounds around the user and the user's own speech.

[1228] Data transmission means: Includes a communication module for encrypting collected data and transmitting it to a server via a communication network.

[1229] Information display means: includes a display for displaying the received recommendation information on a waveguide type display.

[1230] Feedback sending means: Includes an interface for obtaining feedback from the user and sending it to the server.

[1231] 2. Server

[1232] Server configuration:

[1233] Data receiving and decoding means: Includes a module for receiving and decoding data sent from the glasses-type device.

[1234] Image analysis means: Includes an image analysis model (e.g., TensorFlow, Keras) to analyze the received image data and identify the object the user is looking at.

[1235] Speech analysis means, including speech recognition models (e.g., Google Cloud Speech-to-Text) to analyze received audio data and determine what the user is hearing.

[1236] Emotion analysis tools include emotion recognition models (e.g., Emotion API) that analyze facial expressions and vocal tone to identify a user's emotions.

[1237] Recommendation generation means: Includes a generative AI model that combines gaze data, image analysis results, audio analysis results, and emotion data to generate recommendation information to be provided to users.

[1238] Data transmission means: Includes a communication module for encrypting the generated recommendation information and transmitting it to the glasses-type device.

[1239] 3. Emotion Engine

[1240] Emotion engine configuration:

[1241] Facial expression analysis means: includes a module that analyzes the received image data and identifies the user's emotions.

[1242] Voice tone analysis means: Includes a module that analyzes the tone and rhythm of received voice data to identify the user's emotions.

[1243] Emotion data transmission means: includes a module for transmitting the identified emotion data to the server.

[1244] Component Processing

[1245] Glasses-type device

[1246] While the user is wearing the glasses, the gaze detection module captures the user's gaze, image data, and voice data in real time. The emotion engine analyzes facial expressions and voice tone to generate emotion data. This data is temporarily stored in local memory and then encrypted and sent to the server.

[1247] server

[1248] The server decrypts the received data, uses an image analysis model to identify the object the user is looking at, and converts what the user is hearing into text using a voice recognition model. It then analyzes the user's emotions using an emotion recognition model and generates recommendation information by integrating gaze data, image analysis results, voice analysis results, and emotion data. The generated recommendation information is then encrypted and sent to the glasses-type device.

[1249] Recommendation and feedback

[1250] The glasses-type device decodes the received recommendation information, displays it on the display, and presents it to the user. It also collects feedback from the user and sends it back to the server to use in improving the accuracy of the next recommendation.

[1251] Specific examples

[1252] Use in shopping guide apps

[1253] When a user looks at a particular product in a shopping mall, the camera in the glasses-type device captures image data of the product and sends it along with gaze data to a server. The server identifies the user's interests based on image analysis and gaze data, and also analyzes the user's emotions using an emotion recognition model. Based on the results, it generates optimal recommendation information for the user and sends it to the glasses-type device, providing information to increase purchasing motivation in real time.

[1254] Prompt Sentence Examples

[1255] “Every time a user sees a new piece of clothing, analyze their facial expression and gaze data to see if they tend to be interested in that piece of clothing. Based on the analysis, generate a list of similar products and order them by proximity to past purchase history, rather than random sampling.”

[1256] In this way, the present invention improves the user experience in physical stores by acquiring and analyzing user gaze information and emotional data in real time and providing more personalized recommendation information.

[1257] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1258] Step 1:

[1259] A user puts on a glasses-type device and begins to move around in a physical store such as a shopping mall. The glasses-type device uses gaze detection, image capture, and audio capture means to acquire the user's gaze information, image data of products or objects in front of the user's gaze, surrounding audio, and the user's speech in real time. These data are temporarily stored in local memory. The inputs are the user's gaze, images, and audio, which are acquired by the device. The output is raw data stored in local memory.

[1260] Step 2:

[1261] The terminal generates emotional data. An emotion analysis means installed in the glasses-type device analyzes the acquired image data and audio data and identifies the emotion from the user's facial expression and vocal tone. This emotional data is also temporarily stored in local memory. The input is facial expression and vocal tone, and the output is the generated emotional data.

[1262] Step 3:

[1263] The data collected by the terminal (gaze information, image data, voice data, emotion data) is encrypted and sent to the server via a communication network using a data transmission means. The input is data stored in the local memory, and the output is encrypted data.

[1264] Step 4:

[1265] The server receives the encrypted data sent from the glasses-type device using the data receiving and decrypting means and decrypts it. The input is the encrypted data, and the output is the decrypted raw data.

[1266] Step 5:

[1267] The server uses image analysis means to analyze the decoded image data and identify the object or product the user is looking at. The input is the decoded image data, and the output is the identified object or product information. Specifically, image analysis models such as TensorFlow and Keras are used to identify products and objects.

[1268] Step 6:

[1269] The server uses a speech analysis tool to analyze the decoded audio data and convert what the user is listening to or saying into text data. The input is the decoded audio data, and the output is the text-translated audio information. Specifically, Google Cloud Speech-to-Text is used to convert the audio to text.

[1270] Step 7:

[1271] The server uses emotion analysis means to analyze the received emotion data and identify the user's current emotional state. The input is emotion data, and the output is the identified emotional state of the user. Specifically, the emotion data is analyzed using an Emotion API or similar.

[1272] Step 8:

[1273] The server integrates the gaze data and analysis results (objects, voice, emotions) and uses a recommendation generation means to generate recommendation information to be provided to the user. The input is gaze data, object information, voice information, and emotional state, and the output is recommendation information. Specifically, a generative AI model is used to generate personalized recommendation information while matching it with the user's behavioral history.

[1274] Step 9:

[1275] The server encrypts the generated recommendation information and transmits it to the glasses-type device using a data transmission means. The input is the generated recommendation information, and the output is the encrypted recommendation information.

[1276] Step 10:

[1277] The terminal decrypts the received recommendation information and displays the information on the display of the glasses-type device using the information display means. The input is the encrypted recommendation information, and the output is the information displayed on the display.

[1278] Step 11:

[1279] The terminal obtains feedback from the user and transmits it to the server via the feedback transmission means. The input is the user feedback, and the output is the feedback information transmitted to the server. The feedback is used when generating the next recommendation.

[1280] In this way, a system can be constructed that analyzes user behavior and emotions in real time through each step and provides optimal recommendation information.

[1281] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1282] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1283] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1284] [Third embodiment]

[1285] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1286] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1287] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1288] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1289] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1290] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1291] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1292] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1293] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1294] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1295] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1296] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1297] System overview and equipment configuration

[1298] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[1299] 1. Glasses-type device

[1300] The glasses-type device is equipped with the following functions:

[1301] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[1302] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[1303] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[1304] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[1305] Information display means: Received recommendation information is displayed on a waveguide type display.

[1306] Feedback sending means: Obtains feedback from users and sends it to the server.

[1307] 2. Server

[1308] The server has the following functions:

[1309] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[1310] Image analysis means: Analyzes the received image data and identifies the object or scene the user is looking at.

[1311] Audio analysis means: Analyzes received audio data to determine what the user is hearing and saying.

[1312] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[1313] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[1314] Program processing

[1315] Information collection and transmission

[1316] Terminal

[1317] While the user is wearing the glasses-type device and moving around, the gaze detection module, camera, and microphone capture the user's gaze, image data, and audio data in real time.

[1318] The acquired data is encrypted and sent to the server at regular intervals.

[1319] Analysis and recommendation generation

[1320] server

[1321] The received data is decoded and the image data is analyzed using an image analysis model to determine what the user is looking at.

[1322] The voice data is analyzed using a voice recognition model, and what the user is hearing and saying is converted into text data.

[1323] Identify the user's interests based on gaze data, image analysis results, and audio analysis results.

[1324] This is compared with the user's past behavioral history to generate appropriate recommendation information.

[1325] The generated recommendation information is encrypted and sent to the glasses-type device.

[1326] Recommendation and feedback

[1327] Terminal

[1328] The recommendation information sent from the server is received, decrypted, and displayed on the screen.

[1329] Get feedback from the user and send this feedback information back to the server.

[1330] Specific examples

[1331] Example 1: Shopping recommendations

[1332] User

[1333] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[1334] Terminal

[1335] The glasses-type device captures the gaze and images and sends them to a server.

[1336] server

[1337] The received data is analyzed to identify the clothes the user is looking at and recommend accessories and shoes that go well with them.

[1338] Terminal

[1339] Recommendation information is displayed on the screen and presented to the user.

[1340] Example 2: Support for calorie counting

[1341] User

[1342] When a user is looking at a menu at a restaurant, their eyes are drawn to the menu.

[1343] Terminal

[1344] The glasses-type device captures the gaze and images and sends them to a server.

[1345] server

[1346] It analyzes images of menu items, identifies the names of dishes, and retrieves calorie information for the dishes from a database.

[1347] Terminal

[1348] Calorie information is displayed on the display and presented to the user.

[1349] In this way, the present invention is a system that improves the user experience by learning user behavior and interests in real time and providing appropriate recommendation information.

[1350] The processing flow will be explained below.

[1351] Step 1:

[1352] Terminal

[1353] When the user puts on the glasses, the device automatically activates its gaze detection module, camera, and microphone. It detects and collects data on the user's gaze and pupil movements in real time. It also captures image data of the scenery and objects the user is looking at at a frequency of several frames per second, and records surrounding sounds and the user's speech.

[1354] Step 2:

[1355] Terminal

[1356] The acquired gaze data, image data, and audio data are temporarily stored in local memory, and this data is encrypted for later transmission to the server.

[1357] Step 3:

[1358] Terminal

[1359] The stored data is encrypted using an encryption algorithm such as AES-256, and the encrypted data is sent to the server via a secure communication network.

[1360] Step 4:

[1361] server

[1362] The server receives the encrypted data sent over the Internet and decrypts it using a decryption algorithm such as AES-256.

[1363] Step 5:

[1364] server

[1365] The decoded image data is analyzed using an image recognition model (e.g., YOLO, ResNet), which identifies the object or scene the user is looking at.

[1366] Step 6:

[1367] server

[1368] The decoded audio data is analyzed using a speech recognition model (e.g., DeepSpeech, Wav2Vec), which converts the audio data into text data and identifies what the user is hearing and saying.

[1369] Step 7:

[1370] server

[1371] By combining gaze data, image analysis results, and audio analysis results, the system identifies the user's interests. Based on this data, it compares it with the user's behavioral history and generates appropriate recommendation information.

[1372] Step 8:

[1373] server

[1374] The generated recommendation information is encrypted and retransmitted to the glasses-type device via the data transmission means.

[1375] Step 9:

[1376] Terminal

[1377] The glasses-type device decodes the received recommendation information, and the decoded information is displayed on a waveguide display and presented to the user.

[1378] Step 10:

[1379] Terminal

[1380] A function to obtain feedback from the user is activated. For example, feedback is collected in the form of voice commands or touch operations, and this feedback is sent back to the server. The feedback is used to improve the accuracy of future recommendations.

[1381] Example 1

[1382] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1383] Conventional personalized information provision systems have difficulty accurately understanding user behavior and interests in real time. They also have difficulty providing appropriate recommendations in a timely manner. Furthermore, they have been unable to effectively utilize user feedback and reflect it in system improvements.

[1384] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1385] In this invention, the server includes a data receiving / decoding means for receiving and decoding data, an image analysis means for analyzing image data to identify objects that the user is interested in, and an audio analysis means for analyzing audio data to identify what the user is listening to. This makes it possible to identify the user's interests based on the user's gaze information, image analysis results, and audio analysis results, and to dynamically update recommendation information.

[1386] The "gaze detection means" is a device or tool for acquiring information about the user's gaze.

[1387] An "image capture device" is a device or tool for capturing image data of an object that a user is looking at.

[1388] An "audio capture device" is a device or tool used to record a user's speech.

[1389] The "data transmission means" is a device or tool for encrypting the captured data and transmitting it to the server via a communication network.

[1390] "Data receiving and decrypting means" refers to a device or tool for receiving and decrypting encrypted data sent from the glasses-type device.

[1391] "Image analysis means" refers to a device or tool that analyzes received image data and identifies objects in which the user is interested.

[1392] "Audio analysis means" refers to a device or tool that analyzes received audio data and identifies what the user is listening to.

[1393] The "recommendation generating means" is a device or tool for generating recommendation information based on gaze data and analysis results.

[1394] The "information display means" is a device or tool for decoding the received recommendation information and displaying it on a display.

[1395] The "feedback sending means" is a device or tool for obtaining feedback from a user and sending it to the server.

[1396] The "analysis means" is a device or tool for identifying the user's subject of interest based on gaze information, image analysis results, and audio analysis results, and is used to dynamically update recommendation information.

[1397] System overview and equipment configuration

[1398] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[1399] 1. Glasses-type device

[1400] The glasses-type device is equipped with the following functions:

[1401] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[1402] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[1403] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[1404] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[1405] Information display means: Received recommendation information is displayed on a waveguide type display.

[1406] Feedback sending means: Obtains feedback from users and sends it to the server.

[1407] 2. Server

[1408] The server has the following functions:

[1409] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[1410] Image analysis means: Analyzes the received image data and identifies the objects and scenery the user is looking at. Specifically, an image analysis model (e.g., YOLO or ResNet) is used.

[1411] Speech analysis means: Analyzes the received audio data and converts what the user is hearing and saying into text data. A speech recognition model (e.g., Google Speech-to-Text) is used.

[1412] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[1413] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[1414] Program processing

[1415] Information collection and transmission

[1416] User

[1417] The user wears a glasses-type device and acts accordingly.

[1418] Terminal

[1419] The glasses-type device uses a gaze detection module, camera, and microphone to detect the user's gaze and acquires gaze, image data, and audio data in real time.

[1420] The acquired data is encrypted and sent to the server via the communication network. Specifically, encryption technology (e.g., AES-256) is used.

[1421] Analysis and recommendation generation

[1422] server

[1423] Decrypt the received data.

[1424] Image data is analyzed using an image analysis model (e.g., YOLO or ResNet) to identify objects of interest to the user.

[1425] The voice data is analyzed using a voice recognition model (e.g., Google Speech-to-Text) and converted into text data.

[1426] Based on this data, the user's interests are identified and compared with past behavioral history to generate appropriate recommendation information.

[1427] The generated recommendation information is encrypted and sent to the glasses-type device.

[1428] Recommendation and feedback

[1429] Terminal

[1430] The recommendation information received from the server is decoded and displayed on a waveguide display.

[1431] The feedback from the user is taken, re-encrypted and sent to the server.

[1432] Specific examples

[1433] Example 1: Shopping recommendations

[1434] User

[1435] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[1436] Terminal

[1437] The glasses-type device captures the gaze and images and sends them to a server.

[1438] server

[1439] The received data is analyzed to identify the clothes the user is looking at and recommend accessories and shoes that go well with them.

[1440] Terminal

[1441] Recommendation information is displayed on the screen and presented to the user.

[1442] Example 2: Support for calorie counting

[1443] User

[1444] When a user is looking at a menu at a restaurant, their eyes are drawn to the menu.

[1445] Terminal

[1446] The glasses-type device captures the gaze and images and sends them to a server.

[1447] server

[1448] It analyzes images of menu items, identifies the names of dishes, and retrieves calorie information for the dishes from a database.

[1449] Terminal

[1450] Calorie information is displayed on the display and presented to the user.

[1451] In this way, the present invention is a system that improves the user experience by learning user behavior and interests in real time and providing appropriate recommendation information.

[1452] Prompt Sentence Examples

[1453] The prompt to be input to the generative AI model is as follows:

[1454] "Please explain what kind of recommendations will be provided when a user shows interest in a particular product in the shopping mall."

[1455] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1456] Processing Steps

[1457] Step 1: Collect data

[1458] User

[1459] Users wear the glasses-type device and go about their daily lives.

[1460] Input: User gaze, surrounding images, ambient sounds, and speech.

[1461] Output: Raw gaze data, image data, and audio data.

[1462] Terminal

[1463] The glasses-type device tracks the user's gaze using an eye-gaze detection module.

[1464] An image capture means captures the scenery or object the user is looking at in real time.

[1465] An audio capture method records the ambient sounds around the user and the user's own speech.

[1466] Specifically, when a user is looking at a menu in a cafe, the gaze detection means detects that the user is looking at the menu, and the image capture means takes an image of the menu, while the audio capture means records the sounds of the cafe.

[1467] Step 2: Encrypt and send data

[1468] Terminal

[1469] The collected gaze data, image data, and audio data will be encrypted using encryption technology (e.g., AES-256).

[1470] Input: Raw gaze data, image data, audio data.

[1471] Output: Encrypted gaze data, image data, and audio data.

[1472] The encrypted data is sent to the server at regular intervals (e.g., once per second).

[1473] Specifically, the terminal encrypts the gaze data, menu images, and ambient sounds of the cafe, and then transmits the encrypted data to a server via the Internet.

[1474] Step 3: Receiving and Decrypting Data

[1475] server

[1476] The server receives the encrypted data sent from the eyeglass-type device.

[1477] The received data includes line-of-sight data, image data, and audio data.

[1478] The encrypted data is decrypted using a data receiving and decrypting means.

[1479] Input: Encrypted gaze data, image data, and audio data.

[1480] Output: Decoded gaze data, image data, and audio data.

[1481] Specifically, the server receives the encrypted data sent from the terminal and uses encryption / decryption technology to restore the original gaze data, image data, and audio data.

[1482] Step 4: Data analysis

[1483] server

[1484] The received image data is analyzed using an image analysis model (e.g., YOLO or ResNet) to determine what the user is looking at.

[1485] The voice data is analyzed using a speech recognition model (e.g., Google Speech-to-Text) and the user's speech is converted into text data.

[1486] Input: Decoded image and audio data.

[1487] Output: Image analysis results, audio analysis results.

[1488] Specifically, the server analyzes the menu image to determine that the user is looking at a pizza menu, and analyzes the voice data to determine that the user is saying, "How many calories?"

[1489] Step 5: Generate recommendations

[1490] server

[1491] Identify the user's interests based on gaze data, image analysis results, and audio analysis results.

[1492] This is compared with the user's past behavioral history to generate appropriate recommendation information.

[1493] Input: Gaze data, image analysis results, audio analysis results, past behavioral history.

[1494] Output: Recommendation information.

[1495] Specifically, the server retrieves pizza calorie information from a database based on the user's gaze data and analysis results, and generates recommendation information.

[1496] Step 6: Encrypt and send recommendation information

[1497] server

[1498] The generated recommendation information is encrypted.

[1499] The encrypted recommendation information is sent to the glasses-type device.

[1500] Input: Recommendation information.

[1501] Output: Encrypted recommendation information.

[1502] Specifically, the server encrypts the calorie information of the pizza and sends it to the terminal.

[1503] Step 7: Present recommendations

[1504] Terminal

[1505] The glasses-type device receives the recommendation information transmitted from the server.

[1506] The received recommendation information is decoded and displayed on a waveguide display.

[1507] Input: Encrypted recommendation information.

[1508] Output: Display of decoded recommendation information.

[1509] Specifically, the terminal displays the decrypted calorie information of the pizza on the display.

[1510] Step 8: Getting and sending feedback

[1511] User

[1512] Providing feedback from the user, for example, through voice input or gaze input.

[1513] Terminal

[1514] The glasses-type device collects feedback from the user.

[1515] The obtained feedback is encrypted and sent to the server.

[1516] Input: User feedback.

[1517] Output: Encrypted feedback data.

[1518] Specifically, the user says, "This information was helpful," and the feedback is acquired by the device, encrypted, and sent to the server.

[1519] server

[1520] The server decodes and analyzes the received feedback.

[1521] The results of this analysis will be used to generate the next recommendation information.

[1522] Input: Encrypted feedback data.

[1523] Output: Feedback analysis results.

[1524] Specifically, the server decrypts the encrypted feedback and stores the analysis results for future recommendation generation.

[1525] Prompt Sentence Examples

[1526] The prompt to be input to the generative AI model is as follows:

[1527] "Please explain what kind of recommendations will be provided when a user shows interest in a particular product in the shopping mall."

[1528] (Application example 1)

[1529] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1530] Traditional shopping experiences lack personalized product recommendations based on a user's specific interests and behavioral history, resulting in the significant time and effort required for users to select the right products. Even in brick-and-mortar stores, product recommendations often remain general and do not adequately address individual user needs. Furthermore, there is a lack of a system for instantly incorporating user feedback and improving the next recommendation. This prevents an improved user experience and efficient shopping.

[1531] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1532] In this invention, the server includes a data receiving / decoding means, an image analysis means, a voice analysis means, a product recommendation means, a recommendation generation means, and a feedback transmission means. This makes it possible to make personalized product recommendations based on detailed analysis results of the user's gaze, voice, and behavioral history, and to send feedback from the user to the server and use the generation AI model to improve the accuracy of the next recommendation.

[1533] A "glasses-type device" is a device worn by the user that can collect gaze information, image data, audio data, etc. in real time and send it to a server for analysis.

[1534] The "gaze detection means" is a function installed in the glasses-type device that detects the user's gaze and pupil movement to collect gaze information.

[1535] The "image capture means" is a function that captures image data of the scenery or object that the user is viewing in real time.

[1536] "Audio capture means" is a function that records the user's speech and surrounding environmental sounds.

[1537] The "data transmission means" is a function that encrypts collected data and transmits it to a server via a communication network.

[1538] The "data receiving and decrypting means" is a function that receives and decrypts encrypted data sent from the glasses-type device.

[1539] The "image analysis means" is a function that analyzes the received image data and identifies the object or scene that the user is looking at.

[1540] The "voice analysis means" is a function that analyzes received voice data and identifies the content and statements that the user is listening to.

[1541] The "recommendation generation means" is a function that generates appropriate recommendation information based on the analysis results and the user's past behavioral history.

[1542] The "information display means" is a function that decodes the received recommendation information and displays it on a display.

[1543] The "product recommendation means" is a function that provides personalized product recommendations based on the user's past behavioral history when selecting products in a physical store.

[1544] The "feedback sending means" is a function that obtains feedback information from users and sends prompt sentences to the server using a generative AI model in order to improve the accuracy of the next recommendation based on the feedback information.

[1545] A "generative AI model" is an artificial intelligence model that uses machine learning to learn patterns from data and generate output based on input data.

[1546] A "prompt" is a text sentence entered into a generative AI model to instruct it on a specific task.

[1547] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[1548] System overview and equipment configuration

[1549] 1. Glasses-type device

[1550] The glasses-type device is equipped with the following functions:

[1551] Gaze detection means: A means of detecting the user's gaze and pupil movements and collecting that data.

[1552] Image capture means: A means for capturing image data of the scenery or object the user is viewing in real time.

[1553] Audio capture means: A means of recording the user's surroundings and their own speech.

[1554] Data transmission means: A means for encrypting the captured data and transmitting it to a server via a communication network.

[1555] Information display means: A means for decrypting the received recommendation information and displaying it on a display.

[1556] Feedback sending means: A means for obtaining feedback from users and sending it to the server.

[1557] 2. Server

[1558] The server has the following functions:

[1559] Data reception and decryption means: A means for receiving and decrypting encrypted data sent from the glasses-type device.

[1560] Image analysis means: A means for analyzing the received image data and identifying the object or scene the user is looking at. In this case, image analysis software such as TensorFlow or OpenCV is used.

[1561] Speech analysis means: A means of analyzing the received audio data to identify what the user is listening to and saying, such as using voice recognition software like Google Cloud Speech-to-Text API or IBM Watson.

[1562] Recommendation generation method: A method for generating appropriate recommendation information based on analysis results and the user's past behavioral history. In this case, a generative AI model is used.

[1563] Data transmission means: A means for encrypting the generated recommendation information and transmitting it to the glasses-type device.

[1564] Program processing and specific examples

[1565] Information collection and transmission

[1566] The glasses-type device uses a gaze detection module, camera, and microphone to capture the user's gaze, image data, and voice data in real time while the user is walking around the physical store. For example, when a user looks at a pair of sneakers, that information is collected. The collected data is encrypted and sent to a server at regular intervals.

[1567] Analysis and recommendation generation

[1568] The server decrypts the received data and analyzes the image data using an image analysis model to identify the object the user is looking at (e.g., sneakers). It also analyzes the audio data using a voice recognition model to identify what the user is listening to (e.g., product description) and what they are saying (e.g., "Do these sneakers come in other colors?"). Based on the gaze data, image analysis results, and audio analysis results, it identifies the user's interests and compares them with their past behavioral history to generate recommendation information. The recommendation information is encrypted and sent to the glasses-type device.

[1569] Recommendation and feedback

[1570] The glasses-type device receives recommendation information sent from the server, decodes it, and displays it on the screen. If the user is looking at sneakers, it will recommend matching socks and other color variations. Feedback from the user is obtained, and this feedback information is sent back to the server and input into the generative AI model as prompt sentences to improve the accuracy of the next recommendation.

[1571] Examples of prompt statements

[1572] "If a user expresses interest in a blue shirt, offer relevant recommendations, such as matching pants and accessories, or other color variations of the shirt. Take into account the user's previous purchases and style preferences."

[1573] "When a user is listening to information about sneakers, recommend related information, such as socks and other workout apparel to go with the sneakers, or information about current promotions."

[1574] As described above, the present invention is a system that improves the shopping experience in physical stores by learning user behavior and interests in real time and providing appropriate recommendation information.

[1575] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1576] Step 1:

[1577] Collection of information

[1578] The terminal (glasses-type device) acquires the user's gaze information using a gaze detection means, captures image data of the scenery or object the user is looking at in real time using an image capture means, and also collects the user's remarks and surrounding environmental sounds using an audio capture means.

[1579] Input: User gaze data, image data, and audio data.

[1580] Output: Encrypted gaze data, image data, and audio data.

[1581] Step 2:

[1582] Sending data

[1583] The terminal (glasses-type device) encrypts the collected gaze data, image data, and voice data, and transmits them to the server via a communication network using a data transmission means.

[1584] Input: Encrypted gaze data, image data, and audio data.

[1585] Output: The encrypted data sent to the server.

[1586] Step 3:

[1587] Receiving and Decrypting Data

[1588] The server receives the encrypted data sent from the glasses-type device using a data receiving / decrypting means and decrypts it.

[1589] Input: Encrypted gaze data, image data, and audio data.

[1590] Output: Decoded gaze data, image data, and audio data.

[1591] Step 4:

[1592] Image analysis

[1593] The server analyzes the decoded image data using image analysis tools (e.g., TensorFlow or OpenCV) to identify the objects and scenes the user is viewing.

[1594] Input: Decoded image data.

[1595] Output: Analysis results (identification of the object or scene the user is looking at).

[1596] Step 5:

[1597] Audio analysis

[1598] The server then analyzes the decoded audio data using a speech analysis tool (such as Google Cloud Speech-to-Text API or IBM Watson) to determine what the user is hearing and saying.

[1599] Input: Decoded audio data.

[1600] Output: Analysis results (identification of what the user is hearing and saying).

[1601] Step 6:

[1602] Generating recommendation information

[1603] The server identifies the user's interests based on gaze data, image analysis results, and audio analysis results, and generates recommendation information by comparing it with the user's past behavioral history. In this case, a generative AI model is used.

[1604] Input: Gaze data, image analysis results, audio analysis results, past behavioral history.

[1605] Output: Recommendation information (e.g. related products and campaign information).

[1606] Step 7:

[1607] Sending recommendation information

[1608] The server encrypts the generated recommendation information using a data transmission means and transmits it to the glasses-type device.

[1609] Input: Recommendation information.

[1610] Output: Encrypted recommendation information.

[1611] Step 8:

[1612] Displaying Information

[1613] The terminal (glasses-type device) receives the encrypted recommendation information transmitted from the server using the data receiving means and the decrypting means, decrypts it, and displays it to the user using the information displaying means.

[1614] Input: Encrypted recommendation information.

[1615] Output: Recommendation information displayed on the screen.

[1616] Step 9:

[1617] Get feedback

[1618] The terminal (glasses-type device) acquires feedback from the user and transmits the feedback information to the server using the feedback transmission means.

[1619] Input: User feedback information.

[1620] Output: Feedback information sent to the server.

[1621] Step 10:

[1622] Input to generative AI models

[1623] The server creates a prompt sentence for the generative AI model based on the feedback information, which is used to generate the next recommendation information.

[1624] Input: Feedback information.

[1625] Output: The prompt and the generative AI model after retraining.

[1626] Examples of prompt statements

[1627] "If a user expresses interest in a blue shirt, offer relevant recommendations, such as matching pants and accessories, or other color variations of the shirt. Take into account the user's previous purchases and style preferences."

[1628] "When a user is listening to information about sneakers, recommend related information, such as socks and other workout apparel to go with the sneakers, or information about current promotions."

[1629] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1630] System overview and equipment configuration

[1631] This invention is a system that uses a glasses-type device to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that recognizes the user's emotions.

[1632] 1. Glasses-type device

[1633] The glasses-type device is equipped with the following functions:

[1634] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[1635] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[1636] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[1637] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[1638] Information display means: Received recommendation information is displayed on a waveguide type display.

[1639] Feedback sending means: Obtains feedback from users and sends it to the server.

[1640] 2. Server

[1641] The server has the following functions:

[1642] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[1643] Image analysis means: Analyzes the received image data and identifies the object or scene the user is looking at.

[1644] Audio analysis means: Analyzes received audio data to determine what the user is hearing and saying.

[1645] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[1646] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[1647] 3. Emotion Engine

[1648] The emotion engine has the following functions:

[1649] Facial expression analysis means: Analyzes received image data and identifies the user's emotions.

[1650] Voice tone analysis means: Analyzes the tone and rhythm of received voice data to identify the user's emotions.

[1651] Emotion data transmission means: Transmits the identified emotion data to the server.

[1652] Program processing

[1653] Information collection and transmission

[1654] Terminal

[1655] While the user is wearing the glasses-type device and moving around, the gaze detection module, camera, and microphone capture the user's gaze, image data, and audio data in real time.

[1656] As you use it, the emotion engine analyzes your facial expressions and voice tone in real time to generate emotion data, which is also temporarily stored in local memory.

[1657] The acquired data and emotion data are encrypted and sent to the server at regular intervals.

[1658] Analysis and recommendation generation

[1659] server

[1660] The received data is decoded and the image data is analyzed using an image analysis model to determine what the user is looking at.

[1661] The voice data is analyzed using a voice recognition model, and what the user is hearing and saying is converted into text data.

[1662] The emotional data from the emotion engine is analyzed to determine the user's current emotional state.

[1663] Gaze data, image analysis results, audio analysis results, and emotional data are combined to identify the user's interests and emotions.

[1664] By comparing the user's past behavioral history, the system generates appropriate recommendations, especially by prioritizing information that reflects the user's emotional state.

[1665] The generated recommendation information is encrypted and sent to the glasses-type device.

[1666] Recommendation and feedback

[1667] Terminal

[1668] The recommendation information sent from the server is received, decrypted, and displayed on the screen.

[1669] A function to obtain feedback from the user is activated. For example, feedback is collected in the form of voice commands or touch operations, and this information is sent back to the server. The feedback information is used to improve the accuracy of the next recommendation.

[1670] Specific examples

[1671] Example 1: Shopping recommendations

[1672] User

[1673] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[1674] Terminal

[1675] The glasses-type device captures gaze and images, and if the user smiles at the clothes, it sends the data, including their emotions, to the server.

[1676] server

[1677] The system analyzes the received data to identify the clothes the user is looking at, determines that the user has a positive feeling toward the clothes, and recommends accessories and shoes that go well with them.

[1678] Terminal

[1679] Recommendation information is displayed on the screen and presented to the user.

[1680] Example 2: Support for calorie counting

[1681] User

[1682] While looking at a menu at a restaurant, the user glances at the menu and looks a little surprised.

[1683] Terminal

[1684] The glasses-type device captures gaze and images, and sends the user's surprised facial expression as emotional data to the server.

[1685] server

[1686] It analyzes menu images, identifies the names of dishes, and retrieves calorie information from a database, while simultaneously analyzing surprise emotions and generating advice to avoid high-calorie choices.

[1687] Terminal

[1688] Calorie information and advice is displayed on the screen and presented to the user.

[1689] In this way, the present invention is a system that can learn not only a user's behavior and interests but also their emotions in real time, and provide more appropriate and personalized information.

[1690] The processing flow will be explained below.

[1691] Step 1:

[1692] Terminal

[1693] When a user puts on the glasses-type device, the gaze detection module, camera, microphone, and emotion engine are automatically activated. The gaze detection module detects the user's gaze and pupil movements in real time and collects that data. The camera also captures image data of the scenery and objects the user is looking at at a frequency of several frames per second. The microphone records surrounding sounds and what the user is saying.

[1694] Step 2:

[1695] Terminal

[1696] The emotion engine operates, analyzing the user's facial expressions and voice tone in real time to generate emotion data, which is temporarily stored in local memory.

[1697] Step 3:

[1698] Terminal

[1699] The gaze data, image data, voice data, and emotion data are encrypted at regular intervals, and then transmitted to a server via the internet through a data transmission means.

[1700] Step 4:

[1701] server

[1702] The server receives the encrypted data via the Internet, and the received data is decrypted by the data receiving and decrypting means.

[1703] Step 5:

[1704] server

[1705] The decoded image data is analyzed using image analysis means to identify the object or scene the user is looking at.

[1706] Step 6:

[1707] server

[1708] The decoded voice data is analyzed using a voice analysis means, and what the user is hearing or saying is converted into text data.

[1709] Step 7:

[1710] server

[1711] The received emotional data is analyzed by an emotional analysis means to identify the emotional state of the user.

[1712] Step 8:

[1713] server

[1714] Gaze data, image analysis results, audio analysis results, and emotional data are aggregated to identify the user's interests and current emotions.

[1715] Step 9:

[1716] server

[1717] The recommendation generating means generates appropriate recommendation information based on the user's past behavior history and current emotional state, and this recommendation information includes more personalized content.

[1718] Step 10:

[1719] server

[1720] The generated recommendation information is encrypted and transmitted to the glasses-type device via the data transmission means.

[1721] Step 11:

[1722] Terminal

[1723] The glasses-type device decodes the received recommendation information, and the decoded information is displayed on a waveguide display and presented to the user.

[1724] Step 12:

[1725] Terminal

[1726] The user provides some kind of feedback (voice instruction or touch operation) to the presented information. The feedback data is encrypted and sent to the server via the data transmission means.

[1727] Step 13:

[1728] server

[1729] The server receives and analyzes the feedback data, and the results are used to generate the next set of recommendations, helping to improve the accuracy of the system.

[1730] Example 2

[1731] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1732] Conventional information provision systems using glasses-type devices were able to collect data such as the user's gaze, ambient sounds, and speech, but were limited in their ability to provide recommendation information that took the user's emotional state into account. This made it difficult to provide more personalized information based on the user's current emotional state. There was also a need for improvements in the accuracy and real-time nature of the analysis of collected data.

[1733] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1734] In this invention, the server includes a data receiving / decoding means for receiving and decoding the data, an image analysis means for analyzing the image data to identify objects in which the user is interested, a voice analysis means for analyzing the voice data to identify what the user is listening to, an emotion analysis means for analyzing the emotion data to identify the user's emotional state, and a recommendation generation means for integrating the gaze data, image analysis results, voice analysis results, and emotion data to generate recommendation information. This makes it possible to generate appropriate recommendation information that takes into account not only the user's gaze and voice information but also their emotional state.

[1735] A "glasses-type device" is a device that can acquire gaze information, image data, voice data, and emotional data when worn by a user and transmit the data to a server.

[1736] The "gaze detection means" is a mechanism for acquiring the user's gaze information in real time and storing the data internally or transmitting it externally.

[1737] The "image capture means" is a mechanism for capturing and recording image data of the scenery or object that the user is viewing in real time.

[1738] The "audio capture means" is a mechanism for collecting and recording the user's surrounding environmental sounds and the user's speech in real time.

[1739] The "emotion engine" is a mechanism that analyzes the user's facial expressions and tone of voice to identify the user's emotional state in real time.

[1740] The "data transmission means" is a mechanism for encrypting the collected gaze data, image data, voice data, and emotion data and transmitting them to the server.

[1741] The "server" is a central processing unit that receives, decodes, and analyzes data sent from the glasses-type device to generate recommendation information.

[1742] The "data receiving and decrypting means" is a mechanism for receiving and decrypting encrypted data sent from the glasses-type device.

[1743] The "image analysis means" is a mechanism for analyzing received image data and identifying the object or scenery the user is looking at.

[1744] The "voice analysis means" is a mechanism for converting received voice data into text and identifying what the user has said and what they are listening to.

[1745] The "emotion analysis means" is a mechanism for analyzing received emotion data and identifying the user's current emotional state.

[1746] The "recommendation generation means" is a mechanism that generates recommendation information to provide appropriate information to users based on gaze data, image analysis results, audio analysis results, and emotional data.

[1747] The "information display means" is a display device that is installed in the glasses-type device and visually presents the received recommendation information to the user.

[1748] The "feedback transmission means" is a mechanism for obtaining feedback from the user, encrypting it, and transmitting it to the server.

[1749] MODE FOR CARRYING OUT THE INVENTION

[1750] The present invention is a personalized information providing system that uses a glasses-type device and a server. The system configuration and operating principles of the present invention will be described in detail below.

[1751] System Configuration

[1752] This system consists of a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that identifies the user's emotional state.

[1753] Glasses-type device

[1754] The glasses-type device includes the following hardware and software:

[1755] 1. Gaze detection method: Obtaining user gaze information in real time. Specifically, using an infrared sensor or camera to detect gaze and generate gaze tracking data.

[1756] 2. Image capture method: Obtain image data of the scenery or object the user is looking at. Use a high-resolution camera to capture image data in real time.

[1757] 3. Audio capture: Collects the user's surrounding sounds and conversations. Audio data is acquired using the built-in microphone.

[1758] 4. Emotion Engine: Analyzes the user's facial expressions and tone of voice in real time to generate emotion data, for example, using facial expression recognition algorithms and voice analysis algorithms.

[1759] 5. Data transmission method: Collected data and emotion data are encrypted and transmitted to a server via a communication network using AES (Advanced Encryption Standard) via Wi-Fi or mobile networks.

[1760] 6. Information display means: The received recommendation information is displayed on a waveguide type display.

[1761] 7. Feedback sending method: Obtains feedback from the user and sends it to the server. Feedback is collected through voice commands and touch operations.

[1762] server

[1763] The server includes the following software:

[1764] 1. Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[1765] 2. Image analysis: Analyzes the received image data and identifies the object or scene the user is looking at. For image analysis, a deep learning model (such as YOLO or ResNet) is used.

[1766] 3. Voice analysis: Converts received voice data into text and identifies what the user is saying and the surrounding sounds. Google Speech-to-Text API is used for voice recognition.

[1767] 4. Emotion analysis means: Analyzes the received emotion data and identifies the user's current emotional state.

[1768] 5. Recommendation generation method: Integrates gaze data, image analysis results, audio analysis results, and emotion data to generate appropriate recommendations for users, taking into account the user's past behavioral history.

[1769] Specific examples

[1770] For example:

[1771] Example 1: Shopping recommendations

[1772] User

[1773] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[1774] Terminal

[1775] The gaze detection means and image capture means capture images of the user's gaze and clothes, and the emotion engine detects smiling facial expressions. These data are sent to the server.

[1776] server

[1777] The system analyzes the received data, identifies the clothes the user is looking at, determines whether they have a positive emotion, and recommends accessories and shoes that go well with the clothes.

[1778] Terminal

[1779] Recommendation information is displayed on the screen and presented to the user.

[1780] Example 2: Support for calorie counting

[1781] User

[1782] While looking at a menu at a restaurant, the user glances at the menu and looks a little surprised.

[1783] Terminal

[1784] The gaze detection means and image capture means capture the menu image, and the emotion engine detects surprised facial expressions. These data are sent to the server.

[1785] server

[1786] It analyzes menu images, identifies the names of dishes, retrieves calorie information from a database, analyzes surprise emotions, and generates advice to avoid high-calorie choices.

[1787] Terminal

[1788] Calorie information and advice is displayed on the screen and presented to the user.

[1789] Prompt Sentence Examples

[1790] An example of a prompt to be input to a generative AI model is, "Please consider the user's emotional state and generate recommendation information for the object they are looking at."

[1791] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1792] *The following explanation will detail the inputs, outputs, and specific operations at each step of the procedure.

[1793] Step 1: Gather information

[1794] Terminal

[1795] Input: The glasses detect the user's gaze, images of the surrounding scenery and objects, ambient sounds and speech, facial expressions, and tone of voice.

[1796] How it works: The gaze detection means captures the user's gaze data in real time (using infrared sensors and cameras). The image capture means uses a high-resolution camera to capture image data of scenery and objects. The audio capture means uses a built-in microphone to record the surrounding environmental sounds and the user's speech. The emotion engine analyzes facial expressions and voice tone to generate emotion data (using facial expression recognition algorithms and audio analysis algorithms).

[1797] Output: Gaze data, image data, audio data, emotion data.

[1798] Step 2: Sending data

[1799] Terminal

[1800] Input: Gaze data, image data, audio data, and emotion data collected in Step 1.

[1801] How it works: Collected data is encrypted using the AES encryption algorithm and sent to a server via Wi-Fi or mobile network.

[1802] Output: Encrypted gaze data, image data, audio data, and emotion data.

[1803] Step 3: Receiving and Decrypting Data

[1804] server

[1805] Input: Encrypted gaze data, image data, audio data, and emotion data.

[1806] How it works: The server receives encrypted data sent from the device via the HTTP communication protocol and decrypts the data using the same AES encryption algorithm.

[1807] Output: Decoded gaze data, image data, audio data, and emotion data.

[1808] Step 4: Analyze the data

[1809] server

[1810] Input: Decoded gaze data, image data, audio data, and emotion data.

[1811] Operation: Analyzes gaze data to identify the direction and object the user is looking at. The image analysis means analyzes the received image data using deep learning models (YOLO or ResNet) to identify the object the user is looking at. The audio analysis means converts audio data into text using the Google Speech-to-Text API to identify what is being said. The emotion analysis means analyzes emotion data to identify the user's emotional state.

[1812] Output: Gaze analysis results, image analysis results, audio analysis results, emotional state.

[1813] Step 5: Generate recommendations

[1814] server

[1815] Input: Gaze analysis results, image analysis results, audio analysis results, emotional state, and user's past behavior history.

[1816] How it works: Integrates gaze data, image analysis results, audio analysis results, and emotional state to identify the user's interests. References past behavioral history and generates appropriate recommendations based on the user's current interests and emotional state. Creates personalized recommendations using a generative AI model.

[1817] Output: Recommendation information.

[1818] Step 6: Sending recommendations

[1819] server

[1820] Input: Recommendation information.

[1821] How it works: The generated recommendation information is encrypted using the AES encryption algorithm and sent to the glasses-type device via Wi-Fi or mobile network.

[1822] Output: Encrypted recommendation information.

[1823] Step 7: Displaying Recommendations

[1824] Terminal

[1825] Input: Encrypted recommendation information.

[1826] How it works: Decodes received recommendation information and visually displays the information on a waveguide display.

[1827] Output: Recommendation information visually presented to the user.

[1828] Step 8: Collect and send feedback

[1829] Terminal

[1830] Input: User actions and feedback on recommendations.

[1831] How it works: The feedback sending means collects feedback from the user through voice commands and touch operations, encrypts it using the AES encryption algorithm, and sends it to the server.

[1832] Output: Encrypted feedback information.

[1833] As a specific example of the operation of this system, the prompt sentence is "Please consider the user's emotional state and generate recommendation information for the object they are looking at."

[1834] (Application example 2)

[1835] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1836] Today's brick-and-mortar stores are required to respond to diversifying customer needs and quickly and accurately provide each customer with the most appropriate product information. However, conventional methods have difficulty grasping customers' interests and emotional state in real time, limiting the accuracy of personalized recommendation information. Furthermore, customers must consciously search for and obtain product information, which makes it difficult to provide information at the right time to stimulate purchasing motivation.

[1837] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a data receiving / decoding means for receiving and decrypting the data, an image analysis means for analyzing the image data to identify an object in which the user is interested, an audio analysis means for analyzing the audio data to identify what the user is listening to, a recommendation generation means for generating recommendation information based on the gaze data and the analysis results, an emotion analysis means for generating emotion data in the glasses-type device and transmitting it to the server, and a data transmission means for encrypting the recommendation information and transmitting it to the glasses-type device. This makes it possible to provide personalized recommendation information in a timely manner by acquiring and analyzing the user's gaze information and emotion data in real time.

[1838] A "glasses-type device" is an electronic device that can be worn by the user to collect and transmit data such as gaze, images, and audio in real time, and has display functions.

[1839] The term "gaze detection means" refers to a device or method for acquiring information about a user's gaze.

[1840] "Image capture means" means a device or method for capturing image data of an object or scene viewed by a user.

[1841] "Audio capture device" means a device or method for recording a user's speech or surrounding sounds.

[1842] "Data transmission means" refers to a device or method for encrypting the captured data and transmitting it to a server via a communication network.

[1843] "Data receiving and decrypting means" refers to a device or method for receiving data transmitted from the glasses-type device and decrypting encrypted data.

[1844] "Image analysis means" refers to a device or method for analyzing received image data to identify objects of interest to the user.

[1845] "Audio analysis means" means a device or method for analyzing received audio data to determine what the user is hearing.

[1846] "Emotion analysis means" refers to a device or method for analyzing a user's facial expressions and tone of voice to identify the user's emotions.

[1847] The "recommendation generation means" refers to a device or method for generating recommendation information to be provided to a user based on gaze data and analysis results.

[1848] The "information display means" refers to a device or method for decoding the received recommendation information and displaying it on a display.

[1849] "Feedback sending means" refers to a device or method for obtaining feedback from a user and sending it to a server.

[1850] System Overview and Configuration

[1851] The present invention is a system that uses a glasses-type device to learn a user's daily behavior and emotions in real time and provide personalized information. This system includes a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that recognizes the user's emotions.

[1852] 1. Glasses-type device

[1853] Device configuration:

[1854] Gaze detection means: Includes devices that detect the user's gaze and pupil movement and collect that data.

[1855] Image capture means: Includes a camera that captures image data of the scene or object the user is viewing in real time.

[1856] Audio capture means: Includes a microphone that records the ambient sounds around the user and the user's own speech.

[1857] Data transmission means: Includes a communication module for encrypting collected data and transmitting it to a server via a communication network.

[1858] Information display means: includes a display for displaying the received recommendation information on a waveguide type display.

[1859] Feedback sending means: Includes an interface for obtaining feedback from the user and sending it to the server.

[1860] 2. Server

[1861] Server configuration:

[1862] Data receiving and decoding means: Includes a module for receiving and decoding data sent from the glasses-type device.

[1863] Image analysis means: Includes an image analysis model (e.g., TensorFlow, Keras) to analyze the received image data and identify the object the user is looking at.

[1864] Speech analysis means, including speech recognition models (e.g., Google Cloud Speech-to-Text) to analyze received audio data and determine what the user is hearing.

[1865] Emotion analysis tools include emotion recognition models (e.g., Emotion API) that analyze facial expressions and vocal tone to identify a user's emotions.

[1866] Recommendation generation means: Includes a generative AI model that combines gaze data, image analysis results, audio analysis results, and emotion data to generate recommendation information to be provided to users.

[1867] Data transmission means: Includes a communication module for encrypting the generated recommendation information and transmitting it to the glasses-type device.

[1868] 3. Emotion Engine

[1869] Emotion engine configuration:

[1870] Facial expression analysis means: includes a module that analyzes the received image data and identifies the user's emotions.

[1871] Voice tone analysis means: Includes a module that analyzes the tone and rhythm of received voice data to identify the user's emotions.

[1872] Emotion data transmission means: includes a module for transmitting the identified emotion data to the server.

[1873] Component Processing

[1874] Glasses-type device

[1875] While the user is wearing the glasses, the gaze detection module captures the user's gaze, image data, and voice data in real time. The emotion engine analyzes facial expressions and voice tone to generate emotion data. This data is temporarily stored in local memory and then encrypted and sent to the server.

[1876] server

[1877] The server decrypts the received data, uses an image analysis model to identify the object the user is looking at, and converts what the user is hearing into text using a voice recognition model. It then analyzes the user's emotions using an emotion recognition model and generates recommendation information by integrating gaze data, image analysis results, voice analysis results, and emotion data. The generated recommendation information is then encrypted and sent to the glasses-type device.

[1878] Recommendation and feedback

[1879] The glasses-type device decodes the received recommendation information, displays it on the display, and presents it to the user. It also collects feedback from the user and sends it back to the server to use in improving the accuracy of the next recommendation.

[1880] Specific examples

[1881] Use in shopping guide apps

[1882] When a user looks at a particular product in a shopping mall, the camera in the glasses-type device captures image data of the product and sends it along with gaze data to a server. The server identifies the user's interests based on image analysis and gaze data, and also analyzes the user's emotions using an emotion recognition model. Based on the results, it generates optimal recommendation information for the user and sends it to the glasses-type device, providing information to increase purchasing motivation in real time.

[1883] Prompt Sentence Examples

[1884] “Every time a user sees a new piece of clothing, analyze their facial expression and gaze data to see if they tend to be interested in that piece of clothing. Based on the analysis, generate a list of similar products and order them by proximity to past purchase history, rather than random sampling.”

[1885] In this way, the present invention improves the user experience in physical stores by acquiring and analyzing user gaze information and emotional data in real time and providing more personalized recommendation information.

[1886] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1887] Step 1:

[1888] A user puts on a glasses-type device and begins to move around in a physical store such as a shopping mall. The glasses-type device uses gaze detection, image capture, and audio capture means to acquire the user's gaze information, image data of products or objects in front of the user's gaze, surrounding audio, and the user's speech in real time. These data are temporarily stored in local memory. The inputs are the user's gaze, images, and audio, which are acquired by the device. The output is raw data stored in local memory.

[1889] Step 2:

[1890] The terminal generates emotional data. An emotion analysis means installed in the glasses-type device analyzes the acquired image data and audio data and identifies the emotion from the user's facial expression and vocal tone. This emotional data is also temporarily stored in local memory. The input is facial expression and vocal tone, and the output is the generated emotional data.

[1891] Step 3:

[1892] The data collected by the terminal (gaze information, image data, voice data, emotion data) is encrypted and sent to the server via a communication network using a data transmission means. The input is data stored in the local memory, and the output is encrypted data.

[1893] Step 4:

[1894] The server receives the encrypted data sent from the glasses-type device using the data receiving and decrypting means and decrypts it. The input is the encrypted data, and the output is the decrypted raw data.

[1895] Step 5:

[1896] The server uses image analysis means to analyze the decoded image data and identify the object or product the user is looking at. The input is the decoded image data, and the output is the identified object or product information. Specifically, image analysis models such as TensorFlow and Keras are used to identify products and objects.

[1897] Step 6:

[1898] The server uses a speech analysis tool to analyze the decoded audio data and convert what the user is listening to or saying into text data. The input is the decoded audio data, and the output is the text-translated audio information. Specifically, Google Cloud Speech-to-Text is used to convert the audio to text.

[1899] Step 7:

[1900] The server uses emotion analysis means to analyze the received emotion data and identify the user's current emotional state. The input is emotion data, and the output is the identified emotional state of the user. Specifically, the emotion data is analyzed using an Emotion API or similar.

[1901] Step 8:

[1902] The server integrates the gaze data and analysis results (objects, voice, emotions) and uses a recommendation generation means to generate recommendation information to be provided to the user. The input is gaze data, object information, voice information, and emotional state, and the output is recommendation information. Specifically, a generative AI model is used to generate personalized recommendation information while matching it with the user's behavioral history.

[1903] Step 9:

[1904] The server encrypts the generated recommendation information and transmits it to the glasses-type device using a data transmission means. The input is the generated recommendation information, and the output is the encrypted recommendation information.

[1905] Step 10:

[1906] The terminal decrypts the received recommendation information and displays the information on the display of the glasses-type device using the information display means. The input is the encrypted recommendation information, and the output is the information displayed on the display.

[1907] Step 11:

[1908] The terminal obtains feedback from the user and transmits it to the server via the feedback transmission means. The input is the user feedback, and the output is the feedback information transmitted to the server. The feedback is used when generating the next recommendation.

[1909] In this way, a system can be constructed that analyzes user behavior and emotions in real time through each step and provides optimal recommendation information.

[1910] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1911] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1912] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1913] [Fourth embodiment]

[1914] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1915] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1916] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1917] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1918] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1919] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1920] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1921] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1922] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1923] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1924] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1925] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1926] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1927] System overview and equipment configuration

[1928] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[1929] 1. Glasses-type device

[1930] The glasses-type device is equipped with the following functions:

[1931] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[1932] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[1933] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[1934] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[1935] Information display means: Received recommendation information is displayed on a waveguide type display.

[1936] Feedback sending means: Obtains feedback from users and sends it to the server.

[1937] 2. Server

[1938] The server has the following functions:

[1939] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[1940] Image analysis means: Analyzes the received image data and identifies the object or scene the user is looking at.

[1941] Audio analysis means: Analyzes received audio data to determine what the user is hearing and saying.

[1942] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[1943] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[1944] Program processing

[1945] Information collection and transmission

[1946] Terminal

[1947] While the user is wearing the glasses-type device and moving around, the gaze detection module, camera, and microphone capture the user's gaze, image data, and audio data in real time.

[1948] The acquired data is encrypted and sent to the server at regular intervals.

[1949] Analysis and recommendation generation

[1950] server

[1951] The received data is decoded and the image data is analyzed using an image analysis model to determine what the user is looking at.

[1952] The voice data is analyzed using a voice recognition model, and what the user is hearing and saying is converted into text data.

[1953] Identify the user's interests based on gaze data, image analysis results, and audio analysis results.

[1954] This is compared with the user's past behavioral history to generate appropriate recommendation information.

[1955] The generated recommendation information is encrypted and sent to the glasses-type device.

[1956] Recommendation and feedback

[1957] Terminal

[1958] The recommendation information sent from the server is received, decrypted, and displayed on the screen.

[1959] Get feedback from the user and send this feedback information back to the server.

[1960] Specific examples

[1961] Example 1: Shopping recommendations

[1962] User

[1963] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[1964] Terminal

[1965] The glasses-type device captures the gaze and images and sends them to a server.

[1966] server

[1967] The received data is analyzed to identify the clothes the user is looking at and recommend accessories and shoes that go well with them.

[1968] Terminal

[1969] Recommendation information is displayed on the screen and presented to the user.

[1970] Example 2: Support for calorie counting

[1971] User

[1972] When a user is looking at a menu at a restaurant, their eyes are drawn to the menu.

[1973] Terminal

[1974] The glasses-type device captures the gaze and images and sends them to a server.

[1975] server

[1976] It analyzes images of menu items, identifies the names of dishes, and retrieves calorie information for the dishes from a database.

[1977] Terminal

[1978] Calorie information is displayed on the display and presented to the user.

[1979] In this way, the present invention is a system that improves the user experience by learning user behavior and interests in real time and providing appropriate recommendation information.

[1980] The processing flow will be explained below.

[1981] Step 1:

[1982] Terminal

[1983] When the user puts on the glasses, the device automatically activates its gaze detection module, camera, and microphone. It detects and collects data on the user's gaze and pupil movements in real time. It also captures image data of the scenery and objects the user is looking at at a frequency of several frames per second, and records surrounding sounds and the user's speech.

[1984] Step 2:

[1985] Terminal

[1986] The acquired gaze data, image data, and audio data are temporarily stored in local memory, and this data is encrypted for later transmission to the server.

[1987] Step 3:

[1988] Terminal

[1989] The stored data is encrypted using an encryption algorithm such as AES-256, and the encrypted data is sent to the server via a secure communication network.

[1990] Step 4:

[1991] server

[1992] The server receives the encrypted data sent over the Internet and decrypts it using a decryption algorithm such as AES-256.

[1993] Step 5:

[1994] server

[1995] The decoded image data is analyzed using an image recognition model (e.g., YOLO, ResNet), which identifies the object or scene the user is looking at.

[1996] Step 6:

[1997] server

[1998] The decoded audio data is analyzed using a speech recognition model (e.g., DeepSpeech, Wav2Vec), which converts the audio data into text data and identifies what the user is hearing and saying.

[1999] Step 7:

[2000] server

[2001] By combining gaze data, image analysis results, and audio analysis results, the system identifies the user's interests. Based on this data, it compares it with the user's behavioral history and generates appropriate recommendation information.

[2002] Step 8:

[2003] server

[2004] The generated recommendation information is encrypted and retransmitted to the glasses-type device via the data transmission means.

[2005] Step 9:

[2006] Terminal

[2007] The glasses-type device decodes the received recommendation information, and the decoded information is displayed on a waveguide display and presented to the user.

[2008] Step 10:

[2009] Terminal

[2010] A function to obtain feedback from the user is activated. For example, feedback is collected in the form of voice commands or touch operations, and this feedback is sent back to the server. The feedback is used to improve the accuracy of future recommendations.

[2011] Example 1

[2012] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2013] Conventional personalized information provision systems have difficulty accurately understanding user behavior and interests in real time. They also have difficulty providing appropriate recommendations in a timely manner. Furthermore, they have been unable to effectively utilize user feedback and reflect it in system improvements.

[2014] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[2015] In this invention, the server includes a data receiving / decoding means for receiving and decoding data, an image analysis means for analyzing image data to identify objects that the user is interested in, and an audio analysis means for analyzing audio data to identify what the user is listening to. This makes it possible to identify the user's interests based on the user's gaze information, image analysis results, and audio analysis results, and to dynamically update recommendation information.

[2016] The "gaze detection means" is a device or tool for acquiring information about the user's gaze.

[2017] An "image capture device" is a device or tool for capturing image data of an object that a user is looking at.

[2018] An "audio capture device" is a device or tool used to record a user's speech.

[2019] The "data transmission means" is a device or tool for encrypting the captured data and transmitting it to the server via a communication network.

[2020] "Data receiving and decrypting means" refers to a device or tool for receiving and decrypting encrypted data sent from the glasses-type device.

[2021] "Image analysis means" refers to a device or tool that analyzes received image data and identifies objects in which the user is interested.

[2022] "Audio analysis means" refers to a device or tool that analyzes received audio data and identifies what the user is listening to.

[2023] The "recommendation generating means" is a device or tool for generating recommendation information based on gaze data and analysis results.

[2024] The "information display means" is a device or tool for decoding the received recommendation information and displaying it on a display.

[2025] The "feedback sending means" is a device or tool for obtaining feedback from a user and sending it to the server.

[2026] The "analysis means" is a device or tool for identifying the user's subject of interest based on gaze information, image analysis results, and audio analysis results, and is used to dynamically update recommendation information.

[2027] System overview and equipment configuration

[2028] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[2029] 1. Glasses-type device

[2030] The glasses-type device is equipped with the following functions:

[2031] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[2032] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[2033] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[2034] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[2035] Information display means: Received recommendation information is displayed on a waveguide type display.

[2036] Feedback sending means: Obtains feedback from users and sends it to the server.

[2037] 2. Server

[2038] The server has the following functions:

[2039] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[2040] Image analysis means: Analyzes the received image data and identifies the objects and scenery the user is looking at. Specifically, an image analysis model (e.g., YOLO or ResNet) is used.

[2041] Speech analysis means: Analyzes the received audio data and converts what the user is hearing and saying into text data. A speech recognition model (e.g., Google Speech-to-Text) is used.

[2042] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[2043] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[2044] Program processing

[2045] Information collection and transmission

[2046] User

[2047] The user wears a glasses-type device and acts accordingly.

[2048] Terminal

[2049] The glasses-type device uses a gaze detection module, camera, and microphone to detect the user's gaze and acquires gaze, image data, and audio data in real time.

[2050] The acquired data is encrypted and sent to the server via the communication network. Specifically, encryption technology (e.g., AES-256) is used.

[2051] Analysis and recommendation generation

[2052] server

[2053] Decrypt the received data.

[2054] Image data is analyzed using an image analysis model (e.g., YOLO or ResNet) to identify objects of interest to the user.

[2055] The voice data is analyzed using a voice recognition model (e.g., Google Speech-to-Text) and converted into text data.

[2056] Based on this data, the user's interests are identified and compared with past behavioral history to generate appropriate recommendation information.

[2057] The generated recommendation information is encrypted and sent to the glasses-type device.

[2058] Recommendation and feedback

[2059] Terminal

[2060] The recommendation information received from the server is decoded and displayed on a waveguide display.

[2061] The feedback from the user is taken, re-encrypted and sent to the server.

[2062] Specific examples

[2063] Example 1: Shopping recommendations

[2064] User

[2065] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[2066] Terminal

[2067] The glasses-type device captures the gaze and images and sends them to a server.

[2068] server

[2069] The received data is analyzed to identify the clothes the user is looking at and recommend accessories and shoes that go well with them.

[2070] Terminal

[2071] Recommendation information is displayed on the screen and presented to the user.

[2072] Example 2: Support for calorie counting

[2073] User

[2074] When a user is looking at a menu at a restaurant, their eyes are drawn to the menu.

[2075] Terminal

[2076] The glasses-type device captures the gaze and images and sends them to a server.

[2077] server

[2078] It analyzes images of menu items, identifies the names of dishes, and retrieves calorie information for the dishes from a database.

[2079] Terminal

[2080] Calorie information is displayed on the display and presented to the user.

[2081] In this way, the present invention is a system that improves the user experience by learning user behavior and interests in real time and providing appropriate recommendation information.

[2082] Prompt Sentence Examples

[2083] The prompt to be input to the generative AI model is as follows:

[2084] "Please explain what kind of recommendations will be provided when a user shows interest in a particular product in the shopping mall."

[2085] The flow of the identification process in the first embodiment will be described with reference to FIG.

[2086] Processing Steps

[2087] Step 1: Collect data

[2088] User

[2089] Users wear the glasses-type device and go about their daily lives.

[2090] Input: User gaze, surrounding images, ambient sounds, and speech.

[2091] Output: Raw gaze data, image data, and audio data.

[2092] Terminal

[2093] The glasses-type device tracks the user's gaze using an eye-gaze detection module.

[2094] An image capture means captures the scenery or object the user is looking at in real time.

[2095] An audio capture method records the ambient sounds around the user and the user's own speech.

[2096] Specifically, when a user is looking at a menu in a cafe, the gaze detection means detects that the user is looking at the menu, and the image capture means takes an image of the menu, while the audio capture means records the sounds of the cafe.

[2097] Step 2: Encrypt and send data

[2098] Terminal

[2099] The collected gaze data, image data, and audio data will be encrypted using encryption technology (e.g., AES-256).

[2100] Input: Raw gaze data, image data, audio data.

[2101] Output: Encrypted gaze data, image data, and audio data.

[2102] The encrypted data is sent to the server at regular intervals (e.g., once per second).

[2103] Specifically, the terminal encrypts the gaze data, menu images, and ambient sounds of the cafe, and then transmits the encrypted data to a server via the Internet.

[2104] Step 3: Receiving and Decrypting Data

[2105] server

[2106] The server receives the encrypted data sent from the eyeglass-type device.

[2107] The received data includes line-of-sight data, image data, and audio data.

[2108] The encrypted data is decrypted using a data receiving and decrypting means.

[2109] Input: Encrypted gaze data, image data, and audio data.

[2110] Output: Decoded gaze data, image data, and audio data.

[2111] Specifically, the server receives the encrypted data sent from the terminal and uses encryption / decryption technology to restore the original gaze data, image data, and audio data.

[2112] Step 4: Data analysis

[2113] server

[2114] The received image data is analyzed using an image analysis model (e.g., YOLO or ResNet) to determine what the user is looking at.

[2115] The voice data is analyzed using a speech recognition model (e.g., Google Speech-to-Text) and the user's speech is converted into text data.

[2116] Input: Decoded image and audio data.

[2117] Output: Image analysis results, audio analysis results.

[2118] Specifically, the server analyzes the menu image to determine that the user is looking at a pizza menu, and analyzes the voice data to determine that the user is saying, "How many calories?"

[2119] Step 5: Generate recommendations

[2120] server

[2121] Identify the user's interests based on gaze data, image analysis results, and audio analysis results.

[2122] This is compared with the user's past behavioral history to generate appropriate recommendation information.

[2123] Input: Gaze data, image analysis results, audio analysis results, past behavioral history.

[2124] Output: Recommendation information.

[2125] Specifically, the server retrieves pizza calorie information from a database based on the user's gaze data and analysis results, and generates recommendation information.

[2126] Step 6: Encrypt and send recommendation information

[2127] server

[2128] The generated recommendation information is encrypted.

[2129] The encrypted recommendation information is sent to the glasses-type device.

[2130] Input: Recommendation information.

[2131] Output: Encrypted recommendation information.

[2132] Specifically, the server encrypts the calorie information of the pizza and sends it to the terminal.

[2133] Step 7: Present recommendations

[2134] Terminal

[2135] The glasses-type device receives the recommendation information transmitted from the server.

[2136] The received recommendation information is decoded and displayed on a waveguide display.

[2137] Input: Encrypted recommendation information.

[2138] Output: Display of decoded recommendation information.

[2139] Specifically, the terminal displays the decrypted calorie information of the pizza on the display.

[2140] Step 8: Getting and sending feedback

[2141] User

[2142] Providing feedback from the user, for example, through voice input or gaze input.

[2143] Terminal

[2144] The glasses-type device collects feedback from the user.

[2145] The obtained feedback is encrypted and sent to the server.

[2146] Input: User feedback.

[2147] Output: Encrypted feedback data.

[2148] Specifically, the user says, "This information was helpful," and the feedback is acquired by the device, encrypted, and sent to the server.

[2149] server

[2150] The server decodes and analyzes the received feedback.

[2151] The results of this analysis will be used to generate the next recommendation information.

[2152] Input: Encrypted feedback data.

[2153] Output: Feedback analysis results.

[2154] Specifically, the server decrypts the encrypted feedback and stores the analysis results for future recommendation generation.

[2155] Prompt Sentence Examples

[2156] The prompt to be input to the generative AI model is as follows:

[2157] "Please explain what kind of recommendations will be provided when a user shows interest in a particular product in the shopping mall."

[2158] (Application example 1)

[2159] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2160] Traditional shopping experiences lack personalized product recommendations based on a user's specific interests and behavioral history, resulting in the significant time and effort required for users to select the right products. Even in brick-and-mortar stores, product recommendations often remain general and do not adequately address individual user needs. Furthermore, there is a lack of a system for instantly incorporating user feedback and improving the next recommendation. This prevents an improved user experience and efficient shopping.

[2161] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[2162] In this invention, the server includes a data receiving / decoding means, an image analysis means, a voice analysis means, a product recommendation means, a recommendation generation means, and a feedback transmission means. This makes it possible to make personalized product recommendations based on detailed analysis results of the user's gaze, voice, and behavioral history, and to send feedback from the user to the server and use the generation AI model to improve the accuracy of the next recommendation.

[2163] A "glasses-type device" is a device worn by the user that can collect gaze information, image data, audio data, etc. in real time and send it to a server for analysis.

[2164] The "gaze detection means" is a function installed in the glasses-type device that detects the user's gaze and pupil movement to collect gaze information.

[2165] The "image capture means" is a function that captures image data of the scenery or object that the user is viewing in real time.

[2166] "Audio capture means" is a function that records the user's speech and surrounding environmental sounds.

[2167] The "data transmission means" is a function that encrypts collected data and transmits it to a server via a communication network.

[2168] The "data receiving and decrypting means" is a function that receives and decrypts encrypted data sent from the glasses-type device.

[2169] The "image analysis means" is a function that analyzes the received image data and identifies the object or scene that the user is looking at.

[2170] The "voice analysis means" is a function that analyzes received voice data and identifies the content and statements that the user is listening to.

[2171] The "recommendation generation means" is a function that generates appropriate recommendation information based on the analysis results and the user's past behavioral history.

[2172] The "information display means" is a function that decodes the received recommendation information and displays it on a display.

[2173] The "product recommendation means" is a function that provides personalized product recommendations based on the user's past behavioral history when selecting products in a physical store.

[2174] The "feedback sending means" is a function that obtains feedback information from users and sends prompt sentences to the server using a generative AI model in order to improve the accuracy of the next recommendation based on the feedback information.

[2175] A "generative AI model" is an artificial intelligence model that uses machine learning to learn patterns from data and generate output based on input data.

[2176] A "prompt" is a text sentence entered into a generative AI model to instruct it on a specific task.

[2177] The present invention is a system that uses glasses-type devices to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user and a server that analyzes the data and generates recommendation information.

[2178] System overview and equipment configuration

[2179] 1. Glasses-type device

[2180] The glasses-type device is equipped with the following functions:

[2181] Gaze detection means: A means of detecting the user's gaze and pupil movements and collecting that data.

[2182] Image capture means: A means for capturing image data of the scenery or object the user is viewing in real time.

[2183] Audio capture means: A means of recording the user's surroundings and their own speech.

[2184] Data transmission means: A means for encrypting the captured data and transmitting it to a server via a communication network.

[2185] Information display means: A means for decrypting the received recommendation information and displaying it on a display.

[2186] Feedback sending means: A means for obtaining feedback from users and sending it to the server.

[2187] 2. Server

[2188] The server has the following functions:

[2189] Data reception and decryption means: A means for receiving and decrypting encrypted data sent from the glasses-type device.

[2190] Image analysis means: A means for analyzing the received image data and identifying the object or scene the user is looking at. In this case, image analysis software such as TensorFlow or OpenCV is used.

[2191] Speech analysis means: A means of analyzing the received audio data to identify what the user is listening to and saying, such as using voice recognition software like Google Cloud Speech-to-Text API or IBM Watson.

[2192] Recommendation generation method: A method for generating appropriate recommendation information based on analysis results and the user's past behavioral history. In this case, a generative AI model is used.

[2193] Data transmission means: A means for encrypting the generated recommendation information and transmitting it to the glasses-type device.

[2194] Program processing and specific examples

[2195] Information collection and transmission

[2196] The glasses-type device uses a gaze detection module, camera, and microphone to capture the user's gaze, image data, and voice data in real time while the user is walking around the physical store. For example, when a user looks at a pair of sneakers, that information is collected. The collected data is encrypted and sent to a server at regular intervals.

[2197] Analysis and recommendation generation

[2198] The server decrypts the received data and analyzes the image data using an image analysis model to identify the object the user is looking at (e.g., sneakers). It also analyzes the audio data using a voice recognition model to identify what the user is listening to (e.g., product description) and what they are saying (e.g., "Do these sneakers come in other colors?"). Based on the gaze data, image analysis results, and audio analysis results, it identifies the user's interests and compares them with their past behavioral history to generate recommendation information. The recommendation information is encrypted and sent to the glasses-type device.

[2199] Recommendation and feedback

[2200] The glasses-type device receives recommendation information sent from the server, decodes it, and displays it on the screen. If the user is looking at sneakers, it will recommend matching socks and other color variations. Feedback from the user is obtained, and this feedback information is sent back to the server and input into the generative AI model as prompt sentences to improve the accuracy of the next recommendation.

[2201] Examples of prompt statements

[2202] "If a user expresses interest in a blue shirt, offer relevant recommendations, such as matching pants and accessories, or other color variations of the shirt. Take into account the user's previous purchases and style preferences."

[2203] "When a user is listening to information about sneakers, recommend related information, such as socks and other workout apparel to go with the sneakers, or information about current promotions."

[2204] As described above, the present invention is a system that improves the shopping experience in physical stores by learning user behavior and interests in real time and providing appropriate recommendation information.

[2205] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[2206] Step 1:

[2207] Collection of information

[2208] The terminal (glasses-type device) acquires the user's gaze information using a gaze detection means, captures image data of the scenery or object the user is looking at in real time using an image capture means, and also collects the user's remarks and surrounding environmental sounds using an audio capture means.

[2209] Input: User gaze data, image data, and audio data.

[2210] Output: Encrypted gaze data, image data, and audio data.

[2211] Step 2:

[2212] Sending data

[2213] The terminal (glasses-type device) encrypts the collected gaze data, image data, and voice data, and transmits them to the server via a communication network using a data transmission means.

[2214] Input: Encrypted gaze data, image data, and audio data.

[2215] Output: The encrypted data sent to the server.

[2216] Step 3:

[2217] Receiving and Decrypting Data

[2218] The server receives the encrypted data sent from the glasses-type device using a data receiving / decrypting means and decrypts it.

[2219] Input: Encrypted gaze data, image data, and audio data.

[2220] Output: Decoded gaze data, image data, and audio data.

[2221] Step 4:

[2222] Image analysis

[2223] The server analyzes the decoded image data using image analysis tools (e.g., TensorFlow or OpenCV) to identify the objects and scenes the user is viewing.

[2224] Input: Decoded image data.

[2225] Output: Analysis results (identification of the object or scene the user is looking at).

[2226] Step 5:

[2227] Audio analysis

[2228] The server then analyzes the decoded audio data using a speech analysis tool (such as Google Cloud Speech-to-Text API or IBM Watson) to determine what the user is hearing and saying.

[2229] Input: Decoded audio data.

[2230] Output: Analysis results (identification of what the user is hearing and saying).

[2231] Step 6:

[2232] Generating recommendation information

[2233] The server identifies the user's interests based on gaze data, image analysis results, and audio analysis results, and generates recommendation information by comparing it with the user's past behavioral history. In this case, a generative AI model is used.

[2234] Input: Gaze data, image analysis results, audio analysis results, past behavioral history.

[2235] Output: Recommendation information (e.g. related products and campaign information).

[2236] Step 7:

[2237] Sending recommendation information

[2238] The server encrypts the generated recommendation information using a data transmission means and transmits it to the glasses-type device.

[2239] Input: Recommendation information.

[2240] Output: Encrypted recommendation information.

[2241] Step 8:

[2242] Displaying Information

[2243] The terminal (glasses-type device) receives the encrypted recommendation information transmitted from the server using the data receiving means and the decrypting means, decrypts it, and displays it to the user using the information displaying means.

[2244] Input: Encrypted recommendation information.

[2245] Output: Recommendation information displayed on the screen.

[2246] Step 9:

[2247] Get feedback

[2248] The terminal (glasses-type device) acquires feedback from the user and transmits the feedback information to the server using the feedback transmission means.

[2249] Input: User feedback information.

[2250] Output: Feedback information sent to the server.

[2251] Step 10:

[2252] Input to generative AI models

[2253] The server creates a prompt sentence for the generative AI model based on the feedback information, which is used to generate the next recommendation information.

[2254] Input: Feedback information.

[2255] Output: The prompt and the generative AI model after retraining.

[2256] Examples of prompt statements

[2257] "If a user expresses interest in a blue shirt, offer relevant recommendations, such as matching pants and accessories, or other color variations of the shirt. Take into account the user's previous purchases and style preferences."

[2258] "When a user is listening to information about sneakers, recommend related information, such as socks and other workout apparel to go with the sneakers, or information about current promotions."

[2259] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[2260] System overview and equipment configuration

[2261] This invention is a system that uses a glasses-type device to learn a user's daily behavior in real time and provide personalized information. This system includes a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that recognizes the user's emotions.

[2262] 1. Glasses-type device

[2263] The glasses-type device is equipped with the following functions:

[2264] Gaze detection means: Detects the user's gaze and pupil movements and collects that data.

[2265] Image capture means: Captures image data of the scenery or object the user is looking at in real time.

[2266] Audio capture method: Records the ambient sounds around the user and the user's own speech.

[2267] Data transmission means: Collected data is encrypted and sent to a server via a communication network.

[2268] Information display means: Received recommendation information is displayed on a waveguide type display.

[2269] Feedback sending means: Obtains feedback from users and sends it to the server.

[2270] 2. Server

[2271] The server has the following functions:

[2272] Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[2273] Image analysis means: Analyzes the received image data and identifies the object or scene the user is looking at.

[2274] Audio analysis means: Analyzes received audio data to determine what the user is hearing and saying.

[2275] Recommendation generation method: Generate appropriate recommendation information based on the analysis results and the user's past behavioral history.

[2276] Data transmission method: The generated recommendation information is encrypted and sent to the glasses-type device.

[2277] 3. Emotion Engine

[2278] The emotion engine has the following functions:

[2279] Facial expression analysis means: Analyzes received image data and identifies the user's emotions.

[2280] Voice tone analysis means: Analyzes the tone and rhythm of received voice data to identify the user's emotions.

[2281] Emotion data transmission means: Transmits the identified emotion data to the server.

[2282] Program processing

[2283] Information collection and transmission

[2284] Terminal

[2285] While the user is wearing the glasses-type device and moving around, the gaze detection module, camera, and microphone capture the user's gaze, image data, and audio data in real time.

[2286] As you use it, the emotion engine analyzes your facial expressions and voice tone in real time to generate emotion data, which is also temporarily stored in local memory.

[2287] The acquired data and emotion data are encrypted and sent to the server at regular intervals.

[2288] Analysis and recommendation generation

[2289] server

[2290] The received data is decoded and the image data is analyzed using an image analysis model to determine what the user is looking at.

[2291] The voice data is analyzed using a voice recognition model, and what the user is hearing and saying is converted into text data.

[2292] The emotional data from the emotion engine is analyzed to determine the user's current emotional state.

[2293] Gaze data, image analysis results, audio analysis results, and emotional data are combined to identify the user's interests and emotions.

[2294] By comparing the user's past behavioral history, the system generates appropriate recommendations, especially by prioritizing information that reflects the user's emotional state.

[2295] The generated recommendation information is encrypted and sent to the glasses-type device.

[2296] Recommendation and feedback

[2297] Terminal

[2298] The recommendation information sent from the server is received, decrypted, and displayed on the screen.

[2299] A function to obtain feedback from the user is activated. For example, feedback is collected in the form of voice commands or touch operations, and this information is sent back to the server. The feedback information is used to improve the accuracy of the next recommendation.

[2300] Specific examples

[2301] Example 1: Shopping recommendations

[2302] User

[2303] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[2304] Terminal

[2305] The glasses-type device captures gaze and images, and if the user smiles at the clothes, it sends the data, including their emotions, to the server.

[2306] server

[2307] The system analyzes the received data to identify the clothes the user is looking at, determines that the user has a positive feeling toward the clothes, and recommends accessories and shoes that go well with them.

[2308] Terminal

[2309] Recommendation information is displayed on the screen and presented to the user.

[2310] Example 2: Support for calorie counting

[2311] User

[2312] While looking at a menu at a restaurant, the user glances at the menu and looks a little surprised.

[2313] Terminal

[2314] The glasses-type device captures gaze and images, and sends the user's surprised facial expression as emotional data to the server.

[2315] server

[2316] It analyzes menu images, identifies the names of dishes, and retrieves calorie information from a database, while simultaneously analyzing surprise emotions and generating advice to avoid high-calorie choices.

[2317] Terminal

[2318] Calorie information and advice is displayed on the screen and presented to the user.

[2319] In this way, the present invention is a system that can learn not only a user's behavior and interests but also their emotions in real time, and provide more appropriate and personalized information.

[2320] The processing flow will be explained below.

[2321] Step 1:

[2322] Terminal

[2323] When a user puts on the glasses-type device, the gaze detection module, camera, microphone, and emotion engine are automatically activated. The gaze detection module detects the user's gaze and pupil movements in real time and collects that data. The camera also captures image data of the scenery and objects the user is looking at at a frequency of several frames per second. The microphone records surrounding sounds and what the user is saying.

[2324] Step 2:

[2325] Terminal

[2326] The emotion engine operates, analyzing the user's facial expressions and voice tone in real time to generate emotion data, which is temporarily stored in local memory.

[2327] Step 3:

[2328] Terminal

[2329] The gaze data, image data, voice data, and emotion data are encrypted at regular intervals, and then transmitted to a server via the internet through a data transmission means.

[2330] Step 4:

[2331] server

[2332] The server receives the encrypted data via the Internet, and the received data is decrypted by the data receiving and decrypting means.

[2333] Step 5:

[2334] server

[2335] The decoded image data is analyzed using image analysis means to identify the object or scene the user is looking at.

[2336] Step 6:

[2337] server

[2338] The decoded voice data is analyzed using a voice analysis means, and what the user is hearing or saying is converted into text data.

[2339] Step 7:

[2340] server

[2341] The received emotional data is analyzed by an emotional analysis means to identify the emotional state of the user.

[2342] Step 8:

[2343] server

[2344] Gaze data, image analysis results, audio analysis results, and emotional data are aggregated to identify the user's interests and current emotions.

[2345] Step 9:

[2346] server

[2347] The recommendation generating means generates appropriate recommendation information based on the user's past behavior history and current emotional state, and this recommendation information includes more personalized content.

[2348] Step 10:

[2349] server

[2350] The generated recommendation information is encrypted and transmitted to the glasses-type device via the data transmission means.

[2351] Step 11:

[2352] Terminal

[2353] The glasses-type device decodes the received recommendation information, and the decoded information is displayed on a waveguide display and presented to the user.

[2354] Step 12:

[2355] Terminal

[2356] The user provides some kind of feedback (voice instruction or touch operation) to the presented information. The feedback data is encrypted and sent to the server via the data transmission means.

[2357] Step 13:

[2358] server

[2359] The server receives and analyzes the feedback data, and the results are used to generate the next set of recommendations, helping to improve the accuracy of the system.

[2360] Example 2

[2361] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2362] Conventional information provision systems using glasses-type devices were able to collect data such as the user's gaze, ambient sounds, and speech, but were limited in their ability to provide recommendation information that took the user's emotional state into account. This made it difficult to provide more personalized information based on the user's current emotional state. There was also a need for improvements in the accuracy and real-time nature of the analysis of collected data.

[2363] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[2364] In this invention, the server includes a data receiving / decoding means for receiving and decoding the data, an image analysis means for analyzing the image data to identify objects in which the user is interested, a voice analysis means for analyzing the voice data to identify what the user is listening to, an emotion analysis means for analyzing the emotion data to identify the user's emotional state, and a recommendation generation means for integrating the gaze data, image analysis results, voice analysis results, and emotion data to generate recommendation information. This makes it possible to generate appropriate recommendation information that takes into account not only the user's gaze and voice information but also their emotional state.

[2365] A "glasses-type device" is a device that can acquire gaze information, image data, voice data, and emotional data when worn by a user and transmit the data to a server.

[2366] The "gaze detection means" is a mechanism for acquiring the user's gaze information in real time and storing the data internally or transmitting it externally.

[2367] The "image capture means" is a mechanism for capturing and recording image data of the scenery or object that the user is viewing in real time.

[2368] The "audio capture means" is a mechanism for collecting and recording the user's surrounding environmental sounds and the user's speech in real time.

[2369] The "emotion engine" is a mechanism that analyzes the user's facial expressions and tone of voice to identify the user's emotional state in real time.

[2370] The "data transmission means" is a mechanism for encrypting the collected gaze data, image data, voice data, and emotion data and transmitting them to the server.

[2371] The "server" is a central processing unit that receives, decodes, and analyzes data sent from the glasses-type device to generate recommendation information.

[2372] The "data receiving and decrypting means" is a mechanism for receiving and decrypting encrypted data sent from the glasses-type device.

[2373] The "image analysis means" is a mechanism for analyzing received image data and identifying the object or scenery the user is looking at.

[2374] The "voice analysis means" is a mechanism for converting received voice data into text and identifying what the user has said and what they are listening to.

[2375] The "emotion analysis means" is a mechanism for analyzing received emotion data and identifying the user's current emotional state.

[2376] The "recommendation generation means" is a mechanism that generates recommendation information to provide appropriate information to users based on gaze data, image analysis results, audio analysis results, and emotional data.

[2377] The "information display means" is a display device that is installed in the glasses-type device and visually presents the received recommendation information to the user.

[2378] The "feedback transmission means" is a mechanism for obtaining feedback from the user, encrypting it, and transmitting it to the server.

[2379] MODE FOR CARRYING OUT THE INVENTION

[2380] The present invention is a personalized information providing system that uses a glasses-type device and a server. The system configuration and operating principles of the present invention will be described in detail below.

[2381] System Configuration

[2382] This system consists of a glasses-type device worn by the user, a server that analyzes data and generates recommendation information, and an emotion engine that identifies the user's emotional state.

[2383] Glasses-type device

[2384] The glasses-type device includes the following hardware and software:

[2385] 1. Gaze detection method: Obtaining user gaze information in real time. Specifically, using an infrared sensor or camera to detect gaze and generate gaze tracking data.

[2386] 2. Image capture method: Obtain image data of the scenery or object the user is looking at. Use a high-resolution camera to capture image data in real time.

[2387] 3. Audio capture: Collects the user's surrounding sounds and conversations. Audio data is acquired using the built-in microphone.

[2388] 4. Emotion Engine: Analyzes the user's facial expressions and tone of voice in real time to generate emotion data, for example, using facial expression recognition algorithms and voice analysis algorithms.

[2389] 5. Data transmission method: Collected data and emotion data are encrypted and transmitted to a server via a communication network using AES (Advanced Encryption Standard) via Wi-Fi or mobile networks.

[2390] 6. Information display means: The received recommendation information is displayed on a waveguide type display.

[2391] 7. Feedback sending method: Obtains feedback from the user and sends it to the server. Feedback is collected through voice commands and touch operations.

[2392] server

[2393] The server includes the following software:

[2394] 1. Data reception and decryption means: Receives and decrypts encrypted data sent from the glasses-type device.

[2395] 2. Image analysis: Analyzes the received image data and identifies the object or scene the user is looking at. For image analysis, a deep learning model (such as YOLO or ResNet) is used.

[2396] 3. Voice analysis: Converts received voice data into text and identifies what the user is saying and the surrounding sounds. Google Speech-to-Text API is used for voice recognition.

[2397] 4. Emotion analysis means: Analyzes the received emotion data and identifies the user's current emotional state.

[2398] 5. Recommendation generation method: Integrates gaze data, image analysis results, audio analysis results, and emotion data to generate appropriate recommendations for users, taking into account the user's past behavioral history.

[2399] Specific examples

[2400] For example:

[2401] Example 1: Shopping recommendations

[2402] User

[2403] While walking through a shopping mall, a user's eyes are drawn to a particular piece of clothing.

[2404] Terminal

[2405] The gaze detection means and image capture means capture images of the user's gaze and clothes, and the emotion engine detects smiling facial expressions. These data are sent to the server.

[2406] server

[2407] The system analyzes the received data, identifies the clothes the user is looking at, determines whether they have a positive emotion, and recommends accessories and shoes that go well with the clothes.

[2408] Terminal

[2409] Recommendation information is displayed on the screen and presented to the user.

[2410] Example 2: Support for calorie counting

[2411] User

[2412] While looking at a menu at a restaurant, the user glances at the menu and looks a little surprised.

[2413] Terminal

[2414] The gaze detection means and image capture means capture the menu image, and the emotion engine detects surprised facial expressions. These data are sent to the server.

[2415] server

[2416] It analyzes menu images, identifies the names of dishes, retrieves calorie information from a database, analyzes surprise emotions, and generates advice to avoid high-calorie choices.

[2417] Terminal

[2418] Calorie information and advice is displayed on the screen and presented to the user.

[2419] Prompt Sentence Examples

[2420] An example of a prompt to be input to a generative AI model is, "Please consider the user's emotional state and generate recommendation information for the object they are looking at."

[2421] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2422] *The following explanation will detail the inputs, outputs, and specific operations at each step of the procedure.

[2423] Step 1: Gather information

[2424] Terminal

[2425] Input: The glasses detect the user's gaze, images of the surrounding scenery and objects, ambient sounds and speech, facial expressions, and tone of voice.

[2426] How it works: The gaze detection means captures the user's gaze data in real time (using infrared sensors and cameras). The image capture means uses a high-resolution camera to capture image data of scenery and objects. The audio capture means uses a built-in microphone to record the surrounding environmental sounds and the user's speech. The emotion engine analyzes facial expressions and voice tone to generate emotion data (using facial expression recognition algorithms and audio analysis algorithms).

[2427] Output: Gaze data, image data, audio data, emotion data.

[2428] Step 2: Sending data

[2429] Terminal

[2430] Input: Gaze data, image data, audio data, and emotion data collected in Step 1.

[2431] How it works: Collected data is encrypted using the AES encryption algorithm and sent to a server via Wi-Fi or mobile network.

[2432] Output: Encrypted gaze data, image data, audio data, and emotion data.

[2433] Step 3: Receiving and Decrypting Data

[2434] server

[2435] Input: Encrypted gaze data, image data, audio data, and emotion data.

[2436] How it works: The server receives encrypted data sent from the device via the HTTP communication protocol and decrypts the data using the same AES encryption algorithm.

[2437] Output: Decoded gaze data, image data, audio data, and emotion data.

[2438] Step 4: Analyze the data

[2439] server

[2440] Input: Decoded gaze data, image data, audio data, and emotion data.

[2441] Operation: Analyzes gaze data to identify the direction and object the user is looking at. The image analysis means analyzes the received image data using deep learning models (YOLO or ResNet) to identify the object the user is looking at. The audio analysis means converts audio data into text using the Google Speech-to-Text API to identify what is being said. The emotion analysis means analyzes emotion data to identify the user's emotional state.

[2442] Output: Gaze analysis results, image analysis results, audio analysis results, emotional state.

[2443] Step 5: Generate recommendations

[2444] server

[2445] Input: Gaze analysis results, image analysis results, audio analysis results, emotional state, and user's past behavior history.

[2446] How it works: Integrates gaze data, image analysis results, audio analysis results, and emotional state to identify the user's interests. References past behavioral history and generates appropriate recommendations based on the user's current interests and emotional state. Creates personalized recommendations using a generative AI model.

[2447] Output: Recommendation information.

[2448] Step 6: Sending recommendations

[2449] server

[2450] Input: Recommendation information.

[2451] How it works: The generated recommendation information is encrypted using the AES encryption algorithm and sent to the glasses-type device via Wi-Fi or mobile network.

[2452] Output: Encrypted recommendation information.

[2453] Step 7: Displaying Recommendations

[2454] Terminal

[2455] Input: Encrypted recommendation information.

[2456] How it works: Decodes received recommendation information and visually displays the information...

Claims

1. Glasses-type devices A gaze detection means for acquiring gaze information of a user; image capture means for capturing image data of an object viewed by a user; an audio capture means for recording user speech; a data transmission means for encrypting the captured data and transmitting the encrypted data to a server via a communication network; The server data receiving and decoding means for receiving and decoding the data; image analysis means for analyzing the image data to identify objects of interest to the user; a voice analysis means for analyzing the voice data to identify what the user is listening to; a recommendation generating means for generating recommendation information based on the gaze data and analysis results; a data transmission means for encrypting the recommendation information and transmitting the encrypted information to the glasses-type device; Glasses-type devices An information display means for decoding the received recommendation information and displaying it on a display. A system including:

2. The system according to claim 1 , wherein the recommendation generating means generates recommendation information by comparing it with a past user behavior history.

3. 2. The system according to claim 1, wherein said information display means includes feedback transmission means for obtaining feedback from a user and transmitting the feedback to said server.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A