system

The system addresses information overload and environmental impact by converting visual data into digital format, integrating it with user behavior, and providing personalized notifications with feedback-based rewards.

JP2026068306APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Users face challenges in efficiently obtaining relevant information due to information overload, environmental impact from conventional notification methods, and inadequate incentives for information providers and developers.

Method used

A system that converts visual information into digital data using user-worn devices, integrates it with user behavior data, and provides personalized notifications based on user feedback to reward information providers and developers.

Benefits of technology

Enhances user experience by optimizing information delivery, reduces environmental impact, and provides fair incentives for information providers and developers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026068306000001_ABST
    Figure 2026068306000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of converting visual information acquired by a device worn by a user into digital data, A means for analyzing the aforementioned digital data and extracting relevant information, Means by which communication devices collect user behavior data, A means of integrating collected behavioral data and analyzed digital data to generate information optimized for the user, Means for notifying the user of the generated information, A system that includes means for evaluating the aforementioned information based on user feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, with the increase in the amount of information, it has become difficult for users to efficiently obtain useful information they need. Due to information acquisition leakage and schedule forgetting, it is not uncommon to suffer opportunity losses and stress. In addition, conventional analog notification means have become a factor increasing the environmental load. Furthermore, the mechanism for information providers and application developers to obtain a reward corresponding to their contributions is insufficient. It is required to solve such problems.

Means for Solving the Problems

[0005] This invention provides a system that converts visual information into digital data using a device worn by the user and analyzes that digital data. Furthermore, it collects user behavior data using a communication device and integrates this with the analyzed digital data to generate information optimized for the user. The system also includes a mechanism to notify users of the generated information and evaluate its usefulness based on user feedback, thereby providing appropriate rewards to information providers and application developers. This allows users to efficiently acquire necessary information and contributes to reducing environmental impact.

[0006] A "user-worn device" is a device that a user can wear in their daily life to acquire and process information.

[0007] "Visual information" refers to information that users can perceive through their eyes, and includes information displayed on physical media such as books, flyers, and signs.

[0008] "Digital data" refers to data obtained by converting analog information into a format that can be processed electronically, and is binary data that can be handled by a computer.

[0009] "Analyzing" refers to the process of examining, classifying, and evaluating input data to understand its content and extract information relevant to a specific purpose.

[0010] A "communication device" is an electronic device used to send and receive information, and can exchange data with other devices via a network.

[0011] "Behavioral data" refers to data collected based on a user's location, activity history, and digital device usage, and it provides information that indicates a user's interests and activities.

[0012] "Integrating" means combining multiple different pieces of data or information and organizing and processing them into a single, cohesive piece of information.

[0013] "Notifying" refers to a means of conveying specific information or messages to a user, and is a form of communication conducted through electronic devices.

[0014] "Feedback" refers to opinions and evaluations that users express about the information or services they receive, and it is important information for providers to use for improvement and evaluation.

[0015] An "information provider" is an entity that makes information or data publicly available or communicates it to users, and includes companies, organizations, and individuals.

[0016] An "application developer" is a professional or organization that designs and develops software and applications, and whose role is to provide computer programs to users.

[0017] "Reward" refers to monetary or material compensation or gratuity given for an action or contribution, and also serves as a motivator. [Brief explanation of the drawing]

[0018] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7]It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0019] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0020] First, the language used in the following description will be explained.

[0021] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0022] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0026] [First Embodiment]

[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0039] To implement this invention, it is necessary to establish a system in which the user wears a device such as smart glasses and takes in visual information on a daily basis. These smart glasses have a built-in camera that captures the information the user sees. The captured visual information is converted into digital data.

[0040] Smart glasses use OCR (Optical Character Recognition) to convert visual information into text data. This text data is then transmitted wirelessly to a server in the cloud. The server receives this data and analyzes the information using natural language processing algorithms. From the analysis results, it extracts highly relevant keywords and context.

[0041] Meanwhile, the user's smartphone collects behavioral data, including location information and app usage history. This data is also transferred to a server. The server integrates the behavioral data from the smartphone with the analyzed digital data from the smart glasses to build a user profile. Using this profile, the server selects the most relevant information for the user and makes suggestions at the appropriate time.

[0042] For example, suppose a user is walking around town and their smart glasses capture a promotional banner for a new restaurant. The server analyzes this information and, based on the user's current location and past dining preferences, notifies their smartphone of a lunch coupon for that specific restaurant.

[0043] Users who receive a notification can send feedback on the information through the app on their smartphone. This feedback data is also sent to the server and used to reward information providers and application developers.

[0044] In this way, the invention is implemented in a manner that blurs the boundaries between analog and digital information, improves the user experience, and strengthens incentives for information providers.

[0045] The following describes the processing flow.

[0046] Step 1:

[0047] The device (smart glasses) captures analog information seen by the user with a camera and converts that image into digital text data using OCR technology.

[0048] Step 2:

[0049] The device (smart glasses) sends the generated text data to the server via the internet. The transmission is performed using a secure protocol.

[0050] Step 3:

[0051] The server analyzes the received text data and extracts key keywords and context using natural language processing techniques. This process helps to understand the meaning of the data.

[0052] Step 4:

[0053] The device (smartphone) collects the user's current location information and past activity history, and sends this data to the server. Location information is obtained using GPS.

[0054] Step 5:

[0055] The server integrates the analyzed text data and behavioral data sent from the terminal and updates the user profile to estimate the user's interests.

[0056] Step 6:

[0057] The server generates optimal information and suggestions for the user based on integrated data and notifies the user of this information via their device (smartphone). The timing of these notifications is adjusted according to the user's situation.

[0058] Step 7:

[0059] Users provide feedback on the usefulness and interest of the information they receive through notifications on their smartphones. This helps to improve the user's future information delivery.

[0060] Step 8:

[0061] The server collects user feedback and uses it as data to distribute rewards to information providers and application developers. The reward criteria are based on the evaluation of the feedback.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] In recent years, there has been a growing demand for services that digitize users' visual information from their daily lives and activities and use that information to provide detailed, relevant suggestions. However, relying solely on visual information from users makes it difficult to accurately provide highly relevant information, resulting in a lack of improvement in the user experience. Furthermore, if the evaluation of information provision is not fair, incentives for information providers may not function properly. This invention aims to solve these problems and build an information provision system that is more beneficial to users.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for converting visual information acquired by a camera worn by the user into electronic data, natural language processing means for analyzing the electronic data and extracting relevant information, and means for a communication device to collect the user's location data and usage history. This enables the integration and analysis of the user's visual information and behavioral data, allowing for more personalized information provision and fair reward evaluation for information providers.

[0067] "User" refers to an individual or group that uses this system.

[0068] A "photography device" refers to a device that has the function of capturing visual information, and specifically refers to a wearable device with a built-in camera.

[0069] "Electronic data" refers to a data format that includes visual information digitized by a photographic device.

[0070] "Natural language processing means" refers to algorithms and software tools used to analyze electronic data and extract or classify language data.

[0071] "Communication equipment" refers to devices that have the ability to send and receive data via the internet or a local network.

[0072] "Location data" refers to information indicating the user's geographical location, and is obtained through methods such as GPS.

[0073] "Usage history" refers to records of actions a user has taken in the past, applications they have used, places they have visited, and so on.

[0074] "Profile generation means" refers to a method or apparatus for generating detailed information about a user based on collected electronic data and behavioral data.

[0075] "Response" refers to feedback and evaluation from users regarding the generated information.

[0076] "Compensation" refers to the payment given to information providers or application developers that is commensurate with the value of the information they provide.

[0077] To implement this invention, it is fundamental that the user wears smart glasses, which are a camera, to collect everyday visual information. The smart glasses have a built-in camera that acquires information about the outside world as image data. This image data is converted into electronic data using OCR (Optical Character Recognition) technology. An OCR engine such as Tesseract can be used for this conversion.

[0078] The server receives the converted electronic data and performs analysis using natural language processing (NLP) algorithms. This analysis utilizes natural language processing libraries such as spaCy and NLTK to extract relevant information from the electronic data. This yields keywords and sentences relevant to the user's context, preparing the data for information provision.

[0079] Devices, especially smartphones, collect user location data via GPS and also acquire application usage history data. This information is sent to a server, where a profile generation mechanism constructs data about the user. Based on the profile, the server selects the most relevant information for the user and notifies the smartphone.

[0080] As a concrete example, suppose a user is walking through a shopping mall and their smart glasses capture an advertisement for a newly opened store. The server analyzes this advertisement information and, based on the user's current location and shopping preferences, sends a discount coupon for that store to their smartphone.

[0081] This invention enables more personalized information delivery by utilizing users' visual information and behavioral data. An example of a prompt sentence using the generative AI model is, "Please tell me how to suggest products that might interest you based on the advertisements you've seen."

[0082] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0083] Step 1:

[0084] The user wears smart glasses, which are a recording device, to acquire visual information. These glasses have a built-in camera that captures information within the user's field of view as image data. The input is analog visual information, and the output is digital image data. This allows various types of information obtained during the user's daily activities to be digitized.

[0085] Step 2:

[0086] Smart glasses convert acquired image data into electronic data using OCR (Optical Character Recognition) technology. Specifically, they analyze text within images using OCR engines such as Tesseract and extract character information. The input is digital image data, and the output is electronic data in text format. This process makes it possible to effectively extract character information from images.

[0087] Step 3:

[0088] The smart glasses transmit the converted text data to the server via wireless communication. The server takes the received text data as input and analyzes it using natural language processing algorithms. Specifically, it uses libraries such as spaCy and NLTK to perform semantic analysis of the data and extract highly relevant keywords and contexts. The output is the analyzed keyword and context information. This allows the server to obtain detailed information about the user's intent and interests.

[0089] Step 4:

[0090] On the other hand, the smartphone, as the terminal, obtains location information from GPS and collects past app usage history. The input is behavioral data obtained from the terminal, and the output is specific location information and history data that shows the user's behavior patterns. This collected data is transmitted to the server in real time.

[0091] Step 5:

[0092] The server integrates behavioral data from smartphones with electronic data analyzed from smart glasses. The input is the analyzed data obtained in steps 3 and 4, and the output is a user profile. Based on this, the server selects the most relevant information for the user and configures itself to respond quickly.

[0093] Step 6:

[0094] The server notifies the user's smartphone of the generated information. This uses a push notification service, with the user profile and selection information as input, and the output being the specific notification content for the user. The user receives this information and can decide on their next action according to the prompt.

[0095] (Application Example 1)

[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0097] In modern society, users receive vast amounts of information daily, and this information is not always optimized to their interests or circumstances. In particular, in the fields of advertising and marketing, there is a demand for providing information tailored to users' interests and current situations in real time. Furthermore, streamlining the compensation system for information providers and those who create and apply the information is also a crucial challenge.

[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0099] In this invention, the server includes means for converting visual information acquired by a visual device worn by the user into digital data, means for analyzing the digital data and extracting relevant information, and means for a communication device to collect user behavior data. This makes it possible to provide appropriate advertising information in real time based on the user's location information and interest data.

[0100] A "user" refers to a person who wears visual devices to receive information.

[0101] A "visual device" refers to a device worn by a user to acquire visual information, and is capable of converting that information into digital data.

[0102] "Digital data" refers to data obtained by visual devices that has been converted into a format that can be processed.

[0103] "Analysis methods" refer to processes and techniques for extracting relevant information from digital data.

[0104] A "communication device" refers to a device that collects user behavior data and transmits it to a server or other device.

[0105] "Behavioral data" refers to data that shows users' activities, location information, and interests.

[0106] "Notification means" refers to the mechanisms and functions used to deliver generated information to users.

[0107] "Feedback" refers to the evaluations and opinions that users give in response to the information provided.

[0108] "Promotional information" refers to advertisements and promotions provided based on the user's interests and location information.

[0109] "Real-time" refers to information processing and notifications occurring instantly and without delay.

[0110] The system for implementing this invention consists of a visual device worn by the user and a mobile terminal. The visual device is equipped with a camera that captures the user's visual information in real time. This information is converted into text data using built-in optical character recognition (OCR) software. The converted text data is transmitted wirelessly to a server in the cloud.

[0111] The server performs natural language processing on the received text data and analyzes relevant information. During this process, the server builds individual user profiles based on behavioral data transmitted from the user's mobile device, such as location information and past activity history. This allows for the generation of user-optimized information or advertisements, which are then notified to the user's mobile device in real time.

[0112] For example, if a user's visual device captures a sign for a new restaurant while walking around town, that information is immediately analyzed. The server considers the user's location and past dining patterns and sends a special lunch coupon related to that restaurant to the user's mobile device. This notification is timed appropriately so that the user receives the information at the right moment.

[0113] The hardware used includes visual devices such as Google Glass®, and the software used for analysis includes Google Cloud NLP and OCR tools. The terminal is assumed to be a typical smartphone device. An example prompt message is as follows:

[0114] Example of a prompt:

[0115] We have captured promotional information for new exhibitions that the user is interested in at their current location. Based on this information, generate a discount coupon for the relevant exhibition and notify the user on their mobile device.

[0116] This invention allows users to always receive information tailored to their needs, and enables information providers to efficiently reach their target users.

[0117] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0118] Step 1:

[0119] The user wears a visual device to capture visual information from their surroundings. This visual information, as input, is received by a camera sensor and converted into image data. The output is the image data sent to the OCR process.

[0120] Step 2:

[0121] The device uses OCR (Optical Character Recognition) to convert captured visual information into text data. The input is image data from a visual device, which is then converted into text data using optical character recognition software. The output is the text data sent to a cloud server.

[0122] Step 3:

[0123] The server receives text data and performs analysis using natural language processing algorithms. The input is text data sent from the terminal, and the analysis extracts highly relevant keywords. The output consists of keywords and contextual information.

[0124] Step 4:

[0125] The device sends behavioral data (e.g., location information, app usage history) to the server. The input is data showing the user's daily activities, and the output is behavioral data transferred to the server.

[0126] Step 5:

[0127] The server integrates the results of text data analysis with behavioral data to build user profiles. The input consists of analyzed text data and behavioral data, and a generative AI model is used to construct the profiles. The output is a profile tailored to each individual user.

[0128] Step 6:

[0129] The server generates user-optimized information or advertisements and constructs notification prompts. Inputs are user profiles and analytics data, and output is a notification prompt sent to the user's device.

[0130] Step 7:

[0131] The user terminal displays relevant information and advertisements to the user in real time based on the notification prompt message received. The input is a prompt message from the server, and the output is an information notification to the user.

[0132] Step 8:

[0133] Users provide feedback on the information or advertisements they are given. The input is the user's reaction or opinion, and the output is the feedback data sent to the server.

[0134] Step 9:

[0135] The server evaluates the feedback data, calculates rewards for information providers and application developers, and generates new prompt messages as needed. The input is the feedback data, and the output is reward information and improved prompt messages.

[0136] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0137] This invention is based on a system that utilizes a device such as smart glasses as a wearable terminal for the user to capture visual information and process it as digital data. The smart glasses are equipped with cameras and sensors to capture the user's visual information in real time. The captured information is converted into text data using OCR technology.

[0138] This digital data is sent to a server via a secure network. The server analyzes this data using natural language processing technology and extracts highly relevant information. Based on this, the refined information is added to the user profile. In addition, the device (smartphone) collects user behavior data, including location information and behavioral patterns, and transfers it to the server.

[0139] In this embodiment, an emotion engine has been added. The emotion engine analyzes the user's emotional state from their facial expressions, voice data, and behavior to estimate the user's current psychological state. Emotional data is also used for analysis on the server and influences the generation of information.

[0140] The server integrates analyzed digital data, behavioral data, and emotional data to generate the most relevant information and suggestions for the user. In doing so, it adjusts the timing and content of notifications according to the user's emotional state, providing information in a more appropriate format.

[0141] For example, if the emotion engine detects an anxious emotional state while a user is viewing a transportation advertisement, the server will provide suggestions to alleviate the user's anxiety, such as "the status of the next train" or "transfer information." The user will receive this notification on their smartphone and can adjust their behavior accordingly.

[0142] Ultimately, if a user provides feedback on the information they are notified of, that feedback is collected on the server and used to evaluate the rewards given to the provider and app developer. In this way, by linking digital, behavioral, and emotional information, a system is created that maximizes user convenience.

[0143] The following describes the processing flow.

[0144] Step 1:

[0145] The device (smart glasses) captures the user's visual information with a camera. Furthermore, it collects the user's facial expressions and voice using built-in sensors. The captured visual information is converted into text data using OCR (Optical Character Recognition).

[0146] Step 2:

[0147] The device (smart glasses) wirelessly transmits generated text data and sensor data to the server. This transmission is encrypted, ensuring the security of the information.

[0148] Step 3:

[0149] The server analyzes the received text data using natural language processing techniques. This process extracts highly relevant keywords and context.

[0150] Step 4:

[0151] The server analyzes sensor data sent from the terminal using an emotion engine to estimate the user's emotional state. For example, it recognizes joy or anxiety from changes in voice tone and facial expressions.

[0152] Step 5:

[0153] The device (smartphone) collects user behavior data, such as location information and app usage history, and sends it to the server. Location information is obtained using GPS.

[0154] Step 6:

[0155] The server integrates text data, sentiment data, and behavioral data to generate optimal information and suggestions based on the user's interests and needs. In this process, the user's emotional state is a factor that adjusts the priority of information generation.

[0156] Step 7:

[0157] The device (smartphone) will send notifications at the optimal time based on the user's emotions and behavior. These notifications will include information that alleviates the user's anxiety and suggestions that pique their interest.

[0158] Step 8:

[0159] Users can take action based on information received on their smartphones and provide feedback on that information. This feedback is provided within the app.

[0160] Step 9:

[0161] The server aggregates user feedback and evaluates the contributions of information providers and application developers. This evaluation is used to distribute rewards based on user satisfaction and the usefulness of the information.

[0162] (Example 2)

[0163] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0164] In modern society, people are surrounded by a vast amount of information, and are required to appropriately select and utilize the information that is most relevant to them. However, providing information that responds to the user's emotions and the specific situation is difficult. Furthermore, it is necessary to improve the quality of information provision by effectively utilizing feedback obtained from users. Conventional technologies have not been able to adequately address these challenges.

[0165] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0166] In this invention, the server includes means for converting visual information acquired by a visual device worn by the user into digital data, means for analyzing the digital data and extracting relevant information, and emotion analysis means for analyzing and estimating the user's emotional state from facial expressions, voice, and behavior. This makes it possible to provide information that is tailored to the user's emotional state and situation, thereby increasing the relevance and usefulness of the information.

[0167] A "visual device" is a device worn by a user that is equipped with cameras and sensors to acquire visual information in real time.

[0168] "Digital data" refers to data obtained by visual devices, which is electronically represented and converted into an analyzable format.

[0169] "Analysis" is a series of processes that involve processing digital data and extracting meaningful information.

[0170] "Related information" refers to information that is useful or necessary to the user, obtained as a result of the analysis.

[0171] A "communication device" is an integrated system of hardware and software used to collect user behavior and location information and transmit it to a server.

[0172] "Behavioral data" refers to information that includes a user's movement history, location information, and time-dependent behavioral patterns.

[0173] "Emotional analysis" is a process that estimates a user's psychological state based on facial expressions, voice, and behavioral data.

[0174] "Information provision" refers to the act of providing data related to the user's situation, which is generated by the server and notified to the user.

[0175] "Feedback" refers to the evaluations and reactions that users give to the information they receive, and this data is used to improve future information provision.

[0176] This invention provides a system that allows users to efficiently acquire information and utilize its content according to their own situation and psychological state. The system uses a visual device worn by the user, such as a smart device, to convert visual information into digital data in real time.

[0177] The user wears a smart device and acquires visual information using its camera and sensors. The acquired information is converted into text data by the device's built-in OCR (Optical Character Recognition) technology. This digital data is transmitted to a server via a secure communication protocol.

[0178] The server analyzes the received text data using natural language processing techniques to extract highly relevant information. This process also takes into account behavioral data and psychological state. Behavioral data is obtained when the device collects user location information and behavioral patterns and sends them to the server. Furthermore, an emotion analysis engine estimates emotions based on facial expressions, voice, and behavior, and incorporates this into the information generation process.

[0179] The server uses a generative AI model to integrate digital data, behavioral data, and emotional data to generate information and suggestions tailored to the user. This allows the server to notify the user at the optimal time and in the most appropriate way.

[0180] As a concrete example, when a user sees an advertisement in a certain location, the device acquires data at that moment and sends it to the server. The server analyzes the information contained in the advertisement, and if the user feels uneasy about that information, the emotion engine detects this and provides information that alleviates the anxiety, such as "details of the next event" or "relevant traffic information."

[0181] Users can receive notifications on their smartphones, quickly check the information, and adjust their actions as needed. Users can also provide feedback on the information provided, and the server uses this feedback to further improve future information delivery.

[0182] As an example of a prompt, you can enter something like, "Please suggest restaurants based on places I've recently visited," and receive suggestions tailored to your history and situation.

[0183] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0184] Step 1:

[0185] The device captures the user's visual information using its camera. This input data is initially processed within the device and saved as image data in real time. Specifically, the camera automatically adjusts the focus and acquires an image with appropriate exposure.

[0186] Step 2:

[0187] The device converts captured image data into text data using OCR technology. It receives image data as input, the OCR software analyzes the text within the image, and outputs text data. Specifically, the device's processor performs high-precision character recognition.

[0188] Step 3:

[0189] The terminal sends the converted text data to the server using a secure network protocol. The input is text data, and the output is the transfer of data to the server. Specifically, the terminal's communication module packets the data, encrypts it, and transmits it.

[0190] Step 4:

[0191] The server analyzes the received text data using natural language processing and extracts relevant information. In this process, it uses the received text data as input and generates useful information for the user as output. Specifically, the NLP engine parses the data and analyzes keywords and context.

[0192] Step 5:

[0193] The device collects the user's location data and behavioral patterns and sends them to the server. In this step, the device acquires data from location sensors and accelerometers and sends it to the server. Specifically, the device obtains its location from GPS and calculates its speed of movement from the time difference.

[0194] Step 6:

[0195] The server estimates the user's emotional state from facial expressions, voice data, and behavior. It receives emotion-related data as input and generates an estimated emotional state as output. Specifically, the emotion analysis engine uses machine learning algorithms to evaluate the data.

[0196] Step 7:

[0197] The server integrates the analyzed data and uses a generative AI model to generate information best suited to the user. It uses digital data, behavioral data, and emotional data as input and creates notification information as output. Specifically, the AI ​​model integrates the data and selects recommended information.

[0198] Step 8:

[0199] The server adjusts the timing and content of notifications according to the user's emotional state and sends the information to the device. Using the generated information as input, the adjusted notification is delivered to the user as output. Specifically, the server sets the notification priority and determines the sending time to the device.

[0200] Step 9:

[0201] The user modifies their actions based on the received notification and sends necessary feedback back to the server. The system receives notification information as input and sends feedback data as output. Specifically, the user checks the notification on their smartphone and enters their opinion on the evaluation feedback screen.

[0202] (Application Example 2)

[0203] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0204] Existing information delivery systems struggle to tailor information to the user's current emotional state. As a result, the information provided may not align with the user's needs and emotions, potentially leading to a diminished user experience. The challenge lies in resolving this issue and achieving personalized information delivery based on each user's individual circumstances and emotions.

[0205] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0206] In this invention, the server includes means for converting visual information acquired by a device worn by the user into digital data, means for integrating the visual information and behavioral data to detect the user's emotional state, and means for adjusting the information generated based on the user's emotional state. This enables the provision of personalized information that is tailored to the user's emotional state and needs.

[0207] A "user-worn device" is a device that a user can wear and that is equipped with sensors or cameras for acquiring visual information.

[0208] "Methods for converting visual information into digital data" refer to technologies that process images and videos acquired from the user's perspective and transform them into data in a format that can be processed by a computer.

[0209] "Means for extracting relevant information" refers to methods for analyzing acquired digital data and identifying and extracting data that is useful or meaningful to the user.

[0210] A "communication device" is a device that collects information and transfers user behavior data and digital data to a server via a network.

[0211] A "means for detecting emotional state" refers to a system used to analyze a user's facial expressions, voice, and behavioral data to evaluate their current psychological state.

[0212] "Means of adjusting information" refers to the process of changing the content and timing of information provided according to the user's emotions and circumstances.

[0213] "Feedback" refers to the reactions and evaluations that users give to the information and suggestions they provide, and this feedback is used to improve the system and evaluate the information.

[0214] This invention utilizes smart glasses, a smartphone, and a server to realize the entire system. The smart glasses are worn by the user and acquire the user's visual information using the built-in camera and sensors. This visual information is converted into digital data and transmitted to the smartphone for subsequent processing.

[0215] The smartphone, acting as the terminal, converts acquired visual information into text data using OCR technology. In addition, the smartphone functions as a communication device, transmitting the user's location information and behavioral patterns to the server. The server integrates this digital data with the user's behavioral data and performs analysis using natural language processing technology. Based on the analysis results, it extracts highly relevant information and generates information tailored to the user.

[0216] Furthermore, by using an emotion engine, the server can analyze the user's emotional state from their facial expressions and voice data, enabling the provision of emotionally sensitive information. For example, if a user is feeling anxious about a product they are considering purchasing, the server can notify them of reviews and related reassuring information about that product.

[0217] As a concrete example of a prompt message, the generation AI model is instructed to "analyze the emotions the user is experiencing while viewing the product, and create prompts to provide information according to those emotions," thereby supporting the generation of appropriate information.

[0218] In this way, providing personalized information tailored to the user's emotions and situation in real time improves the user experience.

[0219] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0220] Step 1:

[0221] The user wears smart glasses to acquire visual information. The smart glasses use their built-in camera to capture images of the surroundings. This acquired image data becomes the input.

[0222] Step 2:

[0223] The smartphone, acting as the terminal, converts image data received from the smart glasses into text data using OCR technology. The input is image data, and the output is text data based on the OCR processing.

[0224] Step 3:

[0225] The terminal sends text data to the server. The server receives this text data and performs analysis using natural language processing technology. The input is text data, and the analysis outputs highly relevant information.

[0226] Step 4:

[0227] The server analyzes the user's emotional state using an emotion engine based on their visual data and location information. The input consists of visual data and location information, and the emotional state is output as a result of the emotion analysis.

[0228] Step 5:

[0229] The server considers the emotional state, generates relevant information, and adjusts the information it notifies. The input is the analysis result and the emotional state, and the generated appropriate information is output.

[0230] Step 6:

[0231] The user's smartphone receives and displays the adjusted information notification. The input is adjusted information, and the notification is output to the user.

[0232] Step 7:

[0233] Users provide feedback on the information they receive. This feedback is sent to the server and used as input. The server uses this feedback to evaluate the provided information and improve the system. This process improves the quality of the information provided as output.

[0234] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0235] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0236] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0237] [Second Embodiment]

[0238] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0239] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0240] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0241] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0242] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0243] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0244] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0245] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0246] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0247] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0248] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0249] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0250] To implement this invention, it is necessary to establish a system in which the user wears a device such as smart glasses and takes in visual information on a daily basis. These smart glasses have a built-in camera that captures the information the user sees. The captured visual information is converted into digital data.

[0251] Smart glasses use OCR (Optical Character Recognition) to convert visual information into text data. This text data is then transmitted wirelessly to a server in the cloud. The server receives this data and analyzes the information using natural language processing algorithms. From the analysis results, it extracts highly relevant keywords and context.

[0252] Meanwhile, the user's smartphone collects behavioral data, including location information and app usage history. This data is also transferred to a server. The server integrates the behavioral data from the smartphone with the analyzed digital data from the smart glasses to build a user profile. Using this profile, the server selects the most relevant information for the user and makes suggestions at the appropriate time.

[0253] For example, suppose a user is walking around town and their smart glasses capture a promotional banner for a new restaurant. The server analyzes this information and, based on the user's current location and past dining preferences, notifies their smartphone of a lunch coupon for that specific restaurant.

[0254] Users who receive a notification can send feedback on the information through the app on their smartphone. This feedback data is also sent to the server and used to reward information providers and application developers.

[0255] In this way, the invention is implemented in a manner that blurs the boundaries between analog and digital information, improves the user experience, and strengthens incentives for information providers.

[0256] The following describes the processing flow.

[0257] Step 1:

[0258] The device (smart glasses) captures analog information seen by the user with a camera and converts that image into digital text data using OCR technology.

[0259] Step 2:

[0260] The device (smart glasses) sends the generated text data to the server via the internet. The transmission is performed using a secure protocol.

[0261] Step 3:

[0262] The server analyzes the received text data and extracts key keywords and context using natural language processing techniques. This process helps to understand the meaning of the data.

[0263] Step 4:

[0264] The device (smartphone) collects the user's current location information and past activity history, and sends this data to the server. Location information is obtained using GPS.

[0265] Step 5:

[0266] The server integrates the analyzed text data and behavioral data sent from the terminal and updates the user profile to estimate the user's interests.

[0267] Step 6:

[0268] The server generates optimal information and suggestions for the user based on integrated data and notifies the user of this information via their device (smartphone). The timing of these notifications is adjusted according to the user's situation.

[0269] Step 7:

[0270] Users provide feedback on the usefulness and interest of the information they receive through notifications on their smartphones. This helps to improve the user's future information delivery.

[0271] Step 8:

[0272] The server collects user feedback and uses it as data to distribute rewards to information providers and application developers. The reward criteria are based on the evaluation of the feedback.

[0273] (Example 1)

[0274] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0275] In recent years, there has been a growing demand for services that digitize users' visual information from their daily lives and activities and use that information to provide detailed, relevant suggestions. However, relying solely on visual information from users makes it difficult to accurately provide highly relevant information, resulting in a lack of improvement in the user experience. Furthermore, if the evaluation of information provision is not fair, incentives for information providers may not function properly. This invention aims to solve these problems and build an information provision system that is more beneficial to users.

[0276] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0277] In this invention, the server includes means for converting visual information acquired by a camera worn by the user into electronic data, natural language processing means for analyzing the electronic data and extracting relevant information, and means for a communication device to collect the user's location data and usage history. This enables the integration and analysis of the user's visual information and behavioral data, allowing for more personalized information provision and fair reward evaluation for information providers.

[0278] The "user" refers to an individual or organization that uses this system.

[0279] The "imaging device" refers to a device having a function of capturing visual information, specifically a wearable device incorporating a camera.

[0280] The "electronic data" refers to a data format including visual information digitized by an imaging device.

[0281] The "natural language processing means" refers to algorithms and software tools for analyzing electronic data and extracting and classifying language data.

[0282] The "communication device" refers to a device having the ability to transmit and receive data via the Internet or a local network.

[0283] The "location data" is information indicating the geographical location of the user, which is obtained by means such as GPS.

[0284] The "usage history" refers to records of operations performed, applications used, locations visited, etc. by the user in the past.

[0285] The "profile generation means" refers to a method or device for generating detailed information about the user based on the collected electronic data and behavioral data.

[0286] The "response" refers to feedback or evaluation from the user regarding the generated information.

[0287] The "reward" indicates a consideration commensurate with the value of the provided information for the information provider and the developer of the application.

[0288] To implement this invention, it is fundamental that the user wears smart glasses, which are a camera, to collect everyday visual information. The smart glasses have a built-in camera that acquires information about the outside world as image data. This image data is converted into electronic data using OCR (Optical Character Recognition) technology. An OCR engine such as Tesseract can be used for this conversion.

[0289] The server receives the converted electronic data and performs analysis using natural language processing (NLP) algorithms. This analysis utilizes natural language processing libraries such as spaCy and NLTK to extract relevant information from the electronic data. This yields keywords and sentences relevant to the user's context, preparing the data for information provision.

[0290] Devices, especially smartphones, collect user location data via GPS and also acquire application usage history data. This information is sent to a server, where a profile generation mechanism constructs data about the user. Based on the profile, the server selects the most relevant information for the user and notifies the smartphone.

[0291] As a concrete example, suppose a user is walking through a shopping mall and their smart glasses capture an advertisement for a newly opened store. The server analyzes this advertisement information and, based on the user's current location and shopping preferences, sends a discount coupon for that store to their smartphone.

[0292] This invention enables more personalized information delivery by utilizing users' visual information and behavioral data. An example of a prompt sentence using the generative AI model is, "Please tell me how to suggest products that might interest you based on the advertisements you've seen."

[0293] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0294] Step 1:

[0295] The user wears smart glasses, which are a recording device, to acquire visual information. These glasses have a built-in camera that captures information within the user's field of view as image data. The input is analog visual information, and the output is digital image data. This allows various types of information obtained during the user's daily activities to be digitized.

[0296] Step 2:

[0297] Smart glasses convert acquired image data into electronic data using OCR (Optical Character Recognition) technology. Specifically, they analyze text within images using OCR engines such as Tesseract and extract character information. The input is digital image data, and the output is electronic data in text format. This process makes it possible to effectively extract character information from images.

[0298] Step 3:

[0299] The smart glasses transmit the converted text data to the server via wireless communication. The server takes the received text data as input and analyzes it using natural language processing algorithms. Specifically, it uses libraries such as spaCy and NLTK to perform semantic analysis of the data and extract highly relevant keywords and contexts. The output is the analyzed keyword and context information. This allows the server to obtain detailed information about the user's intent and interests.

[0300] Step 4:

[0301] On the other hand, the smartphone, as the terminal, obtains location information from GPS and collects past app usage history. The input is behavioral data obtained from the terminal, and the output is specific location information and history data that shows the user's behavior patterns. This collected data is transmitted to the server in real time.

[0302] Step 5:

[0303] The server integrates the behavior data from the smartphone and the electronic data analyzed from the smart glasses. The input is the analysis data obtained in Step 3 and Step 4, and the output is the user profile. Based on this, the server selects the information most relevant to the user and configures it to enable rapid response.

[0304] Step 6:

[0305] The server notifies the user's smartphone of the generated information. A push notification service is used for this. The input is the user profile and the selected information, and the output is the specific notification content to the user. The user can receive this information and decide on the next action according to the prompt.

[0306] (Application Example 1)

[0307] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0308] In modern society, while users receive a large amount of information daily, that information is not always optimized for the users' interests and situations. Especially in the fields of advertising and marketing, there is a demand to provide information in real time that is in line with the users' interests and current situations. Furthermore, the efficiency of the reward system for information providers and application producers is also an issue that needs to be addressed.

[0309] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0310] In this invention, the server includes means for converting visual information acquired by a visual device worn by the user into digital data, means for analyzing the digital data and extracting relevant information, and means for a communication device to collect user behavior data. This makes it possible to provide appropriate advertising information in real time based on the user's location information and interest data.

[0311] A "user" refers to a person who wears visual devices to receive information.

[0312] A "visual device" refers to a device worn by a user to acquire visual information, and is capable of converting that information into digital data.

[0313] "Digital data" refers to data obtained by visual devices that has been converted into a format that can be processed.

[0314] "Analysis methods" refer to processes and techniques for extracting relevant information from digital data.

[0315] A "communication device" refers to a device that collects user behavior data and transmits it to a server or other device.

[0316] "Behavioral data" refers to data that shows users' activities, location information, and interests.

[0317] "Notification means" refers to the mechanisms and functions used to deliver generated information to users.

[0318] "Feedback" refers to the evaluations and opinions that users give in response to the information provided.

[0319] "Promotional information" refers to advertisements and promotions provided based on the user's interests and location information.

[0320] "Real-time" refers to information processing and notifications occurring instantly and without delay.

[0321] The system for implementing this invention consists of a visual device worn by the user and a mobile terminal. The visual device is equipped with a camera that captures the user's visual information in real time. This information is converted into text data using built-in optical character recognition (OCR) software. The converted text data is transmitted wirelessly to a server in the cloud.

[0322] The server performs natural language processing on the received text data and analyzes relevant information. During this process, the server builds individual user profiles based on behavioral data transmitted from the user's mobile device, such as location information and past activity history. This allows for the generation of user-optimized information or advertisements, which are then notified to the user's mobile device in real time.

[0323] For example, if a user's visual device captures a sign for a new restaurant while walking around town, that information is immediately analyzed. The server considers the user's location and past dining patterns and sends a special lunch coupon related to that restaurant to the user's mobile device. This notification is timed appropriately so that the user receives the information at the right moment.

[0324] The hardware used will include a visual device such as Google Glass, and the software used for analysis will be Google Cloud NLP and OCR tools. The terminal is assumed to be a typical smartphone device. An example of a prompt message would be as follows:

[0325] Example of a prompt:

[0326] We have captured promotional information for new exhibitions that the user is interested in at their current location. Based on this information, generate a discount coupon for the relevant exhibition and notify the user on their mobile device.

[0327] This invention allows users to always receive information tailored to their needs, and enables information providers to efficiently reach their target users.

[0328] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0329] Step 1:

[0330] The user wears a visual device to capture visual information from their surroundings. This visual information, as input, is received by a camera sensor and converted into image data. The output is the image data sent to the OCR process.

[0331] Step 2:

[0332] The device uses OCR (Optical Character Recognition) to convert captured visual information into text data. The input is image data from a visual device, which is then converted into text data using optical character recognition software. The output is the text data sent to a cloud server.

[0333] Step 3:

[0334] The server receives text data and performs analysis using natural language processing algorithms. The input is text data sent from the terminal, and the analysis extracts highly relevant keywords. The output consists of keywords and contextual information.

[0335] Step 4:

[0336] The device sends behavioral data (e.g., location information, app usage history) to the server. The input is data showing the user's daily activities, and the output is behavioral data transferred to the server.

[0337] Step 5:

[0338] The server integrates the results of text data analysis with behavioral data to build user profiles. The input consists of analyzed text data and behavioral data, and a generative AI model is used to construct the profiles. The output is a profile tailored to each individual user.

[0339] Step 6:

[0340] The server generates user-optimized information or advertisements and constructs notification prompts. Inputs are user profiles and analytics data, and output is a notification prompt sent to the user's device.

[0341] Step 7:

[0342] The user terminal displays relevant information and advertisements to the user in real time based on the notification prompt message received. The input is a prompt message from the server, and the output is an information notification to the user.

[0343] Step 8:

[0344] Users provide feedback on the information or advertisements they are given. The input is the user's reaction or opinion, and the output is the feedback data sent to the server.

[0345] Step 9:

[0346] The server evaluates the feedback data, calculates rewards for information providers and application developers, and generates new prompt messages as needed. The input is the feedback data, and the output is reward information and improved prompt messages.

[0347] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0348] This invention is based on a system that utilizes a device such as smart glasses as a wearable terminal for the user to capture visual information and process it as digital data. The smart glasses are equipped with cameras and sensors to capture the user's visual information in real time. The captured information is converted into text data using OCR technology.

[0349] This digital data is sent to a server via a secure network. The server analyzes this data using natural language processing technology and extracts highly relevant information. Based on this, the refined information is added to the user profile. In addition, the device (smartphone) collects user behavior data, including location information and behavioral patterns, and transfers it to the server.

[0350] In this embodiment, an emotion engine has been added. The emotion engine analyzes the user's emotional state from their facial expressions, voice data, and behavior to estimate the user's current psychological state. Emotional data is also used for analysis on the server and influences the generation of information.

[0351] The server integrates analyzed digital data, behavioral data, and emotional data to generate the most relevant information and suggestions for the user. In doing so, it adjusts the timing and content of notifications according to the user's emotional state, providing information in a more appropriate format.

[0352] For example, if the emotion engine detects an anxious emotional state while a user is viewing a transportation advertisement, the server will provide suggestions to alleviate the user's anxiety, such as "the status of the next train" or "transfer information." The user will receive this notification on their smartphone and can adjust their behavior accordingly.

[0353] Ultimately, if a user provides feedback on the information they are notified of, that feedback is collected on the server and used to evaluate the rewards given to the provider and app developer. In this way, by linking digital, behavioral, and emotional information, a system is created that maximizes user convenience.

[0354] The following describes the processing flow.

[0355] Step 1:

[0356] The device (smart glasses) captures the user's visual information with a camera. Furthermore, it collects the user's facial expressions and voice using built-in sensors. The captured visual information is converted into text data using OCR (Optical Character Recognition).

[0357] Step 2:

[0358] The device (smart glasses) wirelessly transmits generated text data and sensor data to the server. This transmission is encrypted, ensuring the security of the information.

[0359] Step 3:

[0360] The server analyzes the received text data using natural language processing techniques. This process extracts highly relevant keywords and context.

[0361] Step 4:

[0362] The server analyzes sensor data sent from the terminal using an emotion engine to estimate the user's emotional state. For example, it recognizes joy or anxiety from changes in voice tone and facial expressions.

[0363] Step 5:

[0364] The device (smartphone) collects user behavior data, such as location information and app usage history, and sends it to the server. Location information is obtained using GPS.

[0365] Step 6:

[0366] The server integrates text data, sentiment data, and behavioral data to generate optimal information and suggestions based on the user's interests and needs. In this process, the user's emotional state is a factor that adjusts the priority of information generation.

[0367] Step 7:

[0368] The device (smartphone) will send notifications at the optimal time based on the user's emotions and behavior. These notifications will include information that alleviates the user's anxiety and suggestions that pique their interest.

[0369] Step 8:

[0370] Users can take action based on information received on their smartphones and provide feedback on that information. This feedback is provided within the app.

[0371] Step 9:

[0372] The server aggregates user feedback and evaluates the contributions of information providers and application developers. This evaluation is used to distribute rewards based on user satisfaction and the usefulness of the information.

[0373] (Example 2)

[0374] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0375] In modern society, people are surrounded by a vast amount of information, and are required to appropriately select and utilize the information that is most relevant to them. However, providing information that responds to the user's emotions and the specific situation is difficult. Furthermore, it is necessary to improve the quality of information provision by effectively utilizing feedback obtained from users. Conventional technologies have not been able to adequately address these challenges.

[0376] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0377] In this invention, the server includes means for converting visual information acquired by a visual device worn by the user into digital data, means for analyzing the digital data and extracting relevant information, and emotion analysis means for analyzing and estimating the user's emotional state from facial expressions, voice, and behavior. This makes it possible to provide information that is tailored to the user's emotional state and situation, thereby increasing the relevance and usefulness of the information.

[0378] A "visual device" is a device worn by a user that is equipped with cameras and sensors to acquire visual information in real time.

[0379] "Digital data" refers to data obtained by visual devices, which is electronically represented and converted into an analyzable format.

[0380] "Analysis" is a series of processes that involve processing digital data and extracting meaningful information.

[0381] "Related information" refers to information that is useful or necessary to the user, obtained as a result of the analysis.

[0382] A "communication device" is an integrated system of hardware and software used to collect user behavior and location information and transmit it to a server.

[0383] "Behavioral data" refers to information that includes a user's movement history, location information, and time-dependent behavioral patterns.

[0384] "Emotional analysis" is a process that estimates a user's psychological state based on facial expressions, voice, and behavioral data.

[0385] "Information provision" refers to the act of providing data related to the user's situation, which is generated by the server and notified to the user.

[0386] "Feedback" refers to the evaluations and reactions that users give to the information they receive, and this data is used to improve future information provision.

[0387] This invention provides a system that allows users to efficiently acquire information and utilize its content according to their own situation and psychological state. The system uses a visual device worn by the user, such as a smart device, to convert visual information into digital data in real time.

[0388] The user wears a smart device and acquires visual information using its camera and sensors. The acquired information is converted into text data by the device's built-in OCR (Optical Character Recognition) technology. This digital data is transmitted to a server via a secure communication protocol.

[0389] The server analyzes the received text data using natural language processing techniques to extract highly relevant information. This process also takes into account behavioral data and psychological state. Behavioral data is obtained when the device collects user location information and behavioral patterns and sends them to the server. Furthermore, an emotion analysis engine estimates emotions based on facial expressions, voice, and behavior, and incorporates this into the information generation process.

[0390] The server uses a generative AI model to integrate digital data, behavioral data, and emotional data to generate information and suggestions tailored to the user. This allows the server to notify the user at the optimal time and in the most appropriate way.

[0391] As a concrete example, when a user sees an advertisement in a certain location, the device acquires data at that moment and sends it to the server. The server analyzes the information contained in the advertisement, and if the user feels uneasy about that information, the emotion engine detects this and provides information that alleviates the anxiety, such as "details of the next event" or "relevant traffic information."

[0392] Users can receive notifications on their smartphones, quickly check the information, and adjust their actions as needed. Users can also provide feedback on the information provided, and the server uses this feedback to further improve future information delivery.

[0393] As an example of a prompt, you can enter something like, "Please suggest restaurants based on places I've recently visited," and receive suggestions tailored to your history and situation.

[0394] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0395] Step 1:

[0396] The device captures the user's visual information using its camera. This input data is initially processed within the device and saved as image data in real time. Specifically, the camera automatically adjusts the focus and acquires an image with appropriate exposure.

[0397] Step 2:

[0398] The device converts captured image data into text data using OCR technology. It receives image data as input, the OCR software analyzes the text within the image, and outputs text data. Specifically, the device's processor performs high-precision character recognition.

[0399] Step 3:

[0400] The terminal sends the converted text data to the server using a secure network protocol. The input is text data, and the output is the transfer of data to the server. Specifically, the terminal's communication module packets the data, encrypts it, and transmits it.

[0401] Step 4:

[0402] The server analyzes the received text data using natural language processing and extracts relevant information. In this process, it uses the received text data as input and generates useful information for the user as output. Specifically, the NLP engine parses the data and analyzes keywords and context.

[0403] Step 5:

[0404] The device collects the user's location data and behavioral patterns and sends them to the server. In this step, the device acquires data from location sensors and accelerometers and sends it to the server. Specifically, the device obtains its location from GPS and calculates its speed of movement from the time difference.

[0405] Step 6:

[0406] The server estimates the user's emotional state from facial expressions, voice data, and behavior. It receives emotion-related data as input and generates an estimated emotional state as output. Specifically, the emotion analysis engine uses machine learning algorithms to evaluate the data.

[0407] Step 7:

[0408] The server integrates the analyzed data and uses a generative AI model to generate information best suited to the user. It uses digital data, behavioral data, and emotional data as input and creates notification information as output. Specifically, the AI ​​model integrates the data and selects recommended information.

[0409] Step 8:

[0410] The server adjusts the timing and content of notifications according to the user's emotional state and sends the information to the device. Using the generated information as input, the adjusted notification is delivered to the user as output. Specifically, the server sets the notification priority and determines the sending time to the device.

[0411] Step 9:

[0412] The user modifies their actions based on the received notification and sends necessary feedback back to the server. The system receives notification information as input and sends feedback data as output. Specifically, the user checks the notification on their smartphone and enters their opinion on the evaluation feedback screen.

[0413] (Application Example 2)

[0414] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0415] Existing information delivery systems struggle to tailor information to the user's current emotional state. As a result, the information provided may not align with the user's needs and emotions, potentially leading to a diminished user experience. The challenge lies in resolving this issue and achieving personalized information delivery based on each user's individual circumstances and emotions.

[0416] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0417] In this invention, the server includes means for converting visual information acquired by a device worn by the user into digital data, means for integrating the visual information and behavioral data to detect the user's emotional state, and means for adjusting the information generated based on the user's emotional state. This enables the provision of personalized information that is tailored to the user's emotional state and needs.

[0418] A "user-worn device" is a device that a user can wear and that is equipped with sensors or cameras for acquiring visual information.

[0419] "Methods for converting visual information into digital data" refer to technologies that process images and videos acquired from the user's perspective and transform them into data in a format that can be processed by a computer.

[0420] "Means for extracting relevant information" refers to methods for analyzing acquired digital data and identifying and extracting data that is useful or meaningful to the user.

[0421] A "communication device" is a device that collects information and transfers user behavior data and digital data to a server via a network.

[0422] A "means for detecting emotional state" refers to a system used to analyze a user's facial expressions, voice, and behavioral data to evaluate their current psychological state.

[0423] "Means of adjusting information" refers to the process of changing the content and timing of information provided according to the user's emotions and circumstances.

[0424] "Feedback" refers to the reactions and evaluations that users give to the information and suggestions they provide, and this feedback is used to improve the system and evaluate the information.

[0425] This invention utilizes smart glasses, a smartphone, and a server to realize the entire system. The smart glasses are worn by the user and acquire the user's visual information using the built-in camera and sensors. This visual information is converted into digital data and transmitted to the smartphone for subsequent processing.

[0426] The smartphone, acting as the terminal, converts acquired visual information into text data using OCR technology. In addition, the smartphone functions as a communication device, transmitting the user's location information and behavioral patterns to the server. The server integrates this digital data with the user's behavioral data and performs analysis using natural language processing technology. Based on the analysis results, it extracts highly relevant information and generates information tailored to the user.

[0427] Furthermore, by using an emotion engine, the server can analyze the user's emotional state from their facial expressions and voice data, enabling the provision of emotionally sensitive information. For example, if a user is feeling anxious about a product they are considering purchasing, the server can notify them of reviews and related reassuring information about that product.

[0428] As a concrete example of a prompt message, the generation AI model is instructed to "analyze the emotions the user is experiencing while viewing the product, and create prompts to provide information according to those emotions," thereby supporting the generation of appropriate information.

[0429] In this way, providing personalized information tailored to the user's emotions and situation in real time improves the user experience.

[0430] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0431] Step 1:

[0432] The user wears smart glasses to acquire visual information. The smart glasses use their built-in camera to capture images of the surroundings. This acquired image data becomes the input.

[0433] Step 2:

[0434] The smartphone, acting as the terminal, converts image data received from the smart glasses into text data using OCR technology. The input is image data, and the output is text data based on the OCR processing.

[0435] Step 3:

[0436] The terminal sends text data to the server. The server receives this text data and performs analysis using natural language processing technology. The input is text data, and the analysis outputs highly relevant information.

[0437] Step 4:

[0438] The server analyzes the user's emotional state using an emotion engine based on their visual data and location information. The input consists of visual data and location information, and the emotional state is output as a result of the emotion analysis.

[0439] Step 5:

[0440] The server considers the emotional state, generates relevant information, and adjusts the information it notifies. The input is the analysis result and the emotional state, and the generated appropriate information is output.

[0441] Step 6:

[0442] The user's smartphone receives and displays the adjusted information notification. The input is adjusted information, and the notification is output to the user.

[0443] Step 7:

[0444] Users provide feedback on the information they receive. This feedback is sent to the server and used as input. The server uses this feedback to evaluate the provided information and improve the system. This process improves the quality of the information provided as output.

[0445] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0446] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0448] [Third Embodiment]

[0449] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0450] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0452] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0456] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0459] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0460] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0461] To implement this invention, it is necessary to establish a system in which the user wears a device such as smart glasses and takes in visual information on a daily basis. These smart glasses have a built-in camera that captures the information the user sees. The captured visual information is converted into digital data.

[0462] Smart glasses use OCR (Optical Character Recognition) to convert visual information into text data. This text data is then transmitted wirelessly to a server in the cloud. The server receives this data and analyzes the information using natural language processing algorithms. From the analysis results, it extracts highly relevant keywords and context.

[0463] Meanwhile, the user's smartphone collects behavioral data, including location information and app usage history. This data is also transferred to a server. The server integrates the behavioral data from the smartphone with the analyzed digital data from the smart glasses to build a user profile. Using this profile, the server selects the most relevant information for the user and makes suggestions at the appropriate time.

[0464] For example, suppose a user is walking around town and their smart glasses capture a promotional banner for a new restaurant. The server analyzes this information and, based on the user's current location and past dining preferences, notifies their smartphone of a lunch coupon for that specific restaurant.

[0465] Users who receive a notification can send feedback on the information through the app on their smartphone. This feedback data is also sent to the server and used to reward information providers and application developers.

[0466] In this way, the invention is implemented in a manner that blurs the boundaries between analog and digital information, improves the user experience, and strengthens incentives for information providers.

[0467] The following describes the processing flow.

[0468] Step 1:

[0469] The device (smart glasses) captures analog information seen by the user with a camera and converts that image into digital text data using OCR technology.

[0470] Step 2:

[0471] The device (smart glasses) sends the generated text data to the server via the internet. The transmission is performed using a secure protocol.

[0472] Step 3:

[0473] The server analyzes the received text data and extracts key keywords and context using natural language processing techniques. This process helps to understand the meaning of the data.

[0474] Step 4:

[0475] The device (smartphone) collects the user's current location information and past activity history, and sends this data to the server. Location information is obtained using GPS.

[0476] Step 5:

[0477] The server integrates the analyzed text data and behavioral data sent from the terminal and updates the user profile to estimate the user's interests.

[0478] Step 6:

[0479] The server generates optimal information and suggestions for the user based on integrated data and notifies the user of this information via their device (smartphone). The timing of these notifications is adjusted according to the user's situation.

[0480] Step 7:

[0481] Users provide feedback on the usefulness and interest of the information they receive through notifications on their smartphones. This helps to improve the user's future information delivery.

[0482] Step 8:

[0483] The server collects user feedback and uses it as data to distribute rewards to information providers and application developers. The reward criteria are based on the evaluation of the feedback.

[0484] (Example 1)

[0485] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0486] In recent years, there has been a growing demand for services that digitize users' visual information from their daily lives and activities and use that information to provide detailed, relevant suggestions. However, relying solely on visual information from users makes it difficult to accurately provide highly relevant information, resulting in a lack of improvement in the user experience. Furthermore, if the evaluation of information provision is not fair, incentives for information providers may not function properly. This invention aims to solve these problems and build an information provision system that is more beneficial to users.

[0487] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0488] In this invention, the server includes means for converting visual information acquired by a camera worn by the user into electronic data, natural language processing means for analyzing the electronic data and extracting relevant information, and means for a communication device to collect the user's location data and usage history. This enables the integration and analysis of the user's visual information and behavioral data, allowing for more personalized information provision and fair reward evaluation for information providers.

[0489] "User" refers to an individual or group that uses this system.

[0490] A "photography device" refers to a device that has the function of capturing visual information, and specifically refers to a wearable device with a built-in camera.

[0491] "Electronic data" refers to a data format that includes visual information digitized by a photographic device.

[0492] "Natural language processing means" refers to algorithms and software tools used to analyze electronic data and extract or classify language data.

[0493] "Communication equipment" refers to devices that have the ability to send and receive data via the internet or a local network.

[0494] "Location data" refers to information indicating the user's geographical location, and is obtained through methods such as GPS.

[0495] "Usage history" refers to records of actions a user has taken in the past, applications they have used, places they have visited, and so on.

[0496] "Profile generation means" refers to a method or apparatus for generating detailed information about a user based on collected electronic data and behavioral data.

[0497] "Response" refers to feedback and evaluation from users regarding the generated information.

[0498] "Compensation" refers to the payment given to information providers or application developers that is commensurate with the value of the information they provide.

[0499] To implement this invention, it is fundamental that the user wears smart glasses, which are a camera, to collect everyday visual information. The smart glasses have a built-in camera that acquires information about the outside world as image data. This image data is converted into electronic data using OCR (Optical Character Recognition) technology. An OCR engine such as Tesseract can be used for this conversion.

[0500] The server receives the converted electronic data and performs analysis using natural language processing (NLP) algorithms. This analysis utilizes natural language processing libraries such as spaCy and NLTK to extract relevant information from the electronic data. This yields keywords and sentences relevant to the user's context, preparing the data for information provision.

[0501] Devices, especially smartphones, collect user location data via GPS and also acquire application usage history data. This information is sent to a server, where a profile generation mechanism constructs data about the user. Based on the profile, the server selects the most relevant information for the user and notifies the smartphone.

[0502] As a concrete example, suppose a user is walking through a shopping mall and their smart glasses capture an advertisement for a newly opened store. The server analyzes this advertisement information and, based on the user's current location and shopping preferences, sends a discount coupon for that store to their smartphone.

[0503] This invention enables more personalized information delivery by utilizing users' visual information and behavioral data. An example of a prompt sentence using the generative AI model is, "Please tell me how to suggest products that might interest you based on the advertisements you've seen."

[0504] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0505] Step 1:

[0506] The user wears smart glasses, which are a recording device, to acquire visual information. These glasses have a built-in camera that captures information within the user's field of view as image data. The input is analog visual information, and the output is digital image data. This allows various types of information obtained during the user's daily activities to be digitized.

[0507] Step 2:

[0508] Smart glasses convert acquired image data into electronic data using OCR (Optical Character Recognition) technology. Specifically, they analyze text within images using OCR engines such as Tesseract and extract character information. The input is digital image data, and the output is electronic data in text format. This process makes it possible to effectively extract character information from images.

[0509] Step 3:

[0510] The smart glasses transmit the converted text data to the server via wireless communication. The server takes the received text data as input and analyzes it using natural language processing algorithms. Specifically, it uses libraries such as spaCy and NLTK to perform semantic analysis of the data and extract highly relevant keywords and contexts. The output is the analyzed keyword and context information. This allows the server to obtain detailed information about the user's intent and interests.

[0511] Step 4:

[0512] On the other hand, the smartphone, as the terminal, obtains location information from GPS and collects past app usage history. The input is behavioral data obtained from the terminal, and the output is specific location information and history data that shows the user's behavior patterns. This collected data is transmitted to the server in real time.

[0513] Step 5:

[0514] The server integrates behavioral data from smartphones with electronic data analyzed from smart glasses. The input is the analyzed data obtained in steps 3 and 4, and the output is a user profile. Based on this, the server selects the most relevant information for the user and configures itself to respond quickly.

[0515] Step 6:

[0516] The server notifies the user's smartphone of the generated information. This uses a push notification service, with the user profile and selection information as input, and the output being the specific notification content for the user. The user receives this information and can decide on their next action according to the prompt.

[0517] (Application Example 1)

[0518] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0519] In modern society, users receive vast amounts of information daily, and this information is not always optimized to their interests or circumstances. In particular, in the fields of advertising and marketing, there is a demand for providing information tailored to users' interests and current situations in real time. Furthermore, streamlining the compensation system for information providers and those who create and apply the information is also a crucial challenge.

[0520] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0521] In this invention, the server includes means for converting visual information acquired by a visual device worn by the user into digital data, means for analyzing the digital data and extracting relevant information, and means for a communication device to collect user behavior data. This makes it possible to provide appropriate advertising information in real time based on the user's location information and interest data.

[0522] A "user" refers to a person who wears visual devices to receive information.

[0523] A "visual device" refers to a device worn by a user to acquire visual information, and is capable of converting that information into digital data.

[0524] "Digital data" refers to data obtained by visual devices that has been converted into a format that can be processed.

[0525] "Analysis methods" refer to processes and techniques for extracting relevant information from digital data.

[0526] A "communication device" refers to a device that collects user behavior data and transmits it to a server or other device.

[0527] "Behavioral data" refers to data that shows users' activities, location information, and interests.

[0528] "Notification means" refers to the mechanisms and functions used to deliver generated information to users.

[0529] "Feedback" refers to the evaluations and opinions that users give in response to the information provided.

[0530] "Promotional information" refers to advertisements and promotions provided based on the user's interests and location information.

[0531] "Real-time" refers to information processing and notifications occurring instantly and without delay.

[0532] The system for implementing this invention consists of a visual device worn by the user and a mobile terminal. The visual device is equipped with a camera that captures the user's visual information in real time. This information is converted into text data using built-in optical character recognition (OCR) software. The converted text data is transmitted wirelessly to a server in the cloud.

[0533] The server performs natural language processing on the received text data and analyzes relevant information. During this process, the server builds individual user profiles based on behavioral data transmitted from the user's mobile device, such as location information and past activity history. This allows for the generation of user-optimized information or advertisements, which are then notified to the user's mobile device in real time.

[0534] For example, if a user's visual device captures a sign for a new restaurant while walking around town, that information is immediately analyzed. The server considers the user's location and past dining patterns and sends a special lunch coupon related to that restaurant to the user's mobile device. This notification is timed appropriately so that the user receives the information at the right moment.

[0535] The hardware used will include a visual device such as Google Glass, and the software used for analysis will be Google Cloud NLP and OCR tools. The terminal is assumed to be a typical smartphone device. An example of a prompt message would be as follows:

[0536] Example of a prompt:

[0537] We have captured promotional information for new exhibitions that the user is interested in at their current location. Based on this information, generate a discount coupon for the relevant exhibition and notify the user on their mobile device.

[0538] This invention allows users to always receive information tailored to their needs, and enables information providers to efficiently reach their target users.

[0539] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0540] Step 1:

[0541] The user wears a visual device to capture visual information from their surroundings. This visual information, as input, is received by a camera sensor and converted into image data. The output is the image data sent to the OCR process.

[0542] Step 2:

[0543] The device uses OCR (Optical Character Recognition) to convert captured visual information into text data. The input is image data from a visual device, which is then converted into text data using optical character recognition software. The output is the text data sent to a cloud server.

[0544] Step 3:

[0545] The server receives text data and performs analysis using natural language processing algorithms. The input is text data sent from the terminal, and the analysis extracts highly relevant keywords. The output consists of keywords and contextual information.

[0546] Step 4:

[0547] The device sends behavioral data (e.g., location information, app usage history) to the server. The input is data showing the user's daily activities, and the output is behavioral data transferred to the server.

[0548] Step 5:

[0549] The server integrates the results of text data analysis with behavioral data to build user profiles. The input consists of analyzed text data and behavioral data, and a generative AI model is used to construct the profiles. The output is a profile tailored to each individual user.

[0550] Step 6:

[0551] The server generates user-optimized information or advertisements and constructs notification prompts. Inputs are user profiles and analytics data, and output is a notification prompt sent to the user's device.

[0552] Step 7:

[0553] The user terminal displays relevant information and advertisements to the user in real time based on the notification prompt message received. The input is a prompt message from the server, and the output is an information notification to the user.

[0554] Step 8:

[0555] Users provide feedback on the information or advertisements they are given. The input is the user's reaction or opinion, and the output is the feedback data sent to the server.

[0556] Step 9:

[0557] The server evaluates the feedback data, calculates rewards for information providers and application developers, and generates new prompt messages as needed. The input is the feedback data, and the output is reward information and improved prompt messages.

[0558] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0559] This invention is based on a system that utilizes a device such as smart glasses as a wearable terminal for the user to capture visual information and process it as digital data. The smart glasses are equipped with cameras and sensors to capture the user's visual information in real time. The captured information is converted into text data using OCR technology.

[0560] This digital data is sent to a server via a secure network. The server analyzes this data using natural language processing technology and extracts highly relevant information. Based on this, the refined information is added to the user profile. In addition, the device (smartphone) collects user behavior data, including location information and behavioral patterns, and transfers it to the server.

[0561] In this embodiment, an emotion engine has been added. The emotion engine analyzes the user's emotional state from their facial expressions, voice data, and behavior to estimate the user's current psychological state. Emotional data is also used for analysis on the server and influences the generation of information.

[0562] The server integrates analyzed digital data, behavioral data, and emotional data to generate the most relevant information and suggestions for the user. In doing so, it adjusts the timing and content of notifications according to the user's emotional state, providing information in a more appropriate format.

[0563] For example, if the emotion engine detects an anxious emotional state while a user is viewing a transportation advertisement, the server will provide suggestions to alleviate the user's anxiety, such as "the status of the next train" or "transfer information." The user will receive this notification on their smartphone and can adjust their behavior accordingly.

[0564] Ultimately, if a user provides feedback on the information they are notified of, that feedback is collected on the server and used to evaluate the rewards given to the provider and app developer. In this way, by linking digital, behavioral, and emotional information, a system is created that maximizes user convenience.

[0565] The following describes the processing flow.

[0566] Step 1:

[0567] The device (smart glasses) captures the user's visual information with a camera. Furthermore, it collects the user's facial expressions and voice using built-in sensors. The captured visual information is converted into text data using OCR (Optical Character Recognition).

[0568] Step 2:

[0569] The device (smart glasses) wirelessly transmits generated text data and sensor data to the server. This transmission is encrypted, ensuring the security of the information.

[0570] Step 3:

[0571] The server analyzes the received text data using natural language processing techniques. This process extracts highly relevant keywords and context.

[0572] Step 4:

[0573] The server analyzes sensor data sent from the terminal using an emotion engine to estimate the user's emotional state. For example, it recognizes joy or anxiety from changes in voice tone and facial expressions.

[0574] Step 5:

[0575] The device (smartphone) collects user behavior data, such as location information and app usage history, and sends it to the server. Location information is obtained using GPS.

[0576] Step 6:

[0577] The server integrates text data, sentiment data, and behavioral data to generate optimal information and suggestions based on the user's interests and needs. In this process, the user's emotional state is a factor that adjusts the priority of information generation.

[0578] Step 7:

[0579] The device (smartphone) will send notifications at the optimal time based on the user's emotions and behavior. These notifications will include information that alleviates the user's anxiety and suggestions that pique their interest.

[0580] Step 8:

[0581] Users can take action based on information received on their smartphones and provide feedback on that information. This feedback is provided within the app.

[0582] Step 9:

[0583] The server aggregates user feedback and evaluates the contributions of information providers and application developers. This evaluation is used to distribute rewards based on user satisfaction and the usefulness of the information.

[0584] (Example 2)

[0585] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0586] In modern society, people are surrounded by a vast amount of information, and are required to appropriately select and utilize the information that is most relevant to them. However, providing information that responds to the user's emotions and the specific situation is difficult. Furthermore, it is necessary to improve the quality of information provision by effectively utilizing feedback obtained from users. Conventional technologies have not been able to adequately address these challenges.

[0587] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0588] In this invention, the server includes means for converting visual information acquired by a visual device worn by the user into digital data, means for analyzing the digital data and extracting relevant information, and emotion analysis means for analyzing and estimating the user's emotional state from facial expressions, voice, and behavior. This makes it possible to provide information that is tailored to the user's emotional state and situation, thereby increasing the relevance and usefulness of the information.

[0589] A "visual device" is a device worn by a user that is equipped with cameras and sensors to acquire visual information in real time.

[0590] "Digital data" refers to data obtained by visual devices, which is electronically represented and converted into an analyzable format.

[0591] "Analysis" is a series of processes that involve processing digital data and extracting meaningful information.

[0592] "Related information" refers to information that is useful or necessary to the user, obtained as a result of the analysis.

[0593] A "communication device" is an integrated system of hardware and software used to collect user behavior and location information and transmit it to a server.

[0594] "Behavioral data" refers to information that includes a user's movement history, location information, and time-dependent behavioral patterns.

[0595] "Emotional analysis" is a process that estimates a user's psychological state based on facial expressions, voice, and behavioral data.

[0596] "Information provision" refers to the act of providing data related to the user's situation, which is generated by the server and notified to the user.

[0597] "Feedback" refers to the evaluations and reactions that users give to the information they receive, and this data is used to improve future information provision.

[0598] This invention provides a system that allows users to efficiently acquire information and utilize its content according to their own situation and psychological state. The system uses a visual device worn by the user, such as a smart device, to convert visual information into digital data in real time.

[0599] The user wears a smart device and acquires visual information using its camera and sensors. The acquired information is converted into text data by the device's built-in OCR (Optical Character Recognition) technology. This digital data is transmitted to a server via a secure communication protocol.

[0600] The server analyzes the received text data using natural language processing techniques to extract highly relevant information. This process also takes into account behavioral data and psychological state. Behavioral data is obtained when the device collects user location information and behavioral patterns and sends them to the server. Furthermore, an emotion analysis engine estimates emotions based on facial expressions, voice, and behavior, and incorporates this into the information generation process.

[0601] The server uses a generative AI model to integrate digital data, behavioral data, and emotional data to generate information and suggestions tailored to the user. This allows the server to notify the user at the optimal time and in the most appropriate way.

[0602] As a concrete example, when a user sees an advertisement in a certain location, the device acquires data at that moment and sends it to the server. The server analyzes the information contained in the advertisement, and if the user feels uneasy about that information, the emotion engine detects this and provides information that alleviates the anxiety, such as "details of the next event" or "relevant traffic information."

[0603] Users can receive notifications on their smartphones, quickly check the information, and adjust their actions as needed. Users can also provide feedback on the information provided, and the server uses this feedback to further improve future information delivery.

[0604] As an example of a prompt, you can enter something like, "Please suggest restaurants based on places I've recently visited," and receive suggestions tailored to your history and situation.

[0605] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0606] Step 1:

[0607] The device captures the user's visual information using its camera. This input data is initially processed within the device and saved as image data in real time. Specifically, the camera automatically adjusts the focus and acquires an image with appropriate exposure.

[0608] Step 2:

[0609] The device converts captured image data into text data using OCR technology. It receives image data as input, the OCR software analyzes the text within the image, and outputs text data. Specifically, the device's processor performs high-precision character recognition.

[0610] Step 3:

[0611] The terminal sends the converted text data to the server using a secure network protocol. The input is text data, and the output is the transfer of data to the server. Specifically, the terminal's communication module packets the data, encrypts it, and transmits it.

[0612] Step 4:

[0613] The server analyzes the received text data using natural language processing and extracts relevant information. In this process, it uses the received text data as input and generates useful information for the user as output. Specifically, the NLP engine parses the data and analyzes keywords and context.

[0614] Step 5:

[0615] The device collects the user's location data and behavioral patterns and sends them to the server. In this step, the device acquires data from location sensors and accelerometers and sends it to the server. Specifically, the device obtains its location from GPS and calculates its speed of movement from the time difference.

[0616] Step 6:

[0617] The server estimates the user's emotional state from facial expressions, voice data, and behavior. It receives emotion-related data as input and generates an estimated emotional state as output. Specifically, the emotion analysis engine uses machine learning algorithms to evaluate the data.

[0618] Step 7:

[0619] The server integrates the analyzed data and uses a generative AI model to generate information best suited to the user. It uses digital data, behavioral data, and emotional data as input and creates notification information as output. Specifically, the AI ​​model integrates the data and selects recommended information.

[0620] Step 8:

[0621] The server adjusts the timing and content of notifications according to the user's emotional state and sends the information to the device. Using the generated information as input, the adjusted notification is delivered to the user as output. Specifically, the server sets the notification priority and determines the sending time to the device.

[0622] Step 9:

[0623] The user modifies their actions based on the received notification and sends necessary feedback back to the server. The system receives notification information as input and sends feedback data as output. Specifically, the user checks the notification on their smartphone and enters their opinion on the evaluation feedback screen.

[0624] (Application Example 2)

[0625] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0626] Existing information delivery systems struggle to tailor information to the user's current emotional state. As a result, the information provided may not align with the user's needs and emotions, potentially leading to a diminished user experience. The challenge lies in resolving this issue and achieving personalized information delivery based on each user's individual circumstances and emotions.

[0627] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0628] In this invention, the server includes means for converting visual information acquired by a device worn by the user into digital data, means for integrating the visual information and behavioral data to detect the user's emotional state, and means for adjusting the information generated based on the user's emotional state. This enables the provision of personalized information that is tailored to the user's emotional state and needs.

[0629] A "user-worn device" is a device that a user can wear and that is equipped with sensors or cameras for acquiring visual information.

[0630] "Methods for converting visual information into digital data" refer to technologies that process images and videos acquired from the user's perspective and transform them into data in a format that can be processed by a computer.

[0631] "Means for extracting relevant information" refers to methods for analyzing acquired digital data and identifying and extracting data that is useful or meaningful to the user.

[0632] A "communication device" is a device that collects information and transfers user behavior data and digital data to a server via a network.

[0633] A "means for detecting emotional state" refers to a system used to analyze a user's facial expressions, voice, and behavioral data to evaluate their current psychological state.

[0634] "Means of adjusting information" refers to the process of changing the content and timing of information provided according to the user's emotions and circumstances.

[0635] "Feedback" refers to the reactions and evaluations that users give to the information and suggestions they provide, and this feedback is used to improve the system and evaluate the information.

[0636] This invention utilizes smart glasses, a smartphone, and a server to realize the entire system. The smart glasses are worn by the user and acquire the user's visual information using the built-in camera and sensors. This visual information is converted into digital data and transmitted to the smartphone for subsequent processing.

[0637] The smartphone, acting as the terminal, converts acquired visual information into text data using OCR technology. In addition, the smartphone functions as a communication device, transmitting the user's location information and behavioral patterns to the server. The server integrates this digital data with the user's behavioral data and performs analysis using natural language processing technology. Based on the analysis results, it extracts highly relevant information and generates information tailored to the user.

[0638] Furthermore, by using an emotion engine, the server can analyze the user's emotional state from their facial expressions and voice data, enabling the provision of emotionally sensitive information. For example, if a user is feeling anxious about a product they are considering purchasing, the server can notify them of reviews and related reassuring information about that product.

[0639] As a concrete example of a prompt message, the generation AI model is instructed to "analyze the emotions the user is experiencing while viewing the product, and create prompts to provide information according to those emotions," thereby supporting the generation of appropriate information.

[0640] In this way, providing personalized information tailored to the user's emotions and situation in real time improves the user experience.

[0641] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0642] Step 1:

[0643] The user wears smart glasses to acquire visual information. The smart glasses use their built-in camera to capture images of the surroundings. This acquired image data becomes the input.

[0644] Step 2:

[0645] The smartphone, acting as the terminal, converts image data received from the smart glasses into text data using OCR technology. The input is image data, and the output is text data based on the OCR processing.

[0646] Step 3:

[0647] The terminal sends text data to the server. The server receives this text data and performs analysis using natural language processing technology. The input is text data, and the analysis outputs highly relevant information.

[0648] Step 4:

[0649] The server analyzes the user's emotional state using an emotion engine based on their visual data and location information. The input consists of visual data and location information, and the emotional state is output as a result of the emotion analysis.

[0650] Step 5:

[0651] The server considers the emotional state, generates relevant information, and adjusts the information it notifies. The input is the analysis result and the emotional state, and the generated appropriate information is output.

[0652] Step 6:

[0653] The user's smartphone receives and displays the adjusted information notification. The input is adjusted information, and the notification is output to the user.

[0654] Step 7:

[0655] Users provide feedback on the information they receive. This feedback is sent to the server and used as input. The server uses this feedback to evaluate the provided information and improve the system. This process improves the quality of the information provided as output.

[0656] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0657] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0658] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0659] [Fourth Embodiment]

[0660] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0661] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0662] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0663] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0664] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0665] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0666] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0667] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0668] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0669] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0670] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0671] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0672] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0673] To implement this invention, it is necessary to establish a system in which the user wears a device such as smart glasses and takes in visual information on a daily basis. These smart glasses have a built-in camera that captures the information the user sees. The captured visual information is converted into digital data.

[0674] Smart glasses use OCR (Optical Character Recognition) to convert visual information into text data. This text data is then transmitted wirelessly to a server in the cloud. The server receives this data and analyzes the information using natural language processing algorithms. From the analysis results, it extracts highly relevant keywords and context.

[0675] Meanwhile, the user's smartphone collects behavioral data, including location information and app usage history. This data is also transferred to a server. The server integrates the behavioral data from the smartphone with the analyzed digital data from the smart glasses to build a user profile. Using this profile, the server selects the most relevant information for the user and makes suggestions at the appropriate time.

[0676] For example, suppose a user is walking around town and their smart glasses capture a promotional banner for a new restaurant. The server analyzes this information and, based on the user's current location and past dining preferences, notifies their smartphone of a lunch coupon for that specific restaurant.

[0677] Users who receive a notification can send feedback on the information through the app on their smartphone. This feedback data is also sent to the server and used to reward information providers and application developers.

[0678] In this way, the invention is implemented in a manner that blurs the boundaries between analog and digital information, improves the user experience, and strengthens incentives for information providers.

[0679] The following describes the processing flow.

[0680] Step 1:

[0681] The device (smart glasses) captures analog information seen by the user with a camera and converts that image into digital text data using OCR technology.

[0682] Step 2:

[0683] The device (smart glasses) sends the generated text data to the server via the internet. The transmission is performed using a secure protocol.

[0684] Step 3:

[0685] The server analyzes the received text data and extracts key keywords and context using natural language processing techniques. This process helps to understand the meaning of the data.

[0686] Step 4:

[0687] The device (smartphone) collects the user's current location information and past activity history, and sends this data to the server. Location information is obtained using GPS.

[0688] Step 5:

[0689] The server integrates the analyzed text data and behavioral data sent from the terminal and updates the user profile to estimate the user's interests.

[0690] Step 6:

[0691] The server generates optimal information and suggestions for the user based on integrated data and notifies the user of this information via their device (smartphone). The timing of these notifications is adjusted according to the user's situation.

[0692] Step 7:

[0693] Users provide feedback on the usefulness and interest of the information they receive through notifications on their smartphones. This helps to improve the user's future information delivery.

[0694] Step 8:

[0695] The server collects user feedback and uses it as data to distribute rewards to information providers and application developers. The reward criteria are based on the evaluation of the feedback.

[0696] (Example 1)

[0697] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0698] In recent years, there has been a growing demand for services that digitize users' visual information from their daily lives and activities and use that information to provide detailed, relevant suggestions. However, relying solely on visual information from users makes it difficult to accurately provide highly relevant information, resulting in a lack of improvement in the user experience. Furthermore, if the evaluation of information provision is not fair, incentives for information providers may not function properly. This invention aims to solve these problems and build an information provision system that is more beneficial to users.

[0699] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0700] In this invention, the server includes means for converting visual information acquired by a camera worn by the user into electronic data, natural language processing means for analyzing the electronic data and extracting relevant information, and means for a communication device to collect the user's location data and usage history. This enables the integration and analysis of the user's visual information and behavioral data, allowing for more personalized information provision and fair reward evaluation for information providers.

[0701] "User" refers to an individual or group that uses this system.

[0702] A "photography device" refers to a device that has the function of capturing visual information, and specifically refers to a wearable device with a built-in camera.

[0703] "Electronic data" refers to a data format that includes visual information digitized by a photographic device.

[0704] "Natural language processing means" refers to algorithms and software tools used to analyze electronic data and extract or classify language data.

[0705] "Communication equipment" refers to devices that have the ability to send and receive data via the internet or a local network.

[0706] "Location data" refers to information indicating the user's geographical location, and is obtained through methods such as GPS.

[0707] "Usage history" refers to records of actions a user has taken in the past, applications they have used, places they have visited, and so on.

[0708] "Profile generation means" refers to a method or apparatus for generating detailed information about a user based on collected electronic data and behavioral data.

[0709] "Response" refers to feedback and evaluation from users regarding the generated information.

[0710] "Compensation" refers to the payment given to information providers or application developers that is commensurate with the value of the information they provide.

[0711] To implement this invention, it is fundamental that the user wears smart glasses, which are a camera, to collect everyday visual information. The smart glasses have a built-in camera that acquires information about the outside world as image data. This image data is converted into electronic data using OCR (Optical Character Recognition) technology. An OCR engine such as Tesseract can be used for this conversion.

[0712] The server receives the converted electronic data and performs analysis using natural language processing (NLP) algorithms. This analysis utilizes natural language processing libraries such as spaCy and NLTK to extract relevant information from the electronic data. This yields keywords and sentences relevant to the user's context, preparing the data for information provision.

[0713] Devices, especially smartphones, collect user location data via GPS and also acquire application usage history data. This information is sent to a server, where a profile generation mechanism constructs data about the user. Based on the profile, the server selects the most relevant information for the user and notifies the smartphone.

[0714] As a concrete example, suppose a user is walking through a shopping mall and their smart glasses capture an advertisement for a newly opened store. The server analyzes this advertisement information and, based on the user's current location and shopping preferences, sends a discount coupon for that store to their smartphone.

[0715] This invention enables more personalized information delivery by utilizing users' visual information and behavioral data. An example of a prompt sentence using the generative AI model is, "Please tell me how to suggest products that might interest you based on the advertisements you've seen."

[0716] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0717] Step 1:

[0718] The user wears smart glasses, which are a recording device, to acquire visual information. These glasses have a built-in camera that captures information within the user's field of view as image data. The input is analog visual information, and the output is digital image data. This allows various types of information obtained during the user's daily activities to be digitized.

[0719] Step 2:

[0720] Smart glasses convert acquired image data into electronic data using OCR (Optical Character Recognition) technology. Specifically, they analyze text within images using OCR engines such as Tesseract and extract character information. The input is digital image data, and the output is electronic data in text format. This process makes it possible to effectively extract character information from images.

[0721] Step 3:

[0722] The smart glasses transmit the converted text data to the server via wireless communication. The server takes the received text data as input and analyzes it using natural language processing algorithms. Specifically, it uses libraries such as spaCy and NLTK to perform semantic analysis of the data and extract highly relevant keywords and contexts. The output is the analyzed keyword and context information. This allows the server to obtain detailed information about the user's intent and interests.

[0723] Step 4:

[0724] On the other hand, the smartphone, as the terminal, obtains location information from GPS and collects past app usage history. The input is behavioral data obtained from the terminal, and the output is specific location information and history data that shows the user's behavior patterns. This collected data is transmitted to the server in real time.

[0725] Step 5:

[0726] The server integrates behavioral data from smartphones with electronic data analyzed from smart glasses. The input is the analyzed data obtained in steps 3 and 4, and the output is a user profile. Based on this, the server selects the most relevant information for the user and configures itself to respond quickly.

[0727] Step 6:

[0728] The server notifies the user's smartphone of the generated information. This uses a push notification service, with the user profile and selection information as input, and the output being the specific notification content for the user. The user receives this information and can decide on their next action according to the prompt.

[0729] (Application Example 1)

[0730] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0731] In modern society, users receive vast amounts of information daily, and this information is not always optimized to their interests or circumstances. In particular, in the fields of advertising and marketing, there is a demand for providing information tailored to users' interests and current situations in real time. Furthermore, streamlining the compensation system for information providers and those who create and apply the information is also a crucial challenge.

[0732] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0733] In this invention, the server includes means for converting visual information acquired by a visual device worn by the user into digital data, means for analyzing the digital data and extracting relevant information, and means for a communication device to collect user behavior data. This makes it possible to provide appropriate advertising information in real time based on the user's location information and interest data.

[0734] A "user" refers to a person who wears visual devices to receive information.

[0735] A "visual device" refers to a device worn by a user to acquire visual information, and is capable of converting that information into digital data.

[0736] "Digital data" refers to data obtained by visual devices that has been converted into a format that can be processed.

[0737] "Analysis methods" refer to processes and techniques for extracting relevant information from digital data.

[0738] A "communication device" refers to a device that collects user behavior data and transmits it to a server or other device.

[0739] "Behavioral data" refers to data that shows users' activities, location information, and interests.

[0740] "Notification means" refers to the mechanisms and functions used to deliver generated information to users.

[0741] "Feedback" refers to the evaluations and opinions that users give in response to the information provided.

[0742] "Promotional information" refers to advertisements and promotions provided based on the user's interests and location information.

[0743] "Real-time" refers to information processing and notifications occurring instantly and without delay.

[0744] The system for implementing this invention consists of a visual device worn by the user and a mobile terminal. The visual device is equipped with a camera that captures the user's visual information in real time. This information is converted into text data using built-in optical character recognition (OCR) software. The converted text data is transmitted wirelessly to a server in the cloud.

[0745] The server performs natural language processing on the received text data and analyzes relevant information. During this process, the server builds individual user profiles based on behavioral data transmitted from the user's mobile device, such as location information and past activity history. This allows for the generation of user-optimized information or advertisements, which are then notified to the user's mobile device in real time.

[0746] For example, if a user's visual device captures a sign for a new restaurant while walking around town, that information is immediately analyzed. The server considers the user's location and past dining patterns and sends a special lunch coupon related to that restaurant to the user's mobile device. This notification is timed appropriately so that the user receives the information at the right moment.

[0747] The hardware used will include a visual device such as Google Glass, and the software used for analysis will be Google Cloud NLP and OCR tools. The terminal is assumed to be a typical smartphone device. An example of a prompt message would be as follows:

[0748] Example of a prompt:

[0749] We have captured promotional information for new exhibitions that the user is interested in at their current location. Based on this information, generate a discount coupon for the relevant exhibition and notify the user on their mobile device.

[0750] This invention allows users to always receive information tailored to their needs, and enables information providers to efficiently reach their target users.

[0751] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0752] Step 1:

[0753] The user wears a visual device to capture visual information from their surroundings. This visual information, as input, is received by a camera sensor and converted into image data. The output is the image data sent to the OCR process.

[0754] Step 2:

[0755] The device uses OCR (Optical Character Recognition) to convert captured visual information into text data. The input is image data from a visual device, which is then converted into text data using optical character recognition software. The output is the text data sent to a cloud server.

[0756] Step 3:

[0757] The server receives text data and performs analysis using natural language processing algorithms. The input is text data sent from the terminal, and the analysis extracts highly relevant keywords. The output consists of keywords and contextual information.

[0758] Step 4:

[0759] The device sends behavioral data (e.g., location information, app usage history) to the server. The input is data showing the user's daily activities, and the output is behavioral data transferred to the server.

[0760] Step 5:

[0761] The server integrates the results of text data analysis with behavioral data to build user profiles. The input consists of analyzed text data and behavioral data, and a generative AI model is used to construct the profiles. The output is a profile tailored to each individual user.

[0762] Step 6:

[0763] The server generates user-optimized information or advertisements and constructs notification prompts. Inputs are user profiles and analytics data, and output is a notification prompt sent to the user's device.

[0764] Step 7:

[0765] The user terminal displays relevant information and advertisements to the user in real time based on the notification prompt message received. The input is a prompt message from the server, and the output is an information notification to the user.

[0766] Step 8:

[0767] Users provide feedback on the information or advertisements they are given. The input is the user's reaction or opinion, and the output is the feedback data sent to the server.

[0768] Step 9:

[0769] The server evaluates the feedback data, calculates rewards for information providers and application developers, and generates new prompt messages as needed. The input is the feedback data, and the output is reward information and improved prompt messages.

[0770] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0771] This invention is based on a system that utilizes a device such as smart glasses as a wearable terminal for the user to capture visual information and process it as digital data. The smart glasses are equipped with cameras and sensors to capture the user's visual information in real time. The captured information is converted into text data using OCR technology.

[0772] This digital data is sent to a server via a secure network. The server analyzes this data using natural language processing technology and extracts highly relevant information. Based on this, the refined information is added to the user profile. In addition, the device (smartphone) collects user behavior data, including location information and behavioral patterns, and transfers it to the server.

[0773] In this embodiment, an emotion engine has been added. The emotion engine analyzes the user's emotional state from their facial expressions, voice data, and behavior to estimate the user's current psychological state. Emotional data is also used for analysis on the server and influences the generation of information.

[0774] The server integrates analyzed digital data, behavioral data, and emotional data to generate the most relevant information and suggestions for the user. In doing so, it adjusts the timing and content of notifications according to the user's emotional state, providing information in a more appropriate format.

[0775] For example, if the emotion engine detects an anxious emotional state while a user is viewing a transportation advertisement, the server will provide suggestions to alleviate the user's anxiety, such as "the status of the next train" or "transfer information." The user will receive this notification on their smartphone and can adjust their behavior accordingly.

[0776] Ultimately, if a user provides feedback on the information they are notified of, that feedback is collected on the server and used to evaluate the rewards given to the provider and app developer. In this way, by linking digital, behavioral, and emotional information, a system is created that maximizes user convenience.

[0777] The following describes the processing flow.

[0778] Step 1:

[0779] The device (smart glasses) captures the user's visual information with a camera. Furthermore, it collects the user's facial expressions and voice using built-in sensors. The captured visual information is converted into text data using OCR (Optical Character Recognition).

[0780] Step 2:

[0781] The device (smart glasses) wirelessly transmits generated text data and sensor data to the server. This transmission is encrypted, ensuring the security of the information.

[0782] Step 3:

[0783] The server analyzes the received text data using natural language processing techniques. This process extracts highly relevant keywords and context.

[0784] Step 4:

[0785] The server analyzes sensor data sent from the terminal using an emotion engine to estimate the user's emotional state. For example, it recognizes joy or anxiety from changes in voice tone and facial expressions.

[0786] Step 5:

[0787] The device (smartphone) collects user behavior data, such as location information and app usage history, and sends it to the server. Location information is obtained using GPS.

[0788] Step 6:

[0789] The server integrates text data, sentiment data, and behavioral data to generate optimal information and suggestions based on the user's interests and needs. In this process, the user's emotional state is a factor that adjusts the priority of information generation.

[0790] Step 7:

[0791] The device (smartphone) will send notifications at the optimal time based on the user's emotions and behavior. These notifications will include information that alleviates the user's anxiety and suggestions that pique their interest.

[0792] Step 8:

[0793] Users can take action based on information received on their smartphones and provide feedback on that information. This feedback is provided within the app.

[0794] Step 9:

[0795] The server aggregates user feedback and evaluates the contributions of information providers and application developers. This evaluation is used to distribute rewards based on user satisfaction and the usefulness of the information.

[0796] (Example 2)

[0797] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0798] In modern society, people are surrounded by a vast amount of information, and are required to appropriately select and utilize the information that is most relevant to them. However, providing information that responds to the user's emotions and the specific situation is difficult. Furthermore, it is necessary to improve the quality of information provision by effectively utilizing feedback obtained from users. Conventional technologies have not been able to adequately address these challenges.

[0799] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0800] In this invention, the server includes means for converting visual information acquired by a visual device worn by the user into digital data, means for analyzing the digital data and extracting relevant information, and emotion analysis means for analyzing and estimating the user's emotional state from facial expressions, voice, and behavior. This makes it possible to provide information that is tailored to the user's emotional state and situation, thereby increasing the relevance and usefulness of the information.

[0801] A "visual device" is a device worn by a user that is equipped with cameras and sensors to acquire visual information in real time.

[0802] "Digital data" refers to data obtained by visual devices, which is electronically represented and converted into an analyzable format.

[0803] "Analysis" is a series of processes that involve processing digital data and extracting meaningful information.

[0804] "Related information" refers to information that is useful or necessary to the user, obtained as a result of the analysis.

[0805] A "communication device" is an integrated system of hardware and software used to collect user behavior and location information and transmit it to a server.

[0806] "Behavioral data" refers to information that includes a user's movement history, location information, and time-dependent behavioral patterns.

[0807] "Emotional analysis" is a process that estimates a user's psychological state based on facial expressions, voice, and behavioral data.

[0808] "Information provision" refers to the act of providing data related to the user's situation, which is generated by the server and notified to the user.

[0809] "Feedback" refers to the evaluations and reactions that users give to the information they receive, and this data is used to improve future information provision.

[0810] This invention provides a system that allows users to efficiently acquire information and utilize its content according to their own situation and psychological state. The system uses a visual device worn by the user, such as a smart device, to convert visual information into digital data in real time.

[0811] The user wears a smart device and acquires visual information using its camera and sensors. The acquired information is converted into text data by the device's built-in OCR (Optical Character Recognition) technology. This digital data is transmitted to a server via a secure communication protocol.

[0812] The server analyzes the received text data using natural language processing techniques to extract highly relevant information. This process also takes into account behavioral data and psychological state. Behavioral data is obtained when the device collects user location information and behavioral patterns and sends them to the server. Furthermore, an emotion analysis engine estimates emotions based on facial expressions, voice, and behavior, and incorporates this into the information generation process.

[0813] The server uses a generative AI model to integrate digital data, behavioral data, and emotional data to generate information and suggestions tailored to the user. This allows the server to notify the user at the optimal time and in the most appropriate way.

[0814] As a concrete example, when a user sees an advertisement in a certain location, the device acquires data at that moment and sends it to the server. The server analyzes the information contained in the advertisement, and if the user feels uneasy about that information, the emotion engine detects this and provides information that alleviates the anxiety, such as "details of the next event" or "relevant traffic information."

[0815] Users can receive notifications on their smartphones, quickly check the information, and adjust their actions as needed. Users can also provide feedback on the information provided, and the server uses this feedback to further improve future information delivery.

[0816] As an example of a prompt, you can enter something like, "Please suggest restaurants based on places I've recently visited," and receive suggestions tailored to your history and situation.

[0817] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0818] Step 1:

[0819] The device captures the user's visual information using its camera. This input data is initially processed within the device and saved as image data in real time. Specifically, the camera automatically adjusts the focus and acquires an image with appropriate exposure.

[0820] Step 2:

[0821] The device converts captured image data into text data using OCR technology. It receives image data as input, the OCR software analyzes the text within the image, and outputs text data. Specifically, the device's processor performs high-precision character recognition.

[0822] Step 3:

[0823] The terminal sends the converted text data to the server using a secure network protocol. The input is text data, and the output is the transfer of data to the server. Specifically, the terminal's communication module packets the data, encrypts it, and transmits it.

[0824] Step 4:

[0825] The server analyzes the received text data using natural language processing and extracts relevant information. In this process, it uses the received text data as input and generates useful information for the user as output. Specifically, the NLP engine parses the data and analyzes keywords and context.

[0826] Step 5:

[0827] The device collects the user's location data and behavioral patterns and sends them to the server. In this step, the device acquires data from location sensors and accelerometers and sends it to the server. Specifically, the device obtains its location from GPS and calculates its speed of movement from the time difference.

[0828] Step 6:

[0829] The server estimates the user's emotional state from facial expressions, voice data, and behavior. It receives emotion-related data as input and generates an estimated emotional state as output. Specifically, the emotion analysis engine uses machine learning algorithms to evaluate the data.

[0830] Step 7:

[0831] The server integrates the analyzed data and uses a generative AI model to generate information best suited to the user. It uses digital data, behavioral data, and emotional data as input and creates notification information as output. Specifically, the AI ​​model integrates the data and selects recommended information.

[0832] Step 8:

[0833] The server adjusts the timing and content of notifications according to the user's emotional state and sends the information to the device. Using the generated information as input, the adjusted notification is delivered to the user as output. Specifically, the server sets the notification priority and determines the sending time to the device.

[0834] Step 9:

[0835] The user modifies their actions based on the received notification and sends necessary feedback back to the server. The system receives notification information as input and sends feedback data as output. Specifically, the user checks the notification on their smartphone and enters their opinion on the evaluation feedback screen.

[0836] (Application Example 2)

[0837] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0838] Existing information delivery systems struggle to tailor information to the user's current emotional state. As a result, the information provided may not align with the user's needs and emotions, potentially leading to a diminished user experience. The challenge lies in resolving this issue and achieving personalized information delivery based on each user's individual circumstances and emotions.

[0839] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0840] In this invention, the server includes means for converting visual information acquired by a device worn by the user into digital data, means for integrating the visual information and behavioral data to detect the user's emotional state, and means for adjusting the information generated based on the user's emotional state. This enables the provision of personalized information that is tailored to the user's emotional state and needs.

[0841] A "user-worn device" is a device that a user can wear and that is equipped with sensors or cameras for acquiring visual information.

[0842] "Methods for converting visual information into digital data" refer to technologies that process images and videos acquired from the user's perspective and transform them into data in a format that can be processed by a computer.

[0843] "Means for extracting relevant information" refers to methods for analyzing acquired digital data and identifying and extracting data that is useful or meaningful to the user.

[0844] A "communication device" is a device that collects information and transfers user behavior data and digital data to a server via a network.

[0845] A "means for detecting emotional state" refers to a system used to analyze a user's facial expressions, voice, and behavioral data to evaluate their current psychological state.

[0846] "Means of adjusting information" refers to the process of changing the content and timing of information provided according to the user's emotions and circumstances.

[0847] "Feedback" refers to the reactions and evaluations that users give to the information and suggestions they provide, and this feedback is used to improve the system and evaluate the information.

[0848] This invention utilizes smart glasses, a smartphone, and a server to realize the entire system. The smart glasses are worn by the user and acquire the user's visual information using the built-in camera and sensors. This visual information is converted into digital data and transmitted to the smartphone for subsequent processing.

[0849] The smartphone, acting as the terminal, converts acquired visual information into text data using OCR technology. In addition, the smartphone functions as a communication device, transmitting the user's location information and behavioral patterns to the server. The server integrates this digital data with the user's behavioral data and performs analysis using natural language processing technology. Based on the analysis results, it extracts highly relevant information and generates information tailored to the user.

[0850] Furthermore, by using an emotion engine, the server can analyze the user's emotional state from their facial expressions and voice data, enabling the provision of emotionally sensitive information. For example, if a user is feeling anxious about a product they are considering purchasing, the server can notify them of reviews and related reassuring information about that product.

[0851] As a concrete example of a prompt message, the generation AI model is instructed to "analyze the emotions the user is experiencing while viewing the product, and create prompts to provide information according to those emotions," thereby supporting the generation of appropriate information.

[0852] In this way, providing personalized information tailored to the user's emotions and situation in real time improves the user experience.

[0853] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0854] Step 1:

[0855] The user wears smart glasses to acquire visual information. The smart glasses use their built-in camera to capture images of the surroundings. This acquired image data becomes the input.

[0856] Step 2:

[0857] The smartphone, acting as the terminal, converts image data received from the smart glasses into text data using OCR technology. The input is image data, and the output is text data based on the OCR processing.

[0858] Step 3:

[0859] The terminal sends text data to the server. The server receives this text data and performs analysis using natural language processing technology. The input is text data, and the analysis outputs highly relevant information.

[0860] Step 4:

[0861] The server analyzes the user's emotional state using an emotion engine based on their visual data and location information. The input consists of visual data and location information, and the emotional state is output as a result of the emotion analysis.

[0862] Step 5:

[0863] The server considers the emotional state, generates relevant information, and adjusts the information it notifies. The input is the analysis result and the emotional state, and the generated appropriate information is output.

[0864] Step 6:

[0865] The user's smartphone receives and displays the adjusted information notification. The input is adjusted information, and the notification is output to the user.

[0866] Step 7:

[0867] Users provide feedback on the information they receive. This feedback is sent to the server and used as input. The server uses this feedback to evaluate the provided information and improve the system. This process improves the quality of the information provided as output.

[0868] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0869] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0870] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0871] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0872] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0873] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0874] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0875] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0876] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0877] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0878] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0879] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0880] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0881] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0882] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0883] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0884] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0885] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0886] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0887] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0888] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0889] The following is further disclosed regarding the embodiments described above.

[0890] (Claim 1)

[0891] A means of converting visual information acquired by a device worn by a user into digital data,

[0892] A means for analyzing the aforementioned digital data and extracting relevant information,

[0893] Means by which communication devices collect user behavior data,

[0894] A means of integrating collected behavioral data and analyzed digital data to generate information optimized for the user,

[0895] Means for notifying the user of the generated information,

[0896] A system that includes means for evaluating the aforementioned information based on user feedback.

[0897] (Claim 2)

[0898] The system according to claim 1, further comprising means for rewarding information providers and application developers based on the aforementioned user feedback.

[0899] (Claim 3)

[0900] The system according to claim 1, comprising means for detecting that a user is in a specific location and including location-based information as part of the generated information.

[0901] "Example 1"

[0902] (Claim 1)

[0903] A means for converting visual information acquired by a camera worn by the user into electronic data,

[0904] A natural language processing means for analyzing the aforementioned electronic data and extracting relevant information,

[0905] A means by which communication devices collect user location data and usage history,

[0906] A profile generation means for integrating collected location data and analyzed electronic data to generate information tailored to the user,

[0907] Means for notifying the user's communication device of the generated information,

[0908] A system that includes means for evaluating the aforementioned information based on responses from users and determining compensation for information providers.

[0909] (Claim 2)

[0910] The system according to claim 1, further comprising means for providing compensation to a software developer based on the response from the user.

[0911] (Claim 3)

[0912] The system according to claim 1, further comprising means for identifying that a user is in a specific geographical area and including information based on that geographical area in the generated information.

[0913] "Application Example 1"

[0914] (Claim 1)

[0915] A means of converting visual information acquired by a visual device worn by a user into digital data,

[0916] A means for analyzing the aforementioned digital data and extracting relevant information,

[0917] A means by which a communication device collects user behavior data,

[0918] A means of integrating collected behavioral data and analyzed digital data to generate information optimized for the user,

[0919] Means for notifying the user of the generated information,

[0920] A means for evaluating the aforementioned information based on user feedback,

[0921] A means of providing appropriate advertising information based on the user's location information and interest data,

[0922] A means of notifying users of special information in real time and

[0923] A system that includes this.

[0924] (Claim 2)

[0925] The system according to claim 1, further comprising means for providing compensation to information providers and application developers based on the aforementioned user feedback.

[0926] (Claim 3)

[0927] The system according to claim 1, comprising means for detecting that a user is in a specific location and including information based on that location as part of the generated information.

[0928] "Example 2 of combining an emotion engine"

[0929] (Claim 1)

[0930] A means of converting visual information acquired by a visual device worn by a user into digital data,

[0931] A means for analyzing the aforementioned digital data and extracting relevant information,

[0932] A means by which a communication device collects behavioral data, including the user's location data and behavioral patterns,

[0933] An emotion analysis method that analyzes and estimates the user's emotional state from facial expressions, voice, and behavior,

[0934] A means of integrating collected behavioral data, analyzed digital data, and emotional data to generate information optimized for the user,

[0935] A means for providing the generated information by adjusting the timing and content of notifications according to the user's emotional state,

[0936] A system that includes means for evaluating the aforementioned information based on user feedback and improving the quality of the information provided.

[0937] (Claim 2)

[0938] The system according to claim 1, further comprising means for rewarding information providers and software developers based on the aforementioned user feedback.

[0939] (Claim 3)

[0940] The system according to claim 1, comprising means for detecting that a user is in a specific geographical location and including location-based information as part of the generated information.

[0941] "Application example 2 when combining with an emotional engine"

[0942] (Claim 1)

[0943] A means of converting visual information acquired by a device worn by a user into digital data,

[0944] A means for analyzing the aforementioned digital data and extracting relevant information,

[0945] A means by which a communication device collects user behavior data,

[0946] A means of integrating collected behavioral data and analyzed digital data to generate information optimized for the user,

[0947] A means for detecting the user's emotional state and adjusting the information generated based on that emotional state,

[0948] Means for notifying the user of the generated information,

[0949] A system that includes means for evaluating the aforementioned information based on user feedback.

[0950] (Claim 2)

[0951] The system according to claim 1, further comprising means for rewarding information providers and application developers based on the aforementioned user feedback.

[0952] (Claim 3)

[0953] The system according to claim 1, further comprising means for detecting the emotional state of a user outside of a specific location and including information based on that emotion as part of the generated information. [Explanation of Symbols]

[0954] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of converting visual information acquired by a device worn by a user into digital data, A means for analyzing the aforementioned digital data and extracting relevant information, Means by which communication devices collect user behavior data, A means of integrating collected behavioral data and analyzed digital data to generate information optimized for the user, Means for notifying the user of the generated information, A system that includes means for evaluating the aforementioned information based on user feedback.

2. The system according to claim 1, further comprising means for rewarding information providers and application developers based on user feedback.

3. The system according to claim 1, further comprising means for detecting that a user is in a specific location and including location-based information as part of the generated information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A