system
A system using generative AI to analyze visual, audio, and location data with emotion recognition provides personalized information overlays, addressing the challenge of information overload and enhancing user experience by adapting to individual needs and emotional states.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Users face challenges in quickly and appropriately obtaining information tailored to their individual needs in information-saturated environments, particularly in fields like shopping, medical care, and tourism, where real-time and personalized services are lacking.
A system utilizing generative artificial intelligence models to analyze visual, audio, and location data, providing information overlays and real-time updates based on user requests, incorporating emotion analysis to optimize user experience.
Enables real-time provision of personalized information and services, enhancing user experience by adapting to individual needs and emotional states, facilitating efficient decision-making.
Smart Images

Figure 2026070997000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern society, users are surrounded by a vast amount of information, and there is a problem that it is difficult to quickly and appropriately obtain information suitable for each individual need. Furthermore, in various fields such as shopping, medical care, and tourism, services and guidance according to individual requests are required. The purpose of the present invention is to realize a system that analyzes visual information, voice information, and location information of users and provides necessary information in real time in order to solve these problems.
Means for Solving the Problems
[0005] This invention relates to a system having means for analyzing visual data, audio data, and location information using a generative artificial intelligence model. This allows for the provision of information tailored to the user's situation and the overlay display of services and information suited to specific needs within the user's field of vision. Furthermore, in the medical field, it improves the user experience by providing diagnostic support based on the patient's condition. This system enables real-time information updates and optimization of the information provided according to user requests.
[0006] A "generative artificial intelligence model" refers to an algorithm that analyzes data and automatically generates new information and suggestions.
[0007] "Visual data" refers to video information acquired using cameras and other imaging devices.
[0008] "Audio data" refers to acoustic information acquired using microphones or other audio acquisition devices.
[0009] "Location information" refers to information about geographical location obtained using GPS or other location measurement technologies.
[0010] "Analysis" refers to the process of processing data and information to extract meaningful patterns and insights from them.
[0011] "Information overlay display" refers to a technique that displays additional data or objects superimposed on the user's field of view.
[0012] "Diagnostic support" in the medical field refers to providing information that helps in treatment planning and diagnosis by analyzing patient data.
[0013] "Real-time" means processing and providing information instantly with minimal delay.
[0014] "User requests" refer to the intentions or instructions of a user who desires specific information or actions. [Brief explanation of the drawing]
[0015] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention relates to an advanced information provision system that utilizes a generative artificial intelligence model, and is operated primarily by analyzing visual data, audio data, and location information. The system is designed to provide information based on the individual needs of users in real time.
[0037] First, the device worn or carried by the user (such as smart glasses) continuously collects data about the surrounding environment and circumstances. This device is equipped with sensors such as a camera, microphone, and GPS, and can acquire visual data, audio data, and location information. This data is analyzed by the device, stored in temporary memory, and then sent to a server.
[0038] The server uses a generative artificial intelligence model to analyze the received data. Specifically, it identifies visual data using image processing technology and understands audio data using natural language processing. It also understands the current location and surrounding environment based on location information. By comprehensively combining this information, it identifies the information and services the user is looking for and generates personalized suggestions.
[0039] For example, when a user is walking through a shopping mall, the device captures visual data of displayed products and sends it to a server. The server uses generated AI to identify the products and provides information on the lowest prices compared to other online platforms. Furthermore, when a user is in a tourist area, the device can analyze information on buildings and artworks captured by the camera and display appropriate explanations in real time.
[0040] In the medical field, healthcare professionals can use terminals to access medical information tailored to the patient's condition. For example, during pre-operative preparation, the system can provide optimal procedures and medication information based on the patient's health status, supporting smooth medical treatment.
[0041] This system provides the most relevant information instantly by continuously updating information in response to user requests. Users can request further details by giving instructions, for example, via voice commands or gestures. As a result, users can obtain the necessary information at the optimal time, enabling efficient decision-making in today's information-saturated society.
[0042] The following describes the processing flow.
[0043] Step 1:
[0044] The device activates sensors when the user begins using it, collecting visual, audio, and location data in real time. The camera captures the surrounding video, the microphone records audio, and GPS determines the user's location.
[0045] Step 2:
[0046] The terminal stores the collected data in temporary memory, compresses the data, and then sends it to the server via a secure protocol. This process minimizes latency during data transfer.
[0047] Step 3:
[0048] The server prepares the received data for analysis. Visual data is identified using image recognition algorithms to identify objects and text, and audio data is converted into language using speech recognition algorithms. Location information is merged with a mapping service to determine the current location.
[0049] Step 4:
[0050] The server applies an artificial intelligence model based on the analysis results to generate information and services tailored to the user. Examples include information on the lowest prices for shopping, historical background and explanations of tourist destinations, and diagnostic information in medical support.
[0051] Step 5:
[0052] The server sends the generated information to the terminal. The terminal receives this information and provides the necessary information by displaying it as an overlay in the user's field of view. The display timing and format are customized according to the user's situation.
[0053] Step 6:
[0054] Users can request additional information or select and manipulate specific data using voice or gestures. The terminal sends these inputs to a server for information updates and further analysis.
[0055] Step 7:
[0056] The server re-analyzes the information based on user feedback and requests, and resends the newly generated information to the terminal. This ensures that the latest information tailored to the user's needs is always provided.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] In modern society, information overload has become the norm, making it difficult for users to quickly and accurately obtain the information they need. This is especially true while on the go or in complex environments, where information acquisition and decision-making become even more challenging. To address this problem, there is a need for systems that provide appropriate information in real time, tailored to the user's situation and needs, and support efficient decision-making.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, the server includes information provision means equipped with a device for collecting data from the surrounding environment, means for transmitting the collected data via a communication line, and means for analyzing visual information using image recognition technology. This makes it possible to provide users with necessary information in real time and support their decision-making.
[0062] A "device that collects data from the surrounding environment" is a device that acquires data such as visual, auditory, and location information using sensors, and understands the state of the user's surroundings in real time.
[0063] "Information provision means" refers to methods and technologies for presenting appropriate information to users based on collected and analyzed data.
[0064] "Means of transmission via communication lines" refers to methods using wireless or wired communication technology to transfer collected data to external devices such as servers.
[0065] "Means of analyzing visual information using image recognition technology" refers to techniques for analyzing acquired visual data for object identification and feature extraction, and utilizes image processing algorithms.
[0066] "Means of analyzing acoustic information using natural language processing technology" refers to technologies that convert audio information into text and analyze its meaning and intent.
[0067] "Methods for understanding spatial information using location data" refers to technologies that utilize location information such as GPS to identify a user's current location and surrounding environment.
[0068] "Means of generating information based on analysis results using generative artificial intelligence models" refers to a method of generating optimal information for users based on analyzed data using AI technology.
[0069] "Means for processing additional information requests from users via prompt messages" refers to interfaces and technologies that understand and provide appropriate responses or information when a user requests additional information via voice or text.
[0070] This invention is a system that provides users with appropriate information in real time using a generative artificial intelligence model. The system primarily utilizes the following devices and technologies.
[0071] The device is equipped with various sensors to collect data from its surroundings. Specifically, these include a camera to acquire visual information, a microphone to acquire acoustic information, and a GPS unit to collect location data. This device is often implemented as smart glasses or other wearable devices. The collected data is initially processed within the device and then transmitted to a server via a communication line.
[0072] The server analyzes the received data. For visual information, image recognition technology is used to recognize objects. In this case, existing image processing libraries such as OpenCV are often used. Acoustic information is analyzed using natural language processing technology, and text conversion is performed using Google Cloud Speech-to-Text, etc. Location information is compared with map data to understand the user's spatial environment.
[0073] A generative AI model is implemented to generate information tailored to the user's needs based on analysis results. This generated information is displayed overlaid on the user's field of view through a visual device. For example, if a user is in a shopping mall, they can view comparison information on the lowest prices of products they have photographed. Furthermore, in tourist areas, it can provide real-time historical background information about buildings. In the medical field, it can also present optimal diagnostic information based on the patient's condition.
[0074] Users can obtain additional information or details by sending prompts via voice commands or text input. Examples of prompts include "Tell me more about this product" or "Tell me the nearest cafe." The server receives these prompts and provides the user with more detailed information or service suggestions.
[0075] As described above, this invention can quickly and efficiently provide users with the information they need in today's information-saturated world.
[0076] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0077] Step 1:
[0078] The device collects data from its surroundings. Inputs include visual data, acoustic data, and location data. This data is acquired through sensors such as cameras, microphones, and GPS. Visual data is saved as image files, and acoustic data is converted into audio files. Location data is acquired as latitude and longitude information. The output is a formatted dataset.
[0079] Step 2:
[0080] The device temporarily stores the collected data and sends it to the server via a communication line. Specifically, the collected data is stored in temporary memory and sent to the server using Wi-Fi or mobile data communication. The input is a formatted dataset, and the output is a data stream received on the server side.
[0081] Step 3:
[0082] The server analyzes the received visual data using image recognition technology. The input is an image file. Specifically, the server uses OpenCV or a similar image processing library to identify objects within the image. The output is a list of recognized objects.
[0083] Step 4:
[0084] The server analyzes audio data using natural language processing techniques. The input is an audio file. The server uses a speech recognition service such as Google Cloud Speech-to-Text to convert the audio to text. It then analyzes the text content to extract meaning. The output is the analyzed text information.
[0085] Step 5:
[0086] The server analyzes spatial information based on location data. The input is latitude and longitude information. Specifically, it combines this with map data to identify the user's current location and surrounding area. The output is location-related environmental information.
[0087] Step 6:
[0088] The server uses a generative AI model to generate information based on the analysis results. The input consists of the analysis results of visual, acoustic, and location information. The server uses generative AI technology to create information and suggestions best suited to the user's situation. The output is personalized information for the user.
[0089] Step 7:
[0090] The device notifies the user of the generated information. The input is personalized information sent from the server. The device displays this information on its screen and provides voice guidance as needed. The output is the information the user receives visually and aurally.
[0091] Step 8:
[0092] The user can request additional information using prompt messages. The input is a prompt message in either voice or text format. The terminal sends this request to the server for further analysis. The output is more detailed information provided to the user.
[0093] (Application Example 1)
[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0095] In today's consumer society, consumers are surrounded by a vast amount of product information, making it difficult to make the best choice. Furthermore, the market contains numerous products with varying prices, performance, and reviews, and comparing them in real time requires considerable time and effort. Moreover, there is a lack of convenient ways for users to obtain this information during their physical shopping experience.
[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0097] In this invention, the server includes means for analyzing visual data, means for handling audio data, means for acquiring location information, and means for recognizing products and presenting detailed information in real time. This enables users to efficiently collect necessary product information when shopping and make optimal purchasing decisions.
[0098] A "generative artificial intelligence model" is an artificial intelligence trained using machine learning techniques to perform a specific task from a vast amount of data.
[0099] "Means for analyzing visual data" refers to devices or software that analyze image or video data using digital image processing technology to recognize objects or scenes.
[0100] "Means for analyzing audio data" refers to devices or software that use speech recognition technology to extract text or intent from audio.
[0101] "Means of acquiring location information" refers to technologies that use GPS or other positioning systems to determine the current location of a device.
[0102] "Means of providing information to users" refers to a system that displays or communicates useful information to users through a user interface based on analysis results.
[0103] A "means for updating information based on user requests" refers to a system that has the ability to dynamically re-evaluate, modify, or update the information it provides in response to user input or requests.
[0104] "Means for identifying products and presenting related information in real time" refers to technology that uses camera data or similar methods to identify products and immediately displays information related to those products to the user.
[0105] The system of this invention operates on the basis of cooperation between a user-carried terminal and a server. The terminal is equipped with sensors such as a camera, microphone, and GPS, which are used to acquire visual data, audio data, and location information of the environment. The terminal has the function of collecting this data in real time and transmitting it to the server.
[0106] The server performs image analysis on visual data using a generative AI model to identify objects and products. Audio data is analyzed by speech recognition software to understand user requests in natural language. Location information is used to determine the current location and understand the context of the environment.
[0107] The server's analysis results are returned to the user's device and displayed as an overlay in the user's field of view. This allows for real-time presentation of detailed product information and price comparisons, supporting the user's decision-making.
[0108] For example, when a user is looking at clothes in a store, the device acquires visual data and sends it to the server. The server uses a generative AI model to identify the clothes and retrieves relevant data by comparing it with registered information. This includes the product price, customer reviews, and price comparisons with other stores. If the user gives a voice command such as "Tell me the reviews for this product," the server retrieves and displays the review information.
[0109] An example of a prompt might be, "Identify the products in the image and tell me their price and description." Based on this prompt, the server uses a generative AI model to select the information to provide to the user.
[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0111] Step 1:
[0112] The device uses its camera and microphone to acquire visual and audio data from its surroundings. The input is raw data from the environment, and the output is this data in digital form. This allows the device to capture objects within the user's field of view and surrounding conversations.
[0113] Step 2:
[0114] The terminal transmits the acquired visual and audio data to the server. The input is the digital data held by the terminal, and the output is a signal to the server. This allows the server to receive the information necessary for analysis.
[0115] Step 3:
[0116] The server uses an artificial intelligence model to analyze visual data and identify objects and products within the image. The input is visual data, and the output is identification information for the identified objects. Specifically, the data is analyzed by an image recognition algorithm.
[0117] Step 4:
[0118] The server uses speech recognition software to analyze voice data and convert user instructions into text. The input is voice data, and the output is the analyzed text information. The specific operation involves extracting voice commands.
[0119] Step 5:
[0120] The server searches for related information about an object it has identified and generates the information in real time. Input is the identification information of the identified object and a voice command, and output is related information that can be displayed. Information is retrieved from a database.
[0121] Step 6:
[0122] The server sends the generated information to the terminal and displays it as an overlay in the user's field of view. The input is a dataset of related information, and the output is a visual presentation on the user's display device. The terminal then reactivates the overlay function.
[0123] Step 7:
[0124] The user requests further details using voice commands or input devices as needed. Input is the user's new instructions, and output is an updated display of information. This cycle is repeated to support the user's decision-making.
[0125] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0126] This invention relates to an information provision system incorporating an emotion engine that recognizes user emotions, thereby enabling the provision of information and services that take into account the user's emotional state. This system has the capability to analyze user emotion data in addition to visual data, audio data, and location information.
[0127] The device is equipped with a camera, microphone, GPS sensor, and emotion engine sensor, which are used to simultaneously collect the user's environmental information and emotional state. The emotion engine analyzes the user's emotions from their facial expressions, voice tone, and body movements. This data collection is performed continuously in the background during the user's natural behavior, and the device sends this data to a server.
[0128] The server comprehensively analyzes various data and generates information using an artificial intelligence model, combining it with the sentiment analysis results from the emotion engine. This system aims to optimize the user experience by selecting the appropriate way to provide information based on the user's emotional state.
[0129] For example, if a user experiences discomfort or fatigue while sightseeing, the system's emotion engine detects this and prioritizes providing information on relaxing places and rest spots rather than typical tourist information. If the user is shopping and their stress level is high, the system will offer suggestions to assist with simple transactions or purchases.
[0130] Furthermore, in the medical field, data obtained from the emotion engine can be used to provide more humane and efficient support when making diagnoses and providing care while considering the patient's emotional responses. This allows healthcare professionals to provide care that is tailored to the patient's psychological state.
[0131] By constantly updating information based on user requests and feedback, and taking appropriate action at each step, the system functions effectively under diverse circumstances. This ensures that users always receive the best information and services, leading to a more comfortable and satisfying experience.
[0132] The following describes the processing flow.
[0133] Step 1:
[0134] The device activates its camera, microphone, GPS sensor, and emotion engine sensors to collect visual data, audio data, location information, and emotion data in real time. The camera captures the user's facial expressions, the microphone records voice tones, and the sensors detect the user's body movements, etc.
[0135] Step 2:
[0136] The terminal sends the collected data to the server. Upon receiving the data, the server prepares it for analysis. This includes data preprocessing and compression as needed.
[0137] Step 3:
[0138] The server uses image recognition algorithms to analyze visual data and identify objects and text in the user's environment.
[0139] Step 4:
[0140] The server applies a speech recognition algorithm, converts the audio data into text format, and performs linguistic analysis. This allows it to understand the content of the user's speech and the nuances of their emotions.
[0141] Step 5:
[0142] The server analyzes emotional data acquired by the emotion engine to identify the user's current emotional state (e.g., happiness, anxiety, excitement, fatigue). This result is then combined with other data to provide a comprehensive assessment of the user's situation.
[0143] Step 6:
[0144] The server uses a generative artificial intelligence model to generate information and services best suited to the user based on collected information and sentiment analysis results. For example, if it detects that the user is tired, it prioritizes providing information about rest spots.
[0145] Step 7:
[0146] The server sends the generated information to the terminal. The terminal receives this information and displays it as an overlay in the user's field of view. The displayed content is customized to include information that responds to the user's emotions.
[0147] Step 8:
[0148] Users can request additional information or manipulate the provided information using voice commands or gestures. The terminal sends this input to the server, which regenerates the information as needed.
[0149] Step 9:
[0150] The server updates information based on user feedback and sends the new information to the device. This ensures that users always receive up-to-date information that reflects their emotional state.
[0151] (Example 2)
[0152] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0153] Conventional information delivery systems have the problem of not optimizing the user experience because they provide information without considering the user's emotional state. Furthermore, there is a lack of technology to automatically provide information based on emotional state, making it difficult to meet the individual needs of users. In particular, there is a need for information delivery that takes users' emotions into account in various situations in daily life, but there is no effective method to achieve this.
[0154] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0155] In this invention, the server includes a device for capturing visual information, a device for acquiring acoustic information, a device for acquiring spatial information, means for analyzing emotional states, means for generating information based on the emotional analysis results using a generative model, means for providing information adjusted according to the user's emotional state, and means for updating the information based on user input. This makes it possible to accurately provide information according to the user's emotional state and optimize the user experience.
[0156] A "device for capturing visual information" is a device that has the function of acquiring the user's visual data, and includes cameras, image sensors, and the like.
[0157] "Devices for acquiring acoustic information" are devices that collect the user's voice and ambient sounds, and include microphones and acoustic sensors.
[0158] "Devices for acquiring spatial information" refer to devices that acquire location-related data such as the user's current location and movement history, and include GPS receivers and location services.
[0159] "Means for analyzing emotional states" refer to mechanisms for analyzing a user's emotions and determining emotional states such as joy or sadness, and these may include software algorithms or emotion analysis engines.
[0160] "Means of generating information based on sentiment analysis results using generative models" refers to methods for generating optimal information while considering the user's emotional state, and includes generative AI models and machine learning algorithms.
[0161] "Means of providing information tailored to the user's emotional state" refers to means of providing information optimized for the user based on analyzed emotional data, and transmitting this information via display devices or notification systems.
[0162] "Means for updating information based on user input" refers to a mechanism for updating information in real time in response to user requests and feedback, and acquires information through interfaces, input devices, etc.
[0163] This system combines advanced data collection and analysis technologies to analyze user emotions in real time and provide appropriate information. Specifically, the device is equipped with a camera to capture visual information, a microphone to acquire acoustic information, and a GPS receiver to acquire spatial information. It also incorporates an emotion analysis engine to analyze emotional states.
[0164] The server receives various data transmitted from the terminal and analyzes the user's emotions via an emotion analysis engine. A generative AI model is used for the analysis, and data processing is performed based on the emotional state. This generative AI model uses algorithms designed to precisely capture the user's emotional state and generates appropriate information.
[0165] Information tailored to the user's emotions is provided to the user through the device. For example, if the system analyzes that the user is emotionally fatigued, the server provides information on nearby rest spots and relaxing facilities. This allows the user to choose appropriate actions based on their emotional state.
[0166] For example, if data acquired by a device while a user is visiting a museum detects "fatigue," the server immediately generates information about nearby cafes and notifies the user. This allows the user to take a break at an appropriate time.
[0167] An example of a prompt message is to input a specific situation, such as "How would you provide information if the user is tired?", into the generating AI model, and then derive the optimal method for providing information from its response.
[0168] This system enables the provision of information that takes user emotions into account, significantly improving the user experience.
[0169] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0170] Step 1:
[0171] The device collects environmental data.
[0172] The device captures visual information of the user's surroundings with a camera and acquires acoustic information with a microphone. Simultaneously, it obtains current location information using a GPS receiver. This allows the user's current environment and activity status to be input as data. This input data serves as the foundational data necessary to analyze the user's emotions.
[0173] Step 2:
[0174] The device analyzes emotional data.
[0175] The system uses an emotion analysis engine to analyze the user's facial expressions and voice tone based on acquired visual and auditory information. This analysis determines the user's emotional state, such as joy, anger, or fatigue. Image recognition and voice analysis technologies are used for data processing. The analysis results are output as data indicating the user's emotional state.
[0176] Step 3:
[0177] The terminal sends the analysis data to the server.
[0178] The device sends analyzed emotion data and location information to the server. The transmitted data includes the user's current emotional state and location information, which then serves as input for the next stage of the information generation process.
[0179] Step 4:
[0180] The server performs an integrated analysis of the data.
[0181] The server analyzes emotional data and location information sent from the terminal. Using a generative AI model, it processes the data to generate appropriate information based on the user's emotional state. For example, based on the emotional state of "tired" and the location information of "tourist spot," it generates recommendations for rest spots.
[0182] Step 5:
[0183] The server provides the adjusted information.
[0184] The generated information is sent from the server to the user's device, which then displays the information on its screen. The information is intended to help the user take actions that align with their current emotional state. For example, it might include directions to a nearby cafe or suggestions for relaxation spots. This allows the user to make decisions that are appropriate to their emotional state.
[0185] (Application Example 2)
[0186] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0187] Traditional information delivery systems often fail to adequately consider the user's emotional state when presenting information, resulting in an unoptimized user experience. For example, in in-store customer service support, it was difficult to provide appropriate responses based on the customer's emotional state, making it difficult to improve customer satisfaction.
[0188] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0189] In this invention, the server includes means for analyzing visual data using a generative artificial intelligence model, means for analyzing audio data, and means for analyzing emotional data to recognize the user's emotional state. This makes it possible to adjust the information provision method based on the user's emotional state and provide customer service support that is tailored to the emotions of the customers.
[0190] A "generative artificial intelligence model" is an artificial intelligence system that has the ability to analyze data and automatically recognize patterns and features.
[0191] "Visual data" refers to information related to images and videos acquired by cameras and sensors.
[0192] "Audio data" refers to sound information collected by audio input devices such as microphones.
[0193] "Location information" refers to information about the current location of a user or device, obtained using GPS or other positioning technologies.
[0194] "Analysis results" refer to specific conclusions or evaluation values obtained after analyzing data.
[0195] "Emotional data" refers to information that indicates the emotional state of a user, estimated from their facial expressions, tone of voice, body movements, etc.
[0196] "Adjusting the method of providing information" means changing the content and format of the information presented to users based on the analyzed sentiment data.
[0197] "Customer service support" refers to providing assistance to sales staff and employees so that they can take appropriate actions and make appropriate suggestions to improve customer satisfaction.
[0198] To realize this system, smart glasses will be used as the main hardware. This will allow the device to acquire visual and audio data and analyze it using an emotion engine. Specifically, the smart glasses are equipped with a camera and microphone, allowing them to collect customers' facial expressions and voice tones in real time. Location information will be acquired using a built-in GPS sensor.
[0199] The server uses open-source facial recognition libraries (e.g., OpenCV) and speech analysis libraries (e.g., librosa) to analyze this data and determine the user's emotional state. Subsequently, a generative AI model (e.g., TENSORFLOW®) is used to comprehensively analyze the data.
[0200] Based on the analysis results, the server will, for example, determine that "the customer is feeling stressed," and then provide specific customer service suggestions to the sales staff, such as "it would be good to suggest products that will help the customer relax." This information is overlaid on the smart glasses' display.
[0201] Users can provide feedback to the smart glasses based on the reactions of customers, and this feedback is sent to the server, where the information is updated. This allows the system to continuously learn, making the customer service assistance provided more precise and effective.
[0202] For example, when a customer is struggling to decide which product to choose from the displayed items, the AI model can be given a prompt message that says, "Determine if the customer is experiencing stress and display more detailed information about the product."
[0203] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0204] Step 1:
[0205] The device uses the camera and microphone of smart glasses to acquire visual and audio data from customers. This data includes facial expressions and voice tone. The input is raw visual and audio data, which is a preparatory step for identifying emotional information based on faces and voices.
[0206] Step 2:
[0207] The device analyzes visual data using an open-source face recognition library (e.g., OpenCV) to estimate emotional states from facial expressions. Similarly, it analyzes audio data using a speech analysis library (e.g., librosa) to estimate emotions from voice tone. Inputs are the acquired visual and audio data, and outputs are the analysis results of the respective emotional states.
[0208] Step 3:
[0209] The results of this emotional state analysis, along with location information including GPS sensor data, are sent to the server. The server receives this data, analyzes it using a generative AI model (e.g., TensorFlow), and determines the overall emotional state. The input is the transmitted emotional analysis results and location information, and the output is the integrated emotional state evaluation result.
[0210] Step 4:
[0211] The server generates informational messages and customer service suggestions using prompts based on the evaluation results of the generated emotional state. These prompts are generated in the format of, for example, "Determine if the customer is stressed and display detailed information about the product." The input is an integrated emotional state, and the output is a prompt that provides specific suggestions or guidelines for action.
[0212] Step 5:
[0213] The generated prompt text is returned to the terminal and overlaid on the smart glasses' display. Based on this, the user can provide quick and appropriate service to customers. The input is the generated prompt text, and the output is the service guidance displayed on the smart glasses' display.
[0214] Step 6:
[0215] Feedback obtained by users through interactions with customers is sent to the server via the terminal, and the system updates the information, which is then used for future analyses and recommendations. In this way, the system is continuously improved, and its accuracy is enhanced. The input is user feedback, and the output is updated information and preparation for the next analysis.
[0216] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0217] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0218] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0219] [Second Embodiment]
[0220] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0221] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0222] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0223] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0224] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0225] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0226] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0227] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0228] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0229] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0230] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0231] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0232] This invention relates to an advanced information provision system that utilizes a generative artificial intelligence model, and is operated primarily by analyzing visual data, audio data, and location information. The system is designed to provide information based on the individual needs of users in real time.
[0233] First, the device worn or carried by the user (such as smart glasses) continuously collects data about the surrounding environment and circumstances. This device is equipped with sensors such as a camera, microphone, and GPS, and can acquire visual data, audio data, and location information. This data is analyzed by the device, stored in temporary memory, and then sent to a server.
[0234] The server uses a generative artificial intelligence model to analyze the received data. Specifically, it identifies visual data using image processing technology and understands audio data using natural language processing. It also understands the current location and surrounding environment based on location information. By comprehensively combining this information, it identifies the information and services the user is looking for and generates personalized suggestions.
[0235] For example, when a user is walking through a shopping mall, the device captures visual data of displayed products and sends it to a server. The server uses generated AI to identify the products and provides information on the lowest prices compared to other online platforms. Furthermore, when a user is in a tourist area, the device can analyze information on buildings and artworks captured by the camera and display appropriate explanations in real time.
[0236] In the medical field, healthcare professionals can use terminals to access medical information tailored to the patient's condition. For example, during pre-operative preparation, the system can provide optimal procedures and medication information based on the patient's health status, supporting smooth medical treatment.
[0237] This system provides the most relevant information instantly by continuously updating information in response to user requests. Users can request further details by giving instructions, for example, via voice commands or gestures. As a result, users can obtain the necessary information at the optimal time, enabling efficient decision-making in today's information-saturated society.
[0238] The following describes the processing flow.
[0239] Step 1:
[0240] The device activates sensors when the user begins using it, collecting visual, audio, and location data in real time. The camera captures the surrounding video, the microphone records audio, and GPS determines the user's location.
[0241] Step 2:
[0242] The terminal stores the collected data in temporary memory, compresses the data, and then sends it to the server via a secure protocol. This process minimizes latency during data transfer.
[0243] Step 3:
[0244] The server prepares the received data for analysis. Visual data is identified using image recognition algorithms to identify objects and text, and audio data is converted into language using speech recognition algorithms. Location information is merged with a mapping service to determine the current location.
[0245] Step 4:
[0246] The server applies an artificial intelligence model based on the analysis results to generate information and services tailored to the user. Examples include information on the lowest prices for shopping, historical background and explanations of tourist destinations, and diagnostic information in medical support.
[0247] Step 5:
[0248] The server sends the generated information to the terminal. The terminal receives this information and provides the necessary information by displaying it as an overlay in the user's field of view. The display timing and format are customized according to the user's situation.
[0249] Step 6:
[0250] Users can request additional information or select and manipulate specific data using voice or gestures. The terminal sends these inputs to a server for information updates and further analysis.
[0251] Step 7:
[0252] The server re-analyzes the information based on user feedback and requests, and resends the newly generated information to the terminal. This ensures that the latest information tailored to the user's needs is always provided.
[0253] (Example 1)
[0254] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0255] In modern society, information overload has become the norm, making it difficult for users to quickly and accurately obtain the information they need. This is especially true while on the go or in complex environments, where information acquisition and decision-making become even more challenging. To address this problem, there is a need for systems that provide appropriate information in real time, tailored to the user's situation and needs, and support efficient decision-making.
[0256] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0257] In this invention, the server includes information provision means equipped with a device for collecting data from the surrounding environment, means for transmitting the collected data via a communication line, and means for analyzing visual information using image recognition technology. This makes it possible to provide users with necessary information in real time and support their decision-making.
[0258] A "device that collects data from the surrounding environment" is a device that acquires data such as visual, auditory, and location information using sensors, and understands the state of the user's surroundings in real time.
[0259] "Information provision means" refers to methods and technologies for presenting appropriate information to users based on collected and analyzed data.
[0260] "Means of transmission via communication lines" refers to methods using wireless or wired communication technology to transfer collected data to external devices such as servers.
[0261] "Means of analyzing visual information using image recognition technology" refers to techniques for analyzing acquired visual data for object identification and feature extraction, and utilizes image processing algorithms.
[0262] "Means of analyzing acoustic information using natural language processing technology" refers to technologies that convert audio information into text and analyze its meaning and intent.
[0263] "Methods for understanding spatial information using location data" refers to technologies that utilize location information such as GPS to identify a user's current location and surrounding environment.
[0264] "Means of generating information based on analysis results using generative artificial intelligence models" refers to a method of generating optimal information for users based on analyzed data using AI technology.
[0265] "Means for processing additional information requests from users via prompt messages" refers to interfaces and technologies that understand and provide appropriate responses or information when a user requests additional information via voice or text.
[0266] This invention is a system that provides users with appropriate information in real time using a generative artificial intelligence model. The system primarily utilizes the following devices and technologies.
[0267] The device is equipped with various sensors to collect data from its surroundings. Specifically, these include a camera to acquire visual information, a microphone to acquire acoustic information, and a GPS unit to collect location data. This device is often implemented as smart glasses or other wearable devices. The collected data is initially processed within the device and then transmitted to a server via a communication line.
[0268] The server analyzes the received data. For visual information, image recognition technology is used to recognize objects. In this case, existing image processing libraries such as OpenCV are often used. Acoustic information is analyzed using natural language processing technology, and text conversion is performed using Google Cloud Speech-to-Text, etc. Location information is compared with map data to understand the user's spatial environment.
[0269] A generative AI model is implemented to generate information tailored to the user's needs based on analysis results. This generated information is displayed overlaid on the user's field of view through a visual device. For example, if a user is in a shopping mall, they can view comparison information on the lowest prices of products they have photographed. Furthermore, in tourist areas, it can provide real-time historical background information about buildings. In the medical field, it can also present optimal diagnostic information based on the patient's condition.
[0270] Users can obtain additional information or details by sending prompts via voice commands or text input. Examples of prompts include "Tell me more about this product" or "Tell me the nearest cafe." The server receives these prompts and provides the user with more detailed information or service suggestions.
[0271] As described above, this invention can quickly and efficiently provide users with the information they need in today's information-saturated world.
[0272] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0273] Step 1:
[0274] The device collects data from its surroundings. Inputs include visual data, acoustic data, and location data. This data is acquired through sensors such as cameras, microphones, and GPS. Visual data is saved as image files, and acoustic data is converted into audio files. Location data is acquired as latitude and longitude information. The output is a formatted dataset.
[0275] Step 2:
[0276] The device temporarily stores the collected data and sends it to the server via a communication line. Specifically, the collected data is stored in temporary memory and sent to the server using Wi-Fi or mobile data communication. The input is a formatted dataset, and the output is a data stream received on the server side.
[0277] Step 3:
[0278] The server analyzes the received visual data using image recognition technology. The input is an image file. Specifically, the server uses OpenCV or a similar image processing library to identify objects within the image. The output is a list of recognized objects.
[0279] Step 4:
[0280] The server analyzes audio data using natural language processing techniques. The input is an audio file. The server uses a speech recognition service such as Google Cloud Speech-to-Text to convert the audio to text. It then analyzes the text content to extract meaning. The output is the analyzed text information.
[0281] Step 5:
[0282] The server analyzes spatial information based on location data. The input is latitude and longitude information. Specifically, it identifies the user's current location and surrounding information in combination with map data. The output is environmental information related to the location.
[0283] Step 6:
[0284] The server uses a generative AI model to generate information based on the analysis results. The input is the analysis results of visual information, acoustic information, and location information. The server uses generative AI technology to create information and proposals optimal for the user's situation. The output is information personalized for the user.
[0285] Step 7:
[0286] The terminal notifies the user of the generated information. The input is the personalized information sent from the server. The terminal displays the information on a display and provides voice guidance if necessary. The output is information received by the user visually and auditorily.
[0287] Step 8:
[0288] The user can make additional information requests using a prompt sentence. The input is a prompt sentence in voice or text format. The terminal sends this request to the server, and additional analysis is performed. The output is more detailed information provided to the user.
[0289] (Application Example 1)
[0290] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0291] In today's consumer society, consumers are surrounded by a vast amount of product information, making it difficult to make the best choice. Furthermore, the market contains numerous products with varying prices, performance, and reviews, and comparing them in real time requires considerable time and effort. Moreover, there is a lack of convenient ways for users to obtain this information during their physical shopping experience.
[0292] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0293] In this invention, the server includes means for analyzing visual data, means for handling audio data, means for acquiring location information, and means for recognizing products and presenting detailed information in real time. This enables users to efficiently collect necessary product information when shopping and make optimal purchasing decisions.
[0294] A "generative artificial intelligence model" is an artificial intelligence trained using machine learning techniques to perform a specific task from a vast amount of data.
[0295] "Means for analyzing visual data" refers to devices or software that analyze image or video data using digital image processing technology to recognize objects or scenes.
[0296] "Means for analyzing audio data" refers to devices or software that use speech recognition technology to extract text or intent from audio.
[0297] "Means of acquiring location information" refers to technologies that use GPS or other positioning systems to determine the current location of a device.
[0298] "Means of providing information to users" refers to a system that displays or communicates useful information to users through a user interface based on analysis results.
[0299] A "means for updating information based on user requests" refers to a system that has the ability to dynamically re-evaluate, modify, or update the information it provides in response to user input or requests.
[0300] "Means for identifying products and presenting related information in real time" refers to technology that uses camera data or similar methods to identify products and immediately displays information related to those products to the user.
[0301] The system of this invention operates on the basis of cooperation between a user-carried terminal and a server. The terminal is equipped with sensors such as a camera, microphone, and GPS, which are used to acquire visual data, audio data, and location information of the environment. The terminal has the function of collecting this data in real time and transmitting it to the server.
[0302] The server performs image analysis on visual data using a generative AI model to identify objects and products. Audio data is analyzed by speech recognition software to understand user requests in natural language. Location information is used to determine the current location and understand the context of the environment.
[0303] The server's analysis results are returned to the user's device and displayed as an overlay in the user's field of view. This allows for real-time presentation of detailed product information and price comparisons, supporting the user's decision-making.
[0304] For example, when a user is looking at clothes in a store, the device acquires visual data and sends it to the server. The server uses a generative AI model to identify the clothes and retrieves relevant data by comparing it with registered information. This includes the product price, customer reviews, and price comparisons with other stores. If the user gives a voice command such as "Tell me the reviews for this product," the server retrieves and displays the review information.
[0305] An example of a prompt might be, "Identify the products in the image and tell me their price and description." Based on this prompt, the server uses a generative AI model to select the information to provide to the user.
[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0307] Step 1:
[0308] The terminal acquires the surrounding visual data and audio data using the camera and microphone. The input is the raw data from the environment, and the output is the digital format of this data. Thereby, the terminal captures the objects within the user's field of view and the surrounding conversations.
[0309] Step 2:
[0310] The terminal transmits the acquired visual data and audio data to the server. The input is the digital data held by the terminal, and the output is the signal to the server. Thereby, the server receives the information necessary for analysis.
[0311] Step 3:
[0312] The server analyzes the visual data using the generated artificial intelligence model to identify the objects and products within the image. The input is the visual data, and the output is the identification information of the identified object. As a specific operation, the data is analyzed by an image recognition algorithm.
[0313] Step 4:
[0314] The server uses voice recognition software to analyze the audio data and textify the user's instructions. The input is the audio data, and the output is the analyzed text information. The extraction of voice commands is the specific operation.
[0315] Step 5:
[0316] The server searches for relevant information for the identified object and generates information in real time. The input is the identification information of the identified object and the voice command, and the output is the displayable relevant information. Information retrieval from the database is performed.
[0317] Step 6:
[0318] The server sends the generated information to the terminal and displays it as an overlay in the user's field of view. The input is a dataset of related information, and the output is a visual presentation on the user's display device. The terminal then reactivates the overlay function.
[0319] Step 7:
[0320] The user requests further details using voice commands or input devices as needed. Input is the user's new instructions, and output is an updated display of information. This cycle is repeated to support the user's decision-making.
[0321] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0322] This invention relates to an information provision system incorporating an emotion engine that recognizes user emotions, thereby enabling the provision of information and services that take into account the user's emotional state. This system has the capability to analyze user emotion data in addition to visual data, audio data, and location information.
[0323] The device is equipped with a camera, microphone, GPS sensor, and emotion engine sensor, which are used to simultaneously collect the user's environmental information and emotional state. The emotion engine analyzes the user's emotions from their facial expressions, voice tone, and body movements. This data collection is performed continuously in the background during the user's natural behavior, and the device sends this data to a server.
[0324] The server comprehensively analyzes various data and generates information using an artificial intelligence model, combining it with the sentiment analysis results from the emotion engine. This system aims to optimize the user experience by selecting the appropriate way to provide information based on the user's emotional state.
[0325] For example, if a user experiences discomfort or fatigue while sightseeing, the system's emotion engine detects this and prioritizes providing information on relaxing places and rest spots rather than typical tourist information. If the user is shopping and their stress level is high, the system will offer suggestions to assist with simple transactions or purchases.
[0326] Furthermore, in the medical field, data obtained from the emotion engine can be used to provide more humane and efficient support when making diagnoses and providing care while considering the patient's emotional responses. This allows healthcare professionals to provide care that is tailored to the patient's psychological state.
[0327] By constantly updating information based on user requests and feedback, and taking appropriate action at each step, the system functions effectively under diverse circumstances. This ensures that users always receive the best information and services, leading to a more comfortable and satisfying experience.
[0328] The following describes the processing flow.
[0329] Step 1:
[0330] The device activates its camera, microphone, GPS sensor, and emotion engine sensors to collect visual data, audio data, location information, and emotion data in real time. The camera captures the user's facial expressions, the microphone records voice tones, and the sensors detect the user's body movements, etc.
[0331] Step 2:
[0332] The terminal sends the collected data to the server. Upon receiving the data, the server prepares it for analysis. This includes data preprocessing and compression as needed.
[0333] Step 3:
[0334] The server uses image recognition algorithms to analyze visual data and identify objects and text in the user's environment.
[0335] Step 4:
[0336] The server applies a speech recognition algorithm, converts the audio data into text format, and performs linguistic analysis. This allows it to understand the content of the user's speech and the nuances of their emotions.
[0337] Step 5:
[0338] The server analyzes emotional data acquired by the emotion engine to identify the user's current emotional state (e.g., happiness, anxiety, excitement, fatigue). This result is then combined with other data to provide a comprehensive assessment of the user's situation.
[0339] Step 6:
[0340] The server uses a generative artificial intelligence model to generate information and services best suited to the user based on collected information and sentiment analysis results. For example, if it detects that the user is tired, it prioritizes providing information about rest spots.
[0341] Step 7:
[0342] The server sends the generated information to the terminal. The terminal receives this information and displays it as an overlay in the user's field of view. The displayed content is customized to include information that responds to the user's emotions.
[0343] Step 8:
[0344] Users can request additional information or manipulate the provided information using voice commands or gestures. The terminal sends this input to the server, which regenerates the information as needed.
[0345] Step 9:
[0346] The server updates information based on user feedback and sends the new information to the device. This ensures that users always receive up-to-date information that reflects their emotional state.
[0347] (Example 2)
[0348] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0349] Conventional information delivery systems have the problem of not optimizing the user experience because they provide information without considering the user's emotional state. Furthermore, there is a lack of technology to automatically provide information based on emotional state, making it difficult to meet the individual needs of users. In particular, there is a need for information delivery that takes users' emotions into account in various situations in daily life, but there is no effective method to achieve this.
[0350] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0351] In this invention, the server includes a device for capturing visual information, a device for acquiring acoustic information, a device for acquiring spatial information, means for analyzing emotional states, means for generating information based on the emotional analysis results using a generative model, means for providing information adjusted according to the user's emotional state, and means for updating the information based on user input. This makes it possible to accurately provide information according to the user's emotional state and optimize the user experience.
[0352] A "device for capturing visual information" is a device that has the function of acquiring the user's visual data, and includes cameras, image sensors, and the like.
[0353] "Devices for acquiring acoustic information" are devices that collect the user's voice and ambient sounds, and include microphones and acoustic sensors.
[0354] "Devices for acquiring spatial information" refer to devices that acquire location-related data such as the user's current location and movement history, and include GPS receivers and location services.
[0355] "Means for analyzing emotional states" refer to mechanisms for analyzing a user's emotions and determining emotional states such as joy or sadness, and these may include software algorithms or emotion analysis engines.
[0356] "Means of generating information based on sentiment analysis results using generative models" refers to methods for generating optimal information while considering the user's emotional state, and includes generative AI models and machine learning algorithms.
[0357] "Means of providing information tailored to the user's emotional state" refers to means of providing information optimized for the user based on analyzed emotional data, and transmitting this information via display devices or notification systems.
[0358] "Means for updating information based on user input" refers to a mechanism for updating information in real time in response to user requests and feedback, and acquires information through interfaces, input devices, etc.
[0359] This system combines advanced data collection and analysis technologies to analyze user emotions in real time and provide appropriate information. Specifically, the device is equipped with a camera to capture visual information, a microphone to acquire acoustic information, and a GPS receiver to acquire spatial information. It also incorporates an emotion analysis engine to analyze emotional states.
[0360] The server receives various data transmitted from the terminal and analyzes the user's emotions via an emotion analysis engine. A generative AI model is used for the analysis, and data processing is performed based on the emotional state. This generative AI model uses algorithms designed to precisely capture the user's emotional state and generates appropriate information.
[0361] Information tailored to the user's emotions is provided to the user through the device. For example, if the system analyzes that the user is emotionally fatigued, the server provides information on nearby rest spots and relaxing facilities. This allows the user to choose appropriate actions based on their emotional state.
[0362] For example, if data acquired by a device while a user is visiting a museum detects "fatigue," the server immediately generates information about nearby cafes and notifies the user. This allows the user to take a break at an appropriate time.
[0363] An example of a prompt message is to input a specific situation, such as "How would you provide information if the user is tired?", into the generating AI model, and then derive the optimal method for providing information from its response.
[0364] This system enables the provision of information that takes user emotions into account, significantly improving the user experience.
[0365] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0366] Step 1:
[0367] The device collects environmental data.
[0368] The device captures visual information of the user's surroundings with a camera and acquires acoustic information with a microphone. Simultaneously, it obtains current location information using a GPS receiver. This allows the user's current environment and activity status to be input as data. This input data serves as the foundational data necessary to analyze the user's emotions.
[0369] Step 2:
[0370] The device analyzes emotional data.
[0371] The system uses an emotion analysis engine to analyze the user's facial expressions and voice tone based on acquired visual and auditory information. This analysis determines the user's emotional state, such as joy, anger, or fatigue. Image recognition and voice analysis technologies are used for data processing. The analysis results are output as data indicating the user's emotional state.
[0372] Step 3:
[0373] The terminal sends the analysis data to the server.
[0374] The device sends analyzed emotion data and location information to the server. The transmitted data includes the user's current emotional state and location information, which then serves as input for the next stage of the information generation process.
[0375] Step 4:
[0376] The server performs an integrated analysis of the data.
[0377] The server analyzes emotional data and location information sent from the terminal. Using a generative AI model, it processes the data to generate appropriate information based on the user's emotional state. For example, based on the emotional state of "tired" and the location information of "tourist spot," it generates recommendations for rest spots.
[0378] Step 5:
[0379] The server provides the adjusted information.
[0380] The generated information is sent from the server to the user's device, which then displays the information on its screen. The information is intended to help the user take actions that align with their current emotional state. For example, it might include directions to a nearby cafe or suggestions for relaxation spots. This allows the user to make decisions that are appropriate to their emotional state.
[0381] (Application Example 2)
[0382] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0383] Traditional information delivery systems often fail to adequately consider the user's emotional state when presenting information, resulting in an unoptimized user experience. For example, in in-store customer service support, it was difficult to provide appropriate responses based on the customer's emotional state, making it difficult to improve customer satisfaction.
[0384] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0385] In this invention, the server includes means for analyzing visual data using a generative artificial intelligence model, means for analyzing audio data, and means for analyzing emotional data to recognize the user's emotional state. This makes it possible to adjust the information provision method based on the user's emotional state and provide customer service support that is tailored to the emotions of the customers.
[0386] A "generative artificial intelligence model" is an artificial intelligence system that has the ability to analyze data and automatically recognize patterns and features.
[0387] "Visual data" refers to information related to images and videos acquired by cameras and sensors.
[0388] "Audio data" refers to sound information collected by audio input devices such as microphones.
[0389] "Location information" refers to information about the current location of a user or device, obtained using GPS or other positioning technologies.
[0390] "Analysis results" refer to specific conclusions or evaluation values obtained after analyzing data.
[0391] "Emotional data" refers to information that indicates the emotional state of a user, estimated from their facial expressions, tone of voice, body movements, etc.
[0392] "Adjusting the method of providing information" means changing the content and format of the information presented to users based on the analyzed sentiment data.
[0393] "Customer service support" refers to providing assistance to sales staff and employees so that they can take appropriate actions and make appropriate suggestions to improve customer satisfaction.
[0394] To realize this system, smart glasses will be used as the main hardware. This will allow the device to acquire visual and audio data and analyze it using an emotion engine. Specifically, the smart glasses are equipped with a camera and microphone, allowing them to collect customers' facial expressions and voice tones in real time. Location information will be acquired using a built-in GPS sensor.
[0395] The server uses open-source facial recognition libraries (e.g., OpenCV) and speech analysis libraries (e.g., librosa) to analyze this data and determine the user's emotional state. Subsequently, a generative AI model (e.g., TensorFlow) is used to comprehensively analyze the data.
[0396] Based on the analysis results, the server will, for example, determine that "the customer is feeling stressed," and then provide specific customer service suggestions to the sales staff, such as "it would be good to suggest products that will help the customer relax." This information is overlaid on the smart glasses' display.
[0397] Users can provide feedback to the smart glasses based on the reactions of customers, and this feedback is sent to the server, where the information is updated. This allows the system to continuously learn, making the customer service assistance provided more precise and effective.
[0398] For example, when a customer is struggling to decide which product to choose from the displayed items, the AI model can be given a prompt message that says, "Determine if the customer is experiencing stress and display more detailed information about the product."
[0399] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0400] Step 1:
[0401] The device uses the camera and microphone of smart glasses to acquire visual and audio data from customers. This data includes facial expressions and voice tone. The input is raw visual and audio data, which is a preparatory step for identifying emotional information based on faces and voices.
[0402] Step 2:
[0403] The device analyzes visual data using an open-source face recognition library (e.g., OpenCV) to estimate emotional states from facial expressions. Similarly, it analyzes audio data using a speech analysis library (e.g., librosa) to estimate emotions from voice tone. Inputs are the acquired visual and audio data, and outputs are the analysis results of the respective emotional states.
[0404] Step 3:
[0405] The results of this emotional state analysis, along with location information including GPS sensor data, are sent to the server. The server receives this data, analyzes it using a generative AI model (e.g., TensorFlow), and determines the overall emotional state. The input is the transmitted emotional analysis results and location information, and the output is the integrated emotional state evaluation result.
[0406] Step 4:
[0407] The server generates informational messages and customer service suggestions using prompts based on the evaluation results of the generated emotional state. These prompts are generated in the format of, for example, "Determine if the customer is stressed and display detailed information about the product." The input is an integrated emotional state, and the output is a prompt that provides specific suggestions or guidelines for action.
[0408] Step 5:
[0409] The generated prompt text is returned to the terminal and overlaid on the smart glasses' display. Based on this, the user can provide quick and appropriate service to customers. The input is the generated prompt text, and the output is the service guidance displayed on the smart glasses' display.
[0410] Step 6:
[0411] Feedback obtained by users through interactions with customers is sent to the server via the terminal, and the system updates the information, which is then used for future analyses and recommendations. In this way, the system is continuously improved, and its accuracy is enhanced. The input is user feedback, and the output is updated information and preparation for the next analysis.
[0412] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0413] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0414] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0415] [Third Embodiment]
[0416] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0417] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0418] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0419] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0420] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0421] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0422] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0423] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0424] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0425] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0426] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0427] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0428] This invention relates to an advanced information provision system that utilizes a generative artificial intelligence model, and is operated primarily by analyzing visual data, audio data, and location information. The system is designed to provide information based on the individual needs of users in real time.
[0429] First, the device worn or carried by the user (such as smart glasses) continuously collects data about the surrounding environment and circumstances. This device is equipped with sensors such as a camera, microphone, and GPS, and can acquire visual data, audio data, and location information. This data is analyzed by the device, stored in temporary memory, and then sent to a server.
[0430] The server uses a generative artificial intelligence model to analyze the received data. Specifically, it identifies visual data using image processing technology and understands audio data using natural language processing. It also understands the current location and surrounding environment based on location information. By comprehensively combining this information, it identifies the information and services the user is looking for and generates personalized suggestions.
[0431] For example, when a user is walking through a shopping mall, the device captures visual data of displayed products and sends it to a server. The server uses generated AI to identify the products and provides information on the lowest prices compared to other online platforms. Furthermore, when a user is in a tourist area, the device can analyze information on buildings and artworks captured by the camera and display appropriate explanations in real time.
[0432] In the medical field, healthcare professionals can use terminals to access medical information tailored to the patient's condition. For example, during pre-operative preparation, the system can provide optimal procedures and medication information based on the patient's health status, supporting smooth medical treatment.
[0433] This system provides the most relevant information instantly by continuously updating information in response to user requests. Users can request further details by giving instructions, for example, via voice commands or gestures. As a result, users can obtain the necessary information at the optimal time, enabling efficient decision-making in today's information-saturated society.
[0434] The following describes the processing flow.
[0435] Step 1:
[0436] The device activates sensors when the user begins using it, collecting visual, audio, and location data in real time. The camera captures the surrounding video, the microphone records audio, and GPS determines the user's location.
[0437] Step 2:
[0438] The terminal stores the collected data in temporary memory, compresses the data, and then sends it to the server via a secure protocol. This process minimizes latency during data transfer.
[0439] Step 3:
[0440] The server prepares the received data for analysis. Visual data is identified using image recognition algorithms to identify objects and text, and audio data is converted into language using speech recognition algorithms. Location information is merged with a mapping service to determine the current location.
[0441] Step 4:
[0442] The server applies an artificial intelligence model based on the analysis results to generate information and services tailored to the user. Examples include information on the lowest prices for shopping, historical background and explanations of tourist destinations, and diagnostic information in medical support.
[0443] Step 5:
[0444] The server sends the generated information to the terminal. The terminal receives this information and provides the necessary information by displaying it as an overlay in the user's field of view. The display timing and format are customized according to the user's situation.
[0445] Step 6:
[0446] Users can request additional information or select and manipulate specific data using voice or gestures. The terminal sends these inputs to a server for information updates and further analysis.
[0447] Step 7:
[0448] The server re-analyzes the information based on user feedback and requests, and resends the newly generated information to the terminal. This ensures that the latest information tailored to the user's needs is always provided.
[0449] (Example 1)
[0450] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0451] In modern society, information overload has become the norm, making it difficult for users to quickly and accurately obtain the information they need. This is especially true while on the go or in complex environments, where information acquisition and decision-making become even more challenging. To address this problem, there is a need for systems that provide appropriate information in real time, tailored to the user's situation and needs, and support efficient decision-making.
[0452] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0453] In this invention, the server includes information provision means equipped with a device for collecting data from the surrounding environment, means for transmitting the collected data via a communication line, and means for analyzing visual information using image recognition technology. This makes it possible to provide users with necessary information in real time and support their decision-making.
[0454] A "device that collects data from the surrounding environment" is a device that acquires data such as visual, auditory, and location information using sensors, and understands the state of the user's surroundings in real time.
[0455] "Information provision means" refers to methods and technologies for presenting appropriate information to users based on collected and analyzed data.
[0456] "Means of transmission via communication lines" refers to methods using wireless or wired communication technology to transfer collected data to external devices such as servers.
[0457] "Means of analyzing visual information using image recognition technology" refers to techniques for analyzing acquired visual data for object identification and feature extraction, and utilizes image processing algorithms.
[0458] "Means of analyzing acoustic information using natural language processing technology" refers to technologies that convert audio information into text and analyze its meaning and intent.
[0459] "Methods for understanding spatial information using location data" refers to technologies that utilize location information such as GPS to identify a user's current location and surrounding environment.
[0460] "Means of generating information based on analysis results using generative artificial intelligence models" refers to a method of generating optimal information for users based on analyzed data using AI technology.
[0461] "Means for processing additional information requests from users via prompt messages" refers to interfaces and technologies that understand and provide appropriate responses or information when a user requests additional information via voice or text.
[0462] This invention is a system that provides users with appropriate information in real time using a generative artificial intelligence model. The system primarily utilizes the following devices and technologies.
[0463] The device is equipped with various sensors to collect data from its surroundings. Specifically, these include a camera to acquire visual information, a microphone to acquire acoustic information, and a GPS unit to collect location data. This device is often implemented as smart glasses or other wearable devices. The collected data is initially processed within the device and then transmitted to a server via a communication line.
[0464] The server analyzes the received data. For visual information, image recognition technology is used to recognize objects. In this case, existing image processing libraries such as OpenCV are often used. Acoustic information is analyzed using natural language processing technology, and text conversion is performed using Google Cloud Speech-to-Text, etc. Location information is compared with map data to understand the user's spatial environment.
[0465] A generative AI model is implemented to generate information tailored to the user's needs based on analysis results. This generated information is displayed overlaid on the user's field of view through a visual device. For example, if a user is in a shopping mall, they can view comparison information on the lowest prices of products they have photographed. Furthermore, in tourist areas, it can provide real-time historical background information about buildings. In the medical field, it can also present optimal diagnostic information based on the patient's condition.
[0466] Users can obtain additional information or details by sending prompts via voice commands or text input. Examples of prompts include "Tell me more about this product" or "Tell me the nearest cafe." The server receives these prompts and provides the user with more detailed information or service suggestions.
[0467] As described above, this invention can quickly and efficiently provide users with the information they need in today's information-saturated world.
[0468] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0469] Step 1:
[0470] The device collects data from its surroundings. Inputs include visual data, acoustic data, and location data. This data is acquired through sensors such as cameras, microphones, and GPS. Visual data is saved as image files, and acoustic data is converted into audio files. Location data is acquired as latitude and longitude information. The output is a formatted dataset.
[0471] Step 2:
[0472] The device temporarily stores the collected data and sends it to the server via a communication line. Specifically, the collected data is stored in temporary memory and sent to the server using Wi-Fi or mobile data communication. The input is a formatted dataset, and the output is a data stream received on the server side.
[0473] Step 3:
[0474] The server analyzes the received visual data using image recognition technology. The input is an image file. Specifically, the server uses OpenCV or a similar image processing library to identify objects within the image. The output is a list of recognized objects.
[0475] Step 4:
[0476] The server analyzes audio data using natural language processing techniques. The input is an audio file. The server uses a speech recognition service such as Google Cloud Speech-to-Text to convert the audio to text. It then analyzes the text content to extract meaning. The output is the analyzed text information.
[0477] Step 5:
[0478] The server analyzes spatial information based on location data. The input is latitude and longitude information. Specifically, it combines this with map data to identify the user's current location and surrounding area. The output is location-related environmental information.
[0479] Step 6:
[0480] The server uses a generative AI model to generate information based on the analysis results. The input consists of the analysis results of visual, acoustic, and location information. The server uses generative AI technology to create information and suggestions best suited to the user's situation. The output is personalized information for the user.
[0481] Step 7:
[0482] The device notifies the user of the generated information. The input is personalized information sent from the server. The device displays this information on its screen and provides voice guidance as needed. The output is the information the user receives visually and aurally.
[0483] Step 8:
[0484] The user can request additional information using prompt messages. The input is a prompt message in either voice or text format. The terminal sends this request to the server for further analysis. The output is more detailed information provided to the user.
[0485] (Application Example 1)
[0486] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0487] In today's consumer society, consumers are surrounded by a vast amount of product information, making it difficult to make the best choice. Furthermore, the market contains numerous products with varying prices, performance, and reviews, and comparing them in real time requires considerable time and effort. Moreover, there is a lack of convenient ways for users to obtain this information during their physical shopping experience.
[0488] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0489] In this invention, the server includes means for analyzing visual data, means for handling audio data, means for acquiring location information, and means for recognizing products and presenting detailed information in real time. This enables users to efficiently collect necessary product information when shopping and make optimal purchasing decisions.
[0490] A "generative artificial intelligence model" is an artificial intelligence trained using machine learning techniques to perform a specific task from a vast amount of data.
[0491] "Means for analyzing visual data" refers to devices or software that analyze image or video data using digital image processing technology to recognize objects or scenes.
[0492] "Means for analyzing audio data" refers to devices or software that use speech recognition technology to extract text or intent from audio.
[0493] "Means of acquiring location information" refers to technologies that use GPS or other positioning systems to determine the current location of a device.
[0494] "Means of providing information to users" refers to a system that displays or communicates useful information to users through a user interface based on analysis results.
[0495] A "means for updating information based on user requests" refers to a system that has the ability to dynamically re-evaluate, modify, or update the information it provides in response to user input or requests.
[0496] "Means for identifying products and presenting related information in real time" refers to technology that uses camera data or similar methods to identify products and immediately displays information related to those products to the user.
[0497] The system of this invention operates on the basis of cooperation between a user-carried terminal and a server. The terminal is equipped with sensors such as a camera, microphone, and GPS, which are used to acquire visual data, audio data, and location information of the environment. The terminal has the function of collecting this data in real time and transmitting it to the server.
[0498] The server performs image analysis on visual data using a generative AI model to identify objects and products. Audio data is analyzed by speech recognition software to understand user requests in natural language. Location information is used to determine the current location and understand the context of the environment.
[0499] The server's analysis results are returned to the user's device and displayed as an overlay in the user's field of view. This allows for real-time presentation of detailed product information and price comparisons, supporting the user's decision-making.
[0500] For example, when a user is looking at clothes in a store, the device acquires visual data and sends it to the server. The server uses a generative AI model to identify the clothes and retrieves relevant data by comparing it with registered information. This includes the product price, customer reviews, and price comparisons with other stores. If the user gives a voice command such as "Tell me the reviews for this product," the server retrieves and displays the review information.
[0501] An example of a prompt might be, "Identify the products in the image and tell me their price and description." Based on this prompt, the server uses a generative AI model to select the information to provide to the user.
[0502] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0503] Step 1:
[0504] The device uses its camera and microphone to acquire visual and audio data from its surroundings. The input is raw data from the environment, and the output is this data in digital form. This allows the device to capture objects within the user's field of view and surrounding conversations.
[0505] Step 2:
[0506] The terminal transmits the acquired visual and audio data to the server. The input is the digital data held by the terminal, and the output is a signal to the server. This allows the server to receive the information necessary for analysis.
[0507] Step 3:
[0508] The server uses an artificial intelligence model to analyze visual data and identify objects and products within the image. The input is visual data, and the output is identification information for the identified objects. Specifically, the data is analyzed by an image recognition algorithm.
[0509] Step 4:
[0510] The server uses speech recognition software to analyze voice data and convert user instructions into text. The input is voice data, and the output is the analyzed text information. The specific operation involves extracting voice commands.
[0511] Step 5:
[0512] The server searches for related information about an object it has identified and generates the information in real time. Input is the identification information of the identified object and a voice command, and output is related information that can be displayed. Information is retrieved from a database.
[0513] Step 6:
[0514] The server sends the generated information to the terminal and displays it as an overlay in the user's field of view. The input is a dataset of related information, and the output is a visual presentation on the user's display device. The terminal then reactivates the overlay function.
[0515] Step 7:
[0516] The user requests further details using voice commands or input devices as needed. Input is the user's new instructions, and output is an updated display of information. This cycle is repeated to support the user's decision-making.
[0517] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0518] This invention relates to an information provision system incorporating an emotion engine that recognizes user emotions, thereby enabling the provision of information and services that take into account the user's emotional state. This system has the capability to analyze user emotion data in addition to visual data, audio data, and location information.
[0519] The device is equipped with a camera, microphone, GPS sensor, and emotion engine sensor, which are used to simultaneously collect the user's environmental information and emotional state. The emotion engine analyzes the user's emotions from their facial expressions, voice tone, and body movements. This data collection is performed continuously in the background during the user's natural behavior, and the device sends this data to a server.
[0520] The server comprehensively analyzes various data and generates information using an artificial intelligence model, combining it with the sentiment analysis results from the emotion engine. This system aims to optimize the user experience by selecting the appropriate way to provide information based on the user's emotional state.
[0521] For example, if a user experiences discomfort or fatigue while sightseeing, the system's emotion engine detects this and prioritizes providing information on relaxing places and rest spots rather than typical tourist information. If the user is shopping and their stress level is high, the system will offer suggestions to assist with simple transactions or purchases.
[0522] Furthermore, in the medical field, data obtained from the emotion engine can be used to provide more humane and efficient support when making diagnoses and providing care while considering the patient's emotional responses. This allows healthcare professionals to provide care that is tailored to the patient's psychological state.
[0523] By constantly updating information based on user requests and feedback, and taking appropriate action at each step, the system functions effectively under diverse circumstances. This ensures that users always receive the best information and services, leading to a more comfortable and satisfying experience.
[0524] The following describes the processing flow.
[0525] Step 1:
[0526] The device activates its camera, microphone, GPS sensor, and emotion engine sensors to collect visual data, audio data, location information, and emotion data in real time. The camera captures the user's facial expressions, the microphone records voice tones, and the sensors detect the user's body movements, etc.
[0527] Step 2:
[0528] The terminal sends the collected data to the server. Upon receiving the data, the server prepares it for analysis. This includes data preprocessing and compression as needed.
[0529] Step 3:
[0530] The server uses image recognition algorithms to analyze visual data and identify objects and text in the user's environment.
[0531] Step 4:
[0532] The server applies a speech recognition algorithm, converts the audio data into text format, and performs linguistic analysis. This allows it to understand the content of the user's speech and the nuances of their emotions.
[0533] Step 5:
[0534] The server analyzes emotional data acquired by the emotion engine to identify the user's current emotional state (e.g., happiness, anxiety, excitement, fatigue). This result is then combined with other data to provide a comprehensive assessment of the user's situation.
[0535] Step 6:
[0536] The server uses a generative artificial intelligence model to generate information and services best suited to the user based on collected information and sentiment analysis results. For example, if it detects that the user is tired, it prioritizes providing information about rest spots.
[0537] Step 7:
[0538] The server sends the generated information to the terminal. The terminal receives this information and displays it as an overlay in the user's field of view. The displayed content is customized to include information that responds to the user's emotions.
[0539] Step 8:
[0540] Users can request additional information or manipulate the provided information using voice commands or gestures. The terminal sends this input to the server, which regenerates the information as needed.
[0541] Step 9:
[0542] The server updates information based on user feedback and sends the new information to the device. This ensures that users always receive up-to-date information that reflects their emotional state.
[0543] (Example 2)
[0544] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0545] Conventional information delivery systems have the problem of not optimizing the user experience because they provide information without considering the user's emotional state. Furthermore, there is a lack of technology to automatically provide information based on emotional state, making it difficult to meet the individual needs of users. In particular, there is a need for information delivery that takes users' emotions into account in various situations in daily life, but there is no effective method to achieve this.
[0546] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0547] In this invention, the server includes a device for capturing visual information, a device for acquiring acoustic information, a device for acquiring spatial information, means for analyzing emotional states, means for generating information based on the emotional analysis results using a generative model, means for providing information adjusted according to the user's emotional state, and means for updating the information based on user input. This makes it possible to accurately provide information according to the user's emotional state and optimize the user experience.
[0548] A "device for capturing visual information" is a device that has the function of acquiring the user's visual data, and includes cameras, image sensors, and the like.
[0549] "Devices for acquiring acoustic information" are devices that collect the user's voice and ambient sounds, and include microphones and acoustic sensors.
[0550] "Devices for acquiring spatial information" refer to devices that acquire location-related data such as the user's current location and movement history, and include GPS receivers and location services.
[0551] "Means for analyzing emotional states" refer to mechanisms for analyzing a user's emotions and determining emotional states such as joy or sadness, and these may include software algorithms or emotion analysis engines.
[0552] "Means of generating information based on sentiment analysis results using generative models" refers to methods for generating optimal information while considering the user's emotional state, and includes generative AI models and machine learning algorithms.
[0553] "Means of providing information tailored to the user's emotional state" refers to means of providing information optimized for the user based on analyzed emotional data, and transmitting this information via display devices or notification systems.
[0554] "Means for updating information based on user input" refers to a mechanism for updating information in real time in response to user requests and feedback, and acquires information through interfaces, input devices, etc.
[0555] This system combines advanced data collection and analysis technologies to analyze user emotions in real time and provide appropriate information. Specifically, the device is equipped with a camera to capture visual information, a microphone to acquire acoustic information, and a GPS receiver to acquire spatial information. It also incorporates an emotion analysis engine to analyze emotional states.
[0556] The server receives various data transmitted from the terminal and analyzes the user's emotions via an emotion analysis engine. A generative AI model is used for the analysis, and data processing is performed based on the emotional state. This generative AI model uses algorithms designed to precisely capture the user's emotional state and generates appropriate information.
[0557] Information tailored to the user's emotions is provided to the user through the device. For example, if the system analyzes that the user is emotionally fatigued, the server provides information on nearby rest spots and relaxing facilities. This allows the user to choose appropriate actions based on their emotional state.
[0558] For example, if data acquired by a device while a user is visiting a museum detects "fatigue," the server immediately generates information about nearby cafes and notifies the user. This allows the user to take a break at an appropriate time.
[0559] An example of a prompt message is to input a specific situation, such as "How would you provide information if the user is tired?", into the generating AI model, and then derive the optimal method for providing information from its response.
[0560] This system enables the provision of information that takes user emotions into account, significantly improving the user experience.
[0561] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0562] Step 1:
[0563] The device collects environmental data.
[0564] The device captures visual information of the user's surroundings with a camera and acquires acoustic information with a microphone. Simultaneously, it obtains current location information using a GPS receiver. This allows the user's current environment and activity status to be input as data. This input data serves as the foundational data necessary to analyze the user's emotions.
[0565] Step 2:
[0566] The device analyzes emotional data.
[0567] The system uses an emotion analysis engine to analyze the user's facial expressions and voice tone based on acquired visual and auditory information. This analysis determines the user's emotional state, such as joy, anger, or fatigue. Image recognition and voice analysis technologies are used for data processing. The analysis results are output as data indicating the user's emotional state.
[0568] Step 3:
[0569] The terminal sends the analysis data to the server.
[0570] The device sends analyzed emotion data and location information to the server. The transmitted data includes the user's current emotional state and location information, which then serves as input for the next stage of the information generation process.
[0571] Step 4:
[0572] The server performs an integrated analysis of the data.
[0573] The server analyzes emotional data and location information sent from the terminal. Using a generative AI model, it processes the data to generate appropriate information based on the user's emotional state. For example, based on the emotional state of "tired" and the location information of "tourist spot," it generates recommendations for rest spots.
[0574] Step 5:
[0575] The server provides the adjusted information.
[0576] The generated information is sent from the server to the user's device, which then displays the information on its screen. The information is intended to help the user take actions that align with their current emotional state. For example, it might include directions to a nearby cafe or suggestions for relaxation spots. This allows the user to make decisions that are appropriate to their emotional state.
[0577] (Application Example 2)
[0578] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0579] Traditional information delivery systems often fail to adequately consider the user's emotional state when presenting information, resulting in an unoptimized user experience. For example, in in-store customer service support, it was difficult to provide appropriate responses based on the customer's emotional state, making it difficult to improve customer satisfaction.
[0580] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0581] In this invention, the server includes means for analyzing visual data using a generative artificial intelligence model, means for analyzing audio data, and means for analyzing emotional data to recognize the user's emotional state. This makes it possible to adjust the information provision method based on the user's emotional state and provide customer service support that is tailored to the emotions of the customers.
[0582] A "generative artificial intelligence model" is an artificial intelligence system that has the ability to analyze data and automatically recognize patterns and features.
[0583] "Visual data" refers to information related to images and videos acquired by cameras and sensors.
[0584] "Audio data" refers to sound information collected by audio input devices such as microphones.
[0585] "Location information" refers to information about the current location of a user or device, obtained using GPS or other positioning technologies.
[0586] "Analysis results" refer to specific conclusions or evaluation values obtained after analyzing data.
[0587] "Emotional data" refers to information that indicates the emotional state of a user, estimated from their facial expressions, tone of voice, body movements, etc.
[0588] "Adjusting the method of providing information" means changing the content and format of the information presented to users based on the analyzed sentiment data.
[0589] "Customer service support" refers to providing assistance to sales staff and employees so that they can take appropriate actions and make appropriate suggestions to improve customer satisfaction.
[0590] To realize this system, smart glasses will be used as the main hardware. This will allow the device to acquire visual and audio data and analyze it using an emotion engine. Specifically, the smart glasses are equipped with a camera and microphone, allowing them to collect customers' facial expressions and voice tones in real time. Location information will be acquired using a built-in GPS sensor.
[0591] The server uses open-source facial recognition libraries (e.g., OpenCV) and speech analysis libraries (e.g., librosa) to analyze this data and determine the user's emotional state. Subsequently, a generative AI model (e.g., TensorFlow) is used to comprehensively analyze the data.
[0592] Based on the analysis results, the server will, for example, determine that "the customer is feeling stressed," and then provide specific customer service suggestions to the sales staff, such as "it would be good to suggest products that will help the customer relax." This information is overlaid on the smart glasses' display.
[0593] Users can provide feedback to the smart glasses based on the reactions of customers, and this feedback is sent to the server, where the information is updated. This allows the system to continuously learn, making the customer service assistance provided more precise and effective.
[0594] For example, when a customer is struggling to decide which product to choose from the displayed items, the AI model can be given a prompt message that says, "Determine if the customer is experiencing stress and display more detailed information about the product."
[0595] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0596] Step 1:
[0597] The device uses the camera and microphone of smart glasses to acquire visual and audio data from customers. This data includes facial expressions and voice tone. The input is raw visual and audio data, which is a preparatory step for identifying emotional information based on faces and voices.
[0598] Step 2:
[0599] The device analyzes visual data using an open-source face recognition library (e.g., OpenCV) to estimate emotional states from facial expressions. Similarly, it analyzes audio data using a speech analysis library (e.g., librosa) to estimate emotions from voice tone. Inputs are the acquired visual and audio data, and outputs are the analysis results of the respective emotional states.
[0600] Step 3:
[0601] The results of this emotional state analysis, along with location information including GPS sensor data, are sent to the server. The server receives this data, analyzes it using a generative AI model (e.g., TensorFlow), and determines the overall emotional state. The input is the transmitted emotional analysis results and location information, and the output is the integrated emotional state evaluation result.
[0602] Step 4:
[0603] The server generates informational messages and customer service suggestions using prompts based on the evaluation results of the generated emotional state. These prompts are generated in the format of, for example, "Determine if the customer is stressed and display detailed information about the product." The input is an integrated emotional state, and the output is a prompt that provides specific suggestions or guidelines for action.
[0604] Step 5:
[0605] The generated prompt text is returned to the terminal and overlaid on the smart glasses' display. Based on this, the user can provide quick and appropriate service to customers. The input is the generated prompt text, and the output is the service guidance displayed on the smart glasses' display.
[0606] Step 6:
[0607] Feedback obtained by users through interactions with customers is sent to the server via the terminal, and the system updates the information, which is then used for future analyses and recommendations. In this way, the system is continuously improved, and its accuracy is enhanced. The input is user feedback, and the output is updated information and preparation for the next analysis.
[0608] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0609] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0610] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0611] [Fourth Embodiment]
[0612] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0613] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0614] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0615] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0616] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0617] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0618] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0619] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0620] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0621] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0622] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0623] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0624] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0625] This invention relates to an advanced information provision system that utilizes a generative artificial intelligence model, and is operated primarily by analyzing visual data, audio data, and location information. The system is designed to provide information based on the individual needs of users in real time.
[0626] First, the device worn or carried by the user (such as smart glasses) continuously collects data about the surrounding environment and circumstances. This device is equipped with sensors such as a camera, microphone, and GPS, and can acquire visual data, audio data, and location information. This data is analyzed by the device, stored in temporary memory, and then sent to a server.
[0627] The server uses a generative artificial intelligence model to analyze the received data. Specifically, it identifies visual data using image processing technology and understands audio data using natural language processing. It also understands the current location and surrounding environment based on location information. By comprehensively combining this information, it identifies the information and services the user is looking for and generates personalized suggestions.
[0628] For example, when a user is walking through a shopping mall, the device captures visual data of displayed products and sends it to a server. The server uses generated AI to identify the products and provides information on the lowest prices compared to other online platforms. Furthermore, when a user is in a tourist area, the device can analyze information on buildings and artworks captured by the camera and display appropriate explanations in real time.
[0629] In the medical field, healthcare professionals can use terminals to access medical information tailored to the patient's condition. For example, during pre-operative preparation, the system can provide optimal procedures and medication information based on the patient's health status, supporting smooth medical treatment.
[0630] This system provides the most relevant information instantly by continuously updating information in response to user requests. Users can request further details by giving instructions, for example, via voice commands or gestures. As a result, users can obtain the necessary information at the optimal time, enabling efficient decision-making in today's information-saturated society.
[0631] The following describes the processing flow.
[0632] Step 1:
[0633] The device activates sensors when the user begins using it, collecting visual, audio, and location data in real time. The camera captures the surrounding video, the microphone records audio, and GPS determines the user's location.
[0634] Step 2:
[0635] The terminal stores the collected data in temporary memory, compresses the data, and then sends it to the server via a secure protocol. This process minimizes latency during data transfer.
[0636] Step 3:
[0637] The server prepares the received data for analysis. Visual data is identified using image recognition algorithms to identify objects and text, and audio data is converted into language using speech recognition algorithms. Location information is merged with a mapping service to determine the current location.
[0638] Step 4:
[0639] The server applies an artificial intelligence model based on the analysis results to generate information and services tailored to the user. Examples include information on the lowest prices for shopping, historical background and explanations of tourist destinations, and diagnostic information in medical support.
[0640] Step 5:
[0641] The server sends the generated information to the terminal. The terminal receives this information and provides the necessary information by displaying it as an overlay in the user's field of view. The display timing and format are customized according to the user's situation.
[0642] Step 6:
[0643] Users can request additional information or select and manipulate specific data using voice or gestures. The terminal sends these inputs to a server for information updates and further analysis.
[0644] Step 7:
[0645] The server re-analyzes the information based on user feedback and requests, and resends the newly generated information to the terminal. This ensures that the latest information tailored to the user's needs is always provided.
[0646] (Example 1)
[0647] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0648] In modern society, information overload has become the norm, making it difficult for users to quickly and accurately obtain the information they need. This is especially true while on the go or in complex environments, where information acquisition and decision-making become even more challenging. To address this problem, there is a need for systems that provide appropriate information in real time, tailored to the user's situation and needs, and support efficient decision-making.
[0649] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0650] In this invention, the server includes information provision means equipped with a device for collecting data from the surrounding environment, means for transmitting the collected data via a communication line, and means for analyzing visual information using image recognition technology. This makes it possible to provide users with necessary information in real time and support their decision-making.
[0651] A "device that collects data from the surrounding environment" is a device that acquires data such as visual, auditory, and location information using sensors, and understands the state of the user's surroundings in real time.
[0652] "Information provision means" refers to methods and technologies for presenting appropriate information to users based on collected and analyzed data.
[0653] "Means of transmission via communication lines" refers to methods using wireless or wired communication technology to transfer collected data to external devices such as servers.
[0654] "Means of analyzing visual information using image recognition technology" refers to techniques for analyzing acquired visual data for object identification and feature extraction, and utilizes image processing algorithms.
[0655] "Means of analyzing acoustic information using natural language processing technology" refers to technologies that convert audio information into text and analyze its meaning and intent.
[0656] "Methods for understanding spatial information using location data" refers to technologies that utilize location information such as GPS to identify a user's current location and surrounding environment.
[0657] "Means of generating information based on analysis results using generative artificial intelligence models" refers to a method of generating optimal information for users based on analyzed data using AI technology.
[0658] "Means for processing additional information requests from users via prompt messages" refers to interfaces and technologies that understand and provide appropriate responses or information when a user requests additional information via voice or text.
[0659] This invention is a system that provides users with appropriate information in real time using a generative artificial intelligence model. The system primarily utilizes the following devices and technologies.
[0660] The device is equipped with various sensors to collect data from its surroundings. Specifically, these include a camera to acquire visual information, a microphone to acquire acoustic information, and a GPS unit to collect location data. This device is often implemented as smart glasses or other wearable devices. The collected data is initially processed within the device and then transmitted to a server via a communication line.
[0661] The server analyzes the received data. For visual information, image recognition technology is used to recognize objects. In this case, existing image processing libraries such as OpenCV are often used. Acoustic information is analyzed using natural language processing technology, and text conversion is performed using Google Cloud Speech-to-Text, etc. Location information is compared with map data to understand the user's spatial environment.
[0662] A generative AI model is implemented to generate information tailored to the user's needs based on analysis results. This generated information is displayed overlaid on the user's field of view through a visual device. For example, if a user is in a shopping mall, they can view comparison information on the lowest prices of products they have photographed. Furthermore, in tourist areas, it can provide real-time historical background information about buildings. In the medical field, it can also present optimal diagnostic information based on the patient's condition.
[0663] Users can obtain additional information or details by sending prompts via voice commands or text input. Examples of prompts include "Tell me more about this product" or "Tell me the nearest cafe." The server receives these prompts and provides the user with more detailed information or service suggestions.
[0664] As described above, this invention can quickly and efficiently provide users with the information they need in today's information-saturated world.
[0665] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0666] Step 1:
[0667] The device collects data from its surroundings. Inputs include visual data, acoustic data, and location data. This data is acquired through sensors such as cameras, microphones, and GPS. Visual data is saved as image files, and acoustic data is converted into audio files. Location data is acquired as latitude and longitude information. The output is a formatted dataset.
[0668] Step 2:
[0669] The device temporarily stores the collected data and sends it to the server via a communication line. Specifically, the collected data is stored in temporary memory and sent to the server using Wi-Fi or mobile data communication. The input is a formatted dataset, and the output is a data stream received on the server side.
[0670] Step 3:
[0671] The server analyzes the received visual data using image recognition technology. The input is an image file. Specifically, the server uses OpenCV or a similar image processing library to identify objects within the image. The output is a list of recognized objects.
[0672] Step 4:
[0673] The server analyzes audio data using natural language processing techniques. The input is an audio file. The server uses a speech recognition service such as Google Cloud Speech-to-Text to convert the audio to text. It then analyzes the text content to extract meaning. The output is the analyzed text information.
[0674] Step 5:
[0675] The server analyzes spatial information based on location data. The input is latitude and longitude information. Specifically, it combines this with map data to identify the user's current location and surrounding area. The output is location-related environmental information.
[0676] Step 6:
[0677] The server uses a generative AI model to generate information based on the analysis results. The input consists of the analysis results of visual, acoustic, and location information. The server uses generative AI technology to create information and suggestions best suited to the user's situation. The output is personalized information for the user.
[0678] Step 7:
[0679] The device notifies the user of the generated information. The input is personalized information sent from the server. The device displays this information on its screen and provides voice guidance as needed. The output is the information the user receives visually and aurally.
[0680] Step 8:
[0681] The user can request additional information using prompt messages. The input is a prompt message in either voice or text format. The terminal sends this request to the server for further analysis. The output is more detailed information provided to the user.
[0682] (Application Example 1)
[0683] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0684] In today's consumer society, consumers are surrounded by a vast amount of product information, making it difficult to make the best choice. Furthermore, the market contains numerous products with varying prices, performance, and reviews, and comparing them in real time requires considerable time and effort. Moreover, there is a lack of convenient ways for users to obtain this information during their physical shopping experience.
[0685] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0686] In this invention, the server includes means for analyzing visual data, means for handling audio data, means for acquiring location information, and means for recognizing products and presenting detailed information in real time. This enables users to efficiently collect necessary product information when shopping and make optimal purchasing decisions.
[0687] A "generative artificial intelligence model" is an artificial intelligence trained using machine learning techniques to perform a specific task from a vast amount of data.
[0688] "Means for analyzing visual data" refers to devices or software that analyze image or video data using digital image processing technology to recognize objects or scenes.
[0689] "Means for analyzing audio data" refers to devices or software that use speech recognition technology to extract text or intent from audio.
[0690] "Means of acquiring location information" refers to technologies that use GPS or other positioning systems to determine the current location of a device.
[0691] "Means of providing information to users" refers to a system that displays or communicates useful information to users through a user interface based on analysis results.
[0692] A "means for updating information based on user requests" refers to a system that has the ability to dynamically re-evaluate, modify, or update the information it provides in response to user input or requests.
[0693] "Means for identifying products and presenting related information in real time" refers to technology that uses camera data or similar methods to identify products and immediately displays information related to those products to the user.
[0694] The system of this invention operates on the basis of cooperation between a user-carried terminal and a server. The terminal is equipped with sensors such as a camera, microphone, and GPS, which are used to acquire visual data, audio data, and location information of the environment. The terminal has the function of collecting this data in real time and transmitting it to the server.
[0695] The server performs image analysis on visual data using a generative AI model to identify objects and products. Audio data is analyzed by speech recognition software to understand user requests in natural language. Location information is used to determine the current location and understand the context of the environment.
[0696] The server's analysis results are returned to the user's device and displayed as an overlay in the user's field of view. This allows for real-time presentation of detailed product information and price comparisons, supporting the user's decision-making.
[0697] For example, when a user is looking at clothes in a store, the device acquires visual data and sends it to the server. The server uses a generative AI model to identify the clothes and retrieves relevant data by comparing it with registered information. This includes the product price, customer reviews, and price comparisons with other stores. If the user gives a voice command such as "Tell me the reviews for this product," the server retrieves and displays the review information.
[0698] An example of a prompt might be, "Identify the products in the image and tell me their price and description." Based on this prompt, the server uses a generative AI model to select the information to provide to the user.
[0699] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0700] Step 1:
[0701] The device uses its camera and microphone to acquire visual and audio data from its surroundings. The input is raw data from the environment, and the output is this data in digital form. This allows the device to capture objects within the user's field of view and surrounding conversations.
[0702] Step 2:
[0703] The terminal transmits the acquired visual and audio data to the server. The input is the digital data held by the terminal, and the output is a signal to the server. This allows the server to receive the information necessary for analysis.
[0704] Step 3:
[0705] The server uses an artificial intelligence model to analyze visual data and identify objects and products within the image. The input is visual data, and the output is identification information for the identified objects. Specifically, the data is analyzed by an image recognition algorithm.
[0706] Step 4:
[0707] The server uses speech recognition software to analyze voice data and convert user instructions into text. The input is voice data, and the output is the analyzed text information. The specific operation involves extracting voice commands.
[0708] Step 5:
[0709] The server searches for related information about an object it has identified and generates the information in real time. Input is the identification information of the identified object and a voice command, and output is related information that can be displayed. Information is retrieved from a database.
[0710] Step 6:
[0711] The server sends the generated information to the terminal and displays it as an overlay in the user's field of view. The input is a dataset of related information, and the output is a visual presentation on the user's display device. The terminal then reactivates the overlay function.
[0712] Step 7:
[0713] The user requests further details using voice commands or input devices as needed. Input is the user's new instructions, and output is an updated display of information. This cycle is repeated to support the user's decision-making.
[0714] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0715] This invention relates to an information provision system incorporating an emotion engine that recognizes user emotions, thereby enabling the provision of information and services that take into account the user's emotional state. This system has the capability to analyze user emotion data in addition to visual data, audio data, and location information.
[0716] The device is equipped with a camera, microphone, GPS sensor, and emotion engine sensor, which are used to simultaneously collect the user's environmental information and emotional state. The emotion engine analyzes the user's emotions from their facial expressions, voice tone, and body movements. This data collection is performed continuously in the background during the user's natural behavior, and the device sends this data to a server.
[0717] The server comprehensively analyzes various data and generates information using an artificial intelligence model, combining it with the sentiment analysis results from the emotion engine. This system aims to optimize the user experience by selecting the appropriate way to provide information based on the user's emotional state.
[0718] For example, if a user experiences discomfort or fatigue while sightseeing, the system's emotion engine detects this and prioritizes providing information on relaxing places and rest spots rather than typical tourist information. If the user is shopping and their stress level is high, the system will offer suggestions to assist with simple transactions or purchases.
[0719] Furthermore, in the medical field, data obtained from the emotion engine can be used to provide more humane and efficient support when making diagnoses and providing care while considering the patient's emotional responses. This allows healthcare professionals to provide care that is tailored to the patient's psychological state.
[0720] By constantly updating information based on user requests and feedback, and taking appropriate action at each step, the system functions effectively under diverse circumstances. This ensures that users always receive the best information and services, leading to a more comfortable and satisfying experience.
[0721] The following describes the processing flow.
[0722] Step 1:
[0723] The device activates its camera, microphone, GPS sensor, and emotion engine sensors to collect visual data, audio data, location information, and emotion data in real time. The camera captures the user's facial expressions, the microphone records voice tones, and the sensors detect the user's body movements, etc.
[0724] Step 2:
[0725] The terminal sends the collected data to the server. Upon receiving the data, the server prepares it for analysis. This includes data preprocessing and compression as needed.
[0726] Step 3:
[0727] The server uses image recognition algorithms to analyze visual data and identify objects and text in the user's environment.
[0728] Step 4:
[0729] The server applies a speech recognition algorithm, converts the audio data into text format, and performs linguistic analysis. This allows it to understand the content of the user's speech and the nuances of their emotions.
[0730] Step 5:
[0731] The server analyzes emotional data acquired by the emotion engine to identify the user's current emotional state (e.g., happiness, anxiety, excitement, fatigue). This result is then combined with other data to provide a comprehensive assessment of the user's situation.
[0732] Step 6:
[0733] The server uses a generative artificial intelligence model to generate information and services best suited to the user based on collected information and sentiment analysis results. For example, if it detects that the user is tired, it prioritizes providing information about rest spots.
[0734] Step 7:
[0735] The server sends the generated information to the terminal. The terminal receives this information and displays it as an overlay in the user's field of view. The displayed content is customized to include information that responds to the user's emotions.
[0736] Step 8:
[0737] Users can request additional information or manipulate the provided information using voice commands or gestures. The terminal sends this input to the server, which regenerates the information as needed.
[0738] Step 9:
[0739] The server updates information based on user feedback and sends the new information to the device. This ensures that users always receive up-to-date information that reflects their emotional state.
[0740] (Example 2)
[0741] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0742] Conventional information delivery systems have the problem of not optimizing the user experience because they provide information without considering the user's emotional state. Furthermore, there is a lack of technology to automatically provide information based on emotional state, making it difficult to meet the individual needs of users. In particular, there is a need for information delivery that takes users' emotions into account in various situations in daily life, but there is no effective method to achieve this.
[0743] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0744] In this invention, the server includes a device for capturing visual information, a device for acquiring acoustic information, a device for acquiring spatial information, means for analyzing emotional states, means for generating information based on the emotional analysis results using a generative model, means for providing information adjusted according to the user's emotional state, and means for updating the information based on user input. This makes it possible to accurately provide information according to the user's emotional state and optimize the user experience.
[0745] A "device for capturing visual information" is a device that has the function of acquiring the user's visual data, and includes cameras, image sensors, and the like.
[0746] "Devices for acquiring acoustic information" are devices that collect the user's voice and ambient sounds, and include microphones and acoustic sensors.
[0747] "Devices for acquiring spatial information" refer to devices that acquire location-related data such as the user's current location and movement history, and include GPS receivers and location services.
[0748] "Means for analyzing emotional states" refer to mechanisms for analyzing a user's emotions and determining emotional states such as joy or sadness, and these may include software algorithms or emotion analysis engines.
[0749] "Means of generating information based on sentiment analysis results using generative models" refers to methods for generating optimal information while considering the user's emotional state, and includes generative AI models and machine learning algorithms.
[0750] "Means of providing information tailored to the user's emotional state" refers to means of providing information optimized for the user based on analyzed emotional data, and transmitting this information via display devices or notification systems.
[0751] "Means for updating information based on user input" refers to a mechanism for updating information in real time in response to user requests and feedback, and acquires information through interfaces, input devices, etc.
[0752] This system combines advanced data collection and analysis technologies to analyze user emotions in real time and provide appropriate information. Specifically, the device is equipped with a camera to capture visual information, a microphone to acquire acoustic information, and a GPS receiver to acquire spatial information. It also incorporates an emotion analysis engine to analyze emotional states.
[0753] The server receives various data transmitted from the terminal and analyzes the user's emotions via an emotion analysis engine. A generative AI model is used for the analysis, and data processing is performed based on the emotional state. This generative AI model uses algorithms designed to precisely capture the user's emotional state and generates appropriate information.
[0754] Information tailored to the user's emotions is provided to the user through the device. For example, if the system analyzes that the user is emotionally fatigued, the server provides information on nearby rest spots and relaxing facilities. This allows the user to choose appropriate actions based on their emotional state.
[0755] For example, if data acquired by a device while a user is visiting a museum detects "fatigue," the server immediately generates information about nearby cafes and notifies the user. This allows the user to take a break at an appropriate time.
[0756] An example of a prompt message is to input a specific situation, such as "How would you provide information if the user is tired?", into the generating AI model, and then derive the optimal method for providing information from its response.
[0757] This system enables the provision of information that takes user emotions into account, significantly improving the user experience.
[0758] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0759] Step 1:
[0760] The device collects environmental data.
[0761] The device captures visual information of the user's surroundings with a camera and acquires acoustic information with a microphone. Simultaneously, it obtains current location information using a GPS receiver. This allows the user's current environment and activity status to be input as data. This input data serves as the foundational data necessary to analyze the user's emotions.
[0762] Step 2:
[0763] The device analyzes emotional data.
[0764] The system uses an emotion analysis engine to analyze the user's facial expressions and voice tone based on acquired visual and auditory information. This analysis determines the user's emotional state, such as joy, anger, or fatigue. Image recognition and voice analysis technologies are used for data processing. The analysis results are output as data indicating the user's emotional state.
[0765] Step 3:
[0766] The terminal sends the analysis data to the server.
[0767] The device sends analyzed emotion data and location information to the server. The transmitted data includes the user's current emotional state and location information, which then serves as input for the next stage of the information generation process.
[0768] Step 4:
[0769] The server performs an integrated analysis of the data.
[0770] The server analyzes emotional data and location information sent from the terminal. Using a generative AI model, it processes the data to generate appropriate information based on the user's emotional state. For example, based on the emotional state of "tired" and the location information of "tourist spot," it generates recommendations for rest spots.
[0771] Step 5:
[0772] The server provides the adjusted information.
[0773] The generated information is sent from the server to the user's device, which then displays the information on its screen. The information is intended to help the user take actions that align with their current emotional state. For example, it might include directions to a nearby cafe or suggestions for relaxation spots. This allows the user to make decisions that are appropriate to their emotional state.
[0774] (Application Example 2)
[0775] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0776] Traditional information delivery systems often fail to adequately consider the user's emotional state when presenting information, resulting in an unoptimized user experience. For example, in in-store customer service support, it was difficult to provide appropriate responses based on the customer's emotional state, making it difficult to improve customer satisfaction.
[0777] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0778] In this invention, the server includes means for analyzing visual data using a generative artificial intelligence model, means for analyzing audio data, and means for analyzing emotional data to recognize the user's emotional state. This makes it possible to adjust the information provision method based on the user's emotional state and provide customer service support that is tailored to the emotions of the customers.
[0779] A "generative artificial intelligence model" is an artificial intelligence system that has the ability to analyze data and automatically recognize patterns and features.
[0780] "Visual data" refers to information related to images and videos acquired by cameras and sensors.
[0781] "Audio data" refers to sound information collected by audio input devices such as microphones.
[0782] "Location information" refers to information about the current location of a user or device, obtained using GPS or other positioning technologies.
[0783] "Analysis results" refer to specific conclusions or evaluation values obtained after analyzing data.
[0784] "Emotional data" refers to information that indicates the emotional state of a user, estimated from their facial expressions, tone of voice, body movements, etc.
[0785] "Adjusting the method of providing information" means changing the content and format of the information presented to users based on the analyzed sentiment data.
[0786] "Customer service support" refers to providing assistance to sales staff and employees so that they can take appropriate actions and make appropriate suggestions to improve customer satisfaction.
[0787] To realize this system, smart glasses will be used as the main hardware. This will allow the device to acquire visual and audio data and analyze it using an emotion engine. Specifically, the smart glasses are equipped with a camera and microphone, allowing them to collect customers' facial expressions and voice tones in real time. Location information will be acquired using a built-in GPS sensor.
[0788] The server uses open-source facial recognition libraries (e.g., OpenCV) and speech analysis libraries (e.g., librosa) to analyze this data and determine the user's emotional state. Subsequently, a generative AI model (e.g., TensorFlow) is used to comprehensively analyze the data.
[0789] Based on the analysis results, the server will, for example, determine that "the customer is feeling stressed," and then provide specific customer service suggestions to the sales staff, such as "it would be good to suggest products that will help the customer relax." This information is overlaid on the smart glasses' display.
[0790] Users can provide feedback to the smart glasses based on the reactions of customers, and this feedback is sent to the server, where the information is updated. This allows the system to continuously learn, making the customer service assistance provided more precise and effective.
[0791] For example, when a customer is struggling to decide which product to choose from the displayed items, the AI model can be given a prompt message that says, "Determine if the customer is experiencing stress and display more detailed information about the product."
[0792] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0793] Step 1:
[0794] The device uses the camera and microphone of smart glasses to acquire visual and audio data from customers. This data includes facial expressions and voice tone. The input is raw visual and audio data, which is a preparatory step for identifying emotional information based on faces and voices.
[0795] Step 2:
[0796] The device analyzes visual data using an open-source face recognition library (e.g., OpenCV) to estimate emotional states from facial expressions. Similarly, it analyzes audio data using a speech analysis library (e.g., librosa) to estimate emotions from voice tone. Inputs are the acquired visual and audio data, and outputs are the analysis results of the respective emotional states.
[0797] Step 3:
[0798] The results of this emotional state analysis, along with location information including GPS sensor data, are sent to the server. The server receives this data, analyzes it using a generative AI model (e.g., TensorFlow), and determines the overall emotional state. The input is the transmitted emotional analysis results and location information, and the output is the integrated emotional state evaluation result.
[0799] Step 4:
[0800] The server generates informational messages and customer service suggestions using prompts based on the evaluation results of the generated emotional state. These prompts are generated in the format of, for example, "Determine if the customer is stressed and display detailed information about the product." The input is an integrated emotional state, and the output is a prompt that provides specific suggestions or guidelines for action.
[0801] Step 5:
[0802] The generated prompt text is returned to the terminal and overlaid on the smart glasses' display. Based on this, the user can provide quick and appropriate service to customers. The input is the generated prompt text, and the output is the service guidance displayed on the smart glasses' display.
[0803] Step 6:
[0804] Feedback obtained by users through interactions with customers is sent to the server via the terminal, and the system updates the information, which is then used for future analyses and recommendations. In this way, the system is continuously improved, and its accuracy is enhanced. The input is user feedback, and the output is updated information and preparation for the next analysis.
[0805] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0806] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0807] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0808] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0809] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0810] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0811] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0812] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0813] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0814] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0815] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0816] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0817] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0818] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0819] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0820] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0821] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0822] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0823] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0824] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0825] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0826] The following is further disclosed regarding the embodiments described above.
[0827] (Claim 1)
[0828] A method for analyzing visual data using a generative artificial intelligence model,
[0829] Methods for analyzing audio data,
[0830] Means of obtaining location information,
[0831] A means of providing information to users based on the analysis results,
[0832] Means for updating the information based on user requests,
[0833] A system that includes this.
[0834] (Claim 2)
[0835] The system according to claim 1, comprising means for displaying information as an overlay in the user's field of view.
[0836] (Claim 3)
[0837] The system according to claim 1, comprising means for providing diagnostic support in accordance with the patient's condition in the medical field.
[0838] "Example 1"
[0839] (Claim 1)
[0840] Information provision means equipped with a device for collecting data from the surrounding environment,
[0841] A means of transmitting the collected data via a communication line,
[0842] A means of analyzing visual information using image recognition technology,
[0843] A method for analyzing acoustic information using natural language processing technology,
[0844] A means of understanding spatial information using location data,
[0845] A means for generating information based on analysis results using a generative artificial intelligence model,
[0846] Means for providing the generated information described above to a display device,
[0847] A means of adaptively updating information based on user settings,
[0848] A means for processing additional information requests from the user via prompt statements,
[0849] A system that includes this.
[0850] (Claim 2)
[0851] The system according to claim 1, comprising a device that overlays and displays information in the user's field of view.
[0852] (Claim 3)
[0853] The system according to claim 1, which provides information support based on an individual's health status in the healthcare field.
[0854] "Application Example 1"
[0855] (Claim 1)
[0856] A method for analyzing visual data using a generative artificial intelligence model,
[0857] Methods for analyzing audio data,
[0858] Means for obtaining location information,
[0859] A means of providing information to users based on the analysis results,
[0860] A means of updating the information based on the user's request,
[0861] A means of identifying products and presenting related information in real time,
[0862] A system that includes this.
[0863] (Claim 2)
[0864] The system according to claim 1, including means for displaying information superimposed from the user's perspective.
[0865] (Claim 3)
[0866] The system according to claim 1, comprising means for providing product information and related detailed information in response to a user's voice instructions.
[0867] "Example 2 of combining an emotion engine"
[0868] (Claim 1)
[0869] A device for capturing visual information,
[0870] A device for acquiring acoustic information,
[0871] A device for acquiring spatial information,
[0872] A means of analyzing emotional states,
[0873] A means of generating information based on sentiment analysis results using a generative model,
[0874] A means of providing information tailored to the user's emotional state,
[0875] Means for updating the information based on user input,
[0876] A system that includes this.
[0877] (Claim 2)
[0878] The system according to claim 1, comprising means for displaying information related to the user's emotional state.
[0879] (Claim 3)
[0880] The system according to claim 1, which includes means for providing support tailored to an individual's psychological state in the field of health management.
[0881] "Application example 2 when combining with an emotional engine"
[0882] (Claim 1)
[0883] A method for analyzing visual data using a generative artificial intelligence model,
[0884] Methods for analyzing audio data,
[0885] Means of obtaining location information,
[0886] A means of providing information to users based on the analysis results,
[0887] Means for updating the information based on user requests,
[0888] A means of analyzing emotional data to recognize the user's emotional state,
[0889] A means of adjusting the method of providing information based on the user's emotional state,
[0890] A system that includes this.
[0891] (Claim 2)
[0892] The system according to claim 1, comprising means for displaying information as an overlay in the user's field of view.
[0893] (Claim 3)
[0894] The system according to claim 1, comprising means for providing customer service support that corresponds to the emotional state of the customer. [Explanation of Symbols]
[0895] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A method for analyzing visual data using a generative artificial intelligence model, Methods for analyzing audio data, Means of obtaining location information, A means of providing information to users based on the analysis results, Means for updating the information based on user requests, A system that includes this.
2. The system according to claim 1, comprising means for displaying information as an overlay in the user's field of view.
3. The system according to claim 1, which includes means for providing diagnostic support in accordance with the patient's condition in the medical field.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A