system

The system addresses the inefficiency of conventional information retrieval by allowing users to capture and analyze images, gather information, and generate understandable explanations, facilitating quick and intuitive learning.

JP2026105513APending Publication Date: 2026-06-26SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-16
Publication Date
2026-06-26

Smart Images

  • Figure 2026105513000001_ABST
    Figure 2026105513000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Image acquisition device, An analysis device that analyzes images received by the aforementioned image acquisition device to identify products, An information collection device that collects information on products identified by the aforementioned analysis device, An explanation generation device that generates an explanation from the information collected by the aforementioned information collection device, A display device that visualizes the explanation generated by the explanation generation device, A display device that displays product information using the aforementioned display device, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In daily life, when people want to immediately know detailed information about a specific object or phenomenon in front of them, the conventional method requires manual search, which has the problem of lacking immediacy and efficiency. In addition, especially for children and people who are not familiar with technology, it is difficult to obtain information, and there is a possibility of missing opportunities for interest and learning.

Means for Solving the Problems

[0005] This invention provides a system in which an image input means acquires an image, an analysis means identifies an object, an information gathering means collects information related to that object, an explanation generation means generates an easy-to-understand explanation, and a display means presents it to the user. With this system, the user can easily and instantly obtain detailed information about an object via a camera, and intuitively expand their knowledge and interests.

[0006] "Image input means" refers to a device or function that acquires images taken by a user and provides them as data in a format usable for subsequent analysis processing.

[0007] "Analysis means" refers to a technology or algorithm for analyzing input image data and recognizing and identifying specific objects within the image.

[0008] "Information gathering means" refers to a function or process for obtaining additional information about an identified object from the internet or other data sources.

[0009] "Explanation generation means" refers to a system or function that organizes collected information into an easily understandable format for the user and generates an explanatory text.

[0010] "Display means" refers to a device or function for visually presenting generated explanatory text and related information to the user. [Brief explanation of the drawing]

[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0013] First, let's explain the terminology used in the following explanation.

[0014] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0015] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0017] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0019] [First Embodiment]

[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0032] In an embodiment of the present invention, the user first launches an application using a terminal. This application incorporates a camera function, allowing the user to point the terminal's camera at an object of interest and take an image.

[0033] The device transmits the captured image to the server via communication. The server uses an AI analysis model to analyze the received image data and identify the object. The server then collects information about the identified object from external databases and information sources via the network and organizes the necessary information.

[0034] Subsequently, the server generates an easy-to-understand explanation based on the collected information. This explanation is sent to the terminal, which visually displays it on the screen for the user. By reading this display, the user can easily understand and learn detailed information about the subject.

[0035] As a concrete example, if a user wants to learn about a flower blooming in a garden, they can use their device to take a picture of the flower. The server analyzes the image and identifies the type of flower. If it is identified as a rose, the server collects information about the rose, such as its characteristics, optimal care, and historical background, and sends it to the device as a concise description. Through this description, the user can gain a deeper understanding of the rose.

[0036] In this way, the present invention enables users to quickly obtain detailed information with simple operations, providing a richer experience and learning opportunity.

[0037] The following describes the processing flow.

[0038] Step 1:

[0039] The user launches an application on their device and uses the camera function to take an image of an object of interest.

[0040] Step 2:

[0041] The device acquires the captured image data, converts it to the appropriate format, and prepares to send it to the server.

[0042] Step 3:

[0043] The terminal sends the converted image data to the server.

[0044] Step 4:

[0045] The server processes the received image data using an analysis tool to recognize and identify objects within the image.

[0046] Step 5:

[0047] The server searches external databases and information sources via the internet to collect information related to the identified object.

[0048] Step 6:

[0049] Based on the information collected by the server, an explanatory text is constructed to generate an easy-to-understand explanation for the user.

[0050] Step 7:

[0051] The server sends the generated explanation to the terminal.

[0052] Step 8:

[0053] The terminal displays the explanatory text received from the server on its screen, presenting it visually to the user.

[0054] Step 9:

[0055] Users can review the displayed description, perform further searches if they want to learn more, and share information as needed.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] In modern times, the means of obtaining information about various objects are limited, and it is particularly difficult for the average user to instantly acquire detailed information about a specific object. There is a need for technology that allows even ordinary users without specialized knowledge to easily identify objects and obtain related information.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes data collection means, analysis means using artificial intelligence, and information acquisition means. This makes it possible for users to quickly and easily obtain identification and related information about various objects simply by acquiring images.

[0061] "Data acquisition means" refers to means consisting of devices and programs for acquiring data such as images.

[0062] "Analysis methods using artificial intelligence" refer to methods that utilize artificial intelligence technology to analyze acquired image data and identify objects.

[0063] "Information acquisition means" refers to means of collecting information related to a specific object via the internet or databases.

[0064] A "generative AI model" is an artificial intelligence model that generates explanatory text in natural language from input information.

[0065] "Display means" refers to devices or programs that visually present generated explanations or information to the user.

[0066] The following processes are performed as embodiments for carrying out the present invention.

[0067] The user launches a dedicated application using their mobile device. This application has a camera function and is used to photograph objects that interest the user. For example, if the user wants to identify a flower they found in a garden, they point the camera of their mobile device at the flower and take a picture.

[0068] The device transmits captured image data to the server via internet communication. This communication uses an encrypted protocol (e.g., HTTPS) to enhance security.

[0069] The server performs artificial intelligence-based analysis on the received images. Specifically, it utilizes machine learning frameworks such as TENSORFLOW® and PyTorch to identify objects within the images. This allows it to recognize what kind of flower is photographed (for example, a rose).

[0070] Next, the server searches online databases or information sources to collect information related to the identified object. During this process, it uses APIs to retrieve detailed information from publicly available sources (e.g., plant identification databases).

[0071] Furthermore, the server uses a generative AI model (e.g., GPT) based on the information obtained to generate explanatory text for the user. For example, it might say, "This flower is a rose, and the best way to grow it is to give it moderate sunlight and keep the soil moist."

[0072] Finally, the server sends the generated description to the terminal, which then displays it to the user. This display allows the user to easily obtain detailed information about the object.

[0073] As a concrete example, an example of a prompt message is as follows: By entering a message in the format "Please tell me the name and characteristics of this plant," you can quickly obtain relevant information.

[0074] In this way, the present invention provides a convenient system that allows users to easily and quickly obtain information about objects of interest.

[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0076] Step 1:

[0077] The user launches the application on their device and uses the camera function to photograph an object. During this process, the device saves the captured image data to a buffer and simultaneously acquires image metadata (e.g., date and time of capture, location information). The input is the image of the object photographed by the user, and the output is the image file stored on the device.

[0078] Step 2:

[0079] The device sends the acquired image data to the server. Specifically, it compresses the image data to an appropriate size using a compression algorithm and sends it to the server via an encrypted protocol (e.g., HTTPS). The input is the captured image data, and the output is the image data transferred to the server.

[0080] Step 3:

[0081] The server analyzes the received image data. This analysis uses an image recognition model (e.g., CNN) equipped with artificial intelligence technology. The server identifies objects from the image and extracts specific features (e.g., color, shape) to determine their type. The input is the received image data, and the output is information about the type of object identified.

[0082] Step 4:

[0083] The server collects information based on the type of object it identifies. Specifically, it searches for and retrieves relevant information from databases and APIs via an internet connection. The input is information about the type of object identified, and the output is detailed information data about the object.

[0084] Step 5:

[0085] The server uses a generative AI model to generate explanatory text for users based on the collected information. The model utilizes natural language processing technology to automatically create appropriate explanations. The input is collected object information, and the output is explanatory text about the object.

[0086] Step 6:

[0087] The server sends the generated description to the terminal. The terminal displays the received description on its screen and presents it to the user. By reading this, the user can obtain detailed information about the object. The input is the generated description, and the output is the visual information displayed on the terminal.

[0088] (Application Example 1)

[0089] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0090] In physical stores, consumers are required to obtain detailed information about products quickly and easily. However, traditional systems have the problem of requiring a lot of effort and time to obtain information, and failing to provide a sufficient purchasing experience.

[0091] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0092] In this invention, the server includes an image acquisition device, an analysis device, and an information collection device. This allows consumers to instantly receive relevant information and make quick and appropriate purchasing decisions simply by taking a picture of products in a store.

[0093] An "image acquisition device" is a device used by users to photograph products they are interested in, and it has the function of acquiring image data.

[0094] An "analysis device" is a device that processes acquired image data and has the function of identifying products. It utilizes artificial intelligence models to extract and analyze product features from images.

[0095] An "information gathering device" is a device that collects information related to a specific product from an external source via an information network, and has the role of accumulating relevant data.

[0096] A "description generation device" is a device that generates product descriptions based on data collected by an information gathering device, providing a function to organize information in a way that is easy for users to understand.

[0097] A "display device" is a device that visually presents generated product descriptions to users, and plays a role in visualizing information.

[0098] A "presentation device" is a device that uses a display device to appropriately provide collected and generated product information to the user.

[0099] This system allows users to easily gather information on products they are interested in within physical stores. Users send images of products they have taken using image acquisition devices such as smartphones to the server. The server processes the received images through an analysis device and identifies the products based on the analysis results. It is desirable to use image recognition software such as Google® Cloud Vision API for this process.

[0100] Once a product is identified, the server uses information gathering devices to collect relevant information from the internet. At this stage, it utilizes databases such as the Amazon Product Review API to obtain product details and review information.

[0101] Next, the server uses an explanation generator to produce an explanation from the collected information in a format that is easy for the user to understand. This process uses generative AI models such as OpenAI's GPT-3 and ChatGPT to generate explanations in natural language.

[0102] The generated descriptions are visualized on a smartphone via a display device. The display device, as part of the display system, accurately presents the product information obtained by this system to the user. This allows the user to instantly check product features, usage instructions, reviews, and more via their smartphone.

[0103] For example, if a user takes a picture of a supplement in a store, its effects, ingredients, and user testimonials will be displayed on their smartphone. Then, by inputting prompts like the following into the AI ​​model, a description will be generated.

[0104] Please describe the following product information: Product name: Supplement, Category: Health food. Briefly describe the product's effects and usage.

[0105] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0106] Step 1:

[0107] The user takes a picture of an item they are interested in using the image acquisition device on their device. The input is image data captured by the camera, and this data is sent to the next processing step.

[0108] Step 2:

[0109] The terminal sends image data to the server. The server inputs the received image data into an analysis device, analyzes the data, extracts product characteristics, and identifies the product. This analysis uses image recognition software such as the Google Cloud Vision API. The output is the identification information of the identified product.

[0110] Step 3:

[0111] The server uses information gathering devices based on the analysis results to obtain product-related information from the internet. It primarily uses the Amazon Product Review API to collect data. The input is product identification information, and the output is detailed product information and review data.

[0112] Step 4:

[0113] The server inputs the collected information into an explanation generation device and generates explanatory text using a generation AI model. This model uses OpenAI's GPT-3 or ChatGPT and provides explanations in natural language using prompts. The output is an explanatory text for the user.

[0114] Step 5:

[0115] The generated explanatory text is sent from the server to the device's display, allowing the user to visually confirm it on their smartphone. The input is the explanatory text, and the output is the visualized information displayed on the device.

[0116] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0117] In an embodiment of the present invention, the user first launches a dedicated application on the terminal and points the camera at an object of interest to take an image. The terminal transmits the image acquired by the image input means to the server via the internet.

[0118] The server first analyzes the received image using an analysis tool to identify the objects contained within the image. This analysis utilizes an artificial intelligence model, enabling accurate identification of the objects. Simultaneously, the server receives user facial expression and voice data from the terminal and analyzes the user's emotional state using an emotion engine.

[0119] Next, the server uses information gathering means to search for and obtain relevant information about the identified object via the internet. Then, based on the collected information, the explanation generation means generates an explanation suitable for the user. At this time, the explanation content is adjusted according to the user's emotional state as determined by the emotion engine, and the explanation is provided in a manner that takes into account the user's interests and level of understanding.

[0120] For example, if a user takes a picture of a painting in a museum, the server identifies the painting and collects information about its history, background, and artist. If the user is excited, it can generate a description that includes many more interesting anecdotes; if the user is calm, it can provide a more detailed and expert explanation.

[0121] The generated description is transmitted to the terminal and displayed on the user's screen by a display device. This information allows the user to gain a deeper understanding of the object and simultaneously experience the system's flexible response. This invention makes it possible not only to obtain information about an object, but also to enjoy optimized information tailored to the user's emotions.

[0122] The following describes the processing flow.

[0123] Step 1:

[0124] The user launches the app on their device, points the camera at an object of interest, and takes a picture.

[0125] Step 2:

[0126] The system acquires image data captured by the device. Simultaneously, it collects the user's facial expressions and voice data.

[0127] Step 3:

[0128] The device sends image data, along with user facial expressions and voice data, to the server.

[0129] Step 4:

[0130] The server analyzes the received image data and uses an AI model to identify the object.

[0131] Step 5:

[0132] The server analyzes facial expressions and voice data through an emotion engine to recognize the user's emotional state.

[0133] Step 6:

[0134] The server collects information related to identified objects via the internet.

[0135] Step 7:

[0136] The server generates a descriptive text based on the information it collects. The descriptive text is adjusted according to the results of the emotion engine, and the information is organized in a format that suits the user's emotional state.

[0137] Step 8:

[0138] The server sends the generated explanation to the terminal.

[0139] Step 9:

[0140] The device displays an explanatory text on the user's screen. The user can review the presented information and, if interested, explore further for more detailed information.

[0141] (Example 2)

[0142] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0143] Conventional systems analyze images and provide information about objects in a uniform manner, making it difficult to provide optimal information tailored to each user's individual interests and emotional state. This resulted in problems such as users receiving either insufficient or excessive information. Furthermore, the lack of a mechanism to adjust information based on the user's emotional state risked a decline in the quality of the user experience.

[0144] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0145] In this invention, the server includes means for acquiring images, means for analyzing the images to identify objects, and means for acquiring and analyzing user emotion data. This enables flexible information provision based on the user's emotional state.

[0146] "Means for acquiring images" refers to a device or function that allows a user to photograph an object of interest and acquire it as image data in digital format.

[0147] "Means for identifying an object" refers to a method or technique for analyzing acquired images to identify specific objects or events present within those images.

[0148] "Means of collecting information from external sources" refers to methods or techniques for obtaining information related to identified objects from databases or websites on the internet.

[0149] "Means for generating explanations" refers to functions and processes that automatically create explanatory text in a human-readable format using a generative AI model based on collected information.

[0150] "Means of display" refers to a device or function for visually presenting the generated explanation on the user's terminal.

[0151] "Means for acquiring and analyzing user emotional data" refers to methods or techniques for acquiring and analyzing user facial expressions, voice, and other emotional expression data to determine the user's emotional state.

[0152] "Means for adjusting the content of explanations" refers to processes and functions for adjusting explanations generated based on the user's sentiment analysis results to be optimized for the user's interests and current emotions.

[0153] To implement this invention, the user, terminal, and server must work together, each fulfilling their respective roles. The user acquires image data on the terminal by taking a picture of an object of interest with the terminal's camera. The terminal then transmits this image data to the server through a dedicated application. A common communication protocol is used for transmitting the image data.

[0154] The server uses artificial intelligence models built with deep learning frameworks such as TensorFlow and PyTorch to analyze images. This analysis identifies objects within the images. Furthermore, the server uses libraries such as OpenCV and Librosa to analyze facial expression and audio data sent from the terminal to evaluate the user's emotional state.

[0155] After identifying the target object, the server uses an information gathering engine to retrieve relevant information from the internet. This process utilizes the Google Custom Search API and web scraping techniques to collect data. Based on the collected information, the server uses a generative AI model to generate a description of the object. This description is adjusted according to the user's emotional state. The generative AI model uses a text generation engine such as GPT-3.

[0156] For example, consider a scenario where a user takes a picture of the painting "Starry Night" in an art museum. The captured image is analyzed on a server, and the subject is identified as "Starry Night." Subsequently, information about the history and background of this painting is collected, and an explanation is generated according to the user's emotional state. A user in an excited state will be provided with an explanation that includes interesting anecdotes about the painter.

[0157] The generated description is sent from the server to the terminal, where the user can visually confirm the information on the terminal's screen. This information allows the user to gain a deeper understanding of the subject matter and enjoy a new experience of receiving information that responds to their own emotions.

[0158] An example of a prompt message would be: "The user has taken an image that includes a painting called 'Starry Night.' The user appears excited. Generate a commentary that includes an interesting anecdote about this painting."

[0159] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0160] Step 1:

[0161] The user launches a dedicated application on their device and takes a picture of an object of interest with the camera. The input is image data of the captured object. The device compresses this image data in JPEG or PNG format and stores it as prepared data.

[0162] Step 2:

[0163] The terminal sends pre-processed image data to the server via internet communication. The input data is a compressed image, and the destination is the server's API endpoint. The output is the image data transferred to the server.

[0164] Step 3:

[0165] The server decompresses the received image data and identifies the object using an image analysis model. The technology used here is a deep learning model based on the TensorFlow or PyTorch framework. The input is image data, and the output is the label of the identified object. In this process, the model analyzes the feature points of the image and determines the most likely object.

[0166] Step 4:

[0167] The server analyzes emotional data (facial expressions and audio data) sent from the terminal. Input includes still images of the user's facial expressions and audio samples. The server evaluates the user's emotional state by analyzing facial expressions with OpenCV and audio with Librosa. The output is the result of the emotional analysis. Specifically, it performs facial feature extraction and audio spectral analysis.

[0168] Step 5:

[0169] The server searches for and collects necessary information via the internet based on the label of the target object. The input is the label of the identified object, and information is collected based on that label. Information is collected using Google Custom Search API and scraping techniques, and structured information data is obtained as output.

[0170] Step 6:

[0171] The server uses the collected information data to generate explanations using a generative AI model. Natural language generation technologies such as GPT-3 are utilized here. The input consists of information data and sentiment analysis results. The output is a user-friendly explanation, with the generated text's style and content adjusted according to the user's sentiment.

[0172] Step 7:

[0173] The server sends the generated explanation back to the terminal. The terminal receives this explanation and displays it visually to the user. The input is the generated explanation, and the output is the display on the user's terminal. Specifically, the text is placed on the terminal screen and presented in a way that the user can easily understand.

[0174] (Application Example 2)

[0175] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0176] While advancements in information technology have led to a vast amount of information available on the internet, it remains difficult for users to efficiently find information that matches their interests and emotions from this enormous volume. In particular, current systems are insufficient in providing information about past works of art and exhibitions tailored to the user's emotional state. Therefore, a system is needed that can provide information that matches the user's interests and emotions.

[0177] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0178] In this invention, the server includes image acquisition means, analysis means, emotion analysis means, information collection means, explanation generation means, and output means. This makes it possible to identify objects of interest to the user and provide information about those objects in an optimized form according to the user's emotional state.

[0179] "Image acquisition means" refers to a function that allows users to photograph objects of interest and acquire the image data.

[0180] "Analysis means" refers to a function that analyzes acquired image data and identifies the target.

[0181] "Emotional analysis means" refers to a function that uses the user's voice and facial expression data to analyze the user's emotional state.

[0182] "Information gathering means" refers to a function for collecting information related to the analyzed subject via a communication network.

[0183] The "explanation generation means" is a function that generates explanations optimized for the user based on collected information and the results of user sentiment analysis.

[0184] "Output means" refers to a function that provides the generated explanation to the user.

[0185] The system that realizes this application example operates through a program installed on a consumer robot, enhancing the user experience in the exhibition space.

[0186] First, the robot uses its built-in camera to acquire images of exhibits that the user is interested in. This image acquisition mechanism allows the robot to process the data of the objects.

[0187] Next, the server uses an analysis tool to analyze the acquired images with a machine learning model (e.g., TensorFlow) to identify the object. After the object is identified, the server activates an emotion analysis tool to collect the user's voice and facial expression data through sensors and uses an emotion analysis engine such as IBM Watson® to determine the user's emotions.

[0188] Subsequently, the server activates information gathering tools via the internet, utilizing the Wikipedia API and other resources to collect relevant information about the target object. Based on this information, the explanation generation tool uses OpenAI's GPT model to generate an appropriate explanation tailored to the user's emotions. This generation AI model uses prompts that reflect the user's interests, derived from the collected data and sentiment analysis.

[0189] The generated explanations are provided to the user via output devices, such as the robot's display or speaker. This allows users to receive customized information tailored to their individual emotional states.

[0190] For example, if the robot recognizes a Renaissance painting and the user is excited, it will provide an interesting explanation including anecdotes about the artist and stories from that time. If the user is calm, it will present detailed information about the technique and historical background of the work.

[0191] An example of a prompt message would be, "Could you tell me about the historical background of this work?" In this way, flexible and advanced information provision tailored to the user's needs is achieved.

[0192] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0193] Step 1:

[0194] The robot uses its built-in camera to capture images of exhibits specified by the user. The input is the image data acquired by the camera, and the output is that same image data. This makes it possible to acquire data on specific objects.

[0195] Step 2:

[0196] The server receives image data and performs image analysis using a machine learning model (such as TensorFlow). The input is the acquired image data, and the output is information about the identified objects. It analyzes the features in the image and identifies the objects.

[0197] Step 3:

[0198] The server acquires user voice and facial expression data through sensors and analyzes it using an emotion analysis engine (e.g., IBM Watson). The input is voice and facial expression data, and the output is the user's emotional state. Data is acquired using sensors, and data calculations are performed to infer the emotional state.

[0199] Step 4:

[0200] The server collects information related to the target object via the internet. It utilizes the Wikipedia API and other search engines for this purpose. The input is information about the identified object, and the output is related information data. The collected data is then searched and retrieved.

[0201] Step 5:

[0202] The server generates explanatory text using a generative AI model (such as GPT) based on collected information and results obtained from sentiment analysis. The input is relevant information data and the user's emotional state, and the output is an optimized explanatory text. The generative AI model is used to generate prompts and appropriate explanations.

[0203] Step 6:

[0204] The generated explanation is provided to the user through the robot's display and speaker. The input is the generated explanation, and the output is achieved in the form of providing information to the user. Explanations are provided to the user using a display and audio output device.

[0205] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0206] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0207] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0208] [Second Embodiment]

[0209] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0210] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0211] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0212] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0213] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0214] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0215] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0216] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0217] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0218] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0219] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0220] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0221] In an embodiment of the present invention, the user first launches an application using a terminal. This application incorporates a camera function, allowing the user to point the terminal's camera at an object of interest and take an image.

[0222] The device transmits the captured image to the server via communication. The server uses an AI analysis model to analyze the received image data and identify the object. The server then collects information about the identified object from external databases and information sources via the network and organizes the necessary information.

[0223] Subsequently, the server generates an easy-to-understand explanation based on the collected information. This explanation is sent to the terminal, which visually displays it on the screen for the user. By reading this display, the user can easily understand and learn detailed information about the subject.

[0224] As a concrete example, if a user wants to learn about a flower blooming in a garden, they can use their device to take a picture of the flower. The server analyzes the image and identifies the type of flower. If it is identified as a rose, the server collects information about the rose, such as its characteristics, optimal care, and historical background, and sends it to the device as a concise description. Through this description, the user can gain a deeper understanding of the rose.

[0225] In this way, the present invention enables users to quickly obtain detailed information with simple operations, providing a richer experience and learning opportunity.

[0226] The following describes the processing flow.

[0227] Step 1:

[0228] The user launches an application on their device and uses the camera function to take an image of an object of interest.

[0229] Step 2:

[0230] The device acquires the captured image data, converts it to the appropriate format, and prepares to send it to the server.

[0231] Step 3:

[0232] The terminal sends the converted image data to the server.

[0233] Step 4:

[0234] The server processes the received image data using an analysis tool to recognize and identify objects within the image.

[0235] Step 5:

[0236] The server searches external databases and information sources via the internet to collect information related to the identified object.

[0237] Step 6:

[0238] Based on the information collected by the server, an explanatory text is constructed to generate an easy-to-understand explanation for the user.

[0239] Step 7:

[0240] The server sends the generated explanation to the terminal.

[0241] Step 8:

[0242] The terminal displays the explanatory text received from the server on its screen, presenting it visually to the user.

[0243] Step 9:

[0244] Users can review the displayed description, perform further searches if they want to learn more, and share information as needed.

[0245] (Example 1)

[0246] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0247] In modern times, the means of obtaining information about a variety of objects are limited, and it is particularly difficult for the average user to instantly acquire detailed information about a specific object. There is a need for technology that allows even ordinary users without specialized knowledge to easily identify objects and obtain related information.

[0248] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0249] In this invention, the server includes data collection means, analysis means using artificial intelligence, and information acquisition means. This makes it possible for users to quickly and easily obtain identification and related information about various objects simply by acquiring images.

[0250] "Data acquisition means" refers to means consisting of devices and programs for acquiring data such as images.

[0251] "Analysis methods using artificial intelligence" refer to methods that utilize artificial intelligence technology to analyze acquired image data and identify objects.

[0252] "Information acquisition means" refers to means of collecting information related to a specific object via the internet or databases.

[0253] A "generative AI model" is an artificial intelligence model that generates explanatory text in natural language from input information.

[0254] "Display means" refers to devices or programs that visually present generated explanations or information to the user.

[0255] The following processes are performed as embodiments for carrying out the present invention.

[0256] The user launches a dedicated application using their mobile device. This application has a camera function and is used to photograph objects that interest the user. For example, if the user wants to identify a flower they found in a garden, they point the camera of their mobile device at the flower and take a picture.

[0257] The device transmits captured image data to the server via internet communication. This communication uses an encrypted protocol (e.g., HTTPS) to enhance security.

[0258] The server performs artificial intelligence-based analysis on the received images. Specifically, it utilizes machine learning frameworks such as TensorFlow and PyTorch to identify objects within the images. This allows it to recognize what kind of flower is photographed (for example, a rose).

[0259] Next, the server searches online databases or information sources to collect information related to the identified object. During this process, it uses APIs to retrieve detailed information from publicly available sources (e.g., plant identification databases).

[0260] Furthermore, the server uses a generative AI model (e.g., GPT) based on the information obtained to generate explanatory text for the user. For example, it might say, "This flower is a rose, and the best way to grow it is to give it moderate sunlight and keep the soil moist."

[0261] Finally, the server sends the generated description to the terminal, which then displays it to the user. This display allows the user to easily obtain detailed information about the object.

[0262] As a concrete example, an example of a prompt message is as follows: By entering a message in the format "Please tell me the name and characteristics of this plant," you can quickly obtain relevant information.

[0263] In this way, the present invention provides a convenient system that allows users to easily and quickly obtain information about objects of interest.

[0264] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0265] Step 1:

[0266] The user launches the application on their device and uses the camera function to photograph an object. During this process, the device saves the captured image data to a buffer and simultaneously acquires image metadata (e.g., date and time of capture, location information). The input is the image of the object photographed by the user, and the output is the image file stored on the device.

[0267] Step 2:

[0268] The device sends the acquired image data to the server. Specifically, it compresses the image data to an appropriate size using a compression algorithm and sends it to the server via an encrypted protocol (e.g., HTTPS). The input is the captured image data, and the output is the image data transferred to the server.

[0269] Step 3:

[0270] The server analyzes the received image data. This analysis uses an image recognition model (e.g., CNN) equipped with artificial intelligence technology. The server identifies objects from the image and extracts specific features (e.g., color, shape) to determine their type. The input is the received image data, and the output is information about the type of object identified.

[0271] Step 4:

[0272] The server collects information based on the type of object it identifies. Specifically, it searches for and retrieves relevant information from databases and APIs via an internet connection. The input is information about the type of object identified, and the output is detailed information data about the object.

[0273] Step 5:

[0274] The server uses a generative AI model to generate explanatory text for users based on the collected information. The model utilizes natural language processing technology to automatically create appropriate explanations. The input is collected object information, and the output is explanatory text about the object.

[0275] Step 6:

[0276] The server sends the generated description to the terminal. The terminal displays the received description on its screen and presents it to the user. By reading this, the user can obtain detailed information about the object. The input is the generated description, and the output is the visual information displayed on the terminal.

[0277] (Application Example 1)

[0278] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0279] In physical stores, consumers are required to obtain detailed information about products quickly and easily. However, traditional systems have the problem of requiring a lot of effort and time to obtain information, and failing to provide a sufficient purchasing experience.

[0280] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0281] In this invention, the server includes an image acquisition device, an analysis device, and an information collection device. This allows consumers to instantly receive relevant information and make quick and appropriate purchasing decisions simply by taking a picture of products in a store.

[0282] An "image acquisition device" is a device used by users to photograph products they are interested in, and it has the function of acquiring image data.

[0283] An "analysis device" is a device that processes acquired image data and has the function of identifying products. It utilizes artificial intelligence models to extract and analyze product features from images.

[0284] The "information collection device" is a device that collects information related to a specified product from the outside via an information network and has the role of integrating related data.

[0285] The "explanation generation device" is a device that generates an explanation about a product based on the data collected by the information collection device and provides a function of organizing information in a form that is easy for users to understand.

[0286] The "display device" is a device for visually presenting the generated product explanation to the user and plays the role of visualizing information.

[0287] The "presentation device" is a device that appropriately provides the collected and generated product information to the user using the display device.

[0288] This system is a mechanism that allows users to easily collect information about products they are interested in within a physical store. The user sends an image of the product taken using an image acquisition device such as a smartphone to the server. The server processes the received image through an analysis device and identifies the product from the analysis results. It is desirable to use image recognition software such as Google Cloud Vision API for this process.

[0289] Once the product is identified, the server uses the information collection device to collect related information from the Internet. At this stage, databases such as the Amazon Product Review API are utilized to obtain product details and review information.

[0290] Next, the server uses the explanation generation device to generate an explanation in a form that is easy for users to understand from the collected information. For this process, generative AI models such as OpenAI's GPT-3 and ChatGPT are used to generate explanations in natural language.

[0291] The generated descriptions are visualized on a smartphone via a display device. The display device, as part of the display system, accurately presents the product information obtained by this system to the user. This allows the user to instantly check product features, usage instructions, reviews, and more via their smartphone.

[0292] For example, if a user takes a picture of a supplement in a store, its effects, ingredients, and user testimonials will be displayed on their smartphone. Then, by inputting prompts like the following into the AI ​​model, a description will be generated.

[0293] Please describe the following product information: Product name: Supplement, Category: Health food. Briefly describe the product's effects and usage.

[0294] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0295] Step 1:

[0296] The user takes a picture of an item they are interested in using the image acquisition device on their device. The input is image data captured by the camera, and this data is sent to the next processing step.

[0297] Step 2:

[0298] The terminal sends image data to the server. The server inputs the received image data into an analysis device, analyzes the data, extracts product characteristics, and identifies the product. This analysis uses image recognition software such as the Google Cloud Vision API. The output is the identification information of the identified product.

[0299] Step 3:

[0300] The server uses the information collection device based on the analysis results to obtain information related to products from the Internet. It mainly collects data using Amazon product review APIs etc. The input is the identification information of the product, and the output is the detailed information of the product and review information.

[0301] Step 4:

[0302] The server inputs the collected information into the explanation generation device and generates an explanatory text using the generation AI model. This model uses OpenAI's GPT-3 or ChatGPT and uses the prompt text to provide explanations in natural language. The output is an explanatory text for the user.

[0303] Step 5:

[0304] The generated explanatory text is sent from the server to the display device of the terminal so that the user can visually confirm it on the smartphone. The input is the explanatory text, and the output is the visualized information displayed on the terminal.

[0305] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion specific model 59 and perform specific processing using the user's emotion.

[0306] In an embodiment of the present invention, the user first launches a dedicated application on the terminal and takes a picture of the object of interest with the camera. The terminal transmits the image acquired by the image input means to the server via the Internet.

[0307] The server first analyzes the received image using the analysis means to identify the object included in the image. An artificial intelligence model is used for this analysis, and what the object is can be accurately identified. At the same time, the server also receives the user's facial expression data and voice data from the terminal and analyzes the user's emotional state using the emotion engine.

[0308] Next, the server uses information gathering means to search for and obtain relevant information about the identified object via the internet. Then, based on the collected information, the explanation generation means generates an explanation suitable for the user. At this time, the explanation content is adjusted according to the user's emotional state as determined by the emotion engine, and the explanation is provided in a manner that takes into account the user's interests and level of understanding.

[0309] For example, if a user takes a picture of a painting in a museum, the server identifies the painting and collects information about its history, background, and artist. If the user is excited, it can generate a description that includes many more interesting anecdotes; if the user is calm, it can provide a more detailed and expert explanation.

[0310] The generated description is transmitted to the terminal and displayed on the user's screen by a display device. This information allows the user to gain a deeper understanding of the object and simultaneously experience the system's flexible response. This invention makes it possible not only to obtain information about an object, but also to enjoy optimized information tailored to the user's emotions.

[0311] The following describes the processing flow.

[0312] Step 1:

[0313] The user launches the app on their device, points the camera at an object of interest, and takes a picture.

[0314] Step 2:

[0315] The system acquires image data captured by the device. Simultaneously, it collects the user's facial expressions and voice data.

[0316] Step 3:

[0317] The device sends image data, along with user facial expressions and voice data, to the server.

[0318] Step 4:

[0319] The server analyzes the received image data and uses an AI model to identify the object.

[0320] Step 5:

[0321] The server analyzes facial expressions and voice data through an emotion engine to recognize the user's emotional state.

[0322] Step 6:

[0323] The server collects information related to identified objects via the internet.

[0324] Step 7:

[0325] The server generates a descriptive text based on the information it collects. The descriptive text is adjusted according to the results of the emotion engine, and the information is organized in a format that suits the user's emotional state.

[0326] Step 8:

[0327] The server sends the generated explanation to the terminal.

[0328] Step 9:

[0329] The device displays an explanatory text on the user's screen. The user can review the presented information and, if interested, explore further for more detailed information.

[0330] (Example 2)

[0331] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0332] Conventional systems analyze images and provide information about objects in a uniform manner, making it difficult to provide optimal information tailored to each user's individual interests and emotional state. This resulted in problems such as users receiving either insufficient or excessive information. Furthermore, the lack of a mechanism to adjust information based on the user's emotional state risked a decline in the quality of the user experience.

[0333] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0334] In this invention, the server includes means for acquiring images, means for analyzing the images to identify objects, and means for acquiring and analyzing user emotion data. This enables flexible information provision based on the user's emotional state.

[0335] "Means for acquiring images" refers to a device or function that allows a user to photograph an object of interest and acquire it as image data in digital format.

[0336] "Means for identifying an object" refers to a method or technique for analyzing an acquired image to identify a specific object or event present within the image.

[0337] "Means of collecting information from external sources" refers to methods or techniques for obtaining information related to identified objects from databases or websites on the internet.

[0338] "Means for generating explanations" refers to functions and processes that use a generative AI model to automatically create explanatory text in a format that is easy for humans to understand, based on the information that has been collected.

[0339] "Means of display" refers to a device or function for visually presenting the generated explanation on the user's terminal.

[0340] "Means for acquiring and analyzing user emotional data" refers to methods or techniques for acquiring and analyzing user facial expressions, voice, and other emotional expression data to determine the user's emotional state.

[0341] "Means for adjusting the content of explanations" refers to processes and functions for adjusting explanations generated based on the user's sentiment analysis results to be optimized for the user's interests and current emotions.

[0342] To implement this invention, the user, terminal, and server must work together, each fulfilling their respective roles. The user acquires image data on the terminal by taking a picture of an object of interest with the terminal's camera. The terminal then transmits this image data to the server through a dedicated application. A common communication protocol is used for transmitting the image data.

[0343] The server uses artificial intelligence models built with deep learning frameworks such as TensorFlow and PyTorch to analyze images. This analysis identifies objects within the images. Furthermore, the server uses libraries such as OpenCV and Librosa to analyze facial expression and audio data sent from the terminal to evaluate the user's emotional state.

[0344] After identifying the target object, the server uses an information gathering engine to retrieve relevant information from the internet. This process utilizes the Google Custom Search API and web scraping techniques to collect data. Based on the collected information, the server uses a generative AI model to generate a description of the object. This description is adjusted according to the user's emotional state. The generative AI model uses a text generation engine such as GPT-3.

[0345] For example, consider a scenario where a user takes a picture of the painting "Starry Night" in an art museum. The captured image is analyzed on a server, and the subject is identified as "Starry Night." Subsequently, information about the history and background of this painting is collected, and an explanation is generated according to the user's emotional state. A user in an excited state will be provided with an explanation that includes interesting anecdotes about the painter.

[0346] The generated description is sent from the server to the terminal, where the user can visually confirm the information on the terminal's screen. This information allows the user to gain a deeper understanding of the subject matter and enjoy a new experience of receiving information that responds to their own emotions.

[0347] An example of a prompt message would be: "The user has taken an image that includes a painting called 'Starry Night.' The user appears excited. Generate a commentary that includes an interesting anecdote about this painting."

[0348] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0349] Step 1:

[0350] The user launches a dedicated application on their device and takes a picture of an object of interest with the camera. The input is image data of the captured object. The device compresses this image data in JPEG or PNG format and stores it as prepared data.

[0351] Step 2:

[0352] The terminal sends pre-processed image data to the server via internet communication. The input data is a compressed image, and the destination is the server's API endpoint. The output is the image data transferred to the server.

[0353] Step 3:

[0354] The server decompresses the received image data and identifies the object using an image analysis model. The technology used here is a deep learning model based on the TensorFlow or PyTorch framework. The input is image data, and the output is the label of the identified object. In this process, the model analyzes the feature points of the image and determines the most likely object.

[0355] Step 4:

[0356] The server analyzes emotional data (facial expressions and audio data) sent from the terminal. Input includes still images of the user's facial expressions and audio samples. The server evaluates the user's emotional state by analyzing facial expressions with OpenCV and audio with Librosa. The output is the result of the emotional analysis. Specifically, it performs facial feature extraction and audio spectral analysis.

[0357] Step 5:

[0358] The server searches for and collects necessary information via the internet based on the label of the target object. The input is the label of the identified object, and information is collected based on that label. Information is collected using Google Custom Search API and scraping techniques, and structured information data is obtained as output.

[0359] Step 6:

[0360] The server uses the collected information data to generate explanations using a generative AI model. Natural language generation technologies such as GPT-3 are utilized here. The input consists of information data and sentiment analysis results. The output is a user-friendly explanation, with the generated text's style and content adjusted according to the user's sentiment.

[0361] Step 7:

[0362] The server sends the generated explanation back to the terminal. The terminal receives this explanation and displays it visually to the user. The input is the generated explanation, and the output is the display on the user's terminal. Specifically, the text is placed on the terminal screen and presented in a way that the user can easily understand.

[0363] (Application Example 2)

[0364] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0365] While advancements in information technology have led to a vast amount of information available on the internet, it remains difficult for users to efficiently find information that matches their interests and emotions from this enormous volume. In particular, current systems are insufficient in providing information about past works of art and exhibitions tailored to the user's emotional state. Therefore, a system is needed that can provide information that matches the user's interests and emotions.

[0366] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0367] In this invention, the server includes image acquisition means, analysis means, emotion analysis means, information collection means, explanation generation means, and output means. This makes it possible to identify objects of interest to the user and provide information about those objects in an optimized form according to the user's emotional state.

[0368] "Image acquisition means" refers to a function that allows users to photograph objects of interest and acquire the image data.

[0369] "Analysis means" refers to a function for analyzing acquired image data and identifying the target.

[0370] "Emotional analysis means" refers to a function that uses the user's voice and facial expression data to analyze the user's emotional state.

[0371] "Information gathering means" refers to a function for collecting information related to the analyzed subject via a communication network.

[0372] The "explanation generation means" is a function that generates explanations optimized for the user based on collected information and the results of user sentiment analysis.

[0373] "Output means" refers to a function that provides the generated explanation to the user.

[0374] The system that realizes this application example operates through a program installed on a consumer robot, enhancing the user experience in the exhibition space.

[0375] First, the robot uses its built-in camera to acquire images of exhibits that the user is interested in. This image acquisition mechanism allows the robot to process the data of the objects.

[0376] Next, the server uses an analysis tool to analyze the acquired images with a machine learning model (e.g., TensorFlow) to identify the object. After the object is identified, the server activates an emotion analysis tool to collect the user's voice and facial expression data through sensors and uses an emotion analysis engine such as IBM Watson to determine the user's emotions.

[0377] Subsequently, the server activates information gathering tools via the internet, utilizing the Wikipedia API and other resources to collect relevant information about the target object. Based on this information, the explanation generation tool uses OpenAI's GPT model to generate an appropriate explanation tailored to the user's emotions. This generation AI model uses prompts that reflect the user's interests, derived from the collected data and sentiment analysis.

[0378] The generated explanations are provided to the user via output devices through the robot's display or speakers. This allows users to receive customized information tailored to their individual emotional states.

[0379] For example, if the robot recognizes a Renaissance painting and the user is excited, it will provide an interesting explanation including anecdotes about the artist and stories from that era. If the user is calm, it will present detailed information about the technique and historical background of the work.

[0380] An example of a prompt message would be, "Could you tell me about the historical background of this work?" In this way, flexible and sophisticated information provision tailored to the user's needs is achieved.

[0381] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0382] Step 1:

[0383] The robot uses its built-in camera to capture images of exhibits specified by the user. The input is the image data acquired by the camera, and the output is that same image data. This makes it possible to acquire data on specific objects.

[0384] Step 2:

[0385] The server receives image data and performs image analysis using a machine learning model (such as TensorFlow). The input is the acquired image data, and the output is information about the identified objects. It analyzes features in the image and identifies the objects.

[0386] Step 3:

[0387] The server acquires user voice and facial expression data through sensors and analyzes it using an emotion analysis engine (e.g., IBM Watson). The input is voice and facial expression data, and the output is the user's emotional state. Data is acquired using sensors, and data calculations are performed to infer the emotional state.

[0388] Step 4:

[0389] The server collects information related to the target object via the internet. It utilizes the Wikipedia API and other search engines for this purpose. The input is information about the identified object, and the output is related information data. The collected data is then searched and retrieved.

[0390] Step 5:

[0391] The server generates explanatory text using a generative AI model (such as GPT) based on collected information and results obtained from sentiment analysis. The input is relevant information data and the user's emotional state, and the output is an optimized explanatory text. The generative AI model is used to generate prompts and appropriate explanations.

[0392] Step 6:

[0393] The generated explanation is provided to the user through the robot's display and speaker. The input is the generated explanation, and the output is achieved in the form of providing information to the user. Explanations are provided to the user using a display and audio output device.

[0394] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0395] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0396] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0397] [Third Embodiment]

[0398] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0399] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0400] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0401] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0402] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0403] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0404] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0405] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0406] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0407] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0408] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0409] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0410] In an embodiment of the present invention, the user first launches an application using a terminal. This application incorporates a camera function, allowing the user to point the terminal's camera at an object of interest and take an image.

[0411] The device transmits the captured image to the server via communication. The server uses an AI analysis model to analyze the received image data and identify the object. The server then collects information about the identified object from external databases and information sources via the network and organizes the necessary information.

[0412] Subsequently, the server generates an easy-to-understand explanation based on the collected information. This explanation is sent to the terminal, which visually displays it on the screen for the user. By reading this display, the user can easily understand and learn detailed information about the subject.

[0413] As a concrete example, if a user wants to learn about a flower blooming in a garden, they can use their device to take a picture of the flower. The server analyzes the image and identifies the type of flower. If it is identified as a rose, the server collects information about the rose, such as its characteristics, optimal care, and historical background, and sends it to the device as a concise description. Through this description, the user can gain a deeper understanding of the rose.

[0414] In this way, the present invention enables users to quickly obtain detailed information with simple operations, providing a richer experience and learning opportunity.

[0415] The following describes the processing flow.

[0416] Step 1:

[0417] The user launches an application on their device and uses the camera function to take an image of an object of interest.

[0418] Step 2:

[0419] The device acquires the captured image data, converts it to the appropriate format, and prepares to send it to the server.

[0420] Step 3:

[0421] The terminal sends the converted image data to the server.

[0422] Step 4:

[0423] The server processes the received image data using an analysis tool to recognize and identify objects within the image.

[0424] Step 5:

[0425] The server searches external databases and information sources via the internet to collect information related to the identified object.

[0426] Step 6:

[0427] Based on the information collected by the server, an explanatory text is constructed to generate an easy-to-understand explanation for the user.

[0428] Step 7:

[0429] The server sends the generated explanation to the terminal.

[0430] Step 8:

[0431] The terminal displays the explanatory text received from the server on its screen, presenting it visually to the user.

[0432] Step 9:

[0433] Users can review the displayed description, perform further searches if they want to learn more, and share information as needed.

[0434] (Example 1)

[0435] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0436] In modern times, the means of obtaining information about a variety of objects are limited, and it is particularly difficult for the average user to instantly acquire detailed information about a specific object. There is a need for technology that allows even ordinary users without specialized knowledge to easily identify objects and obtain related information.

[0437] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0438] In this invention, the server includes data collection means, analysis means using artificial intelligence, and information acquisition means. This makes it possible for users to quickly and easily obtain identification and related information about various objects simply by acquiring images.

[0439] "Data acquisition means" refers to means consisting of devices and programs for acquiring data such as images.

[0440] "Analysis methods using artificial intelligence" refer to methods that utilize artificial intelligence technology to analyze acquired image data and identify objects.

[0441] "Information acquisition means" refers to means of collecting information related to a specific object via the internet or databases.

[0442] A "generative AI model" is an artificial intelligence model that generates explanatory text in natural language from input information.

[0443] "Display means" refers to devices or programs that visually present generated explanations or information to the user.

[0444] The following processes are performed as embodiments for carrying out the present invention.

[0445] The user launches a dedicated application using their mobile device. This application has a camera function and is used to photograph objects that interest the user. For example, if the user wants to identify a flower they found in a garden, they point the camera of their mobile device at the flower and take a picture.

[0446] The device transmits captured image data to the server via internet communication. This communication uses an encrypted protocol (e.g., HTTPS) to enhance security.

[0447] The server performs artificial intelligence-based analysis on the received images. Specifically, it utilizes machine learning frameworks such as TensorFlow and PyTorch to identify objects within the images. This allows it to recognize what kind of flower is photographed (for example, a rose).

[0448] Next, the server searches online databases or information sources to collect information related to the identified object. During this process, it uses APIs to retrieve detailed information from publicly available sources (e.g., plant identification databases).

[0449] Furthermore, the server uses a generative AI model (e.g., GPT) based on the information obtained to generate explanatory text for the user. For example, it might say, "This flower is a rose, and the best way to grow it is to give it moderate sunlight and keep the soil moist."

[0450] Finally, the server sends the generated description to the terminal, which then displays it to the user. This display allows the user to easily obtain detailed information about the object.

[0451] As a concrete example, an example of a prompt message is as follows: By entering a message in the format "Please tell me the name and characteristics of this plant," you can quickly obtain relevant information.

[0452] In this way, the present invention provides a convenient system that allows users to easily and quickly obtain information about objects of interest.

[0453] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0454] Step 1:

[0455] The user launches the application on their device and uses the camera function to photograph an object. During this process, the device saves the captured image data to a buffer and simultaneously acquires image metadata (e.g., date and time of capture, location information). The input is the image of the object photographed by the user, and the output is the image file stored on the device.

[0456] Step 2:

[0457] The device sends the acquired image data to the server. Specifically, it compresses the image data to an appropriate size using a compression algorithm and sends it to the server via an encrypted protocol (e.g., HTTPS). The input is the captured image data, and the output is the image data transferred to the server.

[0458] Step 3:

[0459] The server analyzes the received image data. This analysis uses an image recognition model (e.g., CNN) equipped with artificial intelligence technology. The server identifies objects from the image and extracts specific features (e.g., color, shape) to determine their type. The input is the received image data, and the output is information about the type of object identified.

[0460] Step 4:

[0461] The server collects information based on the type of object it identifies. Specifically, it searches for and retrieves relevant information from databases and APIs via an internet connection. The input is information about the type of object identified, and the output is detailed information data about the object.

[0462] Step 5:

[0463] The server uses a generative AI model to generate explanatory text for users based on the collected information. The model utilizes natural language processing technology to automatically create appropriate explanations. The input is collected object information, and the output is explanatory text about the object.

[0464] Step 6:

[0465] The server sends the generated description to the terminal. The terminal displays the received description on its screen and presents it to the user. By reading this, the user can obtain detailed information about the object. The input is the generated description, and the output is the visual information displayed on the terminal.

[0466] (Application Example 1)

[0467] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0468] In physical stores, consumers are required to obtain detailed information about products quickly and easily. However, traditional systems have the problem of requiring a lot of effort and time to obtain information, and failing to provide a sufficient purchasing experience.

[0469] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0470] In this invention, the server includes an image acquisition device, an analysis device, and an information collection device. This allows consumers to instantly receive relevant information and make quick and appropriate purchasing decisions simply by taking a picture of products in a store.

[0471] An "image acquisition device" is a device used by users to photograph products they are interested in, and it has the function of acquiring image data.

[0472] An "analysis device" is a device that processes acquired image data and has the function of identifying products. It utilizes artificial intelligence models to extract and analyze product features from images.

[0473] An "information gathering device" is a device that collects information related to a specified product from an external source via an information network, and has the role of accumulating relevant data.

[0474] A "description generation device" is a device that generates product descriptions based on data collected by an information gathering device, providing a function to organize information in a way that is easy for users to understand.

[0475] A "display device" is a device that visually presents generated product descriptions to users, and plays a role in visualizing information.

[0476] A "presentation device" is a device that uses a display device to appropriately provide collected and generated product information to the user.

[0477] This system allows users to easily gather information on products they are interested in within a physical store. Users send images of products they have taken using an image acquisition device such as a smartphone to the server. The server processes the received images through an analysis device and identifies the products based on the analysis results. It is desirable to use image recognition software such as Google Cloud Vision API for this process.

[0478] Once a product is identified, the server uses information gathering devices to collect relevant information from the internet. At this stage, it utilizes databases such as the Amazon Product Review API to obtain product details and review information.

[0479] Next, the server uses an explanation generator to produce an explanation from the collected information in a format that is easy for the user to understand. This process uses generative AI models such as OpenAI's GPT-3 and ChatGPT to generate explanations in natural language.

[0480] The generated descriptions are visualized on a smartphone via a display device. The display device, as part of the display system, accurately presents the product information obtained by this system to the user. This allows the user to instantly check product features, usage instructions, reviews, and more via their smartphone.

[0481] For example, if a user takes a picture of a supplement in a store, its effects, ingredients, and user testimonials will be displayed on their smartphone. Then, by inputting prompts like the following into the AI ​​model, a description will be generated.

[0482] Please describe the following product information: Product name: Supplement, Category: Health food. Briefly describe the product's effects and usage.

[0483] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0484] Step 1:

[0485] The user takes a picture of an item they are interested in using the image acquisition device on their device. The input is image data captured by the camera, and this data is sent to the next processing step.

[0486] Step 2:

[0487] The terminal sends image data to the server. The server inputs the received image data into an analysis device, analyzes the data, extracts product characteristics, and identifies the product. This analysis uses image recognition software such as the Google Cloud Vision API. The output is the identification information of the identified product.

[0488] Step 3:

[0489] The server uses information gathering devices based on the analysis results to obtain product-related information from the internet. It primarily uses the Amazon Product Review API to collect data. The input is product identification information, and the output is detailed product information and review data.

[0490] Step 4:

[0491] The server inputs the collected information into an explanation generation device and generates explanatory text using a generation AI model. This model uses OpenAI's GPT-3 or ChatGPT and provides explanations in natural language using prompts. The output is an explanatory text for the user.

[0492] Step 5:

[0493] The generated explanatory text is sent from the server to the device's display, allowing the user to visually confirm it on their smartphone. The input is the explanatory text, and the output is the visualized information displayed on the device.

[0494] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0495] In an embodiment of the present invention, the user first launches a dedicated application on the terminal and points the camera at an object of interest to take an image. The terminal transmits the image acquired by the image input means to the server via the internet.

[0496] The server first analyzes the received image using an analysis tool to identify the objects contained within the image. This analysis utilizes an artificial intelligence model, enabling accurate identification of the objects. Simultaneously, the server receives user facial expression and voice data from the terminal and analyzes the user's emotional state using an emotion engine.

[0497] Next, the server uses information gathering means to search for and obtain relevant information about the identified object via the internet. Then, based on the collected information, the explanation generation means generates an explanation suitable for the user. At this time, the explanation content is adjusted according to the user's emotional state as determined by the emotion engine, and the explanation is provided in a manner that takes into account the user's interests and level of understanding.

[0498] For example, if a user takes a picture of a painting in a museum, the server identifies the painting and collects information about its history, background, and artist. If the user is excited, it can generate a description that includes many more interesting anecdotes; if the user is calm, it can provide a more detailed and expert explanation.

[0499] The generated description is transmitted to the terminal and displayed on the user's screen by a display device. This information allows the user to gain a deeper understanding of the object and simultaneously experience the system's flexible response. This invention makes it possible not only to obtain information about an object, but also to enjoy optimized information tailored to the user's emotions.

[0500] The following describes the processing flow.

[0501] Step 1:

[0502] The user launches the app on their device, points the camera at an object of interest, and takes a picture.

[0503] Step 2:

[0504] The system acquires image data captured by the device. Simultaneously, it collects the user's facial expressions and voice data.

[0505] Step 3:

[0506] The device sends image data, along with user facial expressions and voice data, to the server.

[0507] Step 4:

[0508] The server analyzes the received image data and uses an AI model to identify the object.

[0509] Step 5:

[0510] The server analyzes facial expressions and voice data through an emotion engine to recognize the user's emotional state.

[0511] Step 6:

[0512] The server collects information related to identified objects via the internet.

[0513] Step 7:

[0514] The server generates a descriptive text based on the information it collects. The descriptive text is adjusted according to the results of the emotion engine, and the information is organized in a format that suits the user's emotional state.

[0515] Step 8:

[0516] The server sends the generated explanation to the terminal.

[0517] Step 9:

[0518] The device displays an explanatory text on the user's screen. The user can review the presented information and, if interested, explore further for more detailed information.

[0519] (Example 2)

[0520] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0521] Conventional systems analyze images and provide information about objects in a uniform manner, making it difficult to provide optimal information tailored to each user's individual interests and emotional state. This resulted in problems such as users receiving either insufficient or excessive information. Furthermore, the lack of a mechanism to adjust information based on the user's emotional state risked a decline in the quality of the user experience.

[0522] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0523] In this invention, the server includes means for acquiring images, means for analyzing the images to identify objects, and means for acquiring and analyzing user emotion data. This enables flexible information provision based on the user's emotional state.

[0524] "Means for acquiring images" refers to a device or function that allows a user to photograph an object of interest and acquire it as image data in digital format.

[0525] "Means for identifying an object" refers to a method or technique for analyzing an acquired image to identify a specific object or event present within the image.

[0526] "Means of collecting information from external sources" refers to methods or techniques for obtaining information related to identified objects from databases or websites on the internet.

[0527] "Means for generating explanations" refers to functions and processes that use a generative AI model to automatically create explanatory text in a format that is easy for humans to understand, based on the information that has been collected.

[0528] "Means of display" refers to a device or function for visually presenting the generated explanation on the user's terminal.

[0529] "Means for acquiring and analyzing user emotional data" refers to methods or techniques for acquiring and analyzing user facial expressions, voice, and other emotional expression data to determine the user's emotional state.

[0530] "Means for adjusting the content of explanations" refers to processes and functions for adjusting explanations generated based on the user's sentiment analysis results to be optimized for the user's interests and current emotions.

[0531] To implement this invention, the user, terminal, and server must work together, each fulfilling their respective roles. The user acquires image data on the terminal by taking a picture of an object of interest with the terminal's camera. The terminal then transmits this image data to the server through a dedicated application. A common communication protocol is used for transmitting the image data.

[0532] The server uses artificial intelligence models built with deep learning frameworks such as TensorFlow and PyTorch to analyze images. This analysis identifies objects within the images. Furthermore, the server uses libraries such as OpenCV and Librosa to analyze facial expression and audio data sent from the terminal to evaluate the user's emotional state.

[0533] After identifying the target object, the server uses an information gathering engine to retrieve relevant information from the internet. This process utilizes the Google Custom Search API and web scraping techniques to collect data. Based on the collected information, the server uses a generative AI model to generate a description of the object. This description is adjusted according to the user's emotional state. The generative AI model uses a text generation engine such as GPT-3.

[0534] For example, consider a scenario where a user takes a picture of the painting "Starry Night" in an art museum. The captured image is analyzed on a server, and the subject is identified as "Starry Night." Subsequently, information about the history and background of this painting is collected, and an explanation is generated according to the user's emotional state. A user in an excited state will be provided with an explanation that includes interesting anecdotes about the painter.

[0535] The generated description is sent from the server to the terminal, where the user can visually confirm the information on the terminal's screen. This information allows the user to gain a deeper understanding of the subject matter and enjoy a new experience of receiving information that responds to their own emotions.

[0536] An example of a prompt message would be: "The user has taken an image that includes a painting called 'Starry Night.' The user appears excited. Generate a commentary that includes an interesting anecdote about this painting."

[0537] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0538] Step 1:

[0539] The user launches a dedicated application on their device and takes a picture of an object of interest with the camera. The input is image data of the captured object. The device compresses this image data in JPEG or PNG format and stores it as prepared data.

[0540] Step 2:

[0541] The terminal sends pre-processed image data to the server via internet communication. The input data is a compressed image, and the destination is the server's API endpoint. The output is the image data transferred to the server.

[0542] Step 3:

[0543] The server decompresses the received image data and identifies the object using an image analysis model. The technology used here is a deep learning model based on the TensorFlow or PyTorch framework. The input is image data, and the output is the label of the identified object. In this process, the model analyzes the feature points of the image and determines the most likely object.

[0544] Step 4:

[0545] The server analyzes emotional data (facial expressions and audio data) sent from the terminal. Input includes still images of the user's facial expressions and audio samples. The server evaluates the user's emotional state by analyzing facial expressions with OpenCV and audio with Librosa. The output is the result of the emotional analysis. Specifically, it performs facial feature extraction and audio spectral analysis.

[0546] Step 5:

[0547] The server searches for and collects necessary information via the internet based on the label of the target object. The input is the label of the identified object, and information is collected based on that label. Information is collected using Google Custom Search API and scraping techniques, and structured information data is obtained as output.

[0548] Step 6:

[0549] The server uses the collected information data to generate explanations using a generative AI model. Natural language generation technologies such as GPT-3 are utilized here. The input consists of information data and sentiment analysis results. The output is a user-friendly explanation, with the generated text's style and content adjusted according to the user's sentiment.

[0550] Step 7:

[0551] The server sends the generated explanation back to the terminal. The terminal receives this explanation and displays it visually to the user. The input is the generated explanation, and the output is the display on the user's terminal. Specifically, the text is placed on the terminal screen and presented in a way that the user can easily understand.

[0552] (Application Example 2)

[0553] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0554] While advancements in information technology have led to a vast amount of information available on the internet, it remains difficult for users to efficiently find information that matches their interests and emotions from this enormous volume. In particular, current systems are insufficient in providing information about past works of art and exhibitions tailored to the user's emotional state. Therefore, a system is needed that can provide information that matches the user's interests and emotions.

[0555] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0556] In this invention, the server includes image acquisition means, analysis means, emotion analysis means, information collection means, explanation generation means, and output means. This makes it possible to identify objects of interest to the user and provide information about those objects in an optimized form according to the user's emotional state.

[0557] "Image acquisition means" refers to a function that allows users to photograph objects of interest and acquire the image data.

[0558] "Analysis means" refers to a function for analyzing acquired image data and identifying the target.

[0559] "Emotional analysis means" refers to a function that uses the user's voice and facial expression data to analyze the user's emotional state.

[0560] "Information gathering means" refers to a function for collecting information related to the analyzed subject via a communication network.

[0561] The "explanation generation means" is a function that generates explanations optimized for the user based on collected information and the results of user sentiment analysis.

[0562] "Output means" refers to a function that provides the generated explanation to the user.

[0563] The system that realizes this application example operates through a program installed on a consumer robot, enhancing the user experience in the exhibition space.

[0564] First, the robot uses its built-in camera to acquire images of exhibits that the user is interested in. This image acquisition mechanism allows the robot to process the data of the objects.

[0565] Next, the server uses an analysis tool to analyze the acquired images with a machine learning model (e.g., TensorFlow) to identify the object. After the object is identified, the server activates an emotion analysis tool to collect the user's voice and facial expression data through sensors and uses an emotion analysis engine such as IBM Watson to determine the user's emotions.

[0566] Subsequently, the server activates information gathering tools via the internet, utilizing the Wikipedia API and other resources to collect relevant information about the target object. Based on this information, the explanation generation tool uses OpenAI's GPT model to generate an appropriate explanation tailored to the user's emotions. This generation AI model uses prompts that reflect the user's interests, derived from the collected data and sentiment analysis.

[0567] The generated explanations are provided to the user via output devices through the robot's display or speakers. This allows users to receive customized information tailored to their individual emotional states.

[0568] For example, if the robot recognizes a Renaissance painting and the user is excited, it will provide an interesting explanation including anecdotes about the artist and stories from that era. If the user is calm, it will present detailed information about the technique and historical background of the work.

[0569] An example of a prompt message would be, "Could you tell me about the historical background of this work?" In this way, flexible and sophisticated information provision tailored to the user's needs is achieved.

[0570] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0571] Step 1:

[0572] The robot uses its built-in camera to capture images of exhibits specified by the user. The input is the image data acquired by the camera, and the output is that same image data. This makes it possible to acquire data on specific objects.

[0573] Step 2:

[0574] The server receives image data and performs image analysis using a machine learning model (such as TensorFlow). The input is the acquired image data, and the output is information about the identified objects. It analyzes features in the image and identifies the objects.

[0575] Step 3:

[0576] The server acquires user voice and facial expression data through sensors and analyzes it using an emotion analysis engine (e.g., IBM Watson). The input is voice and facial expression data, and the output is the user's emotional state. Data is acquired using sensors, and data calculations are performed to infer the emotional state.

[0577] Step 4:

[0578] The server collects information related to the target object via the internet. It utilizes the Wikipedia API and other search engines for this purpose. The input is information about the identified object, and the output is related information data. The collected data is then searched and retrieved.

[0579] Step 5:

[0580] The server generates explanatory text using a generative AI model (such as GPT) based on collected information and results obtained from sentiment analysis. The input is relevant information data and the user's emotional state, and the output is an optimized explanatory text. The generative AI model is used to generate prompts and appropriate explanations.

[0581] Step 6:

[0582] The generated explanation is provided to the user through the robot's display and speaker. The input is the generated explanation, and the output is achieved in the form of providing information to the user. Explanations are provided to the user using a display and audio output device.

[0583] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0584] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0585] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0586] [Fourth Embodiment]

[0587] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0588] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0589] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0590] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0591] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0592] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0593] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0594] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0595] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0596] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0597] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0598] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0599] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0600] In an embodiment of the present invention, the user first launches an application using a terminal. This application incorporates a camera function, allowing the user to point the terminal's camera at an object of interest and take an image.

[0601] The device transmits the captured image to the server via communication. The server uses an AI analysis model to analyze the received image data and identify the object. The server then collects information about the identified object from external databases and information sources via the network and organizes the necessary information.

[0602] Subsequently, the server generates an easy-to-understand explanation based on the collected information. This explanation is sent to the terminal, which visually displays it on the screen for the user. By reading this display, the user can easily understand and learn detailed information about the subject.

[0603] As a concrete example, if a user wants to learn about a flower blooming in a garden, they can use their device to take a picture of the flower. The server analyzes the image and identifies the type of flower. If it is identified as a rose, the server collects information about the rose, such as its characteristics, optimal care, and historical background, and sends it to the device as a concise description. Through this description, the user can gain a deeper understanding of the rose.

[0604] In this way, the present invention enables users to quickly obtain detailed information with simple operations, providing a richer experience and learning opportunity.

[0605] The following describes the processing flow.

[0606] Step 1:

[0607] The user launches an application on their device and uses the camera function to take an image of an object of interest.

[0608] Step 2:

[0609] The device acquires the captured image data, converts it to the appropriate format, and prepares to send it to the server.

[0610] Step 3:

[0611] The terminal sends the converted image data to the server.

[0612] Step 4:

[0613] The server processes the received image data using an analysis tool to recognize and identify objects within the image.

[0614] Step 5:

[0615] The server searches external databases and information sources via the internet to collect information related to the identified object.

[0616] Step 6:

[0617] Based on the information collected by the server, an explanatory text is constructed to generate an easy-to-understand explanation for the user.

[0618] Step 7:

[0619] The server sends the generated explanation to the terminal.

[0620] Step 8:

[0621] The terminal displays the explanatory text received from the server on its screen, presenting it visually to the user.

[0622] Step 9:

[0623] Users can review the displayed description, perform further searches if they want to learn more, and share information as needed.

[0624] (Example 1)

[0625] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0626] In modern times, the means of obtaining information about a variety of objects are limited, and it is particularly difficult for the average user to instantly acquire detailed information about a specific object. There is a need for technology that allows even ordinary users without specialized knowledge to easily identify objects and obtain related information.

[0627] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0628] In this invention, the server includes data collection means, analysis means using artificial intelligence, and information acquisition means. This makes it possible for users to quickly and easily obtain identification and related information about various objects simply by acquiring images.

[0629] "Data acquisition means" refers to means consisting of devices and programs for acquiring data such as images.

[0630] "Analysis methods using artificial intelligence" refer to methods that utilize artificial intelligence technology to analyze acquired image data and identify objects.

[0631] "Information acquisition means" refers to means of collecting information related to a specific object via the internet or databases.

[0632] A "generative AI model" is an artificial intelligence model that generates explanatory text in natural language from input information.

[0633] "Display means" refers to devices or programs that visually present generated explanations or information to the user.

[0634] The following processes are performed as embodiments for carrying out the present invention.

[0635] The user launches a dedicated application using their mobile device. This application has a camera function and is used to photograph objects that interest the user. For example, if the user wants to identify a flower they found in a garden, they point the camera of their mobile device at the flower and take a picture.

[0636] The device transmits captured image data to the server via internet communication. This communication uses an encrypted protocol (e.g., HTTPS) to enhance security.

[0637] The server performs artificial intelligence-based analysis on the received images. Specifically, it utilizes machine learning frameworks such as TensorFlow and PyTorch to identify objects within the images. This allows it to recognize what kind of flower is photographed (for example, a rose).

[0638] Next, the server searches online databases or information sources to collect information related to the identified object. During this process, it uses APIs to retrieve detailed information from publicly available sources (e.g., plant identification databases).

[0639] Furthermore, the server uses a generative AI model (e.g., GPT) based on the information obtained to generate explanatory text for the user. For example, it might say, "This flower is a rose, and the best way to grow it is to give it moderate sunlight and keep the soil moist."

[0640] Finally, the server sends the generated description to the terminal, which then displays it to the user. This display allows the user to easily obtain detailed information about the object.

[0641] As a concrete example, an example of a prompt message is as follows: By entering a message in the format "Please tell me the name and characteristics of this plant," you can quickly obtain relevant information.

[0642] In this way, the present invention provides a convenient system that allows users to easily and quickly obtain information about objects of interest.

[0643] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0644] Step 1:

[0645] The user launches the application on their device and uses the camera function to photograph an object. During this process, the device saves the captured image data to a buffer and simultaneously acquires image metadata (e.g., date and time of capture, location information). The input is the image of the object photographed by the user, and the output is the image file stored on the device.

[0646] Step 2:

[0647] The device sends the acquired image data to the server. Specifically, it compresses the image data to an appropriate size using a compression algorithm and sends it to the server via an encrypted protocol (e.g., HTTPS). The input is the captured image data, and the output is the image data transferred to the server.

[0648] Step 3:

[0649] The server analyzes the received image data. This analysis uses an image recognition model (e.g., CNN) equipped with artificial intelligence technology. The server identifies objects from the image and extracts specific features (e.g., color, shape) to determine their type. The input is the received image data, and the output is information about the type of object identified.

[0650] Step 4:

[0651] The server collects information based on the type of object it identifies. Specifically, it searches for and retrieves relevant information from databases and APIs via an internet connection. The input is information about the type of object identified, and the output is detailed information data about the object.

[0652] Step 5:

[0653] The server uses a generative AI model to generate explanatory text for users based on the collected information. The model utilizes natural language processing technology to automatically create appropriate explanations. The input is collected object information, and the output is explanatory text about the object.

[0654] Step 6:

[0655] The server sends the generated description to the terminal. The terminal displays the received description on its screen and presents it to the user. By reading this, the user can obtain detailed information about the object. The input is the generated description, and the output is the visual information displayed on the terminal.

[0656] (Application Example 1)

[0657] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0658] In physical stores, consumers are required to obtain detailed information about products quickly and easily. However, traditional systems have the problem of requiring a lot of effort and time to obtain information, and failing to provide a sufficient purchasing experience.

[0659] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0660] In this invention, the server includes an image acquisition device, an analysis device, and an information collection device. This allows consumers to instantly receive relevant information and make quick and appropriate purchasing decisions simply by taking a picture of products in a store.

[0661] An "image acquisition device" is a device used by users to photograph products they are interested in, and it has the function of acquiring image data.

[0662] An "analysis device" is a device that processes acquired image data and has the function of identifying products. It utilizes artificial intelligence models to extract and analyze product features from images.

[0663] An "information gathering device" is a device that collects information related to a specified product from an external source via an information network, and has the role of accumulating relevant data.

[0664] A "description generation device" is a device that generates product descriptions based on data collected by an information gathering device, providing a function to organize information in a way that is easy for users to understand.

[0665] A "display device" is a device that visually presents generated product descriptions to users, and plays a role in visualizing information.

[0666] A "presentation device" is a device that uses a display device to appropriately provide collected and generated product information to the user.

[0667] This system allows users to easily gather information on products they are interested in within a physical store. Users send images of products they have taken using an image acquisition device such as a smartphone to the server. The server processes the received images through an analysis device and identifies the products based on the analysis results. It is desirable to use image recognition software such as Google Cloud Vision API for this process.

[0668] Once a product is identified, the server uses information gathering devices to collect relevant information from the internet. At this stage, it utilizes databases such as the Amazon Product Review API to obtain product details and review information.

[0669] Next, the server uses an explanation generator to produce an explanation from the collected information in a format that is easy for the user to understand. This process uses generative AI models such as OpenAI's GPT-3 and ChatGPT to generate explanations in natural language.

[0670] The generated descriptions are visualized on a smartphone via a display device. The display device, as part of the display system, accurately presents the product information obtained by this system to the user. This allows the user to instantly check product features, usage instructions, reviews, and more via their smartphone.

[0671] For example, if a user takes a picture of a supplement in a store, its effects, ingredients, and user testimonials will be displayed on their smartphone. Then, by inputting prompts like the following into the AI ​​model, a description will be generated.

[0672] Please describe the following product information: Product name: Supplement, Category: Health food. Briefly describe the product's effects and usage.

[0673] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0674] Step 1:

[0675] The user takes a picture of an item they are interested in using the image acquisition device on their device. The input is image data captured by the camera, and this data is sent to the next processing step.

[0676] Step 2:

[0677] The terminal sends image data to the server. The server inputs the received image data into an analysis device, analyzes the data, extracts product characteristics, and identifies the product. This analysis uses image recognition software such as the Google Cloud Vision API. The output is the identification information of the identified product.

[0678] Step 3:

[0679] The server uses information gathering devices based on the analysis results to obtain product-related information from the internet. It primarily uses the Amazon Product Review API to collect data. The input is product identification information, and the output is detailed product information and review data.

[0680] Step 4:

[0681] The server inputs the collected information into an explanation generation device and generates explanatory text using a generation AI model. This model uses OpenAI's GPT-3 or ChatGPT and provides explanations in natural language using prompts. The output is an explanatory text for the user.

[0682] Step 5:

[0683] The generated explanatory text is sent from the server to the device's display, allowing the user to visually confirm it on their smartphone. The input is the explanatory text, and the output is the visualized information displayed on the device.

[0684] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0685] In an embodiment of the present invention, the user first launches a dedicated application on the terminal and points the camera at an object of interest to take an image. The terminal transmits the image acquired by the image input means to the server via the internet.

[0686] The server first analyzes the received image using an analysis tool to identify the objects contained within the image. This analysis utilizes an artificial intelligence model, enabling accurate identification of the objects. Simultaneously, the server receives user facial expression and voice data from the terminal and analyzes the user's emotional state using an emotion engine.

[0687] Next, the server uses information gathering means to search for and obtain relevant information about the identified object via the internet. Then, based on the collected information, the explanation generation means generates an explanation suitable for the user. At this time, the explanation content is adjusted according to the user's emotional state as determined by the emotion engine, and the explanation is provided in a manner that takes into account the user's interests and level of understanding.

[0688] For example, if a user takes a picture of a painting in a museum, the server identifies the painting and collects information about its history, background, and artist. If the user is excited, it can generate a description that includes many more interesting anecdotes; if the user is calm, it can provide a more detailed and expert explanation.

[0689] The generated description is transmitted to the terminal and displayed on the user's screen by a display device. This information allows the user to gain a deeper understanding of the object and simultaneously experience the system's flexible response. This invention makes it possible not only to obtain information about an object, but also to enjoy optimized information tailored to the user's emotions.

[0690] The following describes the processing flow.

[0691] Step 1:

[0692] The user launches the app on their device, points the camera at an object of interest, and takes a picture.

[0693] Step 2:

[0694] The system acquires image data captured by the device. Simultaneously, it collects the user's facial expressions and voice data.

[0695] Step 3:

[0696] The device sends image data, along with user facial expressions and voice data, to the server.

[0697] Step 4:

[0698] The server analyzes the received image data and uses an AI model to identify the object.

[0699] Step 5:

[0700] The server analyzes facial expressions and voice data through an emotion engine to recognize the user's emotional state.

[0701] Step 6:

[0702] The server collects information related to identified objects via the internet.

[0703] Step 7:

[0704] The server generates a descriptive text based on the information it collects. The descriptive text is adjusted according to the results of the emotion engine, and the information is organized in a format that suits the user's emotional state.

[0705] Step 8:

[0706] The server sends the generated explanation to the terminal.

[0707] Step 9:

[0708] The device displays an explanatory text on the user's screen. The user can review the presented information and, if interested, explore further for more detailed information.

[0709] (Example 2)

[0710] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0711] Conventional systems analyze images and provide information about objects in a uniform manner, making it difficult to provide optimal information tailored to each user's individual interests and emotional state. This resulted in problems such as users receiving either insufficient or excessive information. Furthermore, the lack of a mechanism to adjust information based on the user's emotional state risked a decline in the quality of the user experience.

[0712] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0713] In this invention, the server includes means for acquiring images, means for analyzing the images to identify objects, and means for acquiring and analyzing user emotion data. This enables flexible information provision based on the user's emotional state.

[0714] "Means for acquiring images" refers to a device or function that allows a user to photograph an object of interest and acquire it as image data in digital format.

[0715] "Means for identifying an object" refers to a method or technique for analyzing an acquired image to identify a specific object or event present within the image.

[0716] "Means of collecting information from external sources" refers to methods or techniques for obtaining information related to identified objects from databases or websites on the internet.

[0717] "Means for generating explanations" refers to functions and processes that use a generative AI model to automatically create explanatory text in a format that is easy for humans to understand, based on the information that has been collected.

[0718] "Means of display" refers to a device or function for visually presenting the generated explanation on the user's terminal.

[0719] "Means for acquiring and analyzing user emotional data" refers to methods or techniques for acquiring and analyzing user facial expressions, voice, and other emotional expression data to determine the user's emotional state.

[0720] "Means for adjusting the content of explanations" refers to processes and functions for adjusting explanations generated based on the user's sentiment analysis results to be optimized for the user's interests and current emotions.

[0721] To implement this invention, the user, terminal, and server must work together, each fulfilling their respective roles. The user acquires image data on the terminal by taking a picture of an object of interest with the terminal's camera. The terminal then transmits this image data to the server through a dedicated application. A common communication protocol is used for transmitting the image data.

[0722] The server uses artificial intelligence models built with deep learning frameworks such as TensorFlow and PyTorch to analyze images. This analysis identifies objects within the images. Furthermore, the server uses libraries such as OpenCV and Librosa to analyze facial expression and audio data sent from the terminal to evaluate the user's emotional state.

[0723] After identifying the target object, the server uses an information gathering engine to retrieve relevant information from the internet. This process utilizes the Google Custom Search API and web scraping techniques to collect data. Based on the collected information, the server uses a generative AI model to generate a description of the object. This description is adjusted according to the user's emotional state. The generative AI model uses a text generation engine such as GPT-3.

[0724] For example, consider a scenario where a user takes a picture of the painting "Starry Night" in an art museum. The captured image is analyzed on a server, and the subject is identified as "Starry Night." Subsequently, information about the history and background of this painting is collected, and an explanation is generated according to the user's emotional state. A user in an excited state will be provided with an explanation that includes interesting anecdotes about the painter.

[0725] The generated description is sent from the server to the terminal, where the user can visually confirm the information on the terminal's screen. This information allows the user to gain a deeper understanding of the subject matter and enjoy a new experience of receiving information that responds to their own emotions.

[0726] An example of a prompt message would be: "The user has taken an image that includes a painting called 'Starry Night.' The user appears excited. Generate a commentary that includes an interesting anecdote about this painting."

[0727] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0728] Step 1:

[0729] The user launches a dedicated application on their device and takes a picture of an object of interest with the camera. The input is image data of the captured object. The device compresses this image data in JPEG or PNG format and stores it as prepared data.

[0730] Step 2:

[0731] The terminal sends pre-processed image data to the server via internet communication. The input data is a compressed image, and the destination is the server's API endpoint. The output is the image data transferred to the server.

[0732] Step 3:

[0733] The server decompresses the received image data and identifies the object using an image analysis model. The technology used here is a deep learning model based on the TensorFlow or PyTorch framework. The input is image data, and the output is the label of the identified object. In this process, the model analyzes the feature points of the image and determines the most likely object.

[0734] Step 4:

[0735] The server analyzes emotional data (facial expressions and audio data) sent from the terminal. Input includes still images of the user's facial expressions and audio samples. The server evaluates the user's emotional state by analyzing facial expressions with OpenCV and audio with Librosa. The output is the result of the emotional analysis. Specifically, it performs facial feature extraction and audio spectral analysis.

[0736] Step 5:

[0737] The server searches for and collects necessary information via the internet based on the label of the target object. The input is the label of the identified object, and information is collected based on that label. Information is collected using Google Custom Search API and scraping techniques, and structured information data is obtained as output.

[0738] Step 6:

[0739] The server uses the collected information data to generate explanations using a generative AI model. Natural language generation technologies such as GPT-3 are utilized here. The input consists of information data and sentiment analysis results. The output is a user-friendly explanation, with the generated text's style and content adjusted according to the user's sentiment.

[0740] Step 7:

[0741] The server sends the generated explanation back to the terminal. The terminal receives this explanation and displays it visually to the user. The input is the generated explanation, and the output is the display on the user's terminal. Specifically, the text is placed on the terminal screen and presented in a way that the user can easily understand.

[0742] (Application Example 2)

[0743] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0744] While advancements in information technology have led to a vast amount of information available on the internet, it remains difficult for users to efficiently find information that matches their interests and emotions from this enormous volume. In particular, current systems are insufficient in providing information about past works of art and exhibitions tailored to the user's emotional state. Therefore, a system is needed that can provide information that matches the user's interests and emotions.

[0745] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0746] In this invention, the server includes image acquisition means, analysis means, emotion analysis means, information collection means, explanation generation means, and output means. This makes it possible to identify objects of interest to the user and provide information about those objects in an optimized form according to the user's emotional state.

[0747] "Image acquisition means" refers to a function that allows users to photograph objects of interest and acquire the image data.

[0748] "Analysis means" refers to a function for analyzing acquired image data and identifying the target.

[0749] "Emotional analysis means" refers to a function that uses the user's voice and facial expression data to analyze the user's emotional state.

[0750] "Information gathering means" refers to a function for collecting information related to the analyzed subject via a communication network.

[0751] The "explanation generation means" is a function that generates explanations optimized for the user based on collected information and the results of user sentiment analysis.

[0752] "Output means" refers to a function that provides the generated explanation to the user.

[0753] The system that realizes this application example operates through a program installed on a consumer robot, enhancing the user experience in the exhibition space.

[0754] First, the robot uses its built-in camera to acquire images of exhibits that the user is interested in. This image acquisition mechanism allows the robot to process the data of the objects.

[0755] Next, the server uses an analysis tool to analyze the acquired images with a machine learning model (e.g., TensorFlow) to identify the object. After the object is identified, the server activates an emotion analysis tool to collect the user's voice and facial expression data through sensors and uses an emotion analysis engine such as IBM Watson to determine the user's emotions.

[0756] Subsequently, the server activates information gathering tools via the internet, utilizing the Wikipedia API and other resources to collect relevant information about the target object. Based on this information, the explanation generation tool uses OpenAI's GPT model to generate an appropriate explanation tailored to the user's emotions. This generation AI model uses prompts that reflect the user's interests, derived from the collected data and sentiment analysis.

[0757] The generated explanations are provided to the user via output devices through the robot's display or speakers. This allows users to receive customized information tailored to their individual emotional states.

[0758] For example, if the robot recognizes a Renaissance painting and the user is excited, it will provide an interesting explanation including anecdotes about the artist and stories from that era. If the user is calm, it will present detailed information about the technique and historical background of the work.

[0759] An example of a prompt message would be, "Could you tell me about the historical background of this work?" In this way, flexible and sophisticated information provision tailored to the user's needs is achieved.

[0760] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0761] Step 1:

[0762] The robot uses its built-in camera to capture images of exhibits specified by the user. The input is the image data acquired by the camera, and the output is that same image data. This makes it possible to acquire data on specific objects.

[0763] Step 2:

[0764] The server receives image data and performs image analysis using a machine learning model (such as TensorFlow). The input is the acquired image data, and the output is information about the identified objects. It analyzes features in the image and identifies the objects.

[0765] Step 3:

[0766] The server acquires user voice and facial expression data through sensors and analyzes it using an emotion analysis engine (e.g., IBM Watson). The input is voice and facial expression data, and the output is the user's emotional state. Data is acquired using sensors, and data calculations are performed to infer the emotional state.

[0767] Step 4:

[0768] The server collects information related to the target object via the internet. It utilizes the Wikipedia API and other search engines for this purpose. The input is information about the identified object, and the output is related information data. The collected data is then searched and retrieved.

[0769] Step 5:

[0770] The server generates explanatory text using a generative AI model (such as GPT) based on collected information and results obtained from sentiment analysis. The input is relevant information data and the user's emotional state, and the output is an optimized explanatory text. The generative AI model is used to generate prompts and appropriate explanations.

[0771] Step 6:

[0772] The generated explanation is provided to the user through the robot's display and speaker. The input is the generated explanation, and the output is achieved in the form of providing information to the user. Explanations are provided to the user using a display and audio output device.

[0773] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0774] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0775] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0776] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0777] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0778] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0779] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0780] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0781] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0782] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0783] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0784] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0785] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0786] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0787] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0788] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0789] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0790] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0791] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0792] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0793] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0794] The following is further disclosed regarding the embodiments described above.

[0795] (Claim 1)

[0796] Image input means,

[0797] An analysis means that analyzes the image received by the image input means to identify the object,

[0798] Information gathering means for collecting information related to the object identified by the analysis means,

[0799] An explanation generation means that generates an explanation from the information collected by the aforementioned information collection means,

[0800] A display means for displaying the explanation generated by the explanation generation means,

[0801] A system that includes this.

[0802] (Claim 2)

[0803] The system according to claim 1, wherein the analysis means identifies an object using an artificial intelligence model.

[0804] (Claim 3)

[0805] The system according to claim 1, wherein the information gathering means retrieves information related to the object by searching for external information via the Internet.

[0806] "Example 1"

[0807] (Claim 1)

[0808] Data collection means for acquiring images,

[0809] An analysis means using artificial intelligence to identify objects by analyzing images acquired by the aforementioned data collection means,

[0810] Information acquisition means for acquiring information related to the object identified by the analysis means using the Internet,

[0811] An explanation generation means that generates an explanation using an AI model generated from the information collected by the aforementioned information acquisition means,

[0812] A display means for visually presenting the explanation created by the explanation generation means,

[0813] A system that includes this.

[0814] (Claim 2)

[0815] The system according to claim 1, wherein the analysis means uses a model specifically for image analysis.

[0816] (Claim 3)

[0817] The system according to claim 1, wherein the information acquisition means acquires information related to the object from a database via a network.

[0818] "Application Example 1"

[0819] (Claim 1)

[0820] Image acquisition device,

[0821] An analysis device that analyzes images received by the aforementioned image acquisition device to identify products,

[0822] An information collection device that collects information on products identified by the aforementioned analysis device,

[0823] An explanation generation device that generates an explanation from the information collected by the aforementioned information collection device,

[0824] A display device that visualizes the explanation generated by the explanation generation device,

[0825] A display device that displays product information using the aforementioned display device,

[0826] A system that includes this.

[0827] (Claim 2)

[0828] The system according to claim 1, wherein the analysis device identifies products using an artificial intelligence model.

[0829] (Claim 3)

[0830] The system according to claim 1, wherein the information gathering device retrieves information related to identified products by searching external data through an information network.

[0831] "Example 2 of combining an emotion engine"

[0832] (Claim 1)

[0833] Means of acquiring images,

[0834] A means for analyzing the aforementioned image to identify the object,

[0835] A means for collecting information related to the object of the analyzed image from an external source,

[0836] Means for generating an explanation based on information related to the aforementioned object,

[0837] Means for displaying the generated explanation,

[0838] A means of acquiring and analyzing user sentiment data,

[0839] A means of adjusting the explanatory content based on the results of user sentiment analysis,

[0840] A system that includes this.

[0841] (Claim 2)

[0842] The system according to claim 1, which identifies an object using an artificial intelligence model.

[0843] (Claim 3)

[0844] The system according to claim 1, which searches for external information via the internet and obtains information related to the object.

[0845] "Application example 2 when combining with an emotional engine"

[0846] (Claim 1)

[0847] Image acquisition method,

[0848] An analysis means that analyzes the image received by the image acquisition means to identify the target,

[0849] A means of emotion analysis that collects the user's voice and facial expressions and analyzes their emotions,

[0850] Information gathering means for collecting information related to the object identified by the analysis means,

[0851] An explanation generation means that generates an explanation based on the information collected by the information collection means and the analysis results of the emotion analysis means,

[0852] An output means that provides the explanation generated by the explanation generation means,

[0853] A system that includes this.

[0854] (Claim 2)

[0855] The system according to claim 1, wherein the analysis means identifies an object using a machine learning model.

[0856] (Claim 3)

[0857] The system according to claim 1, wherein the information gathering means retrieves external information via a communication network to obtain information related to the target. [Explanation of Symbols]

[0858] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Image acquisition device, An analysis device that analyzes images received by the aforementioned image acquisition device to identify products, An information collection device that collects information on products identified by the aforementioned analysis device, An explanation generation device that generates an explanation from the information collected by the aforementioned information collection device, A display device that visualizes the explanation generated by the explanation generation device, A display device that displays product information using the aforementioned display device, A system that includes this.

2. The system according to claim 1, wherein the analysis device identifies products using an artificial intelligence model.

3. The system according to claim 1, wherein the information gathering device retrieves information related to identified products by searching external data through an information network.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A