system

The system addresses the limitations of fixed explanatory content in museums by using generative AI for exhibit recognition and personalized responses, enhancing user understanding and engagement through interactive and entertaining explanations.

JP2026063897APending Publication Date: 2026-04-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

Existing methods for providing information about exhibits in art museums and museums are inflexible, requiring fixed explanatory content that does not adapt to user interests, lacks reflection of new information, and necessitates specialized knowledge, limiting user understanding and engagement.

Method used

A system utilizing generative AI for exhibit recognition, user question and feedback collection, and character-based explanations, employing image capture, natural language generation, and generative AI to provide personalized and interactive explanations and responses.

Benefits of technology

Enhances user understanding and engagement by providing multifaceted, interactive, and entertaining explanations and responses to user queries and feedback, leveraging character avatars and voice recordings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063897000001_ABST
    Figure 2026063897000001_ABST
Patent Text Reader

Abstract

We provide a system that utilizes generative AI for recognizing exhibits, providing explanations, and collecting user questions and feedback. [Solution] A system comprising: means for capturing images using a camera installed on a terminal and sending the image data to a server; means for analyzing the received image data and generating an ID corresponding to the identified exhibit; means for obtaining explanatory data from a database based on the identified exhibit ID and converting the explanatory data into a user-friendly format using natural language generation technology; means for sending the converted explanatory data to the user's terminal and displaying or playing it aloud on the terminal; means for inputting user questions into the terminal via voice or text; means for sending the input question data to a server; means for analyzing the received question data and generating an appropriate answer using a generative AI; and means for sending the generated answer to the user's terminal and displaying or playing it aloud on the terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005] ,

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Currently, methods for explaining and providing information about exhibits in art museums and museums generally use pamphlets or fixed explanation plates. However, with these methods, it is difficult for users to flexibly respond to content and questions in which they are individually interested. Also, since the explanatory content is fixed, it is not possible to reflect new information or user impressions. Furthermore, when providing detailed information such as the historical background of the exhibit or the intention of the author, specialized knowledge is required, so it is difficult for users to gain a sufficient understanding. As a result, there is a problem that the user's understanding and interest in the exhibit do not deepen.

Means for Solving the Problems

[0005] To solve the above problems, this invention provides a system that utilizes generative AI for recognizing exhibits, providing explanations, and collecting user questions and feedback. The system includes means for capturing images of exhibits with a camera installed on a terminal and transmitting the image data to a server. The server has means for analyzing the received image data and generating an ID corresponding to the identified exhibit, obtaining explanation data from a database based on the identified exhibit ID, and converting the explanation data into a user-friendly format using natural language generation technology. Furthermore, the system includes means for the user to input questions via voice or text on the terminal and transmit the question data to the server. The server has means for analyzing the received question data using generative AI, generating an appropriate answer and transmitting it to the terminal, and displaying or playing the answer aloud on the terminal. The system also includes means for the user to input their feedback and opinions via voice or text and transmit them to the server. The server collects and organizes the received feedback data and stores it in a database, and generates appropriate feedback based on that feedback data and transmits it to the terminal. Furthermore, the system includes means for obtaining character information set for each exhibit from the server, reflecting the character's tone of voice and personality in the explanations and answers, and providing this to the user through character avatars and voice recordings. In this way, users can gain a deeper understanding of and enjoyment of the exhibits.

[0006] "Exhibits" refer to the diverse works and materials displayed in museums, art galleries, and other similar institutions.

[0007] A "device" refers to a portable information device used by a user, such as a smartphone, tablet, or personal computer.

[0008] A "server" is a computer system that receives requests from terminals via a network and processes the data.

[0009] A "camera" is a device installed in a terminal that receives light and generates image data.

[0010] "Image data" refers to visual information captured by a camera, represented in digital data format.

[0011] "Image recognition" is a method of identifying objects and features from image data using computer vision technology.

[0012] "ID" is an identifier used to uniquely identify an exhibit.

[0013] A "database" is a system for systematically collecting and storing information.

[0014] "Explanatory data" refers to data that includes explanatory information and background information about the exhibits.

[0015] Natural Language Generation (NLG) is an artificial intelligence technology that converts text data into a natural language format.

[0016] "Question data" refers to digital data that includes the content of questions entered by the user.

[0017] "Generative AI" refers to artificial intelligence technology that generates text and information based on input data.

[0018] "Feedback" refers to responses or additional information provided to address users' comments and opinions.

[0019] "Character information" refers to information such as the character's background, speech patterns, and personality related to the exhibit.

[0020] An "avatar" is a digital image or 3D model used to provide a visual representation of a character.

[0021] "Speech" refers to audio data that is produced by reproducing text data as speech using speech synthesis technology. [Brief explanation of the drawing]

[0022] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. < Advocates [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined. [[ID=A2]]

Embodiments for Carrying Out the Invention

[0023] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0024] First, let's explain the terminology used in the following explanation.

[0025] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0026] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0027] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0028] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0030] [First Embodiment]

[0031] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0032] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0033] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0035] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0037] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0038] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0039] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0040] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0041] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0042] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0043] ---

[0044] This invention relates to a system that uses generative AI to provide explanations to users about exhibits in art museums and museums, and to respond to user questions and comments. This system comprehensively performs exhibit recognition, acquisition and display of explanatory data, response to user questions, and collection and feedback of comments.

[0045] System Overview

[0046] First, when a user stands in front of an exhibit, they use a device (such as a smartphone or tablet) to capture an image of the exhibit. The device generates image data using its camera and sends this data to a server. The server uses image recognition technology to identify the exhibit and generates an ID. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent from the server to the device and displayed or played as audio for the user on the device.

[0047] If a user has a question about an exhibit, they enter the question into the terminal via voice or text input. The terminal sends the entered question data to a server, which uses generative AI to generate an appropriate answer to the question. This answer is then sent to the terminal and displayed or played back to the user.

[0048] Furthermore, when users want to input their thoughts and opinions, they do so via voice or text on the device. The device sends the input data to the server, which stores the received data in a database and generates appropriate feedback, which is then sent back to the device. This allows users to receive feedback on their own comments.

[0049] The system also includes features to enhance entertainment value. The server retrieves character information set for each exhibit and reflects the character's tone and personality in the explanations and answers. The generated explanations and answers are played back using the character's avatar and voice, providing users with both visual and auditory experiences.

[0050] Specific example

[0051] For example, suppose a user stands in front of a "famous painting" and launches an app on their device. When the user points their camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses image recognition technology to identify the "famous painting" and generates an ID. Using this ID, the server retrieves explanatory data from a database, generates an explanatory text such as "This painting was painted in [era], and its characteristics are [characteristics]," and sends it to the device. The device then displays or plays the explanatory text for the user.

[0052] Next, when the user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses generative AI to generate an answer to the question and sends the answer, such as "The following events influenced the creation of this painting," to the device. The device then displays or plays the answer aloud to the user.

[0053] Furthermore, when a user enters a comment such as "The use of color in this painting is wonderful," the device sends the comment data to the server. The server stores the comment data in a database and generates feedback such as, "Many experts have praised the use of color. We are pleased that you noticed this point," and sends it back to the device. This allows the user to gain a deeper understanding and greater satisfaction.

[0054] The above is an example of implementing the system according to the present invention. This makes it possible to provide multifaceted and individualized information about exhibits, thereby enhancing the user experience.

[0055] ---

[0056] The following describes the processing flow.

[0057] ---

[0058] Program processing flow

[0059] Recognition processing of exhibits

[0060] Step 1:

[0061] The user stands in front of the exhibit and launches the app on their device (e.g., a smartphone or tablet).

[0062] Step 2:

[0063] The device's camera is used to capture images of the exhibits. The device activates the camera, generates image data, and saves that image data to memory.

[0064] Step 3:

[0065] The device sends the captured image data to the server. Specifically, it sends the image data to the server via an HTTP request.

[0066] Step 4:

[0067] The server executes an image recognition algorithm to identify the received image data. For example, the server uses a Convolutional Neural Network (CNN) model to detect exhibits from the image and generate a corresponding exhibit ID.

[0068] Processing to retrieve explanations for exhibits

[0069] Step 5:

[0070] The server uses the generated exhibit ID to retrieve the corresponding explanatory data from the explanatory database. This involves extracting explanatory information from the database using SQL queries or similar methods.

[0071] Step 6:

[0072] The server uses natural language generation (NLG) technology to convert the acquired explanatory data into a user-friendly format. Specifically, the NLG engine generates the explanatory text.

[0073] Step 7:

[0074] The server sends the generated explanatory data to the terminal. It returns data containing the explanatory text to the terminal as an HTTP response.

[0075] Step 8:

[0076] The device displays the received explanatory data to the user. This is done using a text display UI or a speech synthesis engine to provide the user with audio explanations.

[0077] User question answering process

[0078] Step 9:

[0079] If a user has a question about an exhibit, they can input their question into the terminal either by voice or text. The terminal uses a speech recognition engine to convert the voice to text or accepts the text input.

[0080] Step 10:

[0081] The terminal sends question data in text format to the server. An HTTP request is used for this transmission.

[0082] Step 11:

[0083] The server analyzes the received question data. Natural language understanding (NLU) technology is used for the analysis to understand the user's question.

[0084] Step 12:

[0085] The server uses a generative AI (e.g., GPT-4®) to generate answers to questions. The generated answers are structured as natural language text.

[0086] Step 13:

[0087] The server sends the generated response to the terminal. The response data is returned to the terminal as an HTTP response.

[0088] Step 14:

[0089] The device displays the answer to the user. This display uses a text display UI or a speech synthesis engine.

[0090] User feedback sharing process

[0091] Step 15:

[0092] Users input their impressions and opinions about the exhibits into the device via voice or text. The device uses a speech recognition engine to convert the voice to text or receives the text input.

[0093] Step 16:

[0094] The device sends feedback data in text format to the server. HTTP requests are used for transmission.

[0095] Step 17:

[0096] The server analyzes and organizes the received feedback data and saves it to a database. SQL queries are used to write the data to the database for saving.

[0097] Step 18:

[0098] The server generates feedback using natural language generation technology based on the feedback data. The NLG engine generates appropriate feedback sentences.

[0099] Step 19:

[0100] The server sends the generated feedback to the terminal. The feedback data is returned to the terminal as an HTTP response.

[0101] Step 20:

[0102] The device displays feedback to the user. This is done using a text display UI or a speech synthesis engine.

[0103] Processing to enhance entertainment value

[0104] Step 21:

[0105] The server retrieves character information set for each exhibit. It uses SQL queries to retrieve character information from the database.

[0106] Step 22:

[0107] When the server generates explanatory text and response text based on character information, it reflects the character's speech pattern and personality. An NLG engine is used for generation.

[0108] Step 23:

[0109] The server transmits the character's voice data and avatar information to the terminal. This includes synthesized voice data and image / video data.

[0110] Step 24:

[0111] The device provides explanations and answers to the user using character avatars and voice recordings. The device displays character avatars and plays explanations and answers via voice recordings.

[0112] ---

[0113] The above is the specific processing flow of the system program.

[0114] (Example 1)

[0115] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0116] Museums and art galleries need efficient and effective ways to provide explanations about their exhibits to users. Furthermore, there is a lack of systems that can quickly and appropriately respond to users' questions and comments about the exhibits, thereby providing a richer experience. Traditional systems primarily rely on paper-based descriptions for each exhibit, lacking interactivity and making it difficult to capture user interest. Moreover, the lack of entertainment-oriented exhibit explanations utilizing characters limits the potential for improving the user experience.

[0117] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0118] In this invention, the server includes means for capturing images using a camera device provided on the user terminal to identify exhibits and transmitting the image data to the server; means for analyzing the received image data on the server and generating an identifier corresponding to the identified exhibit; means for obtaining information data from a database based on the identified exhibit identifier on the server and converting the information data into a user-friendly format using natural language generation technology; means for transmitting the converted information data to the user terminal and displaying or playing it aloud on the user terminal; means for inputting user questions into the user terminal via voice or text input; means for transmitting the question data entered on the user terminal to the server; means for analyzing the received question data on the server and generating an appropriate answer using a generation AI model; and means for transmitting the generated answer to the user terminal. The system includes means for displaying or playing audio on the user terminal, means for inputting user feedback and opinions via voice or text into the user terminal, means for sending feedback data entered on the user terminal to a server, means for the server to collect and organize the received feedback data and store it in a database, means for generating appropriate feedback using natural language generation technology based on the feedback data, means for sending the generated feedback to the user terminal and displaying or playing audio on the user terminal, means for obtaining character information set for each exhibit from the server and reflecting the character's tone of voice and personality in the explanations and answers, means for sending data to the user terminal to express the character-based explanations and answers in the character's voice, and means for providing the character to the user as an avatar or voice on the user terminal. This makes it possible to provide users with interactive and rich explanations and responses to exhibits, as well as a highly entertaining experience.

[0119] "Exhibits" refer to works of art and materials that are made available to users in art museums and museums.

[0120] "User terminal" refers to a portable information device owned by a user, primarily such as a smartphone or tablet.

[0121] "Image capture equipment" refers to cameras and other optical instruments used to acquire image data.

[0122] A "server" is a central computing system that processes, stores, and provides data over a network.

[0123] "Image data" refers to data that represents captured visual information in a digital format.

[0124] "Analysis" refers to the process of examining received images and data in detail to extract information.

[0125] An "identifier" is a code or number generated to uniquely identify a particular exhibit.

[0126] A "database" is a system for efficiently storing, searching, and managing large amounts of information.

[0127] "Information data" refers to digital data that includes detailed explanations and information about the exhibits.

[0128] "Natural language generation technology" refers to technology that allows machines to automatically generate sentences that are easy for humans to understand.

[0129] "Question data" refers to data that digitally represents questions about exhibits entered by users.

[0130] A "generative AI model" refers to an algorithm or model that uses artificial intelligence to generate appropriate answers or text.

[0131] "Feedback" refers to the responses or evaluations provided in response to user input.

[0132] "Character information" refers to data that includes specific characteristics and speech patterns of fictional people, animals, and other elements set for each exhibit.

[0133] An "avatar" is an image or model used to visually represent a character.

[0134] "Voice" refers to data used to digitally reproduce the character's voice.

[0135] This invention relates to a system that uses generative AI to provide explanations to users about exhibits in art museums and museums, and to respond to user questions and comments. This system comprehensively performs exhibit recognition, acquisition and display of explanatory data, response to user questions, and collection and feedback of comments.

[0136] System Overview

[0137] This system works as follows:

[0138] First, when a user stands in front of an exhibit, they use their device (e.g., a smartphone or tablet) to capture an image of the exhibit. The device's camera generates image data, which is then sent to a server. The server uses image recognition technology to identify the exhibit and generates an identifier. Based on this identifier, the server retrieves relevant information data from a database and generates an explanatory text using natural language generation (NLG) technology. The generated explanatory text is sent from the server to the device and displayed or played back as audio for the user.

[0139] Hardware and software details

[0140] 1. Terminal (User terminal):

[0141] Hardware used: Smartphone, tablet

[0142] Software used: Dedicated app, camera function, audio recording and playback function

[0143] 2. Server:

[0144] Software used:

[0145] Image recognition technology: AWS® Rekognition, Google® Cloud Vision

[0146] Databases: MySQL (registered trademark), MongoDB

[0147] Natural Language Generation Technology (NLG): OpenAI® GPT-4

[0148] Audio playback technology: Amazon Polly, Google Text-to-Speech

[0149] Processing details

[0150] The processing flow at each step is as follows:

[0151] 1. User-captured images of exhibits:

[0152] Users point their device's camera at the exhibit and capture images using a dedicated app.

[0153] 2. Sending image data from the terminal to the server:

[0154] The terminal encodes the captured image data in BASE64 format and sends it to the server using an HTTP POST request.

[0155] 3. Server-based recognition of exhibits and acquisition of data:

[0156] The server analyzes the received image data, using AWS Rekognition or Google Cloud Vision to recognize exhibits and generate identifiers. Based on the identified exhibit identifiers, the server retrieves relevant information data from MySQL or MongoDB databases.

[0157] 4. Server generates explanatory text and sends it to the terminal:

[0158] The server uses OpenAI's GPT-4 to generate natural language based on the acquired data. The generated explanatory text is sent from the server to the terminal and displayed or played back as audio for the user.

[0159] Specific usage scenarios

[0160] For example, suppose a user stands in front of a "famous painting" and launches a smartphone app. When the user points the camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses AWS Rekognition to identify the "famous painting" and generates an identifier. Using this identifier, the server retrieves explanatory data from a MySQL database, generates an explanatory text using GPT-4, and sends it to the device. The device then displays or plays an audio message stating, "This painting was painted by a famous Renaissance painter, and its distinguishing feature is the meticulous detail."

[0161] Examples of prompts for generative AI models

[0162] "Generate a descriptive text for a famous painting. The text should include the artist, date of creation, and key features."

[0163] The above details the embodiments of the present invention. This system allows users to obtain multifaceted information about exhibits and deepen their understanding of the exhibits interactively.

[0164] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0165] Step 1:

[0166] The user stands in front of the exhibit and uses the camera on their device to capture an image of the exhibit.

[0167] Input: Image of the exhibit

[0168] Output: Captured image data

[0169] Specific operation: The user launches a dedicated app on their smartphone or tablet, points the camera at the exhibit, and presses the shutter button. This causes the camera to capture image data and save it to the device's memory.

[0170] Step 2:

[0171] The terminal encodes the captured image data in BASE64 format and sends it to the server using an HTTP POST request.

[0172] Input: Captured image data

[0173] Output: Image data encoded in BASE64 format

[0174] Specific operation: The device's dedicated app encodes the captured image data into BASE64 format and uploads the image data to the server by sending an HTTP POST request to the specified API endpoint.

[0175] Step 3:

[0176] The server analyzes the received image data and generates identifiers corresponding to the identified exhibits.

[0177] Input: Image data in BASE64 format

[0178] Output: Identifier of the identified exhibit

[0179] Specific operation: The server uses image recognition technology (e.g., AWS Rekognition or Google Cloud Vision) to analyze the received image data. From the analysis results, it identifies the exhibits and generates corresponding identifiers.

[0180] Step 4:

[0181] The server retrieves relevant information data from the database based on the identified exhibit identifier.

[0182] Input: Exhibit identifier

[0183] Output: Information data about the exhibits

[0184] Specific operation: The server uses the generated exhibit identifier to query the database (e.g., MySQL or MongoDB) to retrieve information data about the corresponding exhibit.

[0185] Step 5:

[0186] Based on the acquired information data, the server uses a generative AI model (e.g., GPT-4) to perform natural language generation and generate explanatory text.

[0187] Input: Information data about the exhibits

[0188] Output: Generated explanatory text

[0189] Specific operation: The server analyzes the information data and inputs prompt sentences into the generation AI model to generate appropriate explanatory text.

[0190] Example prompt: "Generate a descriptive text for a famous painting. The text should include the artist, date of creation, and key features."

[0191] Step 6:

[0192] The server sends the generated explanatory text to the user's terminal, where it is displayed or played as audio.

[0193] Input: Generated explanatory text

[0194] Output: Explanatory text sent to the user's terminal

[0195] Specific operation: The server sends the generated explanatory text as an HTTP response to the user's terminal. The user's terminal either displays the received explanatory text on the screen or plays it aloud using speech synthesis technology (e.g., Amazon Polly or Google Text-to-Speech).

[0196] Step 7:

[0197] When a user enters a question about an exhibit, the user's terminal sends the question data to the server.

[0198] Input: Question entered by the user (voice or text)

[0199] Output: Question data sent to the server

[0200] Specific operation: The user asks a question via voice input or text input through the app on their device. In the case of voice input, the device converts the voice to text and sends it to the server as question data.

[0201] Step 8:

[0202] The server analyzes the received question data and generates appropriate answers using a generative AI model.

[0203] Input: Question data sent to the server

[0204] Output: Generated answer

[0205] Specific operation: The server analyzes the question data and prompts the generation AI model to generate appropriate answers. The generated answers are saved in text format.

[0206] Example prompt: "Generate answers to the following question: 'What is the background of this painting?'"

[0207] Step 9:

[0208] The server sends the generated response to the user's terminal, where it is displayed or played back as audio.

[0209] Input: Generated answer

[0210] Output: Response sent to the user's terminal

[0211] Specific operation: The server sends the generated response to the user's terminal as an HTTP response. The user's terminal displays the received response on the screen or plays it back as audio using speech synthesis technology.

[0212] Step 10:

[0213] When a user enters their thoughts and opinions about an exhibit, the user's device sends the feedback data to the server.

[0214] Input: User feedback (voice or text)

[0215] Output: Feedback data sent to the server

[0216] Specific operation: Users provide feedback via voice input or text input through the app on their device. In the case of voice input, the device converts the voice to text and sends it to the server as feedback data.

[0217] Step 11:

[0218] The server collects and organizes the received feedback data and stores it in a database.

[0219] Input: Feedback data sent to the server

[0220] Output: Feedback data stored in the database

[0221] Specific operation: The server collects and organizes the received feedback data, and then saves it to a database (e.g., MongoDB).

[0222] Step 12:

[0223] Based on the feedback data, the server uses a generative AI model to generate appropriate feedback.

[0224] Input: Feedback data stored on the server

[0225] Output: Generated feedback

[0226] Specific operation: The server analyzes the stored feedback data and prompts the AI ​​model to generate appropriate feedback. The generated feedback is saved in text format.

[0227] Step 13:

[0228] The server sends the generated feedback to the user's terminal, where it is displayed or played as audio.

[0229] Input: Generated feedback

[0230] Output: Feedback sent to the user terminal

[0231] Specific operation: The server sends the generated feedback to the user terminal as an HTTP response. The user terminal displays the received feedback on the screen or plays it back as audio using speech synthesis technology.

[0232] Step 14:

[0233] The server retrieves character information set for each exhibit and reflects the character's tone of voice and personality in the explanations and answers.

[0234] Input: Exhibit identifier

[0235] Output: Explanatory text and answer text reflecting character information

[0236] Specific operation: The server uses the exhibit identifier to retrieve character information from the database and reflects the character's tone of voice and personality in the explanatory text and response text.

[0237] Step 15:

[0238] The server sends explanatory text and response text based on character information to the user's terminal, and the terminal plays avatars and voice clips.

[0239] Input: Explanatory text or answer text that reflects character information

[0240] Output: Explanatory text and response text with character information sent to the user's terminal.

[0241] Specific operation: The server sends explanatory text and answer text with character information to the user terminal as an HTTP response. The user terminal displays the received content on the screen and plays the character's voice using speech synthesis technology.

[0242] The above outlines the specific processing steps of this system. It details a comprehensive mechanism for enhancing the user experience, from exhibit recognition and explanation to question-and-answer sessions, feedback, and even character entertainment.

[0243] (Application Example 1)

[0244] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0245] Traditional museum and art gallery systems for explaining exhibits are limited to providing information about the exhibits themselves, and do not adequately address users' questions or comments. Similarly, in physical stores and tourist destinations, there is a lack of detailed information about exhibits and products, as well as insufficient interactive communication with users. As a result, users' understanding and experience are limited, and the appeal of exhibits and products cannot be fully conveyed.

[0246] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0247] This invention includes means for a server to capture images using a camera installed in a terminal to recognize exhibits and transmit the image data to the server; means for the server to analyze the received image data and generate an identifier corresponding to the identified exhibit; means for the server to obtain explanatory data from a database based on the identified exhibit identifier and convert the explanatory data into a user-friendly format using natural language generation technology; means for transmitting the converted explanatory data to the user's terminal and displaying or playing it aloud on the terminal; means for inputting user questions into the terminal via voice or text; means for transmitting the question data entered on the terminal to the server; means for the server to analyze the received question data and generate appropriate answers using a generation AI model; means for transmitting the generated answers to the user's terminal and displaying or playing them aloud on the terminal; and means for providing information via the terminal when the user stands in front of a product or landmark. This allows the user to not only obtain detailed information about exhibits and products, but also to resolve questions in an interactive way and receive feedback on their impressions.

[0248] "Exhibits" is a general term for works of art and historical items displayed in art museums and museums, as well as objects that attract the interest of users in physical stores and tourist destinations.

[0249] A "device" refers to a portable electronic device that a user can carry with them, such as a smartphone, tablet, or personal computer, and is equipped with a camera, microphone, and display.

[0250] "Capturing an image using the camera" refers to saving visual information as a photograph to the device's internal memory using the device's built-in camera.

[0251] A "server" is a computer system that receives and processes data sent from a terminal and provides the necessary information.

[0252] "Image data" refers to digital files containing visual information captured by a camera.

[0253] An "identifier" is a code or number used by the server to uniquely identify a specific exhibit, and is used for database operations.

[0254] A "database" is an information management system for systematically storing explanatory data and other related information about exhibits.

[0255] "Natural language generation technology" refers to the technology that allows computers to generate text in human language that is easy to read and understand.

[0256] "User's device" refers to a device such as a smartphone or tablet used by the user, which displays or plays back information received from the server.

[0257] "Question data" refers to digital data that includes the content of questions entered by the user via voice or text.

[0258] A "generative AI model" is an artificial intelligence technology that automatically generates appropriate answers based on input information.

[0259] "Feedback" refers to responses and reactions to user comments and opinions, and is generated using natural language generation technology.

[0260] An "avatar" is an image used to provide users with a visual representation of a character.

[0261] This invention is a system that provides explanations to users about exhibits and products, and also responds to user questions and feedback. A specific embodiment of this system is shown below.

[0262] System Configuration

[0263] This system consists of terminals, a server, and a database. Terminals include smartphones, tablets, and personal computers. The server handles data processing and storage, as well as the execution of generated AI models. The database stores explanatory data about exhibits and products.

[0264] Hardware and software

[0265] Hardware: Smartphones, tablets, and personal computers (these electronic devices are equipped with cameras, microphones, and displays).

[0266] Software used: Python, OpenCV (for capturing camera images), Requests (for sending and receiving data), Hugging Face's Transformers (using GPT-3® as a generative AI model, etc.).

[0267] Process flow

[0268] 1. Image capture

[0269] The user stands in front of an exhibit or product and captures an image using the device's camera. This image data is then sent from the device to the server.

[0270] 2. Image recognition and acquisition of explanatory data

[0271] The server analyzes the received image data and generates identifiers to identify exhibits and products. Based on these identifiers, the server retrieves explanatory data from a database. This explanatory data is converted into a user-friendly format using natural language generation technology. The converted explanatory data is then sent to the terminal and displayed or played as audio.

[0272] 3. Questions from users

[0273] If a user has questions about the explanation, they can ask them via voice input or text input. The device then sends this question data to the server.

[0274] 4. Generating answers to questions

[0275] The server analyzes the received question data and generates appropriate answers using a generative AI model (e.g., GPT-3). These generated answers are sent to the terminal and displayed or played aloud.

[0276] 5. User feedback and opinions

[0277] Users can input their thoughts and opinions about the exhibits and products. This can be done via voice input or text input. The terminal then sends the inputted feedback data to the server.

[0278] 6. Generating feedback on comments

[0279] The server collects and organizes the received feedback data and stores it in a database. Furthermore, it uses natural language generation technology to generate appropriate feedback based on the feedback. This feedback is sent to the device and displayed or played aloud.

[0280] Specific example

[0281] For example, assume that a user stands in front of a famous spot at a tourist destination and captures an image by turning the camera of the terminal towards it. This image data is sent to the server, and the server identifies the famous spot at the tourist destination and generates explanatory data such as "This famous spot was built in the [century] and is characterized by the following architectural style." When the user inputs a question such as "I want to know more about the history of this famous spot," the server uses the generative AI model to generate an answer such as "The following events had an impact on the history of this famous spot." Also, when the user inputs an impression such as "The design of this building is wonderful," the server generates feedback such as "Many experts also highly evaluate the design."

[0282] Examples of prompt sentences

[0283] Examples of prompt sentences for the generative AI model regarding exhibits include the following:

[0284] "Please tell me about the history of this exhibit."

[0285] "Please provide detailed information about this product."

[0286] In this way, the user can obtain detailed information about the exhibit or product and can also receive an interactive response to the question. As a result, the user's understanding and experience are deepened, and a more abundant information provision is realized.

[0287] ​​​​​​​​​​​​​​​​​​​ Data processing: The terminal's camera takes a picture and saves it as an image file in the internal memory.

[0293] Output: Image data

[0294] Step 2:

[0295] The terminal sends the captured image data to the server.

[0296] Specific operations:

[0297] Input: Image data saved in the terminal

[0298] Data processing: The image data is sent to the server using an HTTP request (e.g., POST request).

[0299] Output: Image data sent to the server

[0300] Step 3:

[0301] The server analyzes the received image data and generates an identifier for identifying the exhibit or product.

[0302] Specific operations:

[0303] Input: Image data sent to the server

[0304] Data processing: Use an image recognition algorithm to identify the exhibit or product in the image and generate its identifier.

[0305] Output: Generated identifier

[0306] Step 4:

[0307] The server retrieves the explanatory data from the database based on the identifier and converts it into a user-friendly format using natural language generation technology.

[0308] Specific operations:

[0309] Input: Generated identifier

[0310] Data processing: Query the database based on identifiers to retrieve relevant descriptive data. Then, use natural language generation techniques (e.g., NLG models) to generate text.

[0311] Output: Converted explanatory data

[0312] Step 5:

[0313] The terminal receives explanatory data sent from the server and displays or plays it as audio for the user.

[0314] Specific operation:

[0315] Input: Converted explanatory data

[0316] Data processing: Display explanatory data on the device's screen, or play it back as audio using speech synthesis technology.

[0317] Output: Displayed or audio-played commentary

[0318] Step 6:

[0319] Users have questions about the explanation and ask them using voice input or text input on their device.

[0320] Specific operation:

[0321] Input: User voice input or text input

[0322] Data processing: If using speech recognition technology, convert the audio to text (in the case of voice input). Then save the question as text data.

[0323] Output: Question data

[0324] Step 7:

[0325] The terminal sends the entered question data to the server.

[0326] Specific operation:

[0327] Input: Entered question data

[0328] Data processing: Send the question data to the server using an HTTP request.

[0329] Output: Question data sent to the server

[0330] Step 8:

[0331] The server analyzes the received question data and uses a generative AI model to generate appropriate answers.

[0332] Specific operation:

[0333] Input: Question data sent to the server

[0334] Data processing: Generative AI models (e.g., GPT-3) are used to generate appropriate answers to questions.

[0335] Output: Generated response data

[0336] Step 9:

[0337] The terminal receives the response data sent from the server and displays it to the user or plays it as audio.

[0338] Specific operation:

[0339] Input: Generated response data

[0340] Data processing: Display the response data on the device's screen, or play it back as audio using speech synthesis technology.

[0341] Output: Displayed or audio-played response

[0342] Step 10:

[0343] Users can input their thoughts and opinions about exhibits and products, providing feedback via voice or text input on their devices.

[0344] Specific operation:

[0345] Input: User voice input or text input

[0346] Data processing: Converting speech to text using speech recognition technology (in the case of voice input). Then, saving the comments as text data.

[0347] Output: Feedback data

[0348] Step 11:

[0349] The terminal sends the entered feedback data to the server.

[0350] Specific operation:

[0351] Input: Entered feedback data

[0352] Data processing: Send feedback data to the server using an HTTP request.

[0353] Output: Feedback data sent to the server

[0354] Step 12:

[0355] The server collects and organizes the feedback data it receives, stores it in a database, and uses natural language generation technology to generate appropriate feedback.

[0356] Specific operation:

[0357] Input: Feedback data sent to the server

[0358] Data processing: Save the feedback data to a database and generate feedback using natural language generation technology (e.g., NLG model).

[0359] Output: Generated feedback data

[0360] Step 13:

[0361] The device receives feedback data sent from the server and displays or plays it aloud for the user.

[0362] Specific operation:

[0363] Input: Generated feedback data

[0364] Data processing: Display feedback data on the device's screen or play it back as audio using speech synthesis technology.

[0365] Output: Displayed or audible feedback

[0366] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0367] ---

[0368] This invention relates to a system that uses generative AI to provide explanations of exhibits and respond to user questions and comments in art museums and museums. Furthermore, by combining this with an emotion engine that recognizes user emotions, the aim is to provide users with personalized explanations and feedback.

[0369] System Overview

[0370] Recognition and explanation of exhibits

[0371] When a user stands in front of an exhibit and launches an app on their device (e.g., a smartphone or tablet), the device's camera is used to capture an image of the exhibit. The captured image data is sent from the device to a server. The server analyzes the received image data and generates an ID to identify the exhibit. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent to the device and displayed or played aloud for the user.

[0372] User Q&A

[0373] If a user has a question about an exhibit, they input their question via voice or text into the terminal. The terminal sends the entered question data to a server, which uses generative AI to generate an appropriate answer to the question. The generated answer is sent to the terminal and displayed or played back to the user.

[0374] User feedback and comments

[0375] When a user inputs their thoughts and opinions into their device via voice or text, the device sends that data to a server. The server analyzes the received data, stores it in a database, and generates feedback based on that data using natural language generation technology. The generated feedback is sent to the device and displayed or played back to the user.

[0376] Enhanced entertainment value

[0377] Furthermore, this system retrieves character information set for each exhibit from a server and reflects the character's tone of voice and personality when generating explanatory texts and answers based on that information. The generated explanatory texts and answers are expressed using character avatars and voice clips, providing users with both visual and auditory experiences.

[0378] Combination of emotional engines

[0379] A distinctive feature of this invention is the integration of an emotion engine. The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. For example, using the user's camera video and audio data, the emotion engine detects emotions such as joy, interest, and dissatisfaction. The terminal sends this emotional data to a server, which analyzes the data to generate explanations, answers, or feedback that correspond to the user's emotions.

[0380] For example, if the emotion engine recognizes that a user is looking at an exhibit with interest, the server generates and sends to the user's device an explanation containing detailed information and interesting anecdotes that will pique the user's interest. On the other hand, if the system recognizes that the user is somewhat tired, it provides a concise and to-the-point explanation. This allows for a more personalized and satisfying user experience.

[0381] Specific example

[0382] For example, suppose a user stands in front of a "famous painting" and launches an app on their device. When the user points their camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses image recognition technology to identify the "famous painting" and generates an ID for it. Based on this ID, the server retrieves explanatory data from its database and generates an explanatory text, such as "This painting was painted in [era], and its characteristics are [characteristics]," and sends it to the device. The device then displays or plays the explanatory text for the user.

[0383] Next, when the user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses generative AI to generate an answer to the question and sends the answer, such as "The following events influenced the creation of this painting," to the device. The device then displays or plays the answer aloud to the user.

[0384] Furthermore, when a user enters a comment such as "The use of color in this painting is wonderful," the device sends the comment data to the server. The server stores the comment data in a database and generates feedback such as "Many experts have praised the use of color. We are pleased that you noticed this point," and sends it to the device. In this way, the user can receive feedback on their own comments.

[0385] Furthermore, by using the emotion engine, if a user says, "This painting is very moving," that emotion can be recognized, and specific feedback such as, "Thank you for being moved. It is said that the artist painted this picture with the same feelings," can be received.

[0386] The above is an example of implementing the system according to the present invention. In this way, it is possible to realize a more personalized user experience by using an emotion engine, in addition to providing explanations for exhibits.

[0387] ---

[0388] The following describes the processing flow.

[0389] ---

[0390] Program processing flow

[0391] Recognition processing of exhibits

[0392] Step 1:

[0393] The user stands in front of the exhibit and launches the app on their device (e.g., a smartphone or tablet).

[0394] Step 2:

[0395] The user uses the device's camera to capture an image of the exhibit. The device activates the camera, generates image data, and saves that image data to memory.

[0396] Step 3:

[0397] The device sends the captured image data to the server. Specifically, it sends the image data to the server via an HTTP request.

[0398] Step 4:

[0399] The server executes an image recognition algorithm to identify the received image data. For example, the server uses a Convolutional Neural Network (CNN) model to detect exhibits from the image and generate a corresponding exhibit ID.

[0400] Processing to retrieve explanations for exhibits

[0401] Step 5:

[0402] The server uses the generated exhibit ID to retrieve the corresponding explanatory data from the explanatory database. This involves extracting explanatory information from the database using SQL queries or similar methods.

[0403] Step 6:

[0404] The server uses natural language generation (NLG) technology to convert the acquired explanatory data into a user-friendly format. Specifically, the NLG engine generates the explanatory text.

[0405] Step 7:

[0406] The server sends the generated explanatory data to the terminal. It returns data containing the explanatory text to the terminal as an HTTP response.

[0407] Step 8:

[0408] The device displays the received explanatory data to the user. This is done using a text display UI or a speech synthesis engine to provide the user with audio explanations.

[0409] User question answering process

[0410] Step 9:

[0411] If a user has a question about an exhibit, they can input their question into the terminal either by voice or text. The terminal uses a speech recognition engine to convert the voice to text or accepts the text input.

[0412] Step 10:

[0413] The terminal sends question data in text format to the server. An HTTP request is used for this transmission.

[0414] Step 11:

[0415] The server analyzes the received question data. Natural language understanding (NLU) technology is used for the analysis to understand the user's question.

[0416] Step 12:

[0417] The server uses a generative AI (e.g., GPT-4) to generate answers to questions. The generated answers are structured as natural language text.

[0418] Step 13:

[0419] The server sends the generated response to the terminal. The response data is returned to the terminal as an HTTP response.

[0420] Step 14:

[0421] The device displays the answer to the user. This display uses a text display UI or a speech synthesis engine.

[0422] User feedback sharing process

[0423] Step 15:

[0424] Users input their impressions and opinions about the exhibits into the device via voice or text. The device uses a speech recognition engine to convert the voice to text or receives the text input.

[0425] Step 16:

[0426] The device sends feedback data in text format to the server. HTTP requests are used for transmission.

[0427] Step 17:

[0428] The server analyzes and organizes the received feedback data and saves it to a database. SQL queries are used to write the data to the database for saving.

[0429] Step 18:

[0430] The server generates feedback using natural language generation technology based on the feedback data. The NLG engine generates appropriate feedback sentences.

[0431] Step 19:

[0432] The server sends the generated feedback to the terminal. The feedback data is returned to the terminal as an HTTP response.

[0433] Step 20:

[0434] The device displays feedback to the user. This is done using a text display UI or a speech synthesis engine.

[0435] Processing to enhance entertainment value

[0436] Step 21:

[0437] The server retrieves character information set for each exhibit. It uses SQL queries to retrieve character information from the database.

[0438] Step 22:

[0439] When the server generates explanatory text and response text based on character information, it reflects the character's speech pattern and personality. An NLG engine is used for generation.

[0440] Step 23:

[0441] The server transmits the character's voice data and avatar information to the terminal. This includes synthesized voice data and image / video data.

[0442] Step 24:

[0443] The device provides explanations and answers to the user using character avatars and voice recordings. The device displays character avatars and plays explanations and answers via voice recordings.

[0444] Combination of emotional engines

[0445] Step 25:

[0446] The device uses its built-in camera and microphone to analyze the user's voice and facial expressions as they view exhibits and express their opinions. The device runs an emotion engine to analyze voice tone and facial expressions.

[0447] Step 26:

[0448] The device converts the analyzed emotion data into text format and sends it to the server. An HTTP request is used for transmission.

[0449] Step 27:

[0450] The server analyzes the received emotional data and generates explanations, responses, or feedback tailored to the user's emotional state. For example, if the user shows interest, it provides a detailed explanation; if the user is tired, it provides a concise explanation.

[0451] Step 28:

[0452] The server sends the generated emotion response explanation, answer, or feedback to the device. It returns the data as an HTTP response.

[0453] Step 29:

[0454] The device displays or plays the generated emotion-responsive data to the user. Text display UI or speech synthesis engines are used for display.

[0455] ---

[0456] (Example 2)

[0457] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0458] Traditional museum and art gallery exhibit explanation systems have been limited to providing static information about the exhibits, failing to address the individual interests and emotions of users. As a result, the user experience was not personalized, resulting in one-sided information delivery and low satisfaction. Furthermore, there was a lack of means to provide appropriate feedback to user questions and comments, often leading to a superficial user experience. In addition to these problems, the lack of entertainment value also needs improvement.

[0459] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing image data transmitted from a terminal to recognize an exhibit and generating an ID corresponding to the identified exhibit; means for converting explanatory data obtained from a database based on the identified exhibit ID into a user-friendly format using natural language generation technology; means for transmitting the generated explanatory data to the user's terminal; means for analyzing questions from the user and generating appropriate answers using a generation AI model; means for analyzing the user's facial expressions and voice tone using an emotion engine to recognize the user's emotions; and means for generating explanations, answers, or feedback that correspond to the user's emotions based on the analyzed emotion data. This makes it possible to provide users with personalized explanations and feedback, thereby achieving higher satisfaction.

[0460] "Exhibits" refer to visual or physical objects such as paintings, sculptures, crafts, and historical artifacts displayed in art museums and museums.

[0461] "Device" refers to a smartphone, tablet, or other portable device used by a user for operation.

[0462] A "camera" is a photographic device attached to a terminal for capturing images of exhibits.

[0463] A "server" refers to a computer system that performs processing such as analyzing received data and generating responses using generative AI.

[0464] "Image data" refers to photographic and video data of exhibits captured using the device's camera.

[0465] An "ID" is a unique identifier generated to identify an exhibit.

[0466] A "database" is a storage area where various types of data, such as explanatory data, character information, and user reviews, are stored.

[0467] "Explanatory data" refers to text data and audio data that include information and explanations related to the exhibits.

[0468] "Natural Language Generation (NLG)" refers to the technology used to convert machine-generated data into human language.

[0469] A "generative AI model" is an artificial intelligence model used to generate appropriate answers or feedback to questions.

[0470] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice tone to detect their emotional state.

[0471] A "prompt message" is text data input to a generative AI model, and it is an instruction message that prompts the model to generate an answer.

[0472] An "avatar" refers to a visual representation generated based on character information, which is typically displayed on the user's device.

[0473] This invention relates to a system that uses generative AI to provide explanations of exhibits and respond to user questions and comments in art museums and museums. Furthermore, by combining this with an emotion engine that recognizes user emotions, the aim is to provide users with personalized explanations and feedback.

[0474] System Overview

[0475] This system consists of the following main elements:

[0476] Recognition and explanation of exhibits

[0477] The user stands in front of an exhibit and uses a smartphone or tablet to capture an image of the exhibit with its camera. The device then sends the captured image data to a server. The server analyzes the received image using image recognition technology (e.g., TENSORFLOW® or OpenCV) and generates an ID to identify the exhibit. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent to the device and displayed or played aloud for the user.

[0478] Specific example

[0479] For example, if a user stands in front of a "famous painting" and points their device's camera at it to capture an image, the device sends the image to a server. The server uses OpenCV to extract features from the image and identify the specific exhibit. Based on this ID, it generates an explanatory text from its database, such as "This painting was painted by [era]," and sends it to the device. The device then plays the explanatory text aloud for the user to listen to.

[0480] User Q&A

[0481] If a user wants to ask a question about an exhibit, they can input it via voice or text into the terminal. The terminal sends the entered question data to a server, which uses a generative AI model (e.g., GPT-3 or GPT-4) to generate an appropriate answer to the question. The generated answer is sent to the terminal and displayed or played back to the user.

[0482] Specific example

[0483] When a user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses a generative AI model to generate an answer, such as "The following events influenced the creation of this painting," and sends it to the device. The device then displays the answer for the user to see.

[0484] User feedback and comments

[0485] When a user inputs their thoughts and opinions into their device via voice or text, the device sends that data to a server. The server analyzes the received data, stores it in a database, and generates feedback based on that data. The generated feedback is sent to the device and displayed or played back to the user.

[0486] Specific example

[0487] When a user enters a comment such as "The colors in this painting are wonderful," the device sends the comment data to the server. The server saves the comment data and generates feedback such as "Many experts have praised this. We are glad you noticed this point," and sends it to the device. The device then plays the feedback text aloud for the user to hear.

[0488] Combination of emotional engines

[0489] The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. Using the user's camera video and audio data, the emotion engine detects emotions such as joy, interest, and dissatisfaction. The device sends this emotional data to a server, which generates explanations, answers, or feedback tailored to the user's emotions based on the analysis results.

[0490] Specific example

[0491] If a user says, "This painting is very moving," the emotion engine recognizes that emotion and generates specific feedback such as, "Thank you for being moved. It is said that the artist painted this picture with the same feelings," and sends it to the device. The device then plays the feedback aloud and delivers it to the user.

[0492] Example of a prompt

[0493] 1. Prompt for generating detailed descriptions of exhibits

[0494] Please generate detailed descriptions for the following exhibits.

[0495] Exhibit ID: ABC123

[0496] Era: Ancient

[0497] Author: Author A

[0498] Features: Vibrant brushwork

[0499] 2. Question and answer prompt sentences

[0500] Please generate answers to the following questions.

[0501] Question: What is the background behind the creation of this painting?

[0502] Exhibit ID: ABC123"

[0503] 3. Prompt for feedback / comments

[0504] Please generate feedback for the following comments.

[0505] Impression: The use of color in this painting is wonderful.

[0506] Exhibit ID: ABC123"

[0507] 4. Feedback prompts based on emotion recognition

[0508] Please generate feedback based on the following emotions:

[0509] Emotion: I'm moved.

[0510] Comment: This painting is very moving.

[0511] Exhibit ID: ABC123"

[0512] In this way, the present invention makes it possible to provide personalized explanations of exhibits in art museums and museums, and to offer an immersive experience that responds to the user's emotions.

[0513] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0514] System program processing flow

[0515] Recognition and explanation of exhibits

[0516] Step 1:

[0517] The user stands in front of the exhibit.

[0518] Input: User's location, exhibits

[0519] Output: Captured camera connection

[0520] Specific actions: The user points their smartphone or tablet camera at the exhibit and launches a dedicated app.

[0521] Step 2:

[0522] The device captures images of the exhibits.

[0523] Input: Exhibit images, camera

[0524] Output: Captured image data

[0525] Specific operation: The device's camera takes a picture of the exhibit and generates image data.

[0526] Step 3:

[0527] The device sends the captured image data to the server.

[0528] Input: Captured image data

[0529] Output: Data sent to the server after completion

[0530] Specific operation: The terminal sends image data to the server using an open communication protocol.

[0531] Step 4:

[0532] The server analyzes the image data.

[0533] Input: Received image data

[0534] Output: Exhibit identification ID

[0535] Specific operation: The server analyzes image data using TensorFlow and OpenCV to extract and identify features of the exhibits.

[0536] Step 5:

[0537] The server retrieves and generates explanatory data from the database based on the identified exhibit ID.

[0538] Input: Exhibit identification ID

[0539] Output: Explanatory data, generated explanatory text

[0540] Specific operation: The server retrieves relevant explanatory data from the explanatory database and uses natural language generation (NLG) technology to generate an easy-to-understand explanatory text for the user.

[0541] Step 6:

[0542] The server sends the generated explanatory text to the terminal.

[0543] Input: Generated explanatory text

[0544] Output: Transmission complete data

[0545] Specific operation: The server sends the generated explanatory text to the terminal.

[0546] Step 7:

[0547] The device displays or plays explanatory text.

[0548] Input: Received explanatory text

[0549] Output: Displayed or audio-played explanatory text

[0550] Specific operation: The device will either display the received explanatory text on the screen or play it as audio.

[0551] User Q&A

[0552] Step 1:

[0553] The user enters a question.

[0554] Input: Questions via voice or text input

[0555] Output: Question data

[0556] Specific operation: The user enters the question using voice input or text input on the device.

[0557] Step 2:

[0558] The device sends the question data to the server.

[0559] Input: Question data

[0560] Output: Data sent to the server after completion

[0561] Specific action: The terminal sends the question data to the server.

[0562] Step 3:

[0563] The server analyzes the question data and generates answers.

[0564] Input: Question data

[0565] Output: Generated answer

[0566] Specific operation: The server analyzes the question data and generates appropriate answers using a generative AI model (e.g., GPT-3 or GPT-4).

[0567] Step 4:

[0568] The server sends the generated response to the terminal.

[0569] Input: Generated answer

[0570] Output: Transmission complete data

[0571] Specific operation: The server sends the generated response to the terminal.

[0572] Step 5:

[0573] The device displays or plays the answer.

[0574] Input: Received response

[0575] Output: Displayed or audio-played answer

[0576] Specific actions: The device will either display the received response on the screen or play it as audio.

[0577] User feedback and comments

[0578] Step 1:

[0579] Users enter their feedback.

[0580] Input: Feedback via voice or text input

[0581] Output: Feedback data

[0582] Specific operation: The user enters their feedback by using voice input or text input on the device.

[0583] Step 2:

[0584] The device sends feedback data to the server.

[0585] Input: Feedback data

[0586] Output: Data sent to the server after completion

[0587] Specific action: The device sends feedback data to the server.

[0588] Step 3:

[0589] The server analyzes and stores the feedback data.

[0590] Input: Received feedback data

[0591] Output: Saved data, analyzed feedback data

[0592] Specific operation: The server analyzes the feedback data and saves it to the database.

[0593] Step 4:

[0594] The server generates feedback based on the feedback data.

[0595] Input: Analyzed feedback data

[0596] Output: Generated feedback

[0597] Specific operation: The server uses natural language generation technology to generate feedback based on the feedback data.

[0598] Step 5:

[0599] The server sends the generated feedback to the terminal.

[0600] Input: Generated feedback

[0601] Output: Transmission complete data

[0602] Specific operation: The server sends the generated feedback to the terminal.

[0603] Step 6:

[0604] The device displays or plays feedback.

[0605] Input: Received feedback

[0606] Output: Displayed or audible feedback

[0607] Specific actions: The device displays the received feedback on the screen or plays it back as audio.

[0608] Combination of emotional engines

[0609] Step 1:

[0610] Capture user sentiment data

[0611] Input: User's facial expressions and voice

[0612] Output: Captured sentiment data

[0613] Specific operation: The device's camera and microphone capture the user's facial expressions and voice.

[0614] Step 2:

[0615] The device sends emotional data to the server.

[0616] Input: Captured emotion data

[0617] Output: Data sent to the server after completion

[0618] Specific operation: The device sends the captured emotion data to the server.

[0619] Step 3:

[0620] The server analyzes the emotional data.

[0621] Input: Received sentiment data

[0622] Output: Analyzed emotional state

[0623] Specific operation: The server analyzes the emotional state using an emotion engine (e.g., Affectiva SDK or Microsoft® Azure® Emotion API).

[0624] Step 4:

[0625] The server generates content based on emotions.

[0626] Input: Analyzed emotional state

[0627] Output: Generated emotion-responsive content

[0628] Specific operation: Based on the analyzed sentiment data, the server generates explanations, responses, or feedback that correspond to the user's emotions.

[0629] Step 5:

[0630] The server sends the generated content to the terminal.

[0631] Input: Generated emotion-responsive content

[0632] Output: Transmission complete data

[0633] Specific operation: The server sends the generated emotion-responsive content to the device.

[0634] Step 6:

[0635] The device displays content or plays audio.

[0636] Input: Received emotion-responsive content

[0637] Output: Displayed or played content

[0638] Specific actions: The device will either display the received content on the screen or play it as audio.

[0639] The above outlines the specific flow of the system's program processing. This allows users to receive personalized explanations and feedback on the exhibits, resulting in a more satisfying experience.

[0640] (Application Example 2)

[0641] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0642] Traditional systems for providing explanations of exhibits and products could only offer the same information at all times, failing to provide an experience tailored to the individual interests and emotions of users. Furthermore, when users asked detailed questions about specific products or exhibits, they often sought more comprehensive information, but the responses were limited to general and consistent answers. As a result, user satisfaction decreased, and problems arose where purchasing intent and learning motivation were not sufficiently stimulated.

[0643] The identification process performed by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing received image data and generating an ID corresponding to the identified exhibit; means for obtaining explanatory data from a database based on the identified exhibit ID and converting the explanatory data into a user-friendly format using natural language generation technology; and means for generating an appropriate response using a generative AI. This makes it possible not only to provide users with detailed information about exhibits and products, but also to provide personalized explanations and feedback that are tailored to the user's emotional state.

[0644] The term "exhibits" refers to all objects that customers are interested in, not only in art museums and museums, but also in physical stores.

[0645] "Terminal" refers to a smartphone, tablet, personal computer, or other portable information processing device.

[0646] A "camera" refers to an optical device used to capture images and videos.

[0647] "Image data" refers to a digital representation of a captured image.

[0648] A "server" refers to a computer system used to receive, process, and transmit data over a network.

[0649] "ID" refers to a code or label used to uniquely identify a specific exhibit or product.

[0650] A "database" refers to a system or software used to organize and store information.

[0651] "Natural language generation technology" refers to the technology that uses computers to generate natural-sounding words and sentences.

[0652] "Generative AI" refers to artificial intelligence technology that generates appropriate responses or new data based on input data.

[0653] "Voice input" refers to a method of recognizing human speech as digital information using a microphone.

[0654] "Text input" refers to the method of entering character data using a keyboard or on-screen keyboard.

[0655] "User" refers to all individuals who use the system.

[0656] An "emotion engine" refers to software or hardware that analyzes a user's facial expressions and tone of voice to detect their psychological state.

[0657] "Feedback" refers to the response or reaction that a system gives to a user's actions or input.

[0658] "Character information" refers to digital data or profiles used to add a specific personality or tone of voice to explanatory texts or answer texts.

[0659] This invention relates to a system that provides customers with detailed information about exhibits and products in physical stores, and further provides personalized explanations and feedback that respond to the customer's emotions. Specific embodiments of this invention are shown below.

[0660] System Overview

[0661] This system includes the following main components:

[0662] terminal

[0663] A terminal is a portable information processing device such as a smartphone or tablet. The terminal is equipped with a camera, microphone, display, and speaker. This allows users to photograph products and exhibits and interact with the system through voice or text input.

[0664] server

[0665] The server communicates with terminals over the network and processes the received data. The server includes multiple modules for image recognition, natural language generation technology, and generative AI processing.

[0666] Image capture and recognition

[0667] The user uses the device's camera to capture images of products or exhibits. The captured image data is sent from the device to a server, where the server's image recognition module analyzes it. An ID corresponding to the identified exhibit is generated from the analysis results. This ID is used to retrieve explanatory data from the database.

[0668] Question and Answer

[0669] When a user sends a question about an exhibit to their device via voice or text input, the question data is transferred to a server. The server uses generative AI to generate an appropriate answer and sends it to the device. The device then displays or plays this answer aloud.

[0670] Emotional engine and feedback

[0671] Furthermore, this system is equipped with an emotion engine that analyzes the user's emotions using camera and audio data. Once the user's emotional state is recognized, that emotional data is sent to the server. The server generates explanations and feedback based on the emotional data, providing a personalized experience. This feedback is generated based on information obtained from the database and the user's emotional state.

[0672] Specific example

[0673] For example, suppose a user is in a physical store and stands before the latest smartphone, then launches an app on the device. The user takes a picture of the smartphone using the camera and sends the image data to a server. The server uses image recognition technology to identify the smartphone and generates an ID. Based on this ID, it retrieves a detailed explanation from a database and uses natural language generation technology to generate an easy-to-understand explanatory text for the user. This explanatory text is sent to the device and is either displayed on the screen or played as audio.

[0674] When a user asks a question by voice, such as "How is the camera function of this smartphone?", the device sends the question data to a server. The server uses generative AI to generate a detailed answer and sends it to the device, such as "This smartphone's camera has a 12-megapixel sensor and excellent night vision capabilities." The device then displays or plays this answer aloud to the user.

[0675] Furthermore, the emotion engine, which analyzes the user's facial expressions and tone of voice, recognizes when the user is excited. For example, when a user says, "This smartphone is amazing!", based on that emotion data, it generates feedback such as, "I'm glad you're pleased! Shall I tell you about other features?"

[0676] Example of a prompt

[0677] Examples of prompt statements are as follows:

[0678] User emotion: Joy

[0679] Question: What features does this washing machine have?

[0680] answer:

[0681] This prompt is input into a generative AI model, which generates a response tailored to the user's emotion of "joy."

[0682] The above describes a specific embodiment of the system according to the present invention. In this way, it becomes possible to provide personalized explanations and feedback that respond to the user's emotions.

[0683] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0684] Step 1:

[0685] The user captures images of products or exhibits using the device's camera. The input is video from the device's camera, and the output is the captured image data. This image data is temporarily stored on the device.

[0686] Step 2:

[0687] The terminal sends the captured image data to the server. Specifically, it uploads the image data to the server via the network connection by including it in the request packet. The captured image data is the input, and the server receives the data as output.

[0688] Step 3:

[0689] The server analyzes the received image data and uses image recognition technology to identify specific exhibits or products. It retrieves the ID of the identified object from the database and generates a dedicated code or label. Image data is the input, and the identified ID is generated as the output.

[0690] Step 4:

[0691] The server retrieves detailed explanatory data from the database based on the identified ID. Next, it uses natural language generation technology to convert the explanatory data into a user-friendly format. The input consists of the identified ID and associated explanatory data, and the output is the converted explanatory text.

[0692] Step 5:

[0693] The server sends the converted explanatory data to the terminal. Specifically, it sends packet data containing the generated explanatory text to the terminal via the network. The input is the explanatory text converted into an easy-to-understand format, and the output is the explanatory text sent to the terminal.

[0694] Step 6:

[0695] The terminal displays or plays explanatory text for the user. The input is explanatory text received from the server, and the output is either text displayed on the screen or audio played through the speaker.

[0696] Step 7:

[0697] Users submit questions about exhibits via voice or text input. The input is the user's voice or text, and the output is the voice recognition or text input data, which is stored on the device.

[0698] Step 8:

[0699] The terminal sends the entered question data to the server. Specifically, it uploads the question data to the server using a network connection. The user's question data is the input, and the server receives the data as output.

[0700] Step 9:

[0701] The server analyzes the received question data and generates appropriate answers using generative AI. The input is the question data, and the output is the generated answers.

[0702] Step 10:

[0703] The server sends the generated response to the terminal. Specifically, it sends packet data containing the generated response to the terminal over the network. The generated response is the input, and the response sent to the terminal is the output.

[0704] Step 11:

[0705] The device displays or plays the answer to the user. The input is the answer received from the server, and the output is either text displayed on the screen or audio played through the speaker.

[0706] Step 12:

[0707] To recognize the user's emotions, the device's camera and microphone are used to capture facial expressions and voice tone. Camera video and audio data are taken as input, and emotion data is generated as output.

[0708] Step 13:

[0709] The device sends emotional data to the server. Specifically, it sends packets containing emotional data to the server over the network. The emotional data is the input, and the server receives the data as the output.

[0710] Step 14:

[0711] The server analyzes emotional data and generates explanations and feedback tailored to the user's emotions. It takes emotional data and related explanation data as input, and generates personalized feedback as output.

[0712] Step 15:

[0713] The server sends personalized feedback to the terminal. Specifically, it sends packet data containing the generated feedback to the terminal over the network. The input is the personalized feedback, and the output is the feedback sent to the terminal.

[0714] Step 16:

[0715] The device displays or plays audio feedback to the user. Input is feedback received from a server, and output is text displayed on the screen or audio played through the speaker.

[0716] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0717] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0718] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0719] [Second Embodiment]

[0720] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0721] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0722] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0723] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0724] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0725] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0726] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0727] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0728] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0729] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0730] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0731] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0732] ---

[0733] This invention relates to a system that uses generative AI to provide explanations to users about exhibits in art museums and museums, and to respond to user questions and comments. This system comprehensively performs exhibit recognition, acquisition and display of explanatory data, response to user questions, and collection and feedback of comments.

[0734] System Overview

[0735] First, when a user stands in front of an exhibit, they use a device (such as a smartphone or tablet) to capture an image of the exhibit. The device generates image data using its camera and sends this data to a server. The server uses image recognition technology to identify the exhibit and generates an ID. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent from the server to the device and displayed or played as audio for the user on the device.

[0736] If a user has a question about an exhibit, they enter the question into the terminal via voice or text input. The terminal sends the entered question data to a server, which uses generative AI to generate an appropriate answer to the question. This answer is then sent to the terminal and displayed or played back to the user.

[0737] Furthermore, when users want to input their thoughts and opinions, they do so via voice or text on the device. The device sends the input data to the server, which stores the received data in a database and generates appropriate feedback, which is then sent back to the device. This allows users to receive feedback on their own comments.

[0738] The system also includes features to enhance entertainment value. The server retrieves character information set for each exhibit and reflects the character's tone and personality in the explanations and answers. The generated explanations and answers are played back using the character's avatar and voice, providing users with both visual and auditory experiences.

[0739] Specific example

[0740] For example, suppose a user stands in front of a "famous painting" and launches an app on their device. When the user points their camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses image recognition technology to identify the "famous painting" and generates an ID. Using this ID, the server retrieves explanatory data from a database, generates an explanatory text such as "This painting was painted in [era], and its characteristics are [characteristics]," and sends it to the device. The device then displays or plays the explanatory text for the user.

[0741] Next, when the user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses generative AI to generate an answer to the question and sends the answer, such as "The following events influenced the creation of this painting," to the device. The device then displays or plays the answer aloud to the user.

[0742] Furthermore, when a user enters a comment such as "The use of color in this painting is wonderful," the device sends the comment data to the server. The server stores the comment data in a database and generates feedback such as, "Many experts have praised the use of color. We are pleased that you noticed this point," and sends it back to the device. This allows the user to gain a deeper understanding and greater satisfaction.

[0743] The above is an example of implementing the system according to the present invention. This makes it possible to provide multifaceted and individualized information about exhibits, thereby enhancing the user experience.

[0744] ---

[0745] The following describes the processing flow.

[0746] ---

[0747] Program processing flow

[0748] Recognition processing of exhibits

[0749] Step 1:

[0750] The user stands in front of the exhibit and launches the app on their device (e.g., a smartphone or tablet).

[0751] Step 2:

[0752] The device's camera is used to capture images of the exhibits. The device activates the camera, generates image data, and saves that image data to memory.

[0753] Step 3:

[0754] The device sends the captured image data to the server. Specifically, it sends the image data to the server via an HTTP request.

[0755] Step 4:

[0756] The server executes an image recognition algorithm to identify the received image data. For example, the server uses a Convolutional Neural Network (CNN) model to detect exhibits from the image and generate a corresponding exhibit ID.

[0757] Processing to retrieve explanations for exhibits

[0758] Step 5:

[0759] The server uses the generated exhibit ID to retrieve the corresponding explanatory data from the explanatory database. This involves extracting explanatory information from the database using SQL queries or similar methods.

[0760] Step 6:

[0761] The server uses natural language generation (NLG) technology to convert the acquired explanatory data into a user-friendly format. Specifically, the NLG engine generates the explanatory text.

[0762] Step 7:

[0763] The server sends the generated explanatory data to the terminal. It returns data containing the explanatory text to the terminal as an HTTP response.

[0764] Step 8:

[0765] The device displays the received explanatory data to the user. This is done using a text display UI or a speech synthesis engine to provide the user with audio explanations.

[0766] User question answering process

[0767] Step 9:

[0768] If a user has a question about an exhibit, they can input their question into the terminal either by voice or text. The terminal uses a speech recognition engine to convert the voice to text or accepts the text input.

[0769] Step 10:

[0770] The terminal sends question data in text format to the server. An HTTP request is used for this transmission.

[0771] Step 11:

[0772] The server analyzes the received question data. Natural language understanding (NLU) technology is used for the analysis to understand the user's question.

[0773] Step 12:

[0774] The server uses a generative AI (e.g., GPT-4) to generate answers to questions. The generated answers are structured as natural language text.

[0775] Step 13:

[0776] The server sends the generated response to the terminal. The response data is returned to the terminal as an HTTP response.

[0777] Step 14:

[0778] The device displays the answer to the user. This display uses a text display UI or a speech synthesis engine.

[0779] User feedback sharing process

[0780] Step 15:

[0781] Users input their impressions and opinions about the exhibits into the device via voice or text. The device uses a speech recognition engine to convert the voice to text or receives the text input.

[0782] Step 16:

[0783] The device sends feedback data in text format to the server. HTTP requests are used for transmission.

[0784] Step 17:

[0785] The server analyzes and organizes the received feedback data and saves it to a database. SQL queries are used to write the data to the database for saving.

[0786] Step 18:

[0787] The server generates feedback using natural language generation technology based on the feedback data. The NLG engine generates appropriate feedback sentences.

[0788] Step 19:

[0789] The server sends the generated feedback to the terminal. The feedback data is returned to the terminal as an HTTP response.

[0790] Step 20:

[0791] The device displays feedback to the user. This is done using a text display UI or a speech synthesis engine.

[0792] Processing to enhance entertainment value

[0793] Step 21:

[0794] The server retrieves character information set for each exhibit. It uses SQL queries to retrieve character information from the database.

[0795] Step 22:

[0796] When the server generates explanatory text and response text based on character information, it reflects the character's speech pattern and personality. An NLG engine is used for generation.

[0797] Step 23:

[0798] The server transmits the character's voice data and avatar information to the terminal. This includes synthesized voice data and image / video data.

[0799] Step 24:

[0800] The device provides explanations and answers to the user using character avatars and voice recordings. The device displays character avatars and plays explanations and answers via voice recordings.

[0801] ---

[0802] The above is the specific processing flow of the system program.

[0803] (Example 1)

[0804] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0805] Museums and art galleries need efficient and effective ways to provide explanations about their exhibits to users. Furthermore, there is a lack of systems that can quickly and appropriately respond to users' questions and comments about the exhibits, thereby providing a richer experience. Traditional systems primarily rely on paper-based descriptions for each exhibit, lacking interactivity and making it difficult to capture user interest. Moreover, the lack of entertainment-oriented exhibit explanations utilizing characters limits the potential for improving the user experience.

[0806] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0807] In this invention, the server includes means for capturing images using a camera device provided on the user terminal to identify exhibits and transmitting the image data to the server; means for analyzing the received image data on the server and generating an identifier corresponding to the identified exhibit; means for obtaining information data from a database based on the identified exhibit identifier on the server and converting the information data into a user-friendly format using natural language generation technology; means for transmitting the converted information data to the user terminal and displaying or playing it aloud on the user terminal; means for inputting user questions into the user terminal via voice or text input; means for transmitting the question data entered on the user terminal to the server; means for analyzing the received question data on the server and generating an appropriate answer using a generation AI model; and means for transmitting the generated answer to the user terminal. The system includes means for displaying or playing audio on the user terminal, means for inputting user feedback and opinions via voice or text into the user terminal, means for sending feedback data entered on the user terminal to a server, means for the server to collect and organize the received feedback data and store it in a database, means for generating appropriate feedback using natural language generation technology based on the feedback data, means for sending the generated feedback to the user terminal and displaying or playing audio on the user terminal, means for obtaining character information set for each exhibit from the server and reflecting the character's tone of voice and personality in the explanations and answers, means for sending data to the user terminal to express the character-based explanations and answers in the character's voice, and means for providing the character to the user as an avatar or voice on the user terminal. This makes it possible to provide users with interactive and rich explanations and responses to exhibits, as well as a highly entertaining experience.

[0808] "Exhibits" refer to works of art and materials that are made available to users in art museums and museums.

[0809] "User terminal" refers to a portable information device owned by a user, primarily such as a smartphone or tablet.

[0810] "Image capture equipment" refers to cameras and other optical instruments used to acquire image data.

[0811] A "server" is a central computing system that processes, stores, and provides data over a network.

[0812] "Image data" refers to data that represents captured visual information in a digital format.

[0813] "Analysis" refers to the process of examining received images and data in detail to extract information.

[0814] An "identifier" is a code or number generated to uniquely identify a particular exhibit.

[0815] A "database" is a system for efficiently storing, searching, and managing large amounts of information.

[0816] "Information data" refers to digital data that includes detailed explanations and information about the exhibits.

[0817] "Natural language generation technology" refers to technology that allows machines to automatically generate sentences that are easy for humans to understand.

[0818] "Question data" refers to data that digitally represents questions about exhibits entered by users.

[0819] A "generative AI model" refers to an algorithm or model that uses artificial intelligence to generate appropriate answers or text.

[0820] "Feedback" refers to the responses or evaluations provided in response to user input.

[0821] "Character information" refers to data that includes specific characteristics and speech patterns of fictional people, animals, and other elements set for each exhibit.

[0822] An "avatar" is an image or model used to visually represent a character.

[0823] "Voice" refers to data used to digitally reproduce the character's voice.

[0824] This invention relates to a system that uses generative AI to provide explanations to users about exhibits in art museums and museums, and to respond to user questions and comments. This system comprehensively performs exhibit recognition, acquisition and display of explanatory data, response to user questions, and collection and feedback of comments.

[0825] System Overview

[0826] This system works as follows:

[0827] First, when a user stands in front of an exhibit, they use their device (e.g., a smartphone or tablet) to capture an image of the exhibit. The device's camera generates image data, which is then sent to a server. The server uses image recognition technology to identify the exhibit and generates an identifier. Based on this identifier, the server retrieves relevant information data from a database and generates an explanatory text using natural language generation (NLG) technology. The generated explanatory text is sent from the server to the device and displayed or played back as audio for the user.

[0828] Hardware and software details

[0829] 1. Terminal (User terminal):

[0830] Hardware used: Smartphone, tablet

[0831] Software used: Dedicated app, camera function, audio recording and playback function

[0832] 2. Server:

[0833] Software used:

[0834] Image recognition technology: AWS Rekognition, Google Cloud Vision

[0835] Database: MySQL, MongoDB

[0836] Natural language generation technology (NLG): OpenAI GPT-4

[0837] Audio playback technology: Amazon Polly, Google Text-to-Speech

[0838] Processing details

[0839] The processing flow at each step is as follows:

[0840] 1. User-captured images of exhibits:

[0841] Users point their device's camera at the exhibit and capture images using a dedicated app.

[0842] 2. Sending image data from the terminal to the server:

[0843] The terminal encodes the captured image data in BASE64 format and sends it to the server using an HTTP POST request.

[0844] 3. Server-based recognition of exhibits and acquisition of data:

[0845] The server analyzes the received image data, using AWS Rekognition or Google Cloud Vision to recognize exhibits and generate identifiers. Based on the identified exhibit identifiers, the server retrieves relevant information data from MySQL or MongoDB databases.

[0846] 4. Server generates explanatory text and sends it to the terminal:

[0847] The server uses OpenAI's GPT-4 to generate natural language based on the acquired data. The generated explanatory text is sent from the server to the terminal and displayed or played back as audio for the user.

[0848] Specific usage scenarios

[0849] For example, suppose a user stands in front of a "famous painting" and launches a smartphone app. When the user points the camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses AWS Rekognition to identify the "famous painting" and generates an identifier. Using this identifier, the server retrieves explanatory data from a MySQL database, generates an explanatory text using GPT-4, and sends it to the device. The device then displays or plays an audio message stating, "This painting was painted by a famous Renaissance painter, and its distinguishing feature is the meticulous detail."

[0850] Examples of prompts for generative AI models

[0851] "Generate a descriptive text for a famous painting. The text should include the artist, date of creation, and key features."

[0852] The above details the embodiments of the present invention. This system allows users to obtain multifaceted information about exhibits and deepen their understanding of the exhibits interactively.

[0853] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0854] Step 1:

[0855] The user stands in front of the exhibit and uses the camera on their device to capture an image of the exhibit.

[0856] Input: Image of the exhibit

[0857] Output: Captured image data

[0858] Specific operation: The user launches a dedicated app on their smartphone or tablet, points the camera at the exhibit, and presses the shutter button. This causes the camera to capture image data and save it to the device's memory.

[0859] Step 2:

[0860] The terminal encodes the captured image data in BASE64 format and sends it to the server using an HTTP POST request.

[0861] Input: Captured image data

[0862] Output: Image data encoded in BASE64 format

[0863] Specific operation: The device's dedicated app encodes the captured image data into BASE64 format and uploads the image data to the server by sending an HTTP POST request to the specified API endpoint.

[0864] Step 3:

[0865] The server analyzes the received image data and generates identifiers corresponding to the identified exhibits.

[0866] Input: Image data in BASE64 format

[0867] Output: Identifier of the identified exhibit

[0868] Specific operation: The server uses image recognition technology (e.g., AWS Rekognition or Google Cloud Vision) to analyze the received image data. From the analysis results, it identifies the exhibits and generates corresponding identifiers.

[0869] Step 4:

[0870] The server retrieves relevant information data from the database based on the identified exhibit identifier.

[0871] Input: Exhibit identifier

[0872] Output: Information data about the exhibits

[0873] Specific operation: The server uses the generated exhibit identifier to query the database (e.g., MySQL or MongoDB) to retrieve information data about the corresponding exhibit.

[0874] Step 5:

[0875] Based on the acquired information data, the server uses a generative AI model (e.g., GPT-4) to perform natural language generation and generate explanatory text.

[0876] Input: Information data about the exhibits

[0877] Output: Generated explanatory text

[0878] Specific operation: The server analyzes the information data and inputs prompt sentences into the generation AI model to generate appropriate explanatory text.

[0879] Example prompt: "Generate a descriptive text for a famous painting. The text should include the artist, date of creation, and key features."

[0880] Step 6:

[0881] The server sends the generated explanatory text to the user's terminal, where it is displayed or played as audio.

[0882] Input: Generated explanatory text

[0883] Output: Explanatory text sent to the user's terminal

[0884] Specific operation: The server sends the generated explanatory text as an HTTP response to the user's terminal. The user's terminal either displays the received explanatory text on the screen or plays it aloud using speech synthesis technology (e.g., Amazon Polly or Google Text-to-Speech).

[0885] Step 7:

[0886] When a user enters a question about an exhibit, the user's terminal sends the question data to the server.

[0887] Input: Question entered by the user (voice or text)

[0888] Output: Question data sent to the server

[0889] Specific operation: The user asks a question via voice input or text input through the app on their device. In the case of voice input, the device converts the voice to text and sends it to the server as question data.

[0890] Step 8:

[0891] The server analyzes the received question data and generates appropriate answers using a generative AI model.

[0892] Input: Question data sent to the server

[0893] Output: Generated answer

[0894] Specific operation: The server analyzes the question data and prompts the generation AI model to generate appropriate answers. The generated answers are saved in text format.

[0895] Example prompt: "Generate answers to the following question: 'What is the background of this painting?'"

[0896] Step 9:

[0897] The server sends the generated response to the user's terminal, where it is displayed or played back as audio.

[0898] Input: Generated answer

[0899] Output: Response sent to the user's terminal

[0900] Specific operation: The server sends the generated response to the user's terminal as an HTTP response. The user's terminal displays the received response on the screen or plays it back as audio using speech synthesis technology.

[0901] Step 10:

[0902] When a user enters their thoughts and opinions about an exhibit, the user's device sends the feedback data to the server.

[0903] Input: User feedback (voice or text)

[0904] Output: Feedback data sent to the server

[0905] Specific operation: Users provide feedback via voice input or text input through the app on their device. In the case of voice input, the device converts the voice to text and sends it to the server as feedback data.

[0906] Step 11:

[0907] The server collects and organizes the received feedback data and stores it in a database.

[0908] Input: Feedback data sent to the server

[0909] Output: Feedback data stored in the database

[0910] Specific operation: The server collects and organizes the received feedback data, and then saves it to a database (e.g., MongoDB).

[0911] Step 12:

[0912] Based on the feedback data, the server uses a generative AI model to generate appropriate feedback.

[0913] Input: Feedback data stored on the server

[0914] Output: Generated feedback

[0915] Specific operation: The server analyzes the stored feedback data and prompts the AI ​​model to generate appropriate feedback. The generated feedback is saved in text format.

[0916] Step 13:

[0917] The server sends the generated feedback to the user's terminal, where it is displayed or played as audio.

[0918] Input: Generated feedback

[0919] Output: Feedback sent to the user terminal

[0920] Specific operation: The server sends the generated feedback to the user terminal as an HTTP response. The user terminal displays the received feedback on the screen or plays it back as audio using speech synthesis technology.

[0921] Step 14:

[0922] The server retrieves character information set for each exhibit and reflects the character's tone of voice and personality in the explanations and answers.

[0923] Input: Exhibit identifier

[0924] Output: Explanatory text and answer text reflecting character information

[0925] Specific operation: The server uses the exhibit identifier to retrieve character information from the database and reflects the character's tone of voice and personality in the explanatory text and response text.

[0926] Step 15:

[0927] The server sends explanatory text and response text based on character information to the user's terminal, and the terminal plays avatars and voice clips.

[0928] Input: Explanatory text or answer text that reflects character information

[0929] Output: Explanatory text and response text with character information sent to the user's terminal.

[0930] Specific operation: The server sends explanatory text and answer text with character information to the user terminal as an HTTP response. The user terminal displays the received content on the screen and plays the character's voice using speech synthesis technology.

[0931] The above outlines the specific processing steps of this system. It details a comprehensive mechanism for enhancing the user experience, from exhibit recognition and explanation to question-and-answer sessions, feedback, and even character entertainment.

[0932] (Application Example 1)

[0933] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0934] Traditional museum and art gallery systems for explaining exhibits are limited to providing information about the exhibits themselves, and do not adequately address users' questions or comments. Similarly, in physical stores and tourist destinations, there is a lack of detailed information about exhibits and products, as well as insufficient interactive communication with users. As a result, users' understanding and experience are limited, and the appeal of exhibits and products cannot be fully conveyed.

[0935] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0936] This invention includes means for a server to capture images using a camera installed in a terminal to recognize exhibits and transmit the image data to the server; means for the server to analyze the received image data and generate an identifier corresponding to the identified exhibit; means for the server to obtain explanatory data from a database based on the identified exhibit identifier and convert the explanatory data into a user-friendly format using natural language generation technology; means for transmitting the converted explanatory data to the user's terminal and displaying or playing it aloud on the terminal; means for inputting user questions into the terminal via voice or text; means for transmitting the question data entered on the terminal to the server; means for the server to analyze the received question data and generate appropriate answers using a generation AI model; means for transmitting the generated answers to the user's terminal and displaying or playing them aloud on the terminal; and means for providing information via the terminal when the user stands in front of a product or landmark. This allows the user to not only obtain detailed information about exhibits and products, but also to resolve questions in an interactive way and receive feedback on their impressions.

[0937] "Exhibits" is a general term for works of art and historical items displayed in art museums and museums, as well as objects that attract the interest of users in physical stores and tourist destinations.

[0938] A "device" refers to a portable electronic device that a user can carry with them, such as a smartphone, tablet, or personal computer, and is equipped with a camera, microphone, and display.

[0939] "Capturing an image using the camera" refers to saving visual information as a photograph to the device's internal memory using the device's built-in camera.

[0940] A "server" is a computer system that receives and processes data sent from a terminal and provides the necessary information.

[0941] "Image data" refers to digital files containing visual information captured by a camera.

[0942] An "identifier" is a code or number used by the server to uniquely identify a specific exhibit, and is used for database operations.

[0943] A "database" is an information management system for systematically storing explanatory data and other related information about exhibits.

[0944] "Natural language generation technology" refers to the technology that allows computers to generate text in human language that is easy to read and understand.

[0945] "User's device" refers to a device such as a smartphone or tablet used by the user, which displays or plays back information received from the server.

[0946] "Question data" refers to digital data that includes the content of questions entered by the user via voice or text.

[0947] A "generative AI model" is an artificial intelligence technology that automatically generates appropriate answers based on input information.

[0948] "Feedback" refers to responses and reactions to user comments and opinions, and is generated using natural language generation technology.

[0949] An "avatar" is an image used to provide users with a visual representation of a character.

[0950] This invention is a system that provides explanations to users about exhibits and products, and also responds to user questions and feedback. A specific embodiment of this system is shown below.

[0951] System Configuration

[0952] This system consists of terminals, a server, and a database. Terminals include smartphones, tablets, and personal computers. The server handles data processing and storage, as well as the execution of generated AI models. The database stores explanatory data about exhibits and products.

[0953] Hardware and software

[0954] Hardware: Smartphones, tablets, and personal computers (these electronic devices are equipped with cameras, microphones, and displays).

[0955] Software: Python, OpenCV (for capturing camera images), Requests (for sending and receiving data), Hugging Face's Transformers (using GPT-3 as a generative AI model).

[0956] Process flow

[0957] 1. Image capture

[0958] The user stands in front of an exhibit or product and captures an image using the device's camera. This image data is then sent from the device to the server.

[0959] 2. Image recognition and acquisition of explanatory data

[0960] The server analyzes the received image data and generates identifiers to identify exhibits and products. Based on these identifiers, the server retrieves explanatory data from a database. This explanatory data is converted into a user-friendly format using natural language generation technology. The converted explanatory data is then sent to the terminal and displayed or played as audio.

[0961] 3. Questions from users

[0962] If a user has questions about the explanation, they can ask them via voice input or text input. The device then sends this question data to the server.

[0963] 4. Generating answers to questions

[0964] The server analyzes the received question data and generates appropriate answers using a generative AI model (e.g., GPT-3). These generated answers are sent to the terminal and displayed or played aloud.

[0965] 5. User feedback and opinions

[0966] Users can input their thoughts and opinions about the exhibits and products. This can be done via voice input or text input. The terminal then sends the inputted feedback data to the server.

[0967] 6. Generating feedback on comments

[0968] The server collects and organizes the received feedback data and stores it in a database. Furthermore, it uses natural language generation technology to generate appropriate feedback based on the feedback. This feedback is sent to the device and displayed or played aloud.

[0969] Specific example

[0970] For example, suppose a user stands in front of a tourist attraction and points their device's camera at it to capture an image. This image data is sent to a server, which identifies the tourist attraction and generates explanatory data such as, "This attraction was built in the 19th century and is characterized by its architectural style." If the user inputs a question like, "I want to know more about the history of this attraction," the server uses a generative AI model to generate an answer such as, "The event of [event name] influenced the history of this attraction." Also, if the user inputs a comment like, "The design of this building is wonderful," the server generates feedback such as, "Many experts have also given its design high marks."

[0971] Example of a prompt

[0972] Examples of prompts for a generative AI model regarding exhibits include the following:

[0973] "Could you tell me about the history of this exhibit?"

[0974] "Please provide detailed information about this product."

[0975] In this way, users can obtain detailed information about exhibits and products, and receive interactive responses to their questions. This deepens the user's understanding and experience, resulting in a richer informational experience.

[0976] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0977] Step 1:

[0978] The user stands in front of the exhibit or product and captures an image using the camera on their device.

[0979] Specific operation:

[0980] Input: The object of the exhibit or product that the user points the camera at.

[0981] Data processing: Take an image with the device's camera and save it as an image file to the internal memory.

[0982] Output: Image data

[0983] Step 2:

[0984] The device sends the captured image data to the server.

[0985] Specific operation:

[0986] Input: Image data stored on the device

[0987] Data processing: Sending image data to a server using an HTTP request (e.g., POST request).

[0988] Output: Image data sent to the server

[0989] Step 3:

[0990] The server analyzes the received image data and generates identifiers to identify exhibits and products.

[0991] Specific operation:

[0992] Input: Image data sent to the server

[0993] Data processing: Use image recognition algorithms to identify exhibits and products within the image and generate their identifiers.

[0994] Output: Generated identifier

[0995] Step 4:

[0996] The server retrieves explanatory data from the database based on an identifier and converts it into a user-friendly format using natural language generation technology.

[0997] Specific operation:

[0998] Input: Generated identifier

[0999] Data processing: Query the database based on identifiers to retrieve relevant descriptive data. Then, use natural language generation techniques (e.g., NLG models) to generate text.

[1000] Output: Converted explanatory data

[1001] Step 5:

[1002] The terminal receives explanatory data sent from the server and displays or plays it as audio for the user.

[1003] Specific operation:

[1004] Input: Converted explanatory data

[1005] Data processing: Display explanatory data on the device's screen, or play it back as audio using speech synthesis technology.

[1006] Output: Displayed or audio-played commentary

[1007] Step 6:

[1008] Users have questions about the explanation and ask them using voice input or text input on their device.

[1009] Specific operation:

[1010] Input: User voice input or text input

[1011] Data processing: If using speech recognition technology, convert the audio to text (in the case of voice input). Then save the question as text data.

[1012] Output: Question data

[1013] Step 7:

[1014] The terminal sends the entered question data to the server.

[1015] Specific operation:

[1016] Input: Entered question data

[1017] Data processing: Send the question data to the server using an HTTP request.

[1018] Output: Question data sent to the server

[1019] Step 8:

[1020] The server analyzes the received question data and uses a generative AI model to generate appropriate answers.

[1021] Specific operation:

[1022] Input: Question data sent to the server

[1023] Data processing: Generative AI models (e.g., GPT-3) are used to generate appropriate answers to questions.

[1024] Output: Generated response data

[1025] Step 9:

[1026] The terminal receives the response data sent from the server and displays it to the user or plays it as audio.

[1027] Specific operation:

[1028] Input: Generated response data

[1029] Data processing: Display the response data on the device's screen, or play it back as audio using speech synthesis technology.

[1030] Output: Displayed or audio-played response

[1031] Step 10:

[1032] Users can input their thoughts and opinions about exhibits and products, providing feedback via voice or text input on their devices.

[1033] Specific operation:

[1034] Input: User voice input or text input

[1035] Data processing: Converting speech to text using speech recognition technology (in the case of voice input). Then, saving the comments as text data.

[1036] Output: Feedback data

[1037] Step 11:

[1038] The terminal sends the entered feedback data to the server.

[1039] Specific operation:

[1040] Input: Entered feedback data

[1041] Data processing: Send feedback data to the server using an HTTP request.

[1042] Output: Feedback data sent to the server

[1043] Step 12:

[1044] The server collects and organizes the feedback data it receives, stores it in a database, and uses natural language generation technology to generate appropriate feedback.

[1045] Specific operation:

[1046] Input: Feedback data sent to the server

[1047] Data processing: Save the feedback data to a database and generate feedback using natural language generation technology (e.g., NLG model).

[1048] Output: Generated feedback data

[1049] Step 13:

[1050] The device receives feedback data sent from the server and displays or plays it aloud for the user.

[1051] Specific operation:

[1052] Input: Generated feedback data

[1053] Data processing: Display feedback data on the device's screen or play it back as audio using speech synthesis technology.

[1054] Output: Displayed or audible feedback

[1055] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1056] ---

[1057] This invention relates to a system that uses generative AI to provide explanations of exhibits and respond to user questions and comments in art museums and museums. Furthermore, by combining this with an emotion engine that recognizes user emotions, the aim is to provide users with personalized explanations and feedback.

[1058] System Overview

[1059] Recognition and explanation of exhibits

[1060] When a user stands in front of an exhibit and launches an app on their device (e.g., a smartphone or tablet), the device's camera is used to capture an image of the exhibit. The captured image data is sent from the device to a server. The server analyzes the received image data and generates an ID to identify the exhibit. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent to the device and displayed or played aloud for the user.

[1061] User Q&A

[1062] If a user has a question about an exhibit, they input their question via voice or text into the terminal. The terminal sends the entered question data to a server, which uses generative AI to generate an appropriate answer to the question. The generated answer is sent to the terminal and displayed or played back to the user.

[1063] User feedback and comments

[1064] When a user inputs their thoughts and opinions into their device via voice or text, the device sends that data to a server. The server analyzes the received data, stores it in a database, and generates feedback based on that data using natural language generation technology. The generated feedback is sent to the device and displayed or played back to the user.

[1065] Enhanced entertainment value

[1066] Furthermore, this system retrieves character information set for each exhibit from a server and reflects the character's tone of voice and personality when generating explanatory texts and answers based on that information. The generated explanatory texts and answers are expressed using character avatars and voice clips, providing users with both visual and auditory experiences.

[1067] Combination of emotional engines

[1068] A distinctive feature of this invention is the integration of an emotion engine. The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. For example, using the user's camera video and audio data, the emotion engine detects emotions such as joy, interest, and dissatisfaction. The terminal sends this emotional data to a server, which analyzes the data to generate explanations, answers, or feedback that correspond to the user's emotions.

[1069] For example, if the emotion engine recognizes that a user is looking at an exhibit with interest, the server generates and sends to the user's device an explanation containing detailed information and interesting anecdotes that will pique the user's interest. On the other hand, if the system recognizes that the user is somewhat tired, it provides a concise and to-the-point explanation. This allows for a more personalized and satisfying user experience.

[1070] Specific example

[1071] For example, suppose a user stands in front of a "famous painting" and launches an app on their device. When the user points their camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses image recognition technology to identify the "famous painting" and generates an ID for it. Based on this ID, the server retrieves explanatory data from its database and generates an explanatory text, such as "This painting was painted in [era], and its characteristics are [characteristics]," and sends it to the device. The device then displays or plays the explanatory text for the user.

[1072] Next, when the user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses generative AI to generate an answer to the question and sends the answer, such as "The following events influenced the creation of this painting," to the device. The device then displays or plays the answer aloud to the user.

[1073] Furthermore, when a user enters a comment such as "The use of color in this painting is wonderful," the device sends the comment data to the server. The server stores the comment data in a database and generates feedback such as "Many experts have praised the use of color. We are pleased that you noticed this point," and sends it to the device. In this way, the user can receive feedback on their own comments.

[1074] Furthermore, by using the emotion engine, if a user says, "This painting is very moving," that emotion can be recognized, and specific feedback such as, "Thank you for being moved. It is said that the artist painted this picture with the same feelings," can be received.

[1075] The above is an example of implementing the system according to the present invention. In this way, it is possible to realize a more personalized user experience by using an emotion engine, in addition to providing explanations for exhibits.

[1076] ---

[1077] The following describes the processing flow.

[1078] ---

[1079] Program processing flow

[1080] Recognition processing of exhibits

[1081] Step 1:

[1082] The user stands in front of the exhibit and launches the app on their device (e.g., a smartphone or tablet).

[1083] Step 2:

[1084] The user uses the device's camera to capture an image of the exhibit. The device activates the camera, generates image data, and saves that image data to memory.

[1085] Step 3:

[1086] The device sends the captured image data to the server. Specifically, it sends the image data to the server via an HTTP request.

[1087] Step 4:

[1088] The server executes an image recognition algorithm to identify the received image data. For example, the server uses a Convolutional Neural Network (CNN) model to detect exhibits from the image and generate a corresponding exhibit ID.

[1089] Processing to retrieve explanations for exhibits

[1090] Step 5:

[1091] The server uses the generated exhibit ID to retrieve the corresponding explanatory data from the explanatory database. This involves extracting explanatory information from the database using SQL queries or similar methods.

[1092] Step 6:

[1093] The server uses natural language generation (NLG) technology to convert the acquired explanatory data into a user-friendly format. Specifically, the NLG engine generates the explanatory text.

[1094] Step 7:

[1095] The server sends the generated explanatory data to the terminal. It returns data containing the explanatory text to the terminal as an HTTP response.

[1096] Step 8:

[1097] The device displays the received explanatory data to the user. This is done using a text display UI or a speech synthesis engine to provide the user with audio explanations.

[1098] User question answering process

[1099] Step 9:

[1100] If a user has a question about an exhibit, they can input their question into the terminal either by voice or text. The terminal uses a speech recognition engine to convert the voice to text or accepts the text input.

[1101] Step 10:

[1102] The terminal sends question data in text format to the server. An HTTP request is used for this transmission.

[1103] Step 11:

[1104] The server analyzes the received question data. Natural language understanding (NLU) technology is used for the analysis to understand the user's question.

[1105] Step 12:

[1106] The server uses a generative AI (e.g., GPT-4) to generate answers to questions. The generated answers are structured as natural language text.

[1107] Step 13:

[1108] The server sends the generated response to the terminal. The response data is returned to the terminal as an HTTP response.

[1109] Step 14:

[1110] The device displays the answer to the user. This display uses a text display UI or a speech synthesis engine.

[1111] User feedback sharing process

[1112] Step 15:

[1113] Users input their impressions and opinions about the exhibits into the device via voice or text. The device uses a speech recognition engine to convert the voice to text or receives the text input.

[1114] Step 16:

[1115] The device sends feedback data in text format to the server. HTTP requests are used for transmission.

[1116] Step 17:

[1117] The server analyzes and organizes the received feedback data and saves it to a database. SQL queries are used to write the data to the database for saving.

[1118] Step 18:

[1119] The server generates feedback using natural language generation technology based on the feedback data. The NLG engine generates appropriate feedback sentences.

[1120] Step 19:

[1121] The server sends the generated feedback to the terminal. The feedback data is returned to the terminal as an HTTP response.

[1122] Step 20:

[1123] The device displays feedback to the user. This is done using a text display UI or a speech synthesis engine.

[1124] Processing to enhance entertainment value

[1125] Step 21:

[1126] The server retrieves character information set for each exhibit. It uses SQL queries to retrieve character information from the database.

[1127] Step 22:

[1128] When the server generates explanatory text and response text based on character information, it reflects the character's speech pattern and personality. An NLG engine is used for generation.

[1129] Step 23:

[1130] The server transmits the character's voice data and avatar information to the terminal. This includes synthesized voice data and image / video data.

[1131] Step 24:

[1132] The device provides explanations and answers to the user using character avatars and voice recordings. The device displays character avatars and plays explanations and answers via voice recordings.

[1133] Combination of emotional engines

[1134] Step 25:

[1135] The device uses its built-in camera and microphone to analyze the user's voice and facial expressions as they view exhibits and express their opinions. The device runs an emotion engine to analyze voice tone and facial expressions.

[1136] Step 26:

[1137] The device converts the analyzed emotion data into text format and sends it to the server. An HTTP request is used for transmission.

[1138] Step 27:

[1139] The server analyzes the received emotional data and generates explanations, responses, or feedback tailored to the user's emotional state. For example, if the user shows interest, it provides a detailed explanation; if the user is tired, it provides a concise explanation.

[1140] Step 28:

[1141] The server sends the generated emotion response explanation, answer, or feedback to the device. It returns the data as an HTTP response.

[1142] Step 29:

[1143] The device displays or plays the generated emotion-responsive data to the user. Text display UI or speech synthesis engines are used for display.

[1144] ---

[1145] (Example 2)

[1146] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[1147] Traditional museum and art gallery exhibit explanation systems have been limited to providing static information about the exhibits, failing to address the individual interests and emotions of users. As a result, the user experience was not personalized, resulting in one-sided information delivery and low satisfaction. Furthermore, there was a lack of means to provide appropriate feedback to user questions and comments, often leading to a superficial user experience. In addition to these problems, the lack of entertainment value also needs improvement.

[1148] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing image data transmitted from a terminal to recognize an exhibit and generating an ID corresponding to the identified exhibit; means for converting explanatory data obtained from a database based on the identified exhibit ID into a user-friendly format using natural language generation technology; means for transmitting the generated explanatory data to the user's terminal; means for analyzing questions from the user and generating appropriate answers using a generation AI model; means for analyzing the user's facial expressions and voice tone using an emotion engine to recognize the user's emotions; and means for generating explanations, answers, or feedback that correspond to the user's emotions based on the analyzed emotion data. This makes it possible to provide users with personalized explanations and feedback, thereby achieving higher satisfaction.

[1149] "Exhibits" refer to visual or physical objects such as paintings, sculptures, crafts, and historical artifacts displayed in art museums and museums.

[1150] "Device" refers to a smartphone, tablet, or other portable device used by a user for operation.

[1151] A "camera" is a photographic device attached to a terminal for capturing images of exhibits.

[1152] A "server" refers to a computer system that performs processing such as analyzing received data and generating responses using generative AI.

[1153] "Image data" refers to photographic and video data of exhibits captured using the device's camera.

[1154] An "ID" is a unique identifier generated to identify an exhibit.

[1155] A "database" is a storage area where various types of data, such as explanatory data, character information, and user reviews, are stored.

[1156] "Explanatory data" refers to text data and audio data that include information and explanations related to the exhibits.

[1157] "Natural Language Generation (NLG)" refers to the technology used to convert machine-generated data into human language.

[1158] A "generative AI model" is an artificial intelligence model used to generate appropriate answers or feedback to questions.

[1159] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice tone to detect their emotional state.

[1160] A "prompt message" is text data input to a generative AI model, and it is an instruction message that prompts the model to generate an answer.

[1161] An "avatar" refers to a visual representation generated based on character information, which is typically displayed on the user's device.

[1162] This invention relates to a system that uses generative AI to provide explanations of exhibits and respond to user questions and comments in art museums and museums. Furthermore, by combining this with an emotion engine that recognizes user emotions, the aim is to provide users with personalized explanations and feedback.

[1163] System Overview

[1164] This system consists of the following main elements:

[1165] Recognition and explanation of exhibits

[1166] The user stands in front of an exhibit and uses a smartphone or tablet to capture an image of the exhibit with its camera. The device then sends the captured image data to a server. The server analyzes the received image using image recognition technology (e.g., TensorFlow or OpenCV) and generates an ID to identify the exhibit. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent to the device and displayed or played aloud for the user.

[1167] Specific example

[1168] For example, if a user stands in front of a "famous painting" and points their device's camera at it to capture an image, the device sends the image to a server. The server uses OpenCV to extract features from the image and identify the specific exhibit. Based on this ID, it generates an explanatory text from its database, such as "This painting was painted by [era]," and sends it to the device. The device then plays the explanatory text aloud for the user to listen to.

[1169] User Q&A

[1170] If a user wants to ask a question about an exhibit, they can input it via voice or text into the terminal. The terminal sends the entered question data to a server, which uses a generative AI model (e.g., GPT-3 or GPT-4) to generate an appropriate answer to the question. The generated answer is sent to the terminal and displayed or played back to the user.

[1171] Specific example

[1172] When a user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses a generative AI model to generate an answer, such as "The following events influenced the creation of this painting," and sends it to the device. The device then displays the answer for the user to see.

[1173] User feedback and comments

[1174] When a user inputs their thoughts and opinions into their device via voice or text, the device sends that data to a server. The server analyzes the received data, stores it in a database, and generates feedback based on that data. The generated feedback is sent to the device and displayed or played back to the user.

[1175] Specific example

[1176] When a user enters a comment such as "The colors in this painting are wonderful," the device sends the comment data to the server. The server saves the comment data and generates feedback such as "Many experts have praised this. We are glad you noticed this point," and sends it to the device. The device then plays the feedback text aloud for the user to hear.

[1177] Combination of emotional engines

[1178] The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. Using the user's camera video and audio data, the emotion engine detects emotions such as joy, interest, and dissatisfaction. The device sends this emotional data to a server, which generates explanations, answers, or feedback tailored to the user's emotions based on the analysis results.

[1179] Specific example

[1180] If a user says, "This painting is very moving," the emotion engine recognizes that emotion and generates specific feedback such as, "Thank you for being moved. It is said that the artist painted this picture with the same feelings," and sends it to the device. The device then plays the feedback aloud and delivers it to the user.

[1181] Example of a prompt

[1182] 1. Prompt for generating detailed descriptions of exhibits

[1183] Please generate detailed descriptions for the following exhibits.

[1184] Exhibit ID: ABC123

[1185] Era: Ancient

[1186] Author: Author A

[1187] Features: Vibrant brushwork

[1188] 2. Question and answer prompt sentences

[1189] Please generate answers to the following questions.

[1190] Question: What is the background behind the creation of this painting?

[1191] Exhibit ID: ABC123"

[1192] 3. Prompt for feedback / comments

[1193] Please generate feedback for the following comments.

[1194] Impression: The use of color in this painting is wonderful.

[1195] Exhibit ID: ABC123"

[1196] 4. Feedback prompts based on emotion recognition

[1197] Please generate feedback based on the following emotions:

[1198] Emotion: I'm moved.

[1199] Comment: This painting is very moving.

[1200] Exhibit ID: ABC123"

[1201] In this way, the present invention makes it possible to provide personalized explanations of exhibits in art museums and museums, and to offer an immersive experience that responds to the user's emotions.

[1202] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1203] System program processing flow

[1204] Recognition and explanation of exhibits

[1205] Step 1:

[1206] The user stands in front of the exhibit.

[1207] Input: User's location, exhibits

[1208] Output: Captured camera connection

[1209] Specific actions: The user points their smartphone or tablet camera at the exhibit and launches a dedicated app.

[1210] Step 2:

[1211] The device captures images of the exhibits.

[1212] Input: Exhibit images, camera

[1213] Output: Captured image data

[1214] Specific operation: The device's camera takes a picture of the exhibit and generates image data.

[1215] Step 3:

[1216] The device sends the captured image data to the server.

[1217] Input: Captured image data

[1218] Output: Data sent to the server after completion

[1219] Specific operation: The terminal sends image data to the server using an open communication protocol.

[1220] Step 4:

[1221] The server analyzes the image data.

[1222] Input: Received image data

[1223] Output: Exhibit identification ID

[1224] Specific operation: The server analyzes image data using TensorFlow and OpenCV to extract and identify features of the exhibits.

[1225] Step 5:

[1226] The server retrieves and generates explanatory data from the database based on the identified exhibit ID.

[1227] Input: Exhibit identification ID

[1228] Output: Explanatory data, generated explanatory text

[1229] Specific operation: The server retrieves relevant explanatory data from the explanatory database and uses natural language generation (NLG) technology to generate an easy-to-understand explanatory text for the user.

[1230] Step 6:

[1231] The server sends the generated explanatory text to the terminal.

[1232] Input: Generated explanatory text

[1233] Output: Transmission complete data

[1234] Specific operation: The server sends the generated explanatory text to the terminal.

[1235] Step 7:

[1236] The device displays or plays explanatory text.

[1237] Input: Received explanatory text

[1238] Output: Displayed or audio-played explanatory text

[1239] Specific operation: The device will either display the received explanatory text on the screen or play it as audio.

[1240] User Q&A

[1241] Step 1:

[1242] The user enters a question.

[1243] Input: Questions via voice or text input

[1244] Output: Question data

[1245] Specific operation: The user enters the question using voice input or text input on the device.

[1246] Step 2:

[1247] The device sends the question data to the server.

[1248] Input: Question data

[1249] Output: Data sent to the server after completion

[1250] Specific action: The terminal sends the question data to the server.

[1251] Step 3:

[1252] The server analyzes the question data and generates answers.

[1253] Input: Question data

[1254] Output: Generated answer

[1255] Specific operation: The server analyzes the question data and generates appropriate answers using a generative AI model (e.g., GPT-3 or GPT-4).

[1256] Step 4:

[1257] The server sends the generated response to the terminal.

[1258] Input: Generated answer

[1259] Output: Transmission complete data

[1260] Specific operation: The server sends the generated response to the terminal.

[1261] Step 5:

[1262] The device displays or plays the answer.

[1263] Input: Received response

[1264] Output: Displayed or audio-played answer

[1265] Specific actions: The device will either display the received response on the screen or play it as audio.

[1266] User feedback and comments

[1267] Step 1:

[1268] Users enter their feedback.

[1269] Input: Feedback via voice or text input

[1270] Output: Feedback data

[1271] Specific operation: The user enters their feedback by using voice input or text input on the device.

[1272] Step 2:

[1273] The device sends feedback data to the server.

[1274] Input: Feedback data

[1275] Output: Data sent to the server after completion

[1276] Specific action: The device sends feedback data to the server.

[1277] Step 3:

[1278] The server analyzes and stores the feedback data.

[1279] Input: Received feedback data

[1280] Output: Saved data, analyzed feedback data

[1281] Specific operation: The server analyzes the feedback data and saves it to the database.

[1282] Step 4:

[1283] The server generates feedback based on the feedback data.

[1284] Input: Analyzed feedback data

[1285] Output: Generated feedback

[1286] Specific operation: The server uses natural language generation technology to generate feedback based on the feedback data.

[1287] Step 5:

[1288] The server sends the generated feedback to the terminal.

[1289] Input: Generated feedback

[1290] Output: Transmission complete data

[1291] Specific operation: The server sends the generated feedback to the terminal.

[1292] Step 6:

[1293] The device displays or plays feedback.

[1294] Input: Received feedback

[1295] Output: Displayed or audible feedback

[1296] Specific actions: The device displays the received feedback on the screen or plays it back as audio.

[1297] Combination of emotional engines

[1298] Step 1:

[1299] Capture user sentiment data

[1300] Input: User's facial expressions and voice

[1301] Output: Captured sentiment data

[1302] Specific operation: The device's camera and microphone capture the user's facial expressions and voice.

[1303] Step 2:

[1304] The device sends emotional data to the server.

[1305] Input: Captured emotion data

[1306] Output: Data sent to the server after completion

[1307] Specific operation: The device sends the captured emotion data to the server.

[1308] Step 3:

[1309] The server analyzes the emotional data.

[1310] Input: Received sentiment data

[1311] Output: Analyzed emotional state

[1312] Specific operation: The server analyzes the emotional state using an emotion engine (e.g., Affectiva SDK or Microsoft Azure Emotion API).

[1313] Step 4:

[1314] The server generates content based on emotions.

[1315] Input: Analyzed emotional state

[1316] Output: Generated emotion-responsive content

[1317] Specific operation: Based on the analyzed sentiment data, the server generates explanations, responses, or feedback that correspond to the user's emotions.

[1318] Step 5:

[1319] The server sends the generated content to the terminal.

[1320] Input: Generated emotion-responsive content

[1321] Output: Transmission complete data

[1322] Specific operation: The server sends the generated emotion-responsive content to the device.

[1323] Step 6:

[1324] The device displays content or plays audio.

[1325] Input: Received emotion-responsive content

[1326] Output: Displayed or played content

[1327] Specific actions: The device will either display the received content on the screen or play it as audio.

[1328] The above outlines the specific flow of the system's program processing. This allows users to receive personalized explanations and feedback on the exhibits, resulting in a more satisfying experience.

[1329] (Application Example 2)

[1330] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[1331] Traditional systems for providing explanations of exhibits and products could only offer the same information at all times, failing to provide an experience tailored to the individual interests and emotions of users. Furthermore, when users asked detailed questions about specific products or exhibits, they often sought more comprehensive information, but the responses were limited to general and consistent answers. As a result, user satisfaction decreased, and problems arose where purchasing intent and learning motivation were not sufficiently stimulated.

[1332] The identification process performed by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing received image data and generating an ID corresponding to the identified exhibit; means for obtaining explanatory data from a database based on the identified exhibit ID and converting the explanatory data into a user-friendly format using natural language generation technology; and means for generating an appropriate response using a generative AI. This makes it possible not only to provide users with detailed information about exhibits and products, but also to provide personalized explanations and feedback that are tailored to the user's emotional state.

[1333] The term "exhibits" refers to all objects that customers are interested in, not only in art museums and museums, but also in physical stores.

[1334] "Terminal" refers to a smartphone, tablet, personal computer, or other portable information processing device.

[1335] A "camera" refers to an optical device used to capture images and videos.

[1336] "Image data" refers to a digital representation of a captured image.

[1337] A "server" refers to a computer system used to receive, process, and transmit data over a network.

[1338] "ID" refers to a code or label used to uniquely identify a specific exhibit or product.

[1339] A "database" refers to a system or software used to organize and store information.

[1340] "Natural language generation technology" refers to the technology that uses computers to generate natural-sounding words and sentences.

[1341] "Generative AI" refers to artificial intelligence technology that generates appropriate responses or new data based on input data.

[1342] "Voice input" refers to a method of recognizing human speech as digital information using a microphone.

[1343] "Text input" refers to the method of entering character data using a keyboard or on-screen keyboard.

[1344] "User" refers to all individuals who use the system.

[1345] An "emotion engine" refers to software or hardware that analyzes a user's facial expressions and tone of voice to detect their psychological state.

[1346] "Feedback" refers to the response or reaction that a system gives to a user's actions or input.

[1347] "Character information" refers to digital data or profiles used to add a specific personality or tone of voice to explanatory texts or answer texts.

[1348] This invention relates to a system that provides customers with detailed information about exhibits and products in physical stores, and further provides personalized explanations and feedback that respond to the customer's emotions. Specific embodiments of this invention are shown below.

[1349] System Overview

[1350] This system includes the following main components:

[1351] terminal

[1352] A terminal is a portable information processing device such as a smartphone or tablet. The terminal is equipped with a camera, microphone, display, and speaker. This allows users to photograph products and exhibits and interact with the system through voice or text input.

[1353] server

[1354] The server communicates with terminals over the network and processes the received data. The server includes multiple modules for image recognition, natural language generation technology, and generative AI processing.

[1355] Image capture and recognition

[1356] The user uses the device's camera to capture images of products or exhibits. The captured image data is sent from the device to a server, where the server's image recognition module analyzes it. An ID corresponding to the identified exhibit is generated from the analysis results. This ID is used to retrieve explanatory data from the database.

[1357] Question and Answer

[1358] When a user sends a question about an exhibit to their device via voice or text input, the question data is transferred to a server. The server uses generative AI to generate an appropriate answer and sends it to the device. The device then displays or plays this answer aloud.

[1359] Emotional engine and feedback

[1360] Furthermore, this system is equipped with an emotion engine that analyzes the user's emotions using camera and audio data. Once the user's emotional state is recognized, that emotional data is sent to the server. The server generates explanations and feedback based on the emotional data, providing a personalized experience. This feedback is generated based on information obtained from the database and the user's emotional state.

[1361] Specific example

[1362] For example, suppose a user is in a physical store and stands before the latest smartphone, then launches an app on the device. The user takes a picture of the smartphone using the camera and sends the image data to a server. The server uses image recognition technology to identify the smartphone and generates an ID. Based on this ID, it retrieves a detailed explanation from a database and uses natural language generation technology to generate an easy-to-understand explanatory text for the user. This explanatory text is sent to the device and is either displayed on the screen or played as audio.

[1363] When a user asks a question by voice, such as "How is the camera function of this smartphone?", the device sends the question data to a server. The server uses generative AI to generate a detailed answer and sends it to the device, such as "This smartphone's camera has a 12-megapixel sensor and excellent night vision capabilities." The device then displays or plays this answer aloud to the user.

[1364] Furthermore, the emotion engine, which analyzes the user's facial expressions and tone of voice, recognizes when the user is excited. For example, when a user says, "This smartphone is amazing!", based on that emotion data, it generates feedback such as, "I'm glad you're pleased! Shall I tell you about other features?"

[1365] Example of a prompt

[1366] Examples of prompt statements are as follows:

[1367] User emotion: Joy

[1368] Question: What features does this washing machine have?

[1369] answer:

[1370] This prompt is input into a generative AI model, which generates a response tailored to the user's emotion of "joy."

[1371] The above describes a specific embodiment of the system according to the present invention. In this way, it becomes possible to provide personalized explanations and feedback that respond to the user's emotions.

[1372] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1373] Step 1:

[1374] The user captures images of products or exhibits using the device's camera. The input is video from the device's camera, and the output is the captured image data. This image data is temporarily stored on the device.

[1375] Step 2:

[1376] The terminal sends the captured image data to the server. Specifically, it uploads the image data to the server via the network connection by including it in the request packet. The captured image data is the input, and the server receives the data as output.

[1377] Step 3:

[1378] The server analyzes the received image data and uses image recognition technology to identify specific exhibits or products. It retrieves the ID of the identified object from the database and generates a dedicated code or label. Image data is the input, and the identified ID is generated as the output.

[1379] Step 4:

[1380] The server retrieves detailed explanatory data from the database based on the identified ID. Next, it uses natural language generation technology to convert the explanatory data into a user-friendly format. The input consists of the identified ID and associated explanatory data, and the output is the converted explanatory text.

[1381] Step 5:

[1382] The server sends the converted explanatory data to the terminal. Specifically, it sends packet data containing the generated explanatory text to the terminal via the network. The input is the explanatory text converted into an easy-to-understand format, and the output is the explanatory text sent to the terminal.

[1383] Step 6:

[1384] The terminal displays or plays explanatory text for the user. The input is explanatory text received from the server, and the output is either text displayed on the screen or audio played through the speaker.

[1385] Step 7:

[1386] Users submit questions about exhibits via voice or text input. The input is the user's voice or text, and the output is the voice recognition or text input data, which is stored on the device.

[1387] Step 8:

[1388] The terminal sends the entered question data to the server. Specifically, it uploads the question data to the server using a network connection. The user's question data is the input, and the server receives the data as output.

[1389] Step 9:

[1390] The server analyzes the received question data and generates appropriate answers using generative AI. The input is the question data, and the output is the generated answers.

[1391] Step 10:

[1392] The server sends the generated response to the terminal. Specifically, it sends packet data containing the generated response to the terminal over the network. The generated response is the input, and the response sent to the terminal is the output.

[1393] Step 11:

[1394] The device displays or plays the answer to the user. The input is the answer received from the server, and the output is either text displayed on the screen or audio played through the speaker.

[1395] Step 12:

[1396] To recognize the user's emotions, the device's camera and microphone are used to capture facial expressions and voice tone. Camera video and audio data are taken as input, and emotion data is generated as output.

[1397] Step 13:

[1398] The device sends emotional data to the server. Specifically, it sends packets containing emotional data to the server over the network. The emotional data is the input, and the server receives the data as the output.

[1399] Step 14:

[1400] The server analyzes emotional data and generates explanations and feedback tailored to the user's emotions. It takes emotional data and related explanation data as input, and generates personalized feedback as output.

[1401] Step 15:

[1402] The server sends personalized feedback to the terminal. Specifically, it sends packet data containing the generated feedback to the terminal over the network. The input is the personalized feedback, and the output is the feedback sent to the terminal.

[1403] Step 16:

[1404] The device displays or plays audio feedback to the user. Input is feedback received from a server, and output is text displayed on the screen or audio played through the speaker.

[1405] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1406] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1407] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1408] [Third Embodiment]

[1409] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1410] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1411] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1412] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1413] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1415] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1416] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1417] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1418] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1419] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1420] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1421] ---

[1422] This invention relates to a system that uses generative AI to provide explanations to users about exhibits in art museums and museums, and to respond to user questions and comments. This system comprehensively performs exhibit recognition, acquisition and display of explanatory data, response to user questions, and collection and feedback of comments.

[1423] System Overview

[1424] First, when a user stands in front of an exhibit, they use a device (such as a smartphone or tablet) to capture an image of the exhibit. The device generates image data using its camera and sends this data to a server. The server uses image recognition technology to identify the exhibit and generates an ID. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent from the server to the device and displayed or played as audio for the user on the device.

[1425] If a user has a question about an exhibit, they enter the question into the terminal via voice or text input. The terminal sends the entered question data to a server, which uses generative AI to generate an appropriate answer to the question. This answer is then sent to the terminal and displayed or played back to the user.

[1426] Furthermore, when users want to input their thoughts and opinions, they do so via voice or text on the device. The device sends the input data to the server, which stores the received data in a database and generates appropriate feedback, which is then sent back to the device. This allows users to receive feedback on their own comments.

[1427] The system also includes features to enhance entertainment value. The server retrieves character information set for each exhibit and reflects the character's tone and personality in the explanations and answers. The generated explanations and answers are played back using the character's avatar and voice, providing users with both visual and auditory experiences.

[1428] Specific example

[1429] For example, suppose a user stands in front of a "famous painting" and launches an app on their device. When the user points their camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses image recognition technology to identify the "famous painting" and generates an ID. Using this ID, the server retrieves explanatory data from a database, generates an explanatory text such as "This painting was painted in [era], and its characteristics are [characteristics]," and sends it to the device. The device then displays or plays the explanatory text for the user.

[1430] Next, when the user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses generative AI to generate an answer to the question and sends the answer, such as "The following events influenced the creation of this painting," to the device. The device then displays or plays the answer aloud to the user.

[1431] Furthermore, when a user enters a comment such as "The use of color in this painting is wonderful," the device sends the comment data to the server. The server stores the comment data in a database and generates feedback such as, "Many experts have praised the use of color. We are pleased that you noticed this point," and sends it back to the device. This allows the user to gain a deeper understanding and greater satisfaction.

[1432] The above is an example of implementing the system according to the present invention. This makes it possible to provide multifaceted and individualized information about exhibits, thereby enhancing the user experience.

[1433] ---

[1434] The following describes the processing flow.

[1435] ---

[1436] Program processing flow

[1437] Recognition processing of exhibits

[1438] Step 1:

[1439] The user stands in front of the exhibit and launches the app on their device (e.g., a smartphone or tablet).

[1440] Step 2:

[1441] The device's camera is used to capture images of the exhibits. The device activates the camera, generates image data, and saves that image data to memory.

[1442] Step 3:

[1443] The device sends the captured image data to the server. Specifically, it sends the image data to the server via an HTTP request.

[1444] Step 4:

[1445] The server executes an image recognition algorithm to identify the received image data. For example, the server uses a Convolutional Neural Network (CNN) model to detect exhibits from the image and generate a corresponding exhibit ID.

[1446] Processing to retrieve explanations for exhibits

[1447] Step 5:

[1448] The server uses the generated exhibit ID to retrieve the corresponding explanatory data from the explanatory database. This involves extracting explanatory information from the database using SQL queries or similar methods.

[1449] Step 6:

[1450] The server uses natural language generation (NLG) technology to convert the acquired explanatory data into a user-friendly format. Specifically, the NLG engine generates the explanatory text.

[1451] Step 7:

[1452] The server sends the generated explanatory data to the terminal. It returns data containing the explanatory text to the terminal as an HTTP response.

[1453] Step 8:

[1454] The device displays the received explanatory data to the user. This is done using a text display UI or a speech synthesis engine to provide the user with audio explanations.

[1455] User question answering process

[1456] Step 9:

[1457] If a user has a question about an exhibit, they can input their question into the terminal either by voice or text. The terminal uses a speech recognition engine to convert the voice to text or accepts the text input.

[1458] Step 10:

[1459] The terminal sends question data in text format to the server. An HTTP request is used for this transmission.

[1460] Step 11:

[1461] The server analyzes the received question data. Natural language understanding (NLU) technology is used for the analysis to understand the user's question.

[1462] Step 12:

[1463] The server uses a generative AI (e.g., GPT-4) to generate answers to questions. The generated answers are structured as natural language text.

[1464] Step 13:

[1465] The server sends the generated response to the terminal. The response data is returned to the terminal as an HTTP response.

[1466] Step 14:

[1467] The device displays the answer to the user. This display uses a text display UI or a speech synthesis engine.

[1468] User feedback sharing process

[1469] Step 15:

[1470] Users input their impressions and opinions about the exhibits into the device via voice or text. The device uses a speech recognition engine to convert the voice to text or receives the text input.

[1471] Step 16:

[1472] The device sends feedback data in text format to the server. HTTP requests are used for transmission.

[1473] Step 17:

[1474] The server analyzes and organizes the received feedback data and saves it to a database. SQL queries are used to write the data to the database for saving.

[1475] Step 18:

[1476] The server generates feedback using natural language generation technology based on the feedback data. The NLG engine generates appropriate feedback sentences.

[1477] Step 19:

[1478] The server sends the generated feedback to the terminal. The feedback data is returned to the terminal as an HTTP response.

[1479] Step 20:

[1480] The device displays feedback to the user. This is done using a text display UI or a speech synthesis engine.

[1481] Processing to enhance entertainment value

[1482] Step 21:

[1483] The server retrieves character information set for each exhibit. It uses SQL queries to retrieve character information from the database.

[1484] Step 22:

[1485] When the server generates explanatory text and response text based on character information, it reflects the character's speech pattern and personality. An NLG engine is used for generation.

[1486] Step 23:

[1487] The server transmits the character's voice data and avatar information to the terminal. This includes synthesized voice data and image / video data.

[1488] Step 24:

[1489] The device provides explanations and answers to the user using character avatars and voice recordings. The device displays character avatars and plays explanations and answers via voice recordings.

[1490] ---

[1491] The above is the specific processing flow of the system program.

[1492] (Example 1)

[1493] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1494] Museums and art galleries need efficient and effective ways to provide explanations about their exhibits to users. Furthermore, there is a lack of systems that can quickly and appropriately respond to users' questions and comments about the exhibits, thereby providing a richer experience. Traditional systems primarily rely on paper-based descriptions for each exhibit, lacking interactivity and making it difficult to capture user interest. Moreover, the lack of entertainment-oriented exhibit explanations utilizing characters limits the potential for improving the user experience.

[1495] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1496] In this invention, the server includes means for capturing images using a camera device provided on the user terminal to identify exhibits and transmitting the image data to the server; means for analyzing the received image data on the server and generating an identifier corresponding to the identified exhibit; means for obtaining information data from a database based on the identified exhibit identifier on the server and converting the information data into a user-friendly format using natural language generation technology; means for transmitting the converted information data to the user terminal and displaying or playing it aloud on the user terminal; means for inputting user questions into the user terminal via voice or text input; means for transmitting the question data entered on the user terminal to the server; means for analyzing the received question data on the server and generating an appropriate answer using a generation AI model; and means for transmitting the generated answer to the user terminal. The system includes means for displaying or playing audio on the user terminal, means for inputting user feedback and opinions via voice or text into the user terminal, means for sending feedback data entered on the user terminal to a server, means for the server to collect and organize the received feedback data and store it in a database, means for generating appropriate feedback using natural language generation technology based on the feedback data, means for sending the generated feedback to the user terminal and displaying or playing audio on the user terminal, means for obtaining character information set for each exhibit from the server and reflecting the character's tone of voice and personality in the explanations and answers, means for sending data to the user terminal to express the character-based explanations and answers in the character's voice, and means for providing the character to the user as an avatar or voice on the user terminal. This makes it possible to provide users with interactive and rich explanations and responses to exhibits, as well as a highly entertaining experience.

[1497] "Exhibits" refer to works of art and materials that are made available to users in art museums and museums.

[1498] "User terminal" refers to a portable information device owned by a user, primarily such as a smartphone or tablet.

[1499] "Image capture equipment" refers to cameras and other optical instruments used to acquire image data.

[1500] A "server" is a central computing system that processes, stores, and provides data over a network.

[1501] "Image data" refers to data that represents captured visual information in a digital format.

[1502] "Analysis" refers to the process of examining received images and data in detail to extract information.

[1503] An "identifier" is a code or number generated to uniquely identify a particular exhibit.

[1504] A "database" is a system for efficiently storing, searching, and managing large amounts of information.

[1505] "Information data" refers to digital data that includes detailed explanations and information about the exhibits.

[1506] "Natural language generation technology" refers to technology that allows machines to automatically generate sentences that are easy for humans to understand.

[1507] "Question data" refers to data that digitally represents questions about exhibits entered by users.

[1508] A "generative AI model" refers to an algorithm or model that uses artificial intelligence to generate appropriate answers or text.

[1509] "Feedback" refers to the responses or evaluations provided in response to user input.

[1510] "Character information" refers to data that includes specific characteristics and speech patterns of fictional people, animals, and other elements set for each exhibit.

[1511] An "avatar" is an image or model used to visually represent a character.

[1512] "Voice" refers to data used to digitally reproduce the character's voice.

[1513] This invention relates to a system that uses generative AI to provide explanations to users about exhibits in art museums and museums, and to respond to user questions and comments. This system comprehensively performs exhibit recognition, acquisition and display of explanatory data, response to user questions, and collection and feedback of comments.

[1514] System Overview

[1515] This system works as follows:

[1516] First, when a user stands in front of an exhibit, they use their device (e.g., a smartphone or tablet) to capture an image of the exhibit. The device's camera generates image data, which is then sent to a server. The server uses image recognition technology to identify the exhibit and generates an identifier. Based on this identifier, the server retrieves relevant information data from a database and generates an explanatory text using natural language generation (NLG) technology. The generated explanatory text is sent from the server to the device and displayed or played back as audio for the user.

[1517] Hardware and software details

[1518] 1. Terminal (User terminal):

[1519] Hardware used: Smartphone, tablet

[1520] Software used: Dedicated app, camera function, audio recording and playback function

[1521] 2. Server:

[1522] Software used:

[1523] Image recognition technology: AWS Rekognition, Google Cloud Vision

[1524] Database: MySQL, MongoDB

[1525] Natural language generation technology (NLG): OpenAI GPT-4

[1526] Audio playback technology: Amazon Polly, Google Text-to-Speech

[1527] Processing details

[1528] The processing flow at each step is as follows:

[1529] 1. User-captured images of exhibits:

[1530] Users point their device's camera at the exhibit and capture images using a dedicated app.

[1531] 2. Sending image data from the terminal to the server:

[1532] The terminal encodes the captured image data in BASE64 format and sends it to the server using an HTTP POST request.

[1533] 3. Server-based recognition of exhibits and acquisition of data:

[1534] The server analyzes the received image data, using AWS Rekognition or Google Cloud Vision to recognize exhibits and generate identifiers. Based on the identified exhibit identifiers, the server retrieves relevant information data from MySQL or MongoDB databases.

[1535] 4. Server generates explanatory text and sends it to the terminal:

[1536] The server uses OpenAI's GPT-4 to generate natural language based on the acquired data. The generated explanatory text is sent from the server to the terminal and displayed or played back as audio for the user.

[1537] Specific usage scenarios

[1538] For example, suppose a user stands in front of a "famous painting" and launches a smartphone app. When the user points the camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses AWS Rekognition to identify the "famous painting" and generates an identifier. Using this identifier, the server retrieves explanatory data from a MySQL database, generates an explanatory text using GPT-4, and sends it to the device. The device then displays or plays an audio message stating, "This painting was painted by a famous Renaissance painter, and its distinguishing feature is the meticulous detail."

[1539] Examples of prompts for generative AI models

[1540] "Generate a descriptive text for a famous painting. The text should include the artist, date of creation, and key features."

[1541] The above details the embodiments of the present invention. This system allows users to obtain multifaceted information about exhibits and deepen their understanding of the exhibits interactively.

[1542] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1543] Step 1:

[1544] The user stands in front of the exhibit and uses the camera on their device to capture an image of the exhibit.

[1545] Input: Image of the exhibit

[1546] Output: Captured image data

[1547] Specific operation: The user launches a dedicated app on their smartphone or tablet, points the camera at the exhibit, and presses the shutter button. This causes the camera to capture image data and save it to the device's memory.

[1548] Step 2:

[1549] The terminal encodes the captured image data in BASE64 format and sends it to the server using an HTTP POST request.

[1550] Input: Captured image data

[1551] Output: Image data encoded in BASE64 format

[1552] Specific operation: The device's dedicated app encodes the captured image data into BASE64 format and uploads the image data to the server by sending an HTTP POST request to the specified API endpoint.

[1553] Step 3:

[1554] The server analyzes the received image data and generates identifiers corresponding to the identified exhibits.

[1555] Input: Image data in BASE64 format

[1556] Output: Identifier of the identified exhibit

[1557] Specific operation: The server uses image recognition technology (e.g., AWS Rekognition or Google Cloud Vision) to analyze the received image data. From the analysis results, it identifies the exhibits and generates corresponding identifiers.

[1558] Step 4:

[1559] The server retrieves relevant information data from the database based on the identified exhibit identifier.

[1560] Input: Exhibit identifier

[1561] Output: Information data about the exhibits

[1562] Specific operation: The server uses the generated exhibit identifier to query the database (e.g., MySQL or MongoDB) to retrieve information data about the corresponding exhibit.

[1563] Step 5:

[1564] Based on the acquired information data, the server uses a generative AI model (e.g., GPT-4) to perform natural language generation and generate explanatory text.

[1565] Input: Information data about the exhibits

[1566] Output: Generated explanatory text

[1567] Specific operation: The server analyzes the information data and inputs prompt sentences into the generation AI model to generate appropriate explanatory text.

[1568] Example prompt: "Generate a descriptive text for a famous painting. The text should include the artist, date of creation, and key features."

[1569] Step 6:

[1570] The server sends the generated explanatory text to the user's terminal, where it is displayed or played as audio.

[1571] Input: Generated explanatory text

[1572] Output: Explanatory text sent to the user's terminal

[1573] Specific operation: The server sends the generated explanatory text as an HTTP response to the user's terminal. The user's terminal either displays the received explanatory text on the screen or plays it aloud using speech synthesis technology (e.g., Amazon Polly or Google Text-to-Speech).

[1574] Step 7:

[1575] When a user enters a question about an exhibit, the user's terminal sends the question data to the server.

[1576] Input: Question entered by the user (voice or text)

[1577] Output: Question data sent to the server

[1578] Specific operation: The user asks a question via voice input or text input through the app on their device. In the case of voice input, the device converts the voice to text and sends it to the server as question data.

[1579] Step 8:

[1580] The server analyzes the received question data and generates appropriate answers using a generative AI model.

[1581] Input: Question data sent to the server

[1582] Output: Generated answer

[1583] Specific operation: The server analyzes the question data and prompts the generation AI model to generate appropriate answers. The generated answers are saved in text format.

[1584] Example prompt: "Generate answers to the following question: 'What is the background of this painting?'"

[1585] Step 9:

[1586] The server sends the generated response to the user's terminal, where it is displayed or played back as audio.

[1587] Input: Generated answer

[1588] Output: Response sent to the user's terminal

[1589] Specific operation: The server sends the generated response to the user's terminal as an HTTP response. The user's terminal displays the received response on the screen or plays it back as audio using speech synthesis technology.

[1590] Step 10:

[1591] When a user enters their thoughts and opinions about an exhibit, the user's device sends the feedback data to the server.

[1592] Input: User feedback (voice or text)

[1593] Output: Feedback data sent to the server

[1594] Specific operation: Users provide feedback via voice input or text input through the app on their device. In the case of voice input, the device converts the voice to text and sends it to the server as feedback data.

[1595] Step 11:

[1596] The server collects and organizes the received feedback data and stores it in a database.

[1597] Input: Feedback data sent to the server

[1598] Output: Feedback data stored in the database

[1599] Specific operation: The server collects and organizes the received feedback data, and then saves it to a database (e.g., MongoDB).

[1600] Step 12:

[1601] Based on the feedback data, the server uses a generative AI model to generate appropriate feedback.

[1602] Input: Feedback data stored on the server

[1603] Output: Generated feedback

[1604] Specific operation: The server analyzes the stored feedback data and prompts the AI ​​model to generate appropriate feedback. The generated feedback is saved in text format.

[1605] Step 13:

[1606] The server sends the generated feedback to the user's terminal, where it is displayed or played as audio.

[1607] Input: Generated feedback

[1608] Output: Feedback sent to the user terminal

[1609] Specific operation: The server sends the generated feedback to the user terminal as an HTTP response. The user terminal displays the received feedback on the screen or plays it back as audio using speech synthesis technology.

[1610] Step 14:

[1611] The server retrieves character information set for each exhibit and reflects the character's tone of voice and personality in the explanations and answers.

[1612] Input: Exhibit identifier

[1613] Output: Explanatory text and answer text reflecting character information

[1614] Specific operation: The server uses the exhibit identifier to retrieve character information from the database and reflects the character's tone of voice and personality in the explanatory text and response text.

[1615] Step 15:

[1616] The server sends explanatory text and response text based on character information to the user's terminal, and the terminal plays avatars and voice clips.

[1617] Input: Explanatory text or answer text that reflects character information

[1618] Output: Explanatory text and response text with character information sent to the user's terminal.

[1619] Specific operation: The server sends explanatory text and answer text with character information to the user terminal as an HTTP response. The user terminal displays the received content on the screen and plays the character's voice using speech synthesis technology.

[1620] The above outlines the specific processing steps of this system. It details a comprehensive mechanism for enhancing the user experience, from exhibit recognition and explanation to question-and-answer sessions, feedback, and even character entertainment.

[1621] (Application Example 1)

[1622] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1623] Traditional museum and art gallery systems for explaining exhibits are limited to providing information about the exhibits themselves, and do not adequately address users' questions or comments. Similarly, in physical stores and tourist destinations, there is a lack of detailed information about exhibits and products, as well as insufficient interactive communication with users. As a result, users' understanding and experience are limited, and the appeal of exhibits and products cannot be fully conveyed.

[1624] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1625] This invention includes means for a server to capture images using a camera installed in a terminal to recognize exhibits and transmit the image data to the server; means for the server to analyze the received image data and generate an identifier corresponding to the identified exhibit; means for the server to obtain explanatory data from a database based on the identified exhibit identifier and convert the explanatory data into a user-friendly format using natural language generation technology; means for transmitting the converted explanatory data to the user's terminal and displaying or playing it aloud on the terminal; means for inputting user questions into the terminal via voice or text; means for transmitting the question data entered on the terminal to the server; means for the server to analyze the received question data and generate appropriate answers using a generation AI model; means for transmitting the generated answers to the user's terminal and displaying or playing them aloud on the terminal; and means for providing information via the terminal when the user stands in front of a product or landmark. This allows the user to not only obtain detailed information about exhibits and products, but also to resolve questions in an interactive way and receive feedback on their impressions.

[1626] "Exhibits" is a general term for works of art and historical items displayed in art museums and museums, as well as objects that attract the interest of users in physical stores and tourist destinations.

[1627] A "device" refers to a portable electronic device that a user can carry with them, such as a smartphone, tablet, or personal computer, and is equipped with a camera, microphone, and display.

[1628] "Capturing an image using the camera" refers to saving visual information as a photograph to the device's internal memory using the device's built-in camera.

[1629] A "server" is a computer system that receives and processes data sent from a terminal and provides the necessary information.

[1630] "Image data" refers to digital files containing visual information captured by a camera.

[1631] An "identifier" is a code or number used by the server to uniquely identify a specific exhibit, and is used for database operations.

[1632] A "database" is an information management system for systematically storing explanatory data and other related information about exhibits.

[1633] "Natural language generation technology" refers to the technology that allows computers to generate text in human language that is easy to read and understand.

[1634] "User's device" refers to a device such as a smartphone or tablet used by the user, which displays or plays back information received from the server.

[1635] "Question data" refers to digital data that includes the content of questions entered by the user via voice or text.

[1636] A "generative AI model" is an artificial intelligence technology that automatically generates appropriate answers based on input information.

[1637] "Feedback" refers to responses and reactions to user comments and opinions, and is generated using natural language generation technology.

[1638] An "avatar" is an image used to provide users with a visual representation of a character.

[1639] This invention is a system that provides explanations to users about exhibits and products, and also responds to user questions and feedback. A specific embodiment of this system is shown below.

[1640] System Configuration

[1641] This system consists of terminals, a server, and a database. Terminals include smartphones, tablets, and personal computers. The server handles data processing and storage, as well as the execution of generated AI models. The database stores explanatory data about exhibits and products.

[1642] Hardware and software

[1643] Hardware: Smartphones, tablets, and personal computers (these electronic devices are equipped with cameras, microphones, and displays).

[1644] Software: Python, OpenCV (for capturing camera images), Requests (for sending and receiving data), Hugging Face's Transformers (using GPT-3 as a generative AI model).

[1645] Process flow

[1646] 1. Image capture

[1647] The user stands in front of an exhibit or product and captures an image using the device's camera. This image data is then sent from the device to the server.

[1648] 2. Image recognition and acquisition of explanatory data

[1649] The server analyzes the received image data and generates identifiers to identify exhibits and products. Based on these identifiers, the server retrieves explanatory data from a database. This explanatory data is converted into a user-friendly format using natural language generation technology. The converted explanatory data is then sent to the terminal and displayed or played as audio.

[1650] 3. Questions from users

[1651] If a user has questions about the explanation, they can ask them via voice input or text input. The device then sends this question data to the server.

[1652] 4. Generating answers to questions

[1653] The server analyzes the received question data and generates appropriate answers using a generative AI model (e.g., GPT-3). These generated answers are sent to the terminal and displayed or played aloud.

[1654] 5. User feedback and opinions

[1655] Users can input their thoughts and opinions about the exhibits and products. This can be done via voice input or text input. The terminal then sends the inputted feedback data to the server.

[1656] 6. Generating feedback on comments

[1657] The server collects and organizes the received feedback data and stores it in a database. Furthermore, it uses natural language generation technology to generate appropriate feedback based on the feedback. This feedback is sent to the device and displayed or played aloud.

[1658] Specific example

[1659] For example, suppose a user stands in front of a tourist attraction and points their device's camera at it to capture an image. This image data is sent to a server, which identifies the tourist attraction and generates explanatory data such as, "This attraction was built in the 19th century and is characterized by its architectural style." If the user inputs a question like, "I want to know more about the history of this attraction," the server uses a generative AI model to generate an answer such as, "The event of [event name] influenced the history of this attraction." Also, if the user inputs a comment like, "The design of this building is wonderful," the server generates feedback such as, "Many experts have also given its design high marks."

[1660] Example of a prompt

[1661] Examples of prompts for a generative AI model regarding exhibits include the following:

[1662] "Could you tell me about the history of this exhibit?"

[1663] "Please provide detailed information about this product."

[1664] In this way, users can obtain detailed information about exhibits and products, and receive interactive responses to their questions. This deepens the user's understanding and experience, resulting in a richer informational experience.

[1665] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1666] Step 1:

[1667] The user stands in front of the exhibit or product and captures an image using the camera on their device.

[1668] Specific operation:

[1669] Input: The object of the exhibit or product that the user points the camera at.

[1670] Data processing: Take an image with the device's camera and save it as an image file to the internal memory.

[1671] Output: Image data

[1672] Step 2:

[1673] The device sends the captured image data to the server.

[1674] Specific operation:

[1675] Input: Image data stored on the device

[1676] Data processing: Sending image data to a server using an HTTP request (e.g., POST request).

[1677] Output: Image data sent to the server

[1678] Step 3:

[1679] The server analyzes the received image data and generates identifiers to identify exhibits and products.

[1680] Specific operation:

[1681] Input: Image data sent to the server

[1682] Data processing: Use image recognition algorithms to identify exhibits and products within the image and generate their identifiers.

[1683] Output: Generated identifier

[1684] Step 4:

[1685] The server retrieves explanatory data from the database based on an identifier and converts it into a user-friendly format using natural language generation technology.

[1686] Specific operation:

[1687] Input: Generated identifier

[1688] Data processing: Query the database based on identifiers to retrieve relevant descriptive data. Then, use natural language generation techniques (e.g., NLG models) to generate text.

[1689] Output: Converted explanatory data

[1690] Step 5:

[1691] The terminal receives explanatory data sent from the server and displays or plays it as audio for the user.

[1692] Specific operation:

[1693] Input: Converted explanatory data

[1694] Data processing: Display explanatory data on the device's screen, or play it back as audio using speech synthesis technology.

[1695] Output: Displayed or audio-played commentary

[1696] Step 6:

[1697] Users have questions about the explanation and ask them using voice input or text input on their device.

[1698] Specific operation:

[1699] Input: User voice input or text input

[1700] Data processing: If using speech recognition technology, convert the audio to text (in the case of voice input). Then save the question as text data.

[1701] Output: Question data

[1702] Step 7:

[1703] The terminal sends the entered question data to the server.

[1704] Specific operation:

[1705] Input: Entered question data

[1706] Data processing: Send the question data to the server using an HTTP request.

[1707] Output: Question data sent to the server

[1708] Step 8:

[1709] The server analyzes the received question data and uses a generative AI model to generate appropriate answers.

[1710] Specific operation:

[1711] Input: Question data sent to the server

[1712] Data processing: Generative AI models (e.g., GPT-3) are used to generate appropriate answers to questions.

[1713] Output: Generated response data

[1714] Step 9:

[1715] The terminal receives the response data sent from the server and displays it to the user or plays it as audio.

[1716] Specific operation:

[1717] Input: Generated response data

[1718] Data processing: Display the response data on the device's screen, or play it back as audio using speech synthesis technology.

[1719] Output: Displayed or audio-played response

[1720] Step 10:

[1721] Users can input their thoughts and opinions about exhibits and products, providing feedback via voice or text input on their devices.

[1722] Specific operation:

[1723] Input: User voice input or text input

[1724] Data processing: Converting speech to text using speech recognition technology (in the case of voice input). Then, saving the comments as text data.

[1725] Output: Feedback data

[1726] Step 11:

[1727] The terminal sends the entered feedback data to the server.

[1728] Specific operation:

[1729] Input: Entered feedback data

[1730] Data processing: Send feedback data to the server using an HTTP request.

[1731] Output: Feedback data sent to the server

[1732] Step 12:

[1733] The server collects and organizes the feedback data it receives, stores it in a database, and uses natural language generation technology to generate appropriate feedback.

[1734] Specific operation:

[1735] Input: Feedback data sent to the server

[1736] Data processing: Save the feedback data to a database and generate feedback using natural language generation technology (e.g., NLG model).

[1737] Output: Generated feedback data

[1738] Step 13:

[1739] The device receives feedback data sent from the server and displays or plays it aloud for the user.

[1740] Specific operation:

[1741] Input: Generated feedback data

[1742] Data processing: Display feedback data on the device's screen or play it back as audio using speech synthesis technology.

[1743] Output: Displayed or audible feedback

[1744] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1745] ---

[1746] This invention relates to a system that uses generative AI to provide explanations of exhibits and respond to user questions and comments in art museums and museums. Furthermore, by combining this with an emotion engine that recognizes user emotions, the aim is to provide users with personalized explanations and feedback.

[1747] System Overview

[1748] Recognition and explanation of exhibits

[1749] When a user stands in front of an exhibit and launches an app on their device (e.g., a smartphone or tablet), the device's camera is used to capture an image of the exhibit. The captured image data is sent from the device to a server. The server analyzes the received image data and generates an ID to identify the exhibit. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent to the device and displayed or played aloud for the user.

[1750] User Q&A

[1751] If a user has a question about an exhibit, they input their question via voice or text into the terminal. The terminal sends the entered question data to a server, which uses generative AI to generate an appropriate answer to the question. The generated answer is sent to the terminal and displayed or played back to the user.

[1752] User feedback and comments

[1753] When a user inputs their thoughts and opinions into their device via voice or text, the device sends that data to a server. The server analyzes the received data, stores it in a database, and generates feedback based on that data using natural language generation technology. The generated feedback is sent to the device and displayed or played back to the user.

[1754] Enhanced entertainment value

[1755] Furthermore, this system retrieves character information set for each exhibit from a server and reflects the character's tone of voice and personality when generating explanatory texts and answers based on that information. The generated explanatory texts and answers are expressed using character avatars and voice clips, providing users with both visual and auditory experiences.

[1756] Combination of emotional engines

[1757] A distinctive feature of this invention is the integration of an emotion engine. The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. For example, using the user's camera video and audio data, the emotion engine detects emotions such as joy, interest, and dissatisfaction. The terminal sends this emotional data to a server, which analyzes the data to generate explanations, answers, or feedback that correspond to the user's emotions.

[1758] For example, if the emotion engine recognizes that a user is looking at an exhibit with interest, the server generates and sends to the user's device an explanation containing detailed information and interesting anecdotes that will pique the user's interest. On the other hand, if the system recognizes that the user is somewhat tired, it provides a concise and to-the-point explanation. This allows for a more personalized and satisfying user experience.

[1759] Specific example

[1760] For example, suppose a user stands in front of a "famous painting" and launches an app on their device. When the user points their camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses image recognition technology to identify the "famous painting" and generates an ID for it. Based on this ID, the server retrieves explanatory data from its database and generates an explanatory text, such as "This painting was painted in [era], and its characteristics are [characteristics]," and sends it to the device. The device then displays or plays the explanatory text for the user.

[1761] Next, when the user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses generative AI to generate an answer to the question and sends the answer, such as "The following events influenced the creation of this painting," to the device. The device then displays or plays the answer aloud to the user.

[1762] Furthermore, when a user enters a comment such as "The use of color in this painting is wonderful," the device sends the comment data to the server. The server stores the comment data in a database and generates feedback such as "Many experts have praised the use of color. We are pleased that you noticed this point," and sends it to the device. In this way, the user can receive feedback on their own comments.

[1763] Furthermore, by using the emotion engine, if a user says, "This painting is very moving," that emotion can be recognized, and specific feedback such as, "Thank you for being moved. It is said that the artist painted this picture with the same feelings," can be received.

[1764] The above is an example of implementing the system according to the present invention. In this way, it is possible to realize a more personalized user experience by using an emotion engine, in addition to providing explanations for exhibits.

[1765] ---

[1766] The following describes the processing flow.

[1767] ---

[1768] Program processing flow

[1769] Recognition processing of exhibits

[1770] Step 1:

[1771] The user stands in front of the exhibit and launches the app on their device (e.g., a smartphone or tablet).

[1772] Step 2:

[1773] The user uses the device's camera to capture an image of the exhibit. The device activates the camera, generates image data, and saves that image data to memory.

[1774] Step 3:

[1775] The device sends the captured image data to the server. Specifically, it sends the image data to the server via an HTTP request.

[1776] Step 4:

[1777] The server executes an image recognition algorithm to identify the received image data. For example, the server uses a Convolutional Neural Network (CNN) model to detect exhibits from the image and generate a corresponding exhibit ID.

[1778] Processing to retrieve explanations for exhibits

[1779] Step 5:

[1780] The server uses the generated exhibit ID to retrieve the corresponding explanatory data from the explanatory database. This involves extracting explanatory information from the database using SQL queries or similar methods.

[1781] Step 6:

[1782] The server uses natural language generation (NLG) technology to convert the acquired explanatory data into a user-friendly format. Specifically, the NLG engine generates the explanatory text.

[1783] Step 7:

[1784] The server sends the generated explanatory data to the terminal. It returns data containing the explanatory text to the terminal as an HTTP response.

[1785] Step 8:

[1786] The device displays the received explanatory data to the user. This is done using a text display UI or a speech synthesis engine to provide the user with audio explanations.

[1787] User question answering process

[1788] Step 9:

[1789] If a user has a question about an exhibit, they can input their question into the terminal either by voice or text. The terminal uses a speech recognition engine to convert the voice to text or accepts the text input.

[1790] Step 10:

[1791] The terminal sends question data in text format to the server. An HTTP request is used for this transmission.

[1792] Step 11:

[1793] The server analyzes the received question data. Natural language understanding (NLU) technology is used for the analysis to understand the user's question.

[1794] Step 12:

[1795] The server uses a generative AI (e.g., GPT-4) to generate answers to questions. The generated answers are structured as natural language text.

[1796] Step 13:

[1797] The server sends the generated response to the terminal. The response data is returned to the terminal as an HTTP response.

[1798] Step 14:

[1799] The device displays the answer to the user. This display uses a text display UI or a speech synthesis engine.

[1800] User feedback sharing process

[1801] Step 15:

[1802] Users input their impressions and opinions about the exhibits into the device via voice or text. The device uses a speech recognition engine to convert the voice to text or receives the text input.

[1803] Step 16:

[1804] The device sends feedback data in text format to the server. HTTP requests are used for transmission.

[1805] Step 17:

[1806] The server analyzes and organizes the received feedback data and saves it to a database. SQL queries are used to write the data to the database for saving.

[1807] Step 18:

[1808] The server generates feedback using natural language generation technology based on the feedback data. The NLG engine generates appropriate feedback sentences.

[1809] Step 19:

[1810] The server sends the generated feedback to the terminal. The feedback data is returned to the terminal as an HTTP response.

[1811] Step 20:

[1812] The device displays feedback to the user. This is done using a text display UI or a speech synthesis engine.

[1813] Processing to enhance entertainment value

[1814] Step 21:

[1815] The server retrieves character information set for each exhibit. It uses SQL queries to retrieve character information from the database.

[1816] Step 22:

[1817] When the server generates explanatory text and response text based on character information, it reflects the character's speech pattern and personality. An NLG engine is used for generation.

[1818] Step 23:

[1819] The server transmits the character's voice data and avatar information to the terminal. This includes synthesized voice data and image / video data.

[1820] Step 24:

[1821] The device provides explanations and answers to the user using character avatars and voice recordings. The device displays character avatars and plays explanations and answers via voice recordings.

[1822] Combination of emotional engines

[1823] Step 25:

[1824] The device uses its built-in camera and microphone to analyze the user's voice and facial expressions as they view exhibits and express their opinions. The device runs an emotion engine to analyze voice tone and facial expressions.

[1825] Step 26:

[1826] The device converts the analyzed emotion data into text format and sends it to the server. An HTTP request is used for transmission.

[1827] Step 27:

[1828] The server analyzes the received emotional data and generates explanations, responses, or feedback tailored to the user's emotional state. For example, if the user shows interest, it provides a detailed explanation; if the user is tired, it provides a concise explanation.

[1829] Step 28:

[1830] The server sends the generated emotion response explanation, answer, or feedback to the device. It returns the data as an HTTP response.

[1831] Step 29:

[1832] The device displays or plays the generated emotion-responsive data to the user. Text display UI or speech synthesis engines are used for display.

[1833] ---

[1834] (Example 2)

[1835] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1836] Traditional museum and art gallery exhibit explanation systems have been limited to providing static information about the exhibits, failing to address the individual interests and emotions of users. As a result, the user experience was not personalized, resulting in one-sided information delivery and low satisfaction. Furthermore, there was a lack of means to provide appropriate feedback to user questions and comments, often leading to a superficial user experience. In addition to these problems, the lack of entertainment value also needs improvement.

[1837] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing image data transmitted from a terminal to recognize an exhibit and generating an ID corresponding to the identified exhibit; means for converting explanatory data obtained from a database based on the identified exhibit ID into a user-friendly format using natural language generation technology; means for transmitting the generated explanatory data to the user's terminal; means for analyzing questions from the user and generating appropriate answers using a generation AI model; means for analyzing the user's facial expressions and voice tone using an emotion engine to recognize the user's emotions; and means for generating explanations, answers, or feedback that correspond to the user's emotions based on the analyzed emotion data. This makes it possible to provide users with personalized explanations and feedback, thereby achieving higher satisfaction.

[1838] "Exhibits" refer to visual or physical objects such as paintings, sculptures, crafts, and historical artifacts displayed in art museums and museums.

[1839] "Device" refers to a smartphone, tablet, or other portable device used by a user for operation.

[1840] A "camera" is a photographic device attached to a terminal for capturing images of exhibits.

[1841] A "server" refers to a computer system that performs processing such as analyzing received data and generating responses using generative AI.

[1842] "Image data" refers to photographic and video data of exhibits captured using the device's camera.

[1843] An "ID" is a unique identifier generated to identify an exhibit.

[1844] A "database" is a storage area where various types of data, such as explanatory data, character information, and user reviews, are stored.

[1845] "Explanatory data" refers to text data and audio data that include information and explanations related to the exhibits.

[1846] "Natural Language Generation (NLG)" refers to the technology used to convert machine-generated data into human language.

[1847] A "generative AI model" is an artificial intelligence model used to generate appropriate answers or feedback to questions.

[1848] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice tone to detect their emotional state.

[1849] A "prompt message" is text data input to a generative AI model, and it is an instruction message that prompts the model to generate an answer.

[1850] An "avatar" refers to a visual representation generated based on character information, which is typically displayed on the user's device.

[1851] This invention relates to a system that uses generative AI to provide explanations of exhibits and respond to user questions and comments in art museums and museums. Furthermore, by combining this with an emotion engine that recognizes user emotions, the aim is to provide users with personalized explanations and feedback.

[1852] System Overview

[1853] This system consists of the following main elements:

[1854] Recognition and explanation of exhibits

[1855] The user stands in front of an exhibit and uses a smartphone or tablet to capture an image of the exhibit with its camera. The device then sends the captured image data to a server. The server analyzes the received image using image recognition technology (e.g., TensorFlow or OpenCV) and generates an ID to identify the exhibit. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent to the device and displayed or played aloud for the user.

[1856] Specific example

[1857] For example, if a user stands in front of a "famous painting" and points their device's camera at it to capture an image, the device sends the image to a server. The server uses OpenCV to extract features from the image and identify the specific exhibit. Based on this ID, it generates an explanatory text from its database, such as "This painting was painted by [era]," and sends it to the device. The device then plays the explanatory text aloud for the user to listen to.

[1858] User Q&A

[1859] If a user wants to ask a question about an exhibit, they can input it via voice or text into the terminal. The terminal sends the entered question data to a server, which uses a generative AI model (e.g., GPT-3 or GPT-4) to generate an appropriate answer to the question. The generated answer is sent to the terminal and displayed or played back to the user.

[1860] Specific example

[1861] When a user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses a generative AI model to generate an answer, such as "The following events influenced the creation of this painting," and sends it to the device. The device then displays the answer for the user to see.

[1862] User feedback and comments

[1863] When a user inputs their thoughts and opinions into their device via voice or text, the device sends that data to a server. The server analyzes the received data, stores it in a database, and generates feedback based on that data. The generated feedback is sent to the device and displayed or played back to the user.

[1864] Specific example

[1865] When a user enters a comment such as "The colors in this painting are wonderful," the device sends the comment data to the server. The server saves the comment data and generates feedback such as "Many experts have praised this. We are glad you noticed this point," and sends it to the device. The device then plays the feedback text aloud for the user to hear.

[1866] Combination of emotional engines

[1867] The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. Using the user's camera video and audio data, the emotion engine detects emotions such as joy, interest, and dissatisfaction. The device sends this emotional data to a server, which generates explanations, answers, or feedback tailored to the user's emotions based on the analysis results.

[1868] Specific example

[1869] If a user says, "This painting is very moving," the emotion engine recognizes that emotion and generates specific feedback such as, "Thank you for being moved. It is said that the artist painted this picture with the same feelings," and sends it to the device. The device then plays the feedback aloud and delivers it to the user.

[1870] Example of a prompt

[1871] 1. Prompt for generating detailed descriptions of exhibits

[1872] Please generate detailed descriptions for the following exhibits.

[1873] Exhibit ID: ABC123

[1874] Era: Ancient

[1875] Author: Author A

[1876] Features: Vibrant brushwork

[1877] 2. Question and answer prompt sentences

[1878] Please generate answers to the following questions.

[1879] Question: What is the background behind the creation of this painting?

[1880] Exhibit ID: ABC123"

[1881] 3. Prompt for feedback / comments

[1882] Please generate feedback for the following comments.

[1883] Impression: The use of color in this painting is wonderful.

[1884] Exhibit ID: ABC123"

[1885] 4. Feedback prompts based on emotion recognition

[1886] Please generate feedback based on the following emotions:

[1887] Emotion: I'm moved.

[1888] Comment: This painting is very moving.

[1889] Exhibit ID: ABC123"

[1890] In this way, the present invention makes it possible to provide personalized explanations of exhibits in art museums and museums, and to offer an immersive experience that responds to the user's emotions.

[1891] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1892] System program processing flow

[1893] Recognition and explanation of exhibits

[1894] Step 1:

[1895] The user stands in front of the exhibit.

[1896] Input: User's location, exhibits

[1897] Output: Captured camera connection

[1898] Specific actions: The user points their smartphone or tablet camera at the exhibit and launches a dedicated app.

[1899] Step 2:

[1900] The device captures images of the exhibits.

[1901] Input: Exhibit images, camera

[1902] Output: Captured image data

[1903] Specific operation: The device's camera takes a picture of the exhibit and generates image data.

[1904] Step 3:

[1905] The device sends the captured image data to the server.

[1906] Input: Captured image data

[1907] Output: Data sent to the server after completion

[1908] Specific operation: The terminal sends image data to the server using an open communication protocol.

[1909] Step 4:

[1910] The server analyzes the image data.

[1911] Input: Received image data

[1912] Output: Exhibit identification ID

[1913] Specific operation: The server analyzes image data using TensorFlow and OpenCV to extract and identify features of the exhibits.

[1914] Step 5:

[1915] The server retrieves and generates explanatory data from the database based on the identified exhibit ID.

[1916] Input: Exhibit identification ID

[1917] Output: Explanatory data, generated explanatory text

[1918] Specific operation: The server retrieves relevant explanatory data from the explanatory database and uses natural language generation (NLG) technology to generate an easy-to-understand explanatory text for the user.

[1919] Step 6:

[1920] The server sends the generated explanatory text to the terminal.

[1921] Input: Generated explanatory text

[1922] Output: Transmission complete data

[1923] Specific operation: The server sends the generated explanatory text to the terminal.

[1924] Step 7:

[1925] The device displays or plays explanatory text.

[1926] Input: Received explanatory text

[1927] Output: Displayed or audio-played explanatory text

[1928] Specific operation: The device will either display the received explanatory text on the screen or play it as audio.

[1929] User Q&A

[1930] Step 1:

[1931] The user enters a question.

[1932] Input: Questions via voice or text input

[1933] Output: Question data

[1934] Specific operation: The user enters the question using voice input or text input on the device.

[1935] Step 2:

[1936] The device sends the question data to the server.

[1937] Input: Question data

[1938] Output: Data sent to the server after completion

[1939] Specific action: The terminal sends the question data to the server.

[1940] Step 3:

[1941] The server analyzes the question data and generates answers.

[1942] Input: Question data

[1943] Output: Generated answer

[1944] Specific operation: The server analyzes the question data and generates appropriate answers using a generative AI model (e.g., GPT-3 or GPT-4).

[1945] Step 4:

[1946] The server sends the generated response to the terminal.

[1947] Input: Generated answer

[1948] Output: Transmission complete data

[1949] Specific operation: The server sends the generated response to the terminal.

[1950] Step 5:

[1951] The device displays or plays the answer.

[1952] Input: Received response

[1953] Output: Displayed or audio-played answer

[1954] Specific actions: The device will either display the received response on the screen or play it as audio.

[1955] User feedback and comments

[1956] Step 1:

[1957] Users enter their feedback.

[1958] Input: Feedback via voice or text input

[1959] Output: Feedback data

[1960] Specific operation: The user enters their feedback by using voice input or text input on the device.

[1961] Step 2:

[1962] The device sends feedback data to the server.

[1963] Input: Feedback data

[1964] Output: Data sent to the server after completion

[1965] Specific action: The device sends feedback data to the server.

[1966] Step 3:

[1967] The server analyzes and stores the feedback data.

[1968] Input: Received feedback data

[1969] Output: Saved data, analyzed feedback data

[1970] Specific operation: The server analyzes the feedback data and saves it to the database.

[1971] Step 4:

[1972] The server generates feedback based on the feedback data.

[1973] Input: Analyzed feedback data

[1974] Output: Generated feedback

[1975] Specific operation: The server uses natural language generation technology to generate feedback based on the feedback data.

[1976] Step 5:

[1977] The server sends the generated feedback to the terminal.

[1978] Input: Generated feedback

[1979] Output: Transmission complete data

[1980] Specific operation: The server sends the generated feedback to the terminal.

[1981] Step 6:

[1982] The device displays or plays feedback.

[1983] Input: Received feedback

[1984] Output: Displayed or audible feedback

[1985] Specific actions: The device displays the received feedback on the screen or plays it back as audio.

[1986] Combination of emotional engines

[1987] Step 1:

[1988] Capture user sentiment data

[1989] Input: User's facial expressions and voice

[1990] Output: Captured sentiment data

[1991] Specific operation: The device's camera and microphone capture the user's facial expressions and voice.

[1992] Step 2:

[1993] The device sends emotional data to the server.

[1994] Input: Captured emotion data

[1995] Output: Data sent to the server after completion

[1996] Specific operation: The device sends the captured emotion data to the server.

[1997] Step 3:

[1998] The server analyzes the emotional data.

[1999] Input: Received sentiment data

[2000] Output: Analyzed emotional state

[2001] Specific operation: The server analyzes the emotional state using an emotion engine (e.g., Affectiva SDK or Microsoft Azure Emotion API).

[2002] Step 4:

[2003] The server generates content based on emotions.

[2004] Input: Analyzed emotional state

[2005] Output: Generated emotion-responsive content

[2006] Specific operation: Based on the analyzed sentiment data, the server generates explanations, responses, or feedback that correspond to the user's emotions.

[2007] Step 5:

[2008] The server sends the generated content to the terminal.

[2009] Input: Generated emotion-responsive content

[2010] Output: Transmission complete data

[2011] Specific operation: The server sends the generated emotion-responsive content to the device.

[2012] Step 6:

[2013] The device displays content or plays audio.

[2014] Input: Received emotion-responsive content

[2015] Output: Displayed or played content

[2016] Specific actions: The device will either display the received content on the screen or play it as audio.

[2017] The above outlines the specific flow of the system's program processing. This allows users to receive personalized explanations and feedback on the exhibits, resulting in a more satisfying experience.

[2018] (Application Example 2)

[2019] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[2020] Traditional systems for providing explanations of exhibits and products could only offer the same information at all times, failing to provide an experience tailored to the individual interests and emotions of users. Furthermore, when users asked detailed questions about specific products or exhibits, they often sought more comprehensive information, but the responses were limited to general and consistent answers. As a result, user satisfaction decreased, and problems arose where purchasing intent and learning motivation were not sufficiently stimulated.

[2021] The identification process performed by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing received image data and generating an ID corresponding to the identified exhibit; means for obtaining explanatory data from a database based on the identified exhibit ID and converting the explanatory data into a user-friendly format using natural language generation technology; and means for generating an appropriate response using a generative AI. This makes it possible not only to provide users with detailed information about exhibits and products, but also to provide personalized explanations and feedback that are tailored to the user's emotional state.

[2022] The term "exhibits" refers to all objects that customers are interested in, not only in art museums and museums, but also in physical stores.

[2023] "Terminal" refers to a smartphone, tablet, personal computer, or other portable information processing device.

[2024] A "camera" refers to an optical device used to capture images and videos.

[2025] "Image data" refers to a digital representation of a captured image.

[2026] A "server" refers to a computer system used to receive, process, and transmit data over a network.

[2027] "ID" refers to a code or label used to uniquely identify a specific exhibit or product.

[2028] A "database" refers to a system or software used to organize and store information.

[2029] "Natural language generation technology" refers to the technology that uses computers to generate natural-sounding words and sentences.

[2030] "Generative AI" refers to artificial intelligence technology that generates appropriate responses or new data based on input data.

[2031] "Voice input" refers to a method of recognizing human speech as digital information using a microphone.

[2032] "Text input" refers to the method of entering character data using a keyboard or on-screen keyboard.

[2033] "User" refers to all individuals who use the system.

[2034] An "emotion engine" refers to software or hardware that analyzes a user's facial expressions and tone of voice to detect their psychological state.

[2035] "Feedback" refers to the response or reaction that a system gives to a user's actions or input.

[2036] "Character information" refers to digital data or profiles used to add a specific personality or tone of voice to explanatory texts or answer texts.

[2037] This invention relates to a system that provides customers with detailed information about exhibits and products in physical stores, and further provides personalized explanations and feedback that respond to the customer's emotions. Specific embodiments of this invention are shown below.

[2038] System Overview

[2039] This system includes the following main components:

[2040] terminal

[2041] A terminal is a portable information processing device such as a smartphone or tablet. The terminal is equipped with a camera, microphone, display, and speaker. This allows users to photograph products and exhibits and interact with the system through voice or text input.

[2042] server

[2043] The server communicates with terminals over the network and processes the received data. The server includes multiple modules for image recognition, natural language generation technology, and generative AI processing.

[2044] Image capture and recognition

[2045] The user uses the device's camera to capture images of products or exhibits. The captured image data is sent from the device to a server, where the server's image recognition module analyzes it. An ID corresponding to the identified exhibit is generated from the analysis results. This ID is used to retrieve explanatory data from the database.

[2046] Question and Answer

[2047] When a user sends a question about an exhibit to their device via voice or text input, the question data is transferred to a server. The server uses generative AI to generate an appropriate answer and sends it to the device. The device then displays or plays this answer aloud.

[2048] Emotional engine and feedback

[2049] Furthermore, this system is equipped with an emotion engine that analyzes the user's emotions using camera and audio data. Once the user's emotional state is recognized, that emotional data is sent to the server. The server generates explanations and feedback based on the emotional data, providing a personalized experience. This feedback is generated based on information obtained from the database and the user's emotional state.

[2050] Specific example

[2051] For example, suppose a user is in a physical store and stands before the latest smartphone, then launches an app on the device. The user takes a picture of the smartphone using the camera and sends the image data to a server. The server uses image recognition technology to identify the smartphone and generates an ID. Based on this ID, it retrieves a detailed explanation from a database and uses natural language generation technology to generate an easy-to-understand explanatory text for the user. This explanatory text is sent to the device and is either displayed on the screen or played as audio.

[2052] When a user asks a question by voice, such as "How is the camera function of this smartphone?", the device sends the question data to a server. The server uses generative AI to generate a detailed answer and sends it to the device, such as "This smartphone's camera has a 12-megapixel sensor and excellent night vision capabilities." The device then displays or plays this answer aloud to the user.

[2053] Furthermore, the emotion engine, which analyzes the user's facial expressions and tone of voice, recognizes when the user is excited. For example, when a user says, "This smartphone is amazing!", based on that emotion data, it generates feedback such as, "I'm glad you're pleased! Shall I tell you about other features?"

[2054] Example of a prompt

[2055] Examples of prompt statements are as follows:

[2056] User emotion: Joy

[2057] Question: What features does this washing machine have?

[2058] answer:

[2059] This prompt is input into a generative AI model, which generates a response tailored to the user's emotion of "joy."

[2060] The above describes a specific embodiment of the system according to the present invention. In this way, it becomes possible to provide personalized explanations and feedback that respond to the user's emotions.

[2061] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[2062] Step 1:

[2063] The user captures images of products or exhibits using the device's camera. The input is video from the device's camera, and the output is the captured image data. This image data is temporarily stored on the device.

[2064] Step 2:

[2065] The terminal sends the captured image data to the server. Specifically, it uploads the image data to the server via the network connection by including it in the request packet. The captured image data is the input, and the server receives the data as output.

[2066] Step 3:

[2067] The server analyzes the received image data and uses image recognition technology to identify specific exhibits or products. It retrieves the ID of the identified object from the database and generates a dedicated code or label. Image data is the input, and the identified ID is generated as the output.

[2068] Step 4:

[2069] The server retrieves detailed explanatory data from the database based on the identified ID. Next, it uses natural language generation technology to convert the explanatory data into a user-friendly format. The input consists of the identified ID and associated explanatory data, and the output is the converted explanatory text.

[2070] Step 5:

[2071] The server sends the converted explanatory data to the terminal. Specifically, it sends packet data containing the generated explanatory text to the terminal via the network. The input is the explanatory text converted into an easy-to-understand format, and the output is the explanatory text sent to the terminal.

[2072] Step 6:

[2073] The terminal displays or plays explanatory text for the user. The input is explanatory text received from the server, and the output is either text displayed on the screen or audio played through the speaker.

[2074] Step 7:

[2075] Users submit questions about exhibits via voice or text input. The input is the user's voice or text, and the output is the voice recognition or text input data, which is stored on the device.

[2076] Step 8:

[2077] The terminal sends the entered question data to the server. Specifically, it uploads the question data to the server using a network connection. The user's question data is the input, and the server receives the data as output.

[2078] Step 9:

[2079] The server analyzes the received question data and generates appropriate answers using generative AI. The input is the question data, and the output is the generated answers.

[2080] Step 10:

[2081] The server sends the generated response to the terminal. Specifically, it sends packet data containing the generated response to the terminal over the network. The generated response is the input, and the response sent to the terminal is the output.

[2082] Step 11:

[2083] The device displays or plays the answer to the user. The input is the answer received from the server, and the output is either text displayed on the screen or audio played through the speaker.

[2084] Step 12:

[2085] To recognize the user's emotions, the device's camera and microphone are used to capture facial expressions and voice tone. Camera video and audio data are taken as input, and emotion data is generated as output.

[2086] Step 13:

[2087] The device sends emotional data to the server. Specifically, it sends packets containing emotional data to the server over the network. The emotional data is the input, and the server receives the data as the output.

[2088] Step 14:

[2089] The server analyzes emotional data and generates explanations and feedback tailored to the user's emotions. It takes emotional data and related explanation data as input, and generates personalized feedback as output.

[2090] Step 15:

[2091] The server sends personalized feedback to the terminal. Specifically, it sends packet data containing the generated feedback to the terminal over the network. The input is the personalized feedback, and the output is the feedback sent to the terminal.

[2092] Step 16:

[2093] The device displays or plays audio feedback to the user. Input is feedback received from a server, and output is text displayed on the screen or audio played through the speaker.

[2094] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[2095] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2096] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[2097] [Fourth Embodiment]

[2098] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[2099] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[2100] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[2101] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[2102] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[2103] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[2104] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[2105] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[2106] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[2107] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[2108] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[2109] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[2110] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2111] ---

[2112] This invention relates to a system that uses generative AI to provide explanations to users about exhibits in art museums and museums, and to respond to user questions and comments. This system comprehensively performs exhibit recognition, acquisition and display of explanatory data, response to user questions, and collection and feedback of comments.

[2113] System Overview

[2114] First, when a user stands in front of an exhibit, they use a device (such as a smartphone or tablet) to capture an image of the exhibit. The device generates image data using its camera and sends this data to a server. The server uses image recognition technology to identify the exhibit and generates an ID. Based on this ID, the server retrieves explanatory data from a database and generates an explanatory text using natural language generation (NLG) technology. This explanatory text is sent from the server to the device and displayed or played as audio for the user on the device.

[2115] If a user has a question about an exhibit, they enter the question into the terminal via voice or text input. The terminal sends the entered question data to a server, which uses generative AI to generate an appropriate answer to the question. This answer is then sent to the terminal and displayed or played back to the user.

[2116] Furthermore, when users want to input their thoughts and opinions, they do so via voice or text on the device. The device sends the input data to the server, which stores the received data in a database and generates appropriate feedback, which is then sent back to the device. This allows users to receive feedback on their own comments.

[2117] The system also includes features to enhance entertainment value. The server retrieves character information set for each exhibit and reflects the character's tone and personality in the explanations and answers. The generated explanations and answers are played back using the character's avatar and voice, providing users with both visual and auditory experiences.

[2118] Specific example

[2119] For example, suppose a user stands in front of a "famous painting" and launches an app on their device. When the user points their camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses image recognition technology to identify the "famous painting" and generates an ID. Using this ID, the server retrieves explanatory data from a database, generates an explanatory text such as "This painting was painted in [era], and its characteristics are [characteristics]," and sends it to the device. The device then displays or plays the explanatory text for the user.

[2120] Next, when the user enters the question, "What is the background behind the creation of this painting?", the device sends the question data to the server. The server uses generative AI to generate an answer to the question and sends the answer, such as "The following events influenced the creation of this painting," to the device. The device then displays or plays the answer aloud to the user.

[2121] Furthermore, when a user enters a comment such as "The use of color in this painting is wonderful," the device sends the comment data to the server. The server stores the comment data in a database and generates feedback such as, "Many experts have praised the use of color. We are pleased that you noticed this point," and sends it back to the device. This allows the user to gain a deeper understanding and greater satisfaction.

[2122] The above is an example of implementing the system according to the present invention. This makes it possible to provide multifaceted and individualized information about exhibits, thereby enhancing the user experience.

[2123] ---

[2124] The following describes the processing flow.

[2125] ---

[2126] Program processing flow

[2127] Recognition processing of exhibits

[2128] Step 1:

[2129] The user stands in front of the exhibit and launches the app on their device (e.g., a smartphone or tablet).

[2130] Step 2:

[2131] The device's camera is used to capture images of the exhibits. The device activates the camera, generates image data, and saves that image data to memory.

[2132] Step 3:

[2133] The device sends the captured image data to the server. Specifically, it sends the image data to the server via an HTTP request.

[2134] Step 4:

[2135] The server executes an image recognition algorithm to identify the received image data. For example, the server uses a Convolutional Neural Network (CNN) model to detect exhibits from the image and generate a corresponding exhibit ID.

[2136] Processing to retrieve explanations for exhibits

[2137] Step 5:

[2138] The server uses the generated exhibit ID to retrieve the corresponding explanatory data from the explanatory database. This involves extracting explanatory information from the database using SQL queries or similar methods.

[2139] Step 6:

[2140] The server uses natural language generation (NLG) technology to convert the acquired explanatory data into a user-friendly format. Specifically, the NLG engine generates the explanatory text.

[2141] Step 7:

[2142] The server sends the generated explanatory data to the terminal. It returns data containing the explanatory text to the terminal as an HTTP response.

[2143] Step 8:

[2144] The device displays the received explanatory data to the user. This is done using a text display UI or a speech synthesis engine to provide the user with audio explanations.

[2145] User question answering process

[2146] Step 9:

[2147] If a user has a question about an exhibit, they can input their question into the terminal either by voice or text. The terminal uses a speech recognition engine to convert the voice to text or accepts the text input.

[2148] Step 10:

[2149] The terminal sends question data in text format to the server. An HTTP request is used for this transmission.

[2150] Step 11:

[2151] The server analyzes the received question data. Natural language understanding (NLU) technology is used for the analysis to understand the user's question.

[2152] Step 12:

[2153] The server uses a generative AI (e.g., GPT-4) to generate answers to questions. The generated answers are structured as natural language text.

[2154] Step 13:

[2155] The server sends the generated response to the terminal. The response data is returned to the terminal as an HTTP response.

[2156] Step 14:

[2157] The device displays the answer to the user. This display uses a text display UI or a speech synthesis engine.

[2158] User feedback sharing process

[2159] Step 15:

[2160] Users input their impressions and opinions about the exhibits into the device via voice or text. The device uses a speech recognition engine to convert the voice to text or receives the text input.

[2161] Step 16:

[2162] The device sends feedback data in text format to the server. HTTP requests are used for transmission.

[2163] Step 17:

[2164] The server analyzes and organizes the received feedback data and saves it to a database. SQL queries are used to write the data to the database for saving.

[2165] Step 18:

[2166] The server generates feedback using natural language generation technology based on the feedback data. The NLG engine generates appropriate feedback sentences.

[2167] Step 19:

[2168] The server sends the generated feedback to the terminal. The feedback data is returned to the terminal as an HTTP response.

[2169] Step 20:

[2170] The device displays feedback to the user. This is done using a text display UI or a speech synthesis engine.

[2171] Processing to enhance entertainment value

[2172] Step 21:

[2173] The server retrieves character information set for each exhibit. It uses SQL queries to retrieve character information from the database.

[2174] Step 22:

[2175] When the server generates explanatory text and response text based on character information, it reflects the character's speech pattern and personality. An NLG engine is used for generation.

[2176] Step 23:

[2177] The server transmits the character's voice data and avatar information to the terminal. This includes synthesized voice data and image / video data.

[2178] Step 24:

[2179] The device provides explanations and answers to the user using character avatars and voice recordings. The device displays character avatars and plays explanations and answers via voice recordings.

[2180] ---

[2181] The above is the specific processing flow of the system program.

[2182] (Example 1)

[2183] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2184] Museums and art galleries need efficient and effective ways to provide explanations about their exhibits to users. Furthermore, there is a lack of systems that can quickly and appropriately respond to users' questions and comments about the exhibits, thereby providing a richer experience. Traditional systems primarily rely on paper-based descriptions for each exhibit, lacking interactivity and making it difficult to capture user interest. Moreover, the lack of entertainment-oriented exhibit explanations utilizing characters limits the potential for improving the user experience.

[2185] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[2186] In this invention, the server includes means for capturing images using a camera device provided on the user terminal to identify exhibits and transmitting the image data to the server; means for analyzing the received image data on the server and generating an identifier corresponding to the identified exhibit; means for obtaining information data from a database based on the identified exhibit identifier on the server and converting the information data into a user-friendly format using natural language generation technology; means for transmitting the converted information data to the user terminal and displaying or playing it aloud on the user terminal; means for inputting user questions into the user terminal via voice or text input; means for transmitting the question data entered on the user terminal to the server; means for analyzing the received question data on the server and generating an appropriate answer using a generation AI model; and means for transmitting the generated answer to the user terminal. The system includes means for displaying or playing audio on the user terminal, means for inputting user feedback and opinions via voice or text into the user terminal, means for sending feedback data entered on the user terminal to a server, means for the server to collect and organize the received feedback data and store it in a database, means for generating appropriate feedback using natural language generation technology based on the feedback data, means for sending the generated feedback to the user terminal and displaying or playing audio on the user terminal, means for obtaining character information set for each exhibit from the server and reflecting the character's tone of voice and personality in the explanations and answers, means for sending data to the user terminal to express the character-based explanations and answers in the character's voice, and means for providing the character to the user as an avatar or voice on the user terminal. This makes it possible to provide users with interactive and rich explanations and responses to exhibits, as well as a highly entertaining experience.

[2187] "Exhibits" refer to works of art and materials that are made available to users in art museums and museums.

[2188] "User terminal" refers to a portable information device owned by a user, primarily such as a smartphone or tablet.

[2189] "Image capture equipment" refers to cameras and other optical instruments used to acquire image data.

[2190] A "server" is a central computing system that processes, stores, and provides data over a network.

[2191] "Image data" refers to data that represents captured visual information in a digital format.

[2192] "Analysis" refers to the process of examining received images and data in detail to extract information.

[2193] An "identifier" is a code or number generated to uniquely identify a particular exhibit.

[2194] A "database" is a system for efficiently storing, searching, and managing large amounts of information.

[2195] "Information data" refers to digital data that includes detailed explanations and information about the exhibits.

[2196] "Natural language generation technology" refers to technology that allows machines to automatically generate sentences that are easy for humans to understand.

[2197] "Question data" refers to data that digitally represents questions about exhibits entered by users.

[2198] A "generative AI model" refers to an algorithm or model that uses artificial intelligence to generate appropriate answers or text.

[2199] "Feedback" refers to the responses or evaluations provided in response to user input.

[2200] "Character information" refers to data that includes specific characteristics and speech patterns of fictional people, animals, and other elements set for each exhibit.

[2201] An "avatar" is an image or model used to visually represent a character.

[2202] "Voice" refers to data used to digitally reproduce the character's voice.

[2203] This invention relates to a system that uses generative AI to provide explanations to users about exhibits in art museums and museums, and to respond to user questions and comments. This system comprehensively performs exhibit recognition, acquisition and display of explanatory data, response to user questions, and collection and feedback of comments.

[2204] System Overview

[2205] This system works as follows:

[2206] First, when a user stands in front of an exhibit, they use their device (e.g., a smartphone or tablet) to capture an image of the exhibit. The device's camera generates image data, which is then sent to a server. The server uses image recognition technology to identify the exhibit and generates an identifier. Based on this identifier, the server retrieves relevant information data from a database and generates an explanatory text using natural language generation (NLG) technology. The generated explanatory text is sent from the server to the device and displayed or played back as audio for the user.

[2207] Hardware and software details

[2208] 1. Terminal (User terminal):

[2209] Hardware used: Smartphone, tablet

[2210] Software used: Dedicated app, camera function, audio recording and playback function

[2211] 2. Server:

[2212] Software used:

[2213] Image recognition technology: AWS Rekognition, Google Cloud Vision

[2214] Database: MySQL, MongoDB

[2215] Natural language generation technology (NLG): OpenAI GPT-4

[2216] Audio playback technology: Amazon Polly, Google Text-to-Speech

[2217] Processing details

[2218] The processing flow at each step is as follows:

[2219] 1. User-captured images of exhibits:

[2220] Users point their device's camera at the exhibit and capture images using a dedicated app.

[2221] 2. Sending image data from the terminal to the server:

[2222] The terminal encodes the captured image data in BASE64 format and sends it to the server using an HTTP POST request.

[2223] 3. Server-based recognition of exhibits and acquisition of data:

[2224] The server analyzes the received image data, using AWS Rekognition or Google Cloud Vision to recognize exhibits and generate identifiers. Based on the identified exhibit identifiers, the server retrieves relevant information data from MySQL or MongoDB databases.

[2225] 4. Server generates explanatory text and sends it to the terminal:

[2226] The server uses OpenAI's GPT-4 to generate natural language based on the acquired data. The generated explanatory text is sent from the server to the terminal and displayed or played back as audio for the user.

[2227] Specific usage scenarios

[2228] For example, suppose a user stands in front of a "famous painting" and launches a smartphone app. When the user points the camera at the "famous painting" and captures an image, the device sends the image data to a server. The server uses AWS Rekognition to identify the "famous painting" and generates an identifier. Using this identifier, the server retrieves explanatory data from a MySQL database, generates an explanatory text using GPT-4, and sends it to the device. The device then displays or plays an audio message stating, "This painting was painted by a famous Renaissance painter, and its distinguishing feature is the meticulous detail."

[2229] Examples of prompts for generative AI models

[2230] "Generate a descriptive text for a famous painting. The text should include the artist, date of creation, and key features."

[2231] The above details the embodiments of the present invention. This system allows users to obtain multifaceted information about exhibits and deepen their understanding of the exhibits interactively.

[2232] The flow of the specific processing in Example 1 will be explained using Figure 11.

[2233] Step 1:

[2234] The user stands in front of the exhibit and uses the camera on their device to capture an image of the exhibit.

[2235] Input: Image of the exhibit

[2236] Output: Captured image data

[2237] Specific operation: The user launches a dedicated app on their smartphone or tablet, points the camera at the exhibit, and presses the shutter button. This causes the camera to capture image data and save it to the device's memory.

[2238] Step 2:

[2239] The terminal encodes the captured image data in BASE64 format and sends it to the server using an HTTP POST request.

[2240] Input: Captured image data

[2241] Output: Image data encoded in BASE64 format

[2242] Specific operation: The device's dedicated app encodes the captured image data into BASE64 format and uploads the image data to the server by sending an HTTP POST request to the specified API endpoint.

[2243] Step 3:

[2244] The server analyzes the received image data and generates identifiers corresponding to the identified exhibits.

[2245] Input: Image data in BASE64 format

[2246] Output: Identifier of the identified exhibit

[2247] Specific operation: The server uses image recognition technology (e.g., AWS Rekognition or Google Cloud Vision) to analyze the received image data. From the analysis results, it identifies the exhibits and generates corresponding identifiers.

[2248] Step 4:

[2249] The server retrieves relevant information data from the database based on the identified exhibit identifier.

[2250] Input: Exhibit identifier

[2251] Output: Information data about the exhibits

[2252] Specific operation: The server uses the generated exhibit identifier to query the database (e.g., MySQL or MongoDB) to retrieve information data about the corresponding exhibit.

[2253] Step 5:

[2254] Based on the acquired information data, the server uses a generative AI model (e.g., GPT-4) to perform natural language generation and generate explanatory text.

[2255] Input: Information data about the exhibits

[2256] Output: Generated explanatory text

[2257] Specific operation: The server analyzes the information data and inputs prompt sentences into the generation AI model to generate appropriate explanatory text.

[2258] Example prompt: "Generate a descriptive text for a famous painting. The text should include the artist, date of creation, and key features."

[2259] Step 6:

[2260] The server sends the generated explanatory text to the user's terminal, where it is displayed or played as audio.

[2261] Input: Generated explanatory text

[2262] Output: Explanatory text sent to the user's terminal

[2263] Specific operation: The server sends the generated explanatory text as an HTTP response to the user's terminal. The user's terminal either displays the received explanatory text on the screen or plays it aloud using speech synthesis technology (e.g., Amazon Polly or Google Text-to-Speech).

[2264] Step 7:

[2265] When a user enters a question about an exhibit, the user's terminal sends the question data to the server.

[2266] Input: Question entered by the user (voice or text)

[2267] Output: Question data sent to the server

[2268] Specific operation: The user asks a question via voice input or text input through the app on their device. In the case of voice input, the device converts the voice to text and sends it to the server as question data.

[2269] Step 8:

[2270] The server analyzes the received question data and generates appropriate answers using a generative AI model.

[2271] Input: Question data sent to the server

[2272] Output: Generated answer

[2273] Specific operation: The server analyzes the question data and prompts the generation AI model to generate appropriate answers. The generated answers are saved in text format.

[2274] Example prompt: "Generate answers to the following question: 'What is the background of this painting?'"

[2275] Step 9:

[2276] The server sends the generated response to the user's terminal, where it is displayed or played back as audio.

[2277] Input: Generated answer

[2278] Output: Response sent to the user's terminal

[2279] Specific operation: The server sends the generated response to the user's terminal as an HTTP response. The user's terminal displays the received response on the screen or plays it back as audio using speech synthesis technology.

[2280] Step 10:

[2281] When a user enters their thoughts and opinions about an exhibit, the user's device sends the feedback data to the server.

[2282] Input: User feedback (voice or text)

[2283] Output: Feedback data sent to the server

[2284] Specific operation: Users provide feedback via voice input or text input through the app on their device. In the case of voice input, the device converts the voice to text and sends it to the server as feedback data.

[2285] Step 11:

[2286] The server collects and organizes the received feedback data and stores it in a database.

[2287] Input: Feedback data sent to the server

[2288] Output: Feedback data stored in the database

[2289] Specific operation: The server collects and organizes the received feedback data, and then saves it to a database (e.g., MongoDB).

[2290] Step 12:

[2291] Based on the feedback data, the server uses a generative AI model to generate appropriate feedback.

[2292] Input: Feedback data stored on the server

[2293] Output: Generated feedback

[2294] Specific operation: The server analyzes the stored feedback data and prompts the AI ​​model to generate appropriate feedback. The generated feedback is saved in text format.

[2295] Step 13:

[2296] The server sends the generated feedback to the user's terminal, where it is displayed or played as audio.

[2297] Input: Generated feedback

[2298] Output: Feedback sent to the user terminal

[2299] Specific operation: The server sends the generated feedback to the user terminal as an HTTP response. The user terminal displays the received feedback on the screen or plays it back as audio using speech synthesis technology.

[2300] Step 14:

[2301] The server retrieves character information set for each exhibit and reflects the character's tone of voice and personality in the explanations and answers.

[2302] Input: Exhibit identifier

[2303] Output: Explanatory text and answer text reflecting character information

[2304] Specific operation: The server uses the exhibit identifier to retrieve character information from the database and reflects the character's tone of voice and personality in the explanatory text and response text.

[2305] Step 15:

[2306] The server sends explanatory text and response text based on character information to the user's terminal, and the terminal plays avatars and voice clips.

[2307] Input: Explanatory text or answer text that reflects character information

[2308] Output: Explanatory text and response text with character information sent to the user's terminal.

[2309] Specific operation: The server sends explanatory text and answer text with character information to the user terminal as an HTTP response. The user terminal displays the received content on the screen and plays the character's voice using speech synthesis technology.

[2310] The above outlines the specific processing steps of this system. It details a comprehensive mechanism for enhancing the user experience, from exhibit recognition and explanation to question-and-answer sessions, feedback, and even character entertainment.

[2311] (Application Example 1)

[2312] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[2313] Traditional museum and art gallery systems for explaining exhibits are limited to providing information about the exhibits themselves, and do not adequately address users' questions or comments. Similarly, in physical stores and tourist destinations, there is a lack of detailed information about exhibits and products, as well as insufficient interactive communication with users. As a result, users' understanding and experience are limited, and the appeal of exhibits and products cannot be fully conveyed.

[2314] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[2315] This invention includes means for a server to capture images using a camera installed in a terminal to recognize exhibits and transmit the image data to the s...

Claims

1. A means of recognizing exhibits by capturing images using a camera installed on the terminal and sending the image data to a server, The server includes means for analyzing received image data and generating an ID corresponding to the identified exhibit, On the server, a means is provided to retrieve explanatory data from a database based on an identified exhibit ID, and to convert that explanatory data into a user-friendly format using natural language generation technology. A means for transmitting the converted explanatory data to the user's terminal and displaying or playing it as audio on the terminal, A means of inputting user questions into the terminal via voice or text, A means for sending question data entered on a terminal to a server, A server provides a means for analyzing received question data and generating appropriate answers using a generative AI, A means for sending the generated response to the user's device and displaying or playing it aloud on the device, A system that includes this.

2. A means of inputting user feedback and opinions into the device via voice or text, A means of sending feedback data entered on the terminal to a server, A means of collecting and organizing the received feedback data on the server and storing it in a database, A means of generating appropriate feedback using natural language generation technology based on feedback data, A means for sending the generated feedback to the user's device and displaying or playing it aloud on the device, The system according to claim 1, including the following:

3. A method for retrieving character information set for each exhibit from a server and reflecting the character's tone of voice and personality in the explanations and answers, A means of sending data to a terminal to express explanatory text and answer text based on the character using the character's voice, A means of providing characters to users on a device, such as avatars and voices, The system according to claim 1, including the following:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A