Artwork information output device, program, artwork information output system, and artwork information registration device

The artwork information output device addresses the limitations of existing systems by using a storage unit, generation AI, and identification processes to provide personalized and relevant responses to user questions, improving user engagement and understanding.

JP2026122432APending Publication Date: 2026-07-28DAI NIPPON PRINTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
DAI NIPPON PRINTING CO LTD
Filing Date
2025-03-26
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing systems for providing information about artworks, such as those using wearable devices, lack the ability to dynamically generate personalized and relevant responses to user questions, requiring significant effort to create multiple guide patterns and often fail to provide users with the information they truly want.

Method used

An artwork information output device that includes an artwork information storage unit, image and question acquisition means, a generation AI for answer generation, and output means, along with identification and truth value determination processes, to provide personalized answers to user questions about viewed artworks.

Benefits of technology

Enables the provision of appropriate and personalized answers to user questions about artworks, enhancing user engagement and understanding by dynamically generating relevant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026122432000001_ABST
    Figure 2026122432000001_ABST
Patent Text Reader

Abstract

This invention provides a work information output device that enables the provision of appropriate answers to a variety of questions from users about the works they are viewing. [Solution] The artwork explanation server 1 includes: an artwork information storage unit 23 that stores artwork information including text about the artwork; a captured image acquisition unit 11 that acquires an image specified by the user; a question acquisition unit 12 that acquires user questions about the image; a text acquisition unit 15 that acquires the text of the artwork shown in the image acquired by the captured image acquisition unit 11 from the artwork information storage unit 23; an answer generation processing unit 16 that acquires the answer output by the generation AI by inputting an answer generation instruction statement that instructs the generation AI to generate an answer to the question acquired by the question acquisition unit 12 based on the text acquired by the text acquisition unit 15; and an answer output processing unit 18 that outputs the answer acquired by the answer generation processing unit 16 to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0002] , ,

[0005] , , , , , , , , , ,

[0004] , , , , , , , ,

[0003]

[0001] The present invention relates to a work information output device, a program, a work information output system, and a work information registration device.

Background Art

[0002] In recent years, in the field of digital archives, not only digitizing and storing objects but also developing means for utilizing digitized archive data has been promoted. In promoting the utilization of digitized data, not limited to artworks, a mechanism is required that allows users who have no knowledge of the digitized archive data to appreciate it while being interested. As a mechanism that allows appreciation while being interested, for example, there are devices for devising ways of showing various work groups in an easy-to-understand manner for users, and devices for providing various viewpoints so that not only individual knowledge of digitized archive data can be acquired but also the relationship between works can be discovered.

[0003] One approach to utilizing archive data is a guide system for exhibits that utilizes wearable devices such as smartphones and smart glasses. For example, an information processing device that estimates the state of a user and the user's appreciation style and presents the most appropriate one from a plurality of pre-prepared guides to the user has been disclosed (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] Therefore, the present invention aims to provide a work information output device, etc., that can provide appropriate answers to a variety of questions from users about the works they are viewing. [Means for solving the problem]

[0007] The present invention solves the above problem by the following means. The first invention is an artwork information output device comprising: an artwork information storage unit that stores artwork information of the artwork, including text relating to the artwork; an image acquisition means that acquires an image specified by a user; a question acquisition means that acquires a question from the user regarding the image; a text acquisition means that acquires the text of the artwork shown in the image acquired by the image acquisition means from the artwork information storage unit; an answer generation means that acquires the answer output by the generation AI by inputting an answer generation instruction sentence that instructs the generation AI to generate an answer to the question acquired by the question acquisition means based on the text acquired by the text acquisition means; and an answer output means that outputs the answer acquired by the answer generation means to the user. The second invention is an artwork information output device of the first invention, wherein the image acquisition means acquires a photographed image that includes at least a portion of the image of the exhibit obtained by photographing the exhibit, the question acquisition means acquires a question from the user regarding the exhibit, the artwork identification means identifies the artwork of the exhibit from the photographed image acquired by the image acquisition means, and the text acquisition means acquires the text of the artwork identified by the artwork identification means from the artwork information storage unit. The third invention is a work information output device of the first or second invention, comprising: a sentence decomposition means for decomposing the answer generated by the answer generation means into individual sentences; and a truth value acquisition means for acquiring the truth value of each sentence output by the generation AI by inputting a truth value determination instruction sentence to the generation AI instructing it to determine the truth value of each sentence after decomposition by the sentence decomposition means based on the text acquired by the text acquisition means, wherein the answer output means outputs the answer generated by the answer generation means to the user based on the truth value determination result of each sentence acquired by the truth value acquisition means. The fourth invention is an artwork information output device of the second invention, further comprising an artwork image storage unit that stores an image of the artwork, wherein the artwork identification means identifies the artwork from the photographed image by comparing the photographed image acquired by the image acquisition means with the image of the artwork in the artwork image storage unit. The fifth invention is an artwork information output device of the fourth invention, wherein the artwork information includes an artwork image vector obtained by converting an image of the artwork into vector information, and the device comprises a captured image vectorization means that converts the captured image acquired by the image acquisition means into vector information, and the artwork identification means compares the captured image vector, which is the vector information converted by the captured image vectorization means, with the artwork image vector contained in the artwork information of the artwork information storage unit, and identifies the artwork of the exhibit based on the degree of similarity. The sixth invention is an artwork information output device of the fifth invention, wherein the artwork information includes a plurality of images relating to the artwork and the artwork image vector of each of the plurality of images, the plurality of images relating to the artwork includes at least one of the images of the artwork taken from different angles and images of a part of the artwork, and the artwork identification means compares the captured image vector converted by the captured image vectorization means with the artwork image vector contained in the artwork information of the artwork information storage unit, and identifies the artwork of the exhibit based on the similarity of each of the plurality of images of the artwork. The seventh invention is an artwork information output device of the second invention, comprising: artwork estimation means for inputting a captured image and an artwork identification instruction statement that instructs the AI ​​to identify the artwork of the exhibit based on the captured image into a generating AI, thereby acquiring artwork information output by the generating AI; and the artwork identification means searches the artwork information storage unit based on the artwork information acquired by the artwork estimation means to identify the artwork of the exhibit. The eighth invention is an artwork information output device of the second invention, comprising: an artwork estimation means that is communicably connected to an image search server, transmits the captured image to the image search server, and estimates the artwork of the exhibit based on information of entities similar to the captured image received from the image search server; and an artwork identification means that searches the artwork information storage unit based on the information of entities corresponding to the artwork of the exhibit estimated by the artwork estimation means to identify the artwork of the exhibit. The ninth invention is an artwork information output device of the fourth invention, further comprising a specific image output means for extracting and outputting the image of the artwork identified by the artwork identification means from the artwork image storage unit. The tenth invention is an artwork information output device of the ninth invention, wherein the specific image output means outputs the images of one or more artworks identified by the artwork identification means together with identification result information, the specific image output means includes a selected artwork acquisition means that acquires the artwork selected by the user from the one or more artworks by the specific image output means, and the text acquisition means acquires the text of the artwork acquired by the selected artwork acquisition means from the artwork information storage unit. The eleventh invention is a work information output device that, in any of the first to tenth inventions, comprises a shooting unit, an audio input unit, and an audio output unit, wherein the image acquisition means acquires the image via the shooting unit, the question acquisition means acquires the question by converting the user's spoken voice acquired via the audio input unit into text, and the answer output means converts the answer into audio and outputs it to the audio output unit. The twelfth invention is a work information output device according to any of the first to tenth inventions, comprising: a display unit; an audio input unit; an audio output unit; a work image storage unit that stores images of the works; and a work image output means that outputs images of at least one work stored in the work image storage unit to the display unit, wherein the image acquisition means acquires the image from the work title obtained by converting the user's spoken voice acquired via the audio input unit into text; the question acquisition means acquires the question by converting the user's spoken voice acquired via the audio input unit into text; and the answer output means converts the answer into voice and outputs it to the audio output unit. The 13th invention is an artwork information output device of the 12th invention, further comprising an input unit, wherein the image acquisition means acquires an image by receiving a selection operation by the user of one of at least one images of artworks displayed on the display unit via the input unit, and the response output means outputs the response to the display unit. The fourteenth invention is a program for causing a computer to function as an output device for artwork information from the first to the tenth inventions. The 15th invention is a work information output system comprising a work information output device from the first to the tenth inventions and a wearable device that is communicatively connected to the work information output device, wherein the wearable device comprises: a captured image transmission means for transmitting a captured image taken by a shooting unit to the work information output device while the wearable device is worn on the user's head; a question transmission means for transmitting the user's spoken voice received by a voice input unit to the work information output device; and a response receiving means for receiving the response from the work information output device and outputting it to a display unit, wherein the question acquisition means acquires the question by converting the spoken voice into text. The sixteenth invention is a work information output system of the fifteenth invention, wherein the answer output means outputs an answer audio converted from the answer to the wearable device, and the answer receiving means receives the answer audio from the work information output device and outputs it to the audio output unit. The 17th invention is a work information output system of the 15th or 16th invention, wherein the mounting device comprises a question display means for outputting a plurality of questions stored in a memory unit to a display unit, and a question receiving means for receiving one question selected by the user from the plurality of questions output to the display unit, and the question transmission means transmits the one question received by the question receiving means to the work information output device in place of the user's spoken voice. The 18th invention is an artwork information registration device comprising: artwork information acquisition means for acquiring an image and artwork attribute information of a work; image vectorization means for converting the image of the work acquired by the artwork information acquisition means into vector information; text vectorization means for converting the artwork attribute information of the work acquired by the artwork information acquisition means into vector information; registration means for registering the image of the work in the artwork image storage unit, and registering the vector information of the image, the artwork attribute information of the work, and the vector information of the artwork attribute information in the artwork information storage unit; related artwork acquisition means for acquiring related works similar to the vector information of the image and the vector information of the artwork attribute information of the work, by referring to the artwork information storage unit based on each vector piece of information; and updating means for associating information relating to the related works acquired by the related artwork acquisition means with the artwork information of the work in the artwork information storage unit. The 19th invention is an artwork information output system comprising: an artwork information output device from the first to the tenth inventions; and an artwork information registration device for registering the artwork information in the artwork information storage unit, wherein the artwork information registration device comprises: artwork information acquisition means for acquiring an image and artwork attribute information of a single artwork; image vectorization means for converting the image of the single artwork acquired by the artwork information acquisition means into vector information; text vectorization means for converting the artwork attribute information of the single artwork acquired by the artwork information acquisition means into vector information; registration means for registering the image of the single artwork in the artwork image storage unit and registering the vector information of the image, the artwork attribute information of the single artwork, and the vector information of the artwork attribute information in the artwork information storage unit; related artwork acquisition means for acquiring related artworks similar to the vector information of the image and the vector information of the artwork attribute information of the single artwork by referring to the artwork information storage unit based on each vector piece of information; and updating means for associating the information relating to the related artworks acquired by the related artwork acquisition means with the artwork information of the single artwork in the artwork information storage unit. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide a work information output device, etc., that can provide appropriate answers to a variety of questions from users about the works they are viewing. [Brief explanation of the drawing]

[0009] [Figure 1] This is an overall configuration diagram of the artwork information output system according to the first embodiment. [Figure 2] This is a functional block diagram of the artwork description server according to the first embodiment. [Figure 3] This figure shows an example of the artwork information storage unit of the artwork commentary server according to the first embodiment. [Figure 4] This is a functional block diagram of the mounting device and information registration terminal according to the first embodiment. [Figure 5] This is a flowchart showing the artwork information registration process of the information registration terminal according to the first embodiment. [Figure 6] It is a diagram for explaining the work reference information stored in the work information storage unit according to the first embodiment. [Figure 7] It is a flowchart showing the work explanation output process of the work information output system according to the first embodiment. [Figure 8] It is a flowchart showing the response acquisition determination process of the work explanation server according to the first embodiment. [Figure 9] It is a diagram showing a specific example in the wearing device related to the work explanation output process of the work information output system according to the first embodiment. [Figure 10] It is a diagram showing a specific example in the work identification process of the work explanation server according to the first embodiment. [Figure 11] It is a diagram showing an example of an instruction sentence generated by the work explanation server according to the first embodiment. [Figure 12] It is a diagram showing an example of an instruction sentence generated by the work explanation server according to the first embodiment. [Figure 13] It is a diagram showing a display example of the wearing device according to a modification example of the first embodiment. [Figure 14] It is a diagram showing a display example of the wearing device according to a modification example of the first embodiment. [Figure 15] It is a flowchart showing the work explanation output process of the work information output system according to a modification example of the first embodiment. [Figure 16] It is a flowchart showing the work explanation acquisition process of the work explanation server according to a modification example of the first embodiment. [Figure 17] It is a diagram showing an example of an instruction sentence generated by the work explanation server according to a modification example of the first embodiment. [Figure 18] It is a functional block diagram of the work information output device according to the second embodiment. [Figure 19] It is a flowchart showing the work information output process of the work information output device according to the second embodiment. [Figure 20] It is a diagram showing a display example of the work information output device according to the second embodiment. [Figure 21] It is a diagram showing an operation example in the work information output device according to the second embodiment. [Modes for carrying out the invention]

[0010] The following describes embodiments for carrying out the present invention with reference to the figures. However, this is merely an example, and the technical scope of the present invention is not limited thereto. (First Embodiment) <Work Information Output System 100> Figure 1 is an overall schematic diagram of the artwork information output system 100 according to the first embodiment. The artwork information output system 100 shown in Figure 1 is a system comprising an artwork commentary server 1 (artwork information output device), a mounting device 3, an information registration terminal 5 (artwork information registration device), and a generation AI server 7.

[0011] In the artwork information output system 100, for example, a glasses-shaped wearable device 3 worn by the user takes pictures of exhibits displayed in a museum or the like, and receives questions about the exhibits. Then, in the artwork information output system 100, the artwork explanation server 1, which receives data from the wearable device 3, identifies the artwork from the image of the exhibit, generates answers to the questions using the artwork information of the identified artwork, and outputs them to the wearable device 3. In other words, the artwork information output system 100 is a system that can predict the artwork that the user is viewing and then provide appropriate answers to a variety of questions from the user about the artwork.

[0012] In the following, the artwork information output system 100 will be described using paintings displayed in art museums and museums as examples of artworks. However, this is just one example, and other exhibits (works) such as sculptures or books may also be used. The artwork explanation server 1, the wearable device 3, the information registration terminal 5, and the generation AI server 7 are each connected to each other via a communication network N. The communication network N is, for example, an internet connection or a mobile device communication network. However, the communication network N is not limited to the above as the communication line connecting each device. For example, it may include a LAN (Local Area Network) in part, and it does not matter whether it is wired or wireless. Furthermore, although Figure 1 shows only one mounting device 3, multiple mounting devices 3 may be connected to the communication network N.

[0013] <Work Explanation Server 1> Next, I will explain the artwork explanation server 1. Figure 2 is a functional block diagram of the artwork commentary server 1 according to the first embodiment. Figure 3 shows an example of the artwork information storage unit 23 of the artwork commentary server 1 according to the first embodiment.

[0014] The artwork explanation server 1 obtains the captured image and the user's question from the wearable device 3, and then identifies the artwork from the captured image. The artwork explanation server 1 then generates an answer to the user's question using the artwork information of the identified artwork and outputs it to the wearable device 3. As shown in Figure 2, the artwork commentary server 1 comprises a control unit 10, a storage unit 20, and a communication interface unit 29. The control unit 10 is the CPU (Central Processing Unit) that controls the entire artwork commentary server 1. The control unit 10 works in cooperation with the aforementioned hardware to perform various functions by appropriately reading and executing the OS (Operating System) and application programs stored in the memory unit 20.

[0015] First, let me explain the memory unit 20. The storage unit 20 is a storage area such as a hard disk or semiconductor memory element for storing programs, data, etc., necessary for the control unit 10 to perform various processes. The memory unit 20 comprises a program memory unit 21, a work information memory unit 23 (work information memory unit, work image memory unit), and an instruction statement memory unit 24. The program storage unit 21 is a memory area that stores various programs. The program storage unit 21 stores the work explanation program 21a. The work explanation program 21a is a program for executing each function of the control unit 10, which will be described later.

[0016] The artwork information storage unit 23 is a memory area that stores artwork information, including images of the artwork and text related to the artwork. Figure 3 shows an example of items in the artwork information storage unit 23 and an example of registered data. As shown in Figure 3, the artwork information storage unit 23 stores the artwork ID (IDentification), artwork title, artwork image (thumbnail) name, artist name, artwork reference information, image vector, text vector, etc.

[0017] The artwork ID is identification information used to uniquely identify an artwork. The artwork ID is uniquely set, for example, when registering artwork information in the artwork information storage unit 23. The title of a work is one of the attributes of the work and is the name of the work. The artwork image (thumbnail) name is, for example, the data name of the artwork's thumbnail image. The thumbnail image may be stored in, for example, the memory unit 20, or in another memory device (not shown). The author's name is one of the attributes of the work and is the name of the person who wrote the work. The reference information for a work consists of various pieces of information about the work, including text related to the work. For example, it includes attribute information such as the work title and author's name. It also includes information about related works. The reference information for a work is used in the processing of the work description server 1 and is in the form of structured data.

[0018] The image vector is the artwork image vector obtained by converting the image of the artwork into vector information. The image vector includes the artwork image vector obtained by converting at least a portion of the image of the artwork into vector information. It is desirable that the artwork information storage unit 23 has artwork image vectors from multiple viewpoints (angles) registered. A text vector is a work text vector obtained by converting work attribute information into vector information. The items stored in the work information storage unit 23 are not limited to those listed above. For example, the work attribute information is not limited to those listed above.

[0019] The instruction statement storage unit 24 is a storage area that stores templates for response generation instruction statements and truth / false determination instruction statements. The answer generation instruction is a prompt that instructs the generating AI to generate an answer to a question based on reference information for the work. The truth / false determination instruction is a prompt that instructs the AI ​​to determine whether the answer generated based on the reference information for the work is inferable or not. Here, a prompt refers to instructions or other information that a user inputs to a conversational system, such as a generating AI, and is written in text.

[0020] Next, the control unit 10 will be described. The control unit 10 includes a captured image acquisition unit 11 (image acquisition means), a question acquisition unit 12 (question acquisition means), an image vectorization unit 13 (captured image vectorization means), a work identification unit 14 (work identification means), a text acquisition unit 15 (text acquisition means), an answer generation processing unit 16 (answer generation means), an answer truth value determination processing unit 17 (sentence decomposition means, truth value acquisition means), and an answer output processing unit 18 (answer output means).

[0021] The image acquisition unit 11 acquires images obtained by photographing the exhibit from the mounting device 3. Ideally, the images should be of the entire exhibit, but images of only a part of the exhibit are not excluded. Furthermore, ideally, the images should be of the exhibit photographed from the front, but images of the exhibit photographed from an oblique angle are not excluded. The question acquisition unit 12 acquires user questions about the exhibits from the wearable device 3. For example, the question acquisition unit 12 acquires the spoken audio of the question uttered by the user. In this case, the question acquisition unit 12 converts the acquired spoken audio into text to obtain the question in text form. Various known techniques can be used to convert audio to text.

[0022] The image vectorization unit 13 converts the captured image acquired by the captured image acquisition unit 11 into vector information. This conversion of an image into vector information can be performed using known technologies, such as Microsoft's Azure Computer Vision service (Vectorize Image). The captured image vector, which is the vector information of the captured image converted by the image vectorization unit 13, is represented as a multidimensional vector. The artwork identification unit 14 identifies the exhibited artwork from the captured image. For example, the artwork identification unit 14 compares the captured image vector obtained by the image vectorization unit 13 with the artwork image vector stored in the artwork information storage unit 23, and identifies the exhibited artwork based on the degree of similarity.

[0023] The text acquisition unit 15 acquires reference information about the work identified by the work identification unit 14 from the work information storage unit 23. The answer generation processing unit 16 uses the work reference information in the work information storage unit 23 related to the work identified by the work identification unit 14 to generate an answer to the question acquired and converted into text by the question acquisition unit 12. More specifically, the answer generation processing unit 16 inserts the converted question and the work reference information into an answer generation instruction statement to instruct the generation AI, and sends the answer generation instruction statement to the generation AI server 7. The answer generation processing unit 16 then receives and acquires the answer statement from the generation AI server 7.

[0024] The answer truthfulness determination processing unit 17 determines the truthfulness of the answer text obtained by the answer generation processing unit 16. The answer truthfulness determination processing unit 17 divides the answer text obtained by the answer generation processing unit 16 into sentence units. Then, the answer truthfulness determination processing unit 17 uses the work reference information in the work information storage unit 23 related to the work identified by the work identification unit 14 to determine the truthfulness of the answer text generated by the answer generation processing unit 16. More specifically, the answer truthfulness determination processing unit 17 inserts each divided sentence and the work reference information into a truthfulness determination instruction statement to instruct the generation AI, and sends the truthfulness determination instruction statement to the generation AI server 7. Then, the answer truthfulness determination processing unit 17 receives and obtains the truthfulness determination result from the generation AI server 7.

[0025] The answer output processing unit 18 outputs the answer to the user's question to the wearable device 3. More specifically, the answer output processing unit 18 outputs the answer generated by the answer generation processing unit 16 to the wearable device 3 if the truthfulness determination result by the answer truthfulness determination processing unit 17 is true. At that time, the answer output processing unit 18 converts the answer text generated by the answer generation processing unit 16 into speech and outputs the answer audio to the wearable device 3. Various known techniques can be used to convert the answer text into speech.

[0026] The communication interface unit 29 is an interface for communicating with the mounting device 3, the information registration terminal 5, and the generation AI server 7, etc. Here, "computer" refers to an information processing device equipped with a control unit, memory device, etc., and the artwork commentary server 1 is an information processing device equipped with a control unit 10, a memory unit 20, etc., and is included in the concept of a computer. Furthermore, there is no limit to the number of hardware components that make up the artwork explanation server 1; it may be configured with one or more components as needed. Also, the artwork explanation server 1 may, for example, be a cloud service.

[0027] <Mounting device 3> Next, I will explain the mounting device 3. Figure 4 is a functional block diagram of the mounting device 3 and information registration terminal 5 according to the first embodiment. The wearable device 3 is, for example, a pair of smart glasses. The mounting device 3 comprises a control unit 30, a storage unit 40, an imaging unit 45, a display unit 46, an audio input unit 47, an audio output unit 48, and a communication interface unit 49. The control unit 30 is a CPU that controls the entire mounting device 3. The control unit 30 works in cooperation with the aforementioned hardware to perform various functions by appropriately reading and executing the OS and application programs stored in the memory unit 40.

[0028] The control unit 30 includes a captured image acquisition unit 31, a captured image transmission unit 32 (captured image transmission means), a spoken voice acquisition unit 33, a spoken voice transmission unit 34 (question transmission means), a response voice reception unit 35 (response reception means), and a response voice output unit 36 ​​(response reception means). The image acquisition unit 31 acquires images of the exhibits that the user is viewing via the shooting unit 45. The image acquisition unit 31 may acquire images by triggering them with user speech, or by accepting user input to start shooting. The captured image transmission unit 32 transmits the captured image acquired by the captured image acquisition unit 31 to the artwork explanation server 1.

[0029] The speech voice acquisition unit 33 acquires the user's speech via the voice input unit 47. The speech voice acquisition unit 33, for example, starts recording the voice when the user speaks and stops recording the voice after detecting a certain period of silence following the end of speech. The speech voice acquisition unit 33 can recognize the user's speech by, for example, pre-registering the user's voice in the storage unit 40. The speech voice transmission unit 34 transmits the speech voice acquired by the speech voice acquisition unit 33 to the work commentary server 1. The processing by the captured image transmission unit 32 and the spoken voice transmission unit 34 may be performed simultaneously.

[0030] The response audio receiving unit 35 receives the response audio from the artwork explanation server 1. The response audio output unit 36 ​​outputs the response audio received by the response audio receiving unit 35 to the audio output unit 48, thereby outputting the response audio from the audio output unit 48. The user can hear the answer to the question they asked through the processing of the response audio output unit 36.

[0031] The memory unit 40 is a memory area such as a semiconductor memory element for storing programs, data, etc., necessary for the control unit 30 to perform various processes. The memory unit 40 includes a program memory unit 41. The program storage unit 41 is a storage area that stores various programs, including programs for performing the various functions executed by the control unit 30 described above.

[0032] The imaging unit 45 is a camera that captures the user's field of view while the user is wearing the device. The imaging unit 45 is located in a position that allows it to capture the surroundings, for example, on the outer surface opposite to the side facing the user's head when the device 3 is attached to the user's head, and near the positions corresponding to the user's left and right eyes. Therefore, the image captured by the imaging unit 45 is an image of approximately the same range as what the user's left and right eyes can see.

[0033] The display unit 46 is a display device such as an LCD (Liquid Crystal Display) or an organic EL display. The display unit 46 is provided on the inner surface that faces the head of the user wearing the mounting device 3. The audio input unit 47 is a microphone. The audio input unit 47 is provided, for example, on the frame of the mounting device 3. The audio output unit 48 is a speaker. The audio output unit 48 is provided, for example, on the part of the vine at the end of the wearable device 3 that rests on the ear. The communication interface unit 49 is an interface for communicating with the artwork commentary server 1 and the like.

[0034] <Information Registration Terminal 5> Next, we will explain the information registration terminal 5. The information registration terminal 5 is used by the operator to register artwork information in the artwork information storage unit 23, a process performed before processing by the artwork explanation server 1. As shown in Figure 4(B), the information registration terminal 5 comprises a control unit 50, a storage unit 60, an input unit 64, a display unit 66, and a communication interface unit 69. The control unit 50 is a CPU that controls the entire information registration terminal 5. The control unit 50 works in cooperation with the aforementioned hardware to perform various functions by appropriately reading and executing the OS and application programs stored in the memory unit 60.

[0035] The control unit 50 includes a work information acquisition unit 51 (work information acquisition means), an image vectorization unit 52 (image vectorization means), a text vectorization unit 53 (text vectorization means), an information registration processing unit 54 (registration means), a related work acquisition unit 55 (related work acquisition means), and an information update processing unit 56 (update means). The artwork information acquisition unit 51 acquires an image and artwork attribute information for a single artwork. For example, the artwork information acquisition unit 51 acquires an image and artwork attribute information for each exhibit or artwork in a museum.

[0036] The image vectorization unit 52 converts the image acquired by the artwork information acquisition unit 51 into vector information to obtain an artwork image vector. Here, the image vectorization unit 52, for example, shapes the image acquired by the artwork information acquisition unit 51 into images from various viewpoints through image processing. The image vectorization unit 52 can pseudo-shape each viewpoint image from a single artwork image by using image processing such as clipping, padding, affine transformation, and projection transformation. Then, the image vectorization unit 52 converts each of the multiple images into vector information.

[0037] The text vectorization unit 53 converts the artwork attribute information acquired by the artwork information acquisition unit 51 into vector information to obtain the artwork text vector. The technique for converting text into vector information can be done using publicly known technologies such as Text Embedding provided by OpenAI. The text vector information converted by the text vectorization unit 53 is represented as a multidimensional vector. The information registration processing unit 54 registers the artwork attribute information of one artwork, the artwork reference information including the artwork attribute information of one artwork, the artwork image vector of one artwork converted by the image vectorization unit 52, and the artwork text vector of one artwork converted by the text vectorization unit 53 in the artwork information storage unit 23 of the artwork explanation server 1.

[0038] The related works acquisition unit 55 acquires related works that are similar to the work image vector and work text vector of a given work by referring to the work information storage unit 23 based on the respective vector information. The information update processing unit 56 associates the information about related works acquired by the related works acquisition unit 55 with the information about one of the works stored in the work information storage unit 23. Here, it is desirable that the information registration terminal 5 performs the processing of the information registration processing unit 54 for each work acquired by the work information acquisition unit 51 and registers it in the work information storage unit 23, and then performs the processing of the related work acquisition unit 55 and the information update processing unit 56 to add information related to related works to the work reference information in the work information storage unit 23.

[0039] The storage unit 60 is a storage area such as a hard disk or semiconductor memory element for storing programs, data, etc., necessary for the control unit 50 to perform various processes. The storage unit 60 includes a program storage unit 61. The program storage unit 61 is a storage area that stores various programs, including programs for performing the various functions executed by the control unit 50 described above.

[0040] The input unit 64 is an input device such as a keyboard or mouse. The display unit 66 is a display device such as an LCD or an organic EL display. Alternatively, the input unit 64 and the display unit 66 may be integrated into a single touch panel display. The communication interface unit 69 is an interface for communicating with the artwork commentary server 1 and the like.

[0041] Here, "computer" refers to an information processing device equipped with a control unit, memory device, etc., and the mounting device 3 and the information registration terminal 5 are information processing devices equipped with a control unit, memory unit, etc., respectively, and both are included in the concept of a computer.

[0042] <Generating AI Server 7> Next, we will explain the generation AI server 7. The generative AI server 7 shown in Figure 1 is a generative AI server capable of text input. The generative AI server 7 may also be a multimodal generative AI server capable of text and image input. The generative AI includes a large-scale language model (LLM). A large-scale language model is a language model constructed using a large amount of text data and deep learning technology. Examples of LLMs include GPT-3.5 and GPT-4.0, but are not limited to these. The generative AI performs inference according to instructions indicated by the input prompt and outputs the inference results as text data. Here, inference refers to, for example, analysis, classification, prediction, summarization, etc. The generation AI server 7, although not shown in the diagram, includes a control unit, a storage unit, a communication interface unit, and the like. The generation AI server 7 may, for example, be a cloud service.

[0043] <Work Information Registration Process> Next, we will explain the processing in the artwork information output system 100. First, we will explain the process of registering artwork information in the artwork information storage unit 23 of the artwork commentary server 1. Figure 5 is a flowchart showing the artwork information registration process of the information registration terminal 5 according to the first embodiment. Figure 6 is a diagram illustrating the work reference information stored in the work information storage unit 23 according to the first embodiment.

[0044] In step S (hereinafter referred to simply as "S") 11 of Figure 5, the control unit 50 (artwork information acquisition unit 51) of the information registration terminal 5 acquires an image and artwork attribute information of one artwork. The control unit 50 may acquire the image of one artwork from a storage device (not shown) in which the image is stored, or it may acquire the image from a camera or other shooting device by connecting to the shooting device in a communicative manner. The control unit 50 may also acquire the artwork attribute information, for example, from a storage device (not shown) in which the artwork attribute information is stored, or it may acquire it via the input unit 64.

[0045] In S12, the control unit 50 (image vectorization unit 52) ​​converts the image of the artwork into vector information to obtain the artwork image vector. Here, in addition to converting the image of the artwork into vector information, the control unit 50 also shapes the image of the artwork into images from various viewpoints using image processing, and converts each of the multiple shaped images into vector information. In S13, the control unit 50 (text vectorization unit 53) converts the work attribute information into vector information to obtain the work text vector. In S14, the control unit 50 (information registration processing unit 54) registers various types of information in the work information storage unit 23. Through this process, information for each item except for the work reference information is registered in the work information storage unit 23 (see Figure 3).

[0046] In S15, the control unit 50 (related work acquisition unit 55) acquires related works that are similar to the work image vector and work text vector by referring to the work information storage unit 23 based on the respective vector information. Here, I will explain in detail how to acquire related works. In this example, the similarity between a work and other registered works is calculated numerically, and works with high scores are designated as related works. For example, cosine similarity can be used to calculate the similarity.

[0047] The formula for calculating cosine similarity is expressed in Equation 1 below.

number

[0048] In the above-mentioned (Equation 1), the control unit 50 calculates the similarity by using vector a as the image vector of one work and vector b as the image vector of another work registered in the work information storage unit 23. Similarly, in the above-mentioned (Equation 1), the control unit 50 calculates the similarity by using vector a as the text vector of one work and vector b as the text vector of another work registered in the work information storage unit 23. If the calculated similarity is equal to or greater than a threshold, the control unit 50 determines that the other work is a related work similar to one work.

[0049] In the above method, each vector of information is output as a multidimensional (over 1000 dimensions) vector. However, calculating cosine similarity in a high-dimensional state can generally lead to unreasonable results, such as the cosine similarity value converging to 0. Therefore, the values ​​to be calculated may be reduced by performing Principal Component Analysis on the above numerical vectors to reduce the number of dimensions before calculating the similarity. Furthermore, while the above example uses cosine similarity to calculate similarity, it is not limited to this. Other methods for calculating similarity include using Euclidean distance or Mahalanobis distance.

[0050] In S16, the control unit 50 (information update processing unit 56) registers related work information for a work registered in the work information storage unit 23. For example, the control unit 50 extracts the work attribute information of related works from the work information storage unit 23 and adds it to the work reference information for related works that have similar work image vectors, treating them as related works with similar appearances. The control unit 50 also extracts the work attribute information of related works from the work information storage unit 23 and adds it to the work reference information for related works that have similar work text vectors, treating them as related works with similar work attributes.

[0051] Here, we will explain the reference information for the work using an example. Figure 6 shows an example 70 of the work reference information stored in the work information storage unit 23. Example 70 shows artwork reference information presented as structured data, and includes an attribute information area 71, a text vector area 72, and an image vector area 73. The attribute information area 71 is an area that contains work attribute information as structured data. Text vector area 72 is an area that contains structured data containing information about related works that have similar text vectors. Image vector region 73 is a region that contains structured data containing information about related works whose image vectors are similar to those of the original work.

[0052] Please note that the information registered in the artwork reference information is not limited to that shown in Example 70. For example, other artwork information such as emotional data (warm impression, soft image quality) may also be registered. Returning to Figure 5, the control unit 50 then terminates this process. The information registration terminal 5 processes the registration of artwork information, allowing various types of artwork-related information to be registered in the artwork information storage unit 23 of the artwork commentary server 1.

[0053] <Product Description Output Processing> Next, we will explain the process of outputting answers to questions about the exhibits the user is viewing, for users wearing the device 3. Figure 7 is a flowchart showing the artwork description output process of the artwork information output system 100 according to the first embodiment. Figure 8 is a flowchart showing the response acquisition decision process of the artwork commentary server 1 according to the first embodiment. Figure 9 shows a specific example of the mounting device 3 related to the output processing of artwork descriptions in the artwork information output system 100 according to the first embodiment. Figure 10 shows a specific example of the artwork identification process of the artwork description server 1 according to the first embodiment. Figures 11 and 12 show examples of instruction texts generated by the artwork explanation server 1 according to the first embodiment.

[0054] As shown in Figure 9(A), user P wears a pair of glasses-type devices 3 while viewing exhibits in the museum. In S21 of Figure 7, the control unit 30 (image acquisition unit 31) of the mounting device 3 acquires a captured image via the imaging unit 45. Here, the control unit 30 acquires a captured image via the imaging unit 45, for example, triggered by a speech uttered by the user. The captured image includes at least a portion of the exhibit that the user is looking at. In S22, the control unit 30 (speech voice acquisition unit 33) acquires the speech voice of the question uttered by the user. For example, the control unit 30 records the speech voice triggered by the utterance made by the user and continues recording the speech voice until there is a certain period of silence. As shown in Figure 9(A), when user P says "Explain this picture," the wearable device 3 acquires a captured image 75 of the exhibit D1 within the frame F that the camera unit 45 of the wearable device 3 can capture, and the spoken audio 76.

[0055] In step S23 of Figure 7, the control unit 30 (captured image transmission unit 32, spoken audio transmission unit 34) transmits the captured image and spoken audio acquired in steps S21 and S23, respectively, to the artwork commentary server 1. In the artwork commentary server 1, when the control unit 10 (image acquisition unit 11, question acquisition unit 12) receives the captured image and spoken audio from the wearable device 3, in S31, the control unit 10 (image vectorization unit 13) converts the captured image into vector information to obtain the captured image vector. In S32, the control unit 10 (artwork identification unit 14) refers to the artwork information storage unit 23 and identifies the artwork using the captured image vector.

[0056] In the S32 process, the control unit 10 calculates the similarity between the captured image vector obtained in the S31 process and the artwork image vector (see Figure 3) in the artwork information storage unit 23, and identifies the artwork with the highest similarity. The control unit 10 can use the cosine similarity shown in (Equation 1) above when calculating the similarity. The similarity score results 80 shown in Figure 10 represent the case where multiple image vectors are registered for each work, and show the similarity score between each work's image vector and the captured image vector. Image 81 has the highest similarity score. The control unit 10 can identify the artwork with the image 81 that has the highest similarity score as "the artwork the user is viewing."

[0057] The control unit 10 may, for example, use the average similarity score for each artwork to identify the artwork with the highest average score. Alternatively, if the similarity score does not meet a predetermined threshold, the control unit 1 may terminate processing and send an audio message to the wearable device 3 stating, "No artwork was found." In step S33 of Figure 7, the control unit 10 (question acquisition unit 12) converts the spoken audio into text to obtain a question in text form. In S34, the control unit 10 performs the response acquisition decision process.

[0058] Here, the process for determining the acquisition of a response will be explained based on Figure 8. In S41 of Figure 8, the control unit 10 (text acquisition unit 15, answer generation processing unit 16) of the artwork explanation server 1 acquires artwork reference information for the identified artwork from the artwork information storage unit 23, and inserts the question converted into text and the acquired artwork reference information into the answer generation instruction template to generate the answer generation instruction. Here, an example of an instruction statement for generating an answer is shown in Figure 11. The response generation instruction statement 90 shown in Figure 11 includes a system instruction area 91, a work information area 92, and a user instruction area 93. The system instruction area 91 is an area that contains instructions for the generating AI, and these instructions are common to all questions. The artwork information area 92 is the area where acquired artwork reference information is inserted. User instruction area 93 is the area where the question converted into text is inserted.

[0059] In step S42 of Figure 8, the control unit 10 (answer generation processing unit 16) transmits the generated answer generation instruction to the generation AI server 7. In the generation AI server 7, when a response generation instruction is input to the generation AI, the generation AI processes the response generation instruction, generates an answer to the question, and outputs the answer text. Then, the generation AI server 7 sends the answer text output by the generation AI to the work explanation server 1, and in S43 of Figure 8, the control unit 10 (answer generation processing unit 16) obtains the answer text from the generation AI server 7.

[0060] In S44, the control unit 10 (answer truth value determination processing unit 17) divides the acquired answer sentence into sentence units. In S45, the control unit 10 (answer truth / false determination processing unit 17) inserts each of the divided sentences of the answer and the reference information of the identified work into the template of the truth / false determination instruction to generate the truth / false determination instruction. Here, an example of a truth / false instruction is shown in Figure 12.

[0061] The truth / false determination instruction statement 95 shown in Figure 12 includes a system instruction area 96, a work information area 97, a prerequisite instruction area 98, and a user instruction area 99. The system instruction area 91 is an area that contains instructions for the generating AI, and these instructions are common to all questions. The artwork information area 97 is the area where acquired artwork reference information is inserted. The prerequisite instruction area 98 is an area that describes the instructions for the generating AI, and these instructions are common to all questions. The user instruction area 93 is an area where each sentence obtained by dividing the response text acquired from the generating AI is inserted as a separate sentence.

[0062] In S46 of Figure 8, the control unit 10 (answer truth / false determination processing unit 17) sends the generated truth / false determination instruction to the generation AI server 7. When a truth / false judgment instruction is input to the generation AI on the generation AI server 7, the generation AI processes the truth / false judgment instruction, generates truth / false values ​​for each sentence, and outputs the truth / false judgment result. The generation AI server 7 transmits the truth value determination result for each sentence output by the generation AI to the work explanation server 1. In S47 of Figure 8, the control unit 10 (answer truth value determination processing unit 17) obtains the truth value determination result for each sentence from the generation AI server 7.

[0063] In S48, the control unit 10 (answer truth value determination processing unit 17) determines the answer to output based on the truth value determination result of each acquired sentence. The control unit 10 may, for example, determine the reliability score as the number of true statements divided by the total number of sentences, set a certain threshold, and decide to send the generated response text to the device 3 if the reliability score is higher than the threshold, and decide not to send the generated response text to the device 3 if it is lower than the threshold. The control unit 10 may also decide to send the response text after excluding sentences that have been determined to be false. This process ensures the reliability of the answer to the question. If it is decided not to send the generated response text, the control unit 10 may decide to send a standard message such as, for example, "I'm sorry, but I cannot answer that question."

[0064] Furthermore, the control unit 10 may appropriately change the reliability score threshold depending on the intended use. For example, if the device is used in an art museum or museum, it is necessary to have a higher level of reliability than usual. Therefore, even if the normal threshold is set to, for example, less than 70, when the reliability score ranges from 0 to 100 and higher reliability corresponds to a higher reliability score, the threshold may be set to, for example, less than 100 when the device is used in an art museum or museum. The control unit 10 moves the processing to S35 in Figure 7.

[0065] In step S35 of Figure 7, the control unit 10 (answer output processing unit 18) transmits the answer determined by the answer acquisition and determination process to the wearable device 3. At that time, the control unit 10 (answer output processing unit 18) converts the answer into audio and transmits the audio answer to the wearable device 3. After that, the control unit 10 terminates this process.

[0066] In the mounting device 3, when the control unit 30 (answer audio receiving unit 35) receives the answer audio from the work explanation server 1, in S24, the control unit 30 (answer audio output unit 36) outputs the received answer audio to the audio output unit 48, and the answer audio is output from the audio output unit 48. After that, the control unit 30 terminates this process. As shown in Figure 9(B), when user P says "Explain this painting," the artwork explanation server 1 identifies the artwork image 77, obtains a question answer based on the artwork reference information of the identified artwork image 77, and generates the answer as audio data 78 after a truth / false judgment. Therefore, the audio output unit 48 of the wearable device 3 outputs the audio "This is an oil painting called 'Thread Spool'..."

[0067] <Example 1> Variation 1 describes a system that displays the answer to a question. Figure 13(A) shows an example of a display on the mounting device 3 according to Modification 1 of the First Embodiment, illustrating an example of what is displayed on the display unit 46 of the mounting device 3 when a user looks at an exhibit and asks a question while wearing the mounting device 3. The control unit 10 (specific image output means) of the artwork commentary server 1 transmits the thumbnail of the identified artwork and the acquired response to the wearable device 3, causing the control unit 30 (response receiving means) of the wearable device 3 to display information about the artwork (image and text) on the display unit 46, as shown in Figure 13(A).

[0068] Furthermore, in order to avoid interfering with the user's appreciation of the artwork, it is desirable to display information about the artwork to one eye, as exemplified in Figure 13(A), or to display it in a location that does not interfere with the artwork in the user's line of sight. Furthermore, if a user asks a question about an exhibit, the answer may be output as audio along with the display shown in Figure 13(A). In this way, by outputting information about the identified artwork, the identified artwork can be presented to the user, allowing the user to confirm whether or not it is the same artwork as the one on display.

[0069] <Modification 2> Variation 2 describes a method for displaying anticipated questions. Figure 13(B) shows an example of a display on the mounting device 3 according to Modification 2 of the First Embodiment, illustrating an example where a hypothetical question is displayed on the display unit 46 of the mounting device 3 when the user views an exhibit while wearing the mounting device 3. The anticipated questions displayed on the wearable device 3 are pre-set in the storage unit of the wearable device 3, for example. Then, when the user views an exhibit while wearing the wearable device 3, the control unit 30 (question display means) of the wearable device 3 displays the anticipated questions on the display unit 46 of the wearable device 3. This is useful in situations where verbal questioning is not possible. The user may select the displayed question using, for example, a controller (not shown), or they may select the question by arbitrarily recognizing gestures using hand tracking technology. The control unit 30 (question receiving means, question transmission means) of the mounting device 3 receives the question selected by the user and transmits it to the artwork explanation server 1. The subsequent processing by the artwork explanation server 1 is the same as in the embodiment described above.

[0070] <Variation 3> Modification 3 describes a method for displaying a guide frame for shooting. Figure 14(A) shows an example of display on the mounting device 3 according to modification 3 of the first embodiment, and shows an example in which a guide frame G is displayed on the display unit 46 of the mounting device 3 when the user views an exhibit while the mounting device 3 is attached. In this way, by displaying the guide frame G on the display unit 46 of the mounting device 3, the user adjusts the position of the exhibit so that it fits within the guide frame G. Then, for example, the control unit 30 of the mounting device 3 can detect when the exhibit is within the guide frame G and acquire a captured image, thereby more accurately inferring the location of the artwork and, as a result, providing appropriate answers to questions about the artwork.

[0071] <Modification 4> Modification 4 describes a method that allows the user to identify the work. Figure 14(B) shows an example of display on the mounting device 3 according to modification 4 of the first embodiment, illustrating an example where, when the user views an exhibit while the mounting device 3 is attached, the display unit 46 of the mounting device 3 displays candidate works. In the embodiment described above, the artwork explanation server 1 calculates the similarity between the captured image vector, which is obtained by converting the captured image into vector information, and the artwork image vector of the artwork information storage unit 23, and identifies the artwork with the highest similarity. In the modified example 4, the artwork explanation server 1 (specific image output means) identifies several artworks with high similarity and outputs thumbnails of the identified artworks to the mounting device 3. As shown in Figure 14(B), the control unit 30 of the mounting device 3 displays the candidate artworks on the display unit 46.

[0072] The user may select from the displayed artwork candidates using, for example, a controller (not shown), or they may select one artwork from the candidates by arbitrarily recognizing gestures using hand tracking technology. The control unit 30 of the mounting device 3 transmits the artwork selected by the user to the artwork explanation server 1, and the control unit 10 (selected artwork acquisition means) of the artwork explanation server 1 acquires the artwork selected by the user. Then, the control unit 10 (answer generation means) of the artwork explanation server 1 generates answers to questions about the acquired artwork. This approach makes it possible to provide users with accurate and reliable explanations of the artwork. Furthermore, along with the suggested works presented to the user, the similarity score value (specific result information) may also be output. This allows the user to be provided with both the suggested works and the similarity score value.

[0073] <Modification 5> Modification 5 describes a system that provides the user with information about the artwork before any questions are taken from the user, when the user is wearing the attachment device 3 and viewing the exhibit. Figure 15 is a flowchart showing the artwork description output process of the artwork information output system 100 according to Modification 5 of the first embodiment. Figure 16 is a flowchart showing the process of acquiring artwork descriptions from the artwork description server 1 according to Modification 5 of the first embodiment. Figure 17 shows an example of an instruction text generated by the work explanation server 1 according to modification 5 of the first embodiment.

[0074] First, the user is wearing glasses-type wearable device 3 while viewing exhibits in the museum. In S121 of Figure 15, the control unit 30 (image acquisition unit 31) of the mounting device 3 acquires an image via the shooting unit 45, provided that the shooting unit 45 has been photographing one exhibit for a predetermined time (for example, 5 seconds). In S122, the control unit 30 (captured image transmission unit 32) transmits the captured image to the artwork commentary server 1. In the artwork explanation server 1, when the control unit 10 (image acquisition unit 11) receives an image from the mounting device 3, in S131, the control unit 10 performs the artwork explanation acquisition process.

[0075] Here, the process for obtaining artwork descriptions will be explained based on Figure 16. The processes S151 and S152 in Figure 16 are the same as the processes S31 and S32 in the first embodiment (Figure 7). In S153, the control unit 10 obtains the explanatory text for the identified work. Here, various methods can be considered for obtaining the explanatory text for the artwork, but this can be achieved by using one or more of these methods.

[0076] One example is to pre-register work description information (not shown) that explains the work in the work information storage unit 23. In this case, the control unit 10 can retrieve the work description information of the work stored in the work information storage unit 23 as the work description text. As another example, images that supplement the work may be provided to the user along with the explanatory text. For example, a timeline image (not shown) showing the "production date" of the work and the dates of works closely related to this work that were produced around that time may be pre-registered in the work information storage unit 23. Then, the control unit 10 retrieves the timeline image along with the work explanation information. In this case, for example, a timeline image can be presented to the user along with audio (or text) explanatory text, thereby promoting a deeper understanding of the work.

[0077] Furthermore, as an example, the control unit 10 can obtain an explanatory text for a work using the generation AI server 7. In this case, the control unit 10 first obtains reference information for the identified work from the work information storage unit 23, inserts the reference information into the template for the work explanation generation instruction, and generates the work explanation generation instruction. Here, an example of the instruction text for generating a work description is shown in Figure 17. The artwork description generation instruction text 190 shown in Figure 17 includes an instruction area 191 and an artwork information area 192. Instruction area 191 is an area that contains instructions for the generating AI. In this example, it shows a timeline. The artwork information area 192 is the area where acquired artwork reference information is inserted.

[0078] Next, the control unit 10 sends the generated artwork description generation instruction text to an image generation AI server (not shown). The image generation AI server is, for example, a server having an image generation model such as a deep learning text-to-image model represented by Stable Diffusion. The image generation AI server takes a text instruction for generating a work description as input to the image generation model, processes the text instruction, and generates and outputs a description of the work, including the image. The image generation AI server then sends the artwork description to the artwork description server 1, and the control unit 10 of the artwork description server 1 retrieves the artwork description from the image generation AI server. The instructions to be written in instruction area 191 are not limited to those listed above. Other instructions may be used, such as generating a map showing the production location, as long as they provide information that helps the user understand the work.

[0079] When the explanatory text for the artwork is obtained, in S132 of Figure 15, the control unit 10 converts the explanatory text into audio and transmits the explanatory audio to the wearable device 3. In the mounting device 3, when the control unit 30 receives commentary audio from the artwork commentary server 1, in S123, the control unit 30 outputs the received commentary audio via the audio output unit 48. In the above explanation, if an image is generated, the control unit 10 of the artwork explanation server 1 will also transmit the image, and the control unit 30 of the mounting device 3 will output the received image to the display unit 46.

[0080] The process in S124 is the same as the process in S22 in the first embodiment (Figure 7). The processing in S125 is the same as the processing related to spoken audio in S23 of the first embodiment (Figure 7). The processes from S133 to S135 and the process in S126 are the same as the processes from S33 to S35 and the process in S24 in the first embodiment (Figure 7). Furthermore, after processing S123, the process may be terminated without performing the processes from S124 onward. In this case, the user will only receive explanations about the exhibits they have photographed.

[0081] Thus, the artwork information output system 100 of the first embodiment has the following effects. (1) The artwork explanation server 1 acquires photographic images that include at least part of the images of the exhibits obtained by taking pictures of the exhibits, acquires user questions about the exhibits, identifies the exhibits from the acquired photographic images, generates answers to the questions using the artwork information about the identified artwork stored in the artwork information storage unit 23, and outputs the generated answers. Therefore, by inferring the works that the user is watching, we can provide answers to the user's questions.

[0082] (2) The artwork explanation server 1 obtains text related to the identified artwork from the artwork information storage unit 23, and inputs an answer generation instruction statement to the generation AI instructing it to generate an answer based on the text obtained for the question, thereby obtaining the answer output by the generation AI. Therefore, by using generative AI, we can provide appropriate answers to a variety of user questions regarding the predicted works.

[0083] (3) The work explanation server 1 inputs a truth value judgment instruction to the generation AI, which instructs the AI ​​to break down the generated answer into individual sentences and determine the truth value of each sentence based on the acquired text. The server then obtains the truth value of each sentence output by the generation AI and outputs the generated answer based on the truth value judgment results of each acquired sentence. Therefore, by verifying the truthfulness of the answers output by the generating AI and outputting only those answers that have been confirmed to be true, the accuracy of the answers can be guaranteed.

[0084] (4) The artwork explanation server 1 identifies the exhibited artwork from the captured image by comparing the acquired captured image with the image of the artwork stored in the artwork information storage unit 23. Therefore, the artwork can be estimated and identified by comparing images.

[0085] (5) The artwork information includes an artwork image vector obtained by converting an image of the artwork into vector information. The artwork explanation server 1 converts the acquired photographed image into vector information, compares the resulting photographed image vector with the artwork image vector contained in the artwork information storage unit 23, and identifies the exhibited artwork based on the degree of similarity. Therefore, by using vector information to calculate similarity, and estimating and identifying works based on the calculated similarity, a uniform judgment can be made.

[0086] (6) The artwork information includes multiple images of the artwork and the artwork image vector for each of the multiple images, and the multiple images of the artwork include at least one of the images of the artwork taken from different angles and images of a part of the artwork, and the artwork explanation server 1 compares the captured image vector with the artwork image vector held in the artwork information storage unit 23 and identifies the exhibited artwork based on the similarity of each of the multiple images of the artwork. Therefore, even if the photographed image shows only a part of the exhibit or is taken from an oblique angle, the artwork can be estimated and identified with greater sensitivity.

[0087] (7) The mounting device 3, when mounted on the user's head, transmits the captured image taken by the shooting unit 45 to the work commentary server 1, and the voice input unit 47 transmits the user's spoken voice received to the work commentary server 1, and the work commentary server 1 converts the spoken voice into text to obtain the question. Furthermore, the artwork explanation server 1 converts the answers into audio and transmits them to the wearable device 3, and the wearable device 3 outputs the received audio answers to the audio output unit 48. Therefore, by using the wearable device 3 and the artwork explanation server 1, it is possible to infer the artwork the user is viewing and provide the user with appropriate answers to a variety of questions.

[0088] (8) The information registration terminal 5 acquires an image and attribute information of a work, converts the acquired image of the work into vector information to obtain a work image vector, and converts the work attribute information of the work into vector information to obtain a work text vector, and registers these in the work information storage unit 23. Related works similar to the work image vector and work text vector of the work are acquired based on the respective vector information, and the work information related to the acquired related works is associated with the information of the work in the work information storage unit 23. Therefore, when constructing the artwork information storage unit 23 used for inferring and identifying artworks in advance, related artworks similar to the artwork can be determined based on vector information.

[0089] (Second Embodiment) In the first embodiment, the system processing was described using a specific example of a user wearing the attachment device asking questions while viewing exhibits in a museum or similar setting. In the second embodiment, the processing of an artwork information output device installed in a museum or similar setting that enables the user to ask questions about the artwork and manipulate the artwork is described. In the following description, parts that perform the same functions as in the first embodiment described above will be denoted with the same reference numerals or the same reference numerals at the end, and redundant explanations will be omitted as appropriate.

[0090] <Artwork Information Output Device 201> Figure 18 is a functional block diagram of the artwork information output device 201 according to the second embodiment. The artwork information output device 201 is installed, for example, in a museum, and allows users to operate it by touching a touch panel display 204 or to ask questions about the artwork by voice using a voice recognition device 206. The artwork information output device 201 generates and outputs answers based on the user's operations and inputs. In other words, the artwork information output device 201 is a device that enables users to deepen their knowledge of the artwork they have viewed by answering questions from the user about the artwork they have viewed.

[0091] The artwork information output device 201 is a device that integrates a user interface and a processing unit for processing. The artwork information output device 201 shown in Figure 18 comprises a control unit 210, a storage unit 220, a touch panel display 204 (display unit, input unit), and a voice recognition device 206.

[0092] First, let me explain the memory unit 220. The memory unit 220 includes a program memory unit 221, a work information memory unit 223 (work information memory unit, work image memory unit), an instruction statement memory unit 24, and a generation AI memory unit 226. The program storage unit 221 stores various programs, including the work commentary program 221a, which is a program for executing each function of the control unit 210, which will be described later.

[0093] The artwork information storage unit 223 is a storage area that stores artwork information, including, for example, an image (thumbnail) of the artwork and text related to the artwork such as the artwork title, using the artwork ID as the key. The artwork information storage unit 223 may be the same as the example items and example of registered data of the artwork information storage unit 23 in the first embodiment (Figure 3), and there may be some missing information, such as not having an image vector, but at a minimum, artwork attribute information such as the artwork title and author name, and the artwork image are required.

[0094] The generation AI memory unit 226, provided in the artwork information output device 201, offers the functionality of the generation AI. The generation AI memory unit 226 includes a local LLM, which is the generation AI, and RAG (Retrieval Augmented Generation), which is reference data for the generation AI. The generation AI performs inference according to instructions indicated by the input prompt and outputs the inference result as text data. In doing so, the generation AI adds supplementary information to the input prompt using the mechanism of RAG. RAG contains information specific to the artwork, as well as information about the museum where the artwork information output device 201 is installed.

[0095] The artwork information output device 201, equipped with a generation AI memory unit 226, does not require data transmission or reception with external sources, allowing for the secure handling of highly confidential information. Furthermore, because the generation AI does not use external data, it can avoid, for example, learning using data related to unreliable information, and can provide answers based on highly reliable information.

[0096] Next, the control unit 210 will be described. The control unit 210 includes a work image output unit 219 (work image output means), an image acquisition unit 211 (image acquisition means), a question acquisition unit 212 (question acquisition means), a text acquisition unit 215 (text acquisition means), an answer generation processing unit 216 (answer generation means), and an answer output processing unit 218 (answer output means).

[0097] The artwork image output unit 219 outputs images of at least one artwork stored in the artwork information storage unit 223 to the touch panel display 204. The image acquisition unit 211 acquires an image specified by the user. For example, the image acquisition unit 211 acquires an image by receiving a user's selection operation of one image from at least one image of artworks displayed on the touch panel display 204 via the touch panel display 204. Alternatively, the image acquisition unit 211 may acquire an image from artwork text obtained by converting the user's spoken voice acquired via the speech recognition device 206 into text using known technology. In this case, the image acquisition unit 211 acquires the artwork ID of the artwork corresponding to the image.

[0098] The question acquisition unit 212 acquires, for example, the spoken audio of a question uttered by the user via the speech recognition device 206. The question acquisition unit 212 converts the user's spoken audio acquired via the speech recognition device 206 into text to obtain the question. The text acquisition unit 215 acquires the artwork text (text) of the artwork shown in the image acquired by the image acquisition unit 211 from the artwork information storage unit 223, for example, based on the artwork ID.

[0099] The answer generation processing unit 216 generates an answer to the question that the question acquisition unit 212 has acquired and converted into text. More specifically, the answer generation processing unit 216 inserts the converted question and the work text, which is text acquired by the text acquisition unit 215, into an answer generation instruction statement to instruct the generation AI, and inputs the answer generation instruction statement into the generation AI in the generation AI storage unit 226. The answer generation processing unit 216 uses the answer generation instruction statement and RAG to have the generation AI generate an answer. The answer generation processing unit 216 acquires the answer output by the generation AI.

[0100] The response output processing unit 218 outputs the response generated by the response generation processing unit 216 to the user. The response output processing unit 218 may, for example, convert the response into speech and output it to the speech recognition device 206. Alternatively, the response output processing unit 218 may output the response to, for example, the touch panel display 204.

[0101] The touch panel display 204 has the function of a display unit composed of a liquid crystal panel or the like, and the function of an input unit that detects touch input from the user, such as a finger. The voice recognition device 206 is, for example, a handset-type device and includes a voice input unit 247 and a voice output unit 248. The audio input unit 247 is, for example, a microphone. The audio output unit 248 is, for example, a speaker. Here, "computer" refers to an information processing device equipped with a control unit, memory device, etc., and the artwork information output device 201 is an information processing device equipped with a control unit 210, a memory unit 220, etc., and is included in the concept of a computer.

[0102] <Product Information Output Processing> Next, we will explain the process of outputting artwork information based on user actions. Figure 19 is a flowchart showing the artwork information output process of the artwork information output device 201 according to the second embodiment. Figure 20 shows an example of the display of the artwork information output device 201 according to the second embodiment. Figure 21 shows an example of operation in the artwork information output device 201 according to the second embodiment.

[0103] In step S221 of Figure 19, the control unit 210 (artwork image output unit 219) of the artwork information output device 201 outputs the images of each artwork stored in the artwork information storage unit 223 to the touch panel display 204. Figure 20 shows the screen 240 displayed on the touch panel display 204 of the artwork information output device 201. Screen 240 consists of a speech display area 241 that outputs the content of the speech, and an artwork display area 242 that outputs an image of the artwork. The speech display area 241 is an area that outputs speech from the work information output device 201 and speech from the user in text format. In the example shown on screen 240, the speech display area 241 displays the text "Please tell me ○×" as speech from the work information output device 201. The artwork display area 242 is an area in which thumbnail images 244 of each artwork stored in the artwork information storage unit 223 are displayed on cube-shaped objects 243 based on the year of production of each artwork.

[0104] In step S222 of Figure 19, the control unit 210 (image acquisition unit 211) acquires the artwork ID of the image corresponding to the user's input. With screen 240 displayed, for example, when the user touches (selects) the position of a thumbnail image 244 of a single artwork, the control unit 210 acquires the artwork ID associated with the touched thumbnail image 244. Alternatively, when the user speaks, for example, the title of an artwork using the voice recognition device 206, the control unit 210 acquires the artwork ID associated with the title of the artwork.

[0105] In S223, the control unit 210 (image acquisition unit 211) extracts and outputs the image corresponding to the acquired artwork ID from the artwork information storage unit 223. Figure 21 shows a user P configuration 250, including a screen 251 displayed on the touch panel display 204 of the artwork information output device 201. Screen 251 consists of a speech display area 241 that outputs the content of the speech, and an artwork display area 252 that outputs an image of the artwork. User P is using the speech recognition device 206 to utter a conversation that includes the title of the work (in this example, a conversation about the work title "Thread Spool"). In the example shown on screen 251, the utterance display area 241 displays the text "Tell me ○×" as the user's utterance, and "Tell me about thread spool" as the user's subsequent utterance. The artwork display area 252 is an area that displays the artwork image 253 of the artwork specified by the user's instructions. In the state of screen 251 shown in embodiment 250, user P can ask questions about the artwork image 253. Furthermore, in S223, the control unit 210 (text acquisition unit 215) acquires the title of the work from the work information storage unit 223.

[0106] In S224 of Figure 19, the control unit 210 (question acquisition unit 212) acquires the spoken audio of the question uttered by the user and converts it into text. In S225, the control unit 210 (answer generation processing unit 216) inserts the transcribed question and the acquired work title into a template for an answer generation instruction sentence to instruct the generation AI, thereby generating an answer generation instruction sentence. In S226, the control unit 210 (answer generation processing unit 216) inputs an answer generation instruction to the generation AI in the generation AI storage unit 226 and obtains the answer. In S227, the control unit 210 (answer output processing unit 218) converts the acquired answer into speech and outputs the answer speech to the speech recognition device 206. The control unit 210 (answer output processing unit 218) also outputs the acquired answer to the speech display area 241 of the screen 251.

[0107] In S228, the control unit 210 (question acquisition unit 212) determines whether or not it has received a question. If the user continues to ask questions, the control unit 210 determines that it has received a question. If a question has been received (S228: YES), the control unit 210 moves the process to S224. On the other hand, if no question has been received, for example, if the voice recognition device 206 is returned to the receiving side (S228: NO), the control unit 210 terminates this process. In this way, users can ask various questions about the artwork, and the answers output by the artwork information output device 201 allow them to delve deeper into the artwork and promote their understanding of it.

[0108] Thus, the artwork information output device 201 of the second embodiment has the following effects. (1) The generation AI memory unit 226 is provided in the artwork information output device 201, and an answer is obtained by inputting an answer generation instruction sentence, which includes a question about the artwork, into the generation AI of the generation AI memory unit 226. Therefore, since the response is generated using data stored in the generation AI memory unit 226, it does not use external data of unknown reliability, and can output a response based on highly reliable information.

[0109] (2) The system includes a touch panel display 204 and a voice recognition device 206, which allows the system to identify artworks in response to touch operations or voice input from the user, and to provide answers when the user enters questions about the identified artwork. Therefore, it is possible to build a system that allows users to obtain information about the artwork using a user-friendly interface.

[0110] Although embodiments of the present invention have been described above, the present invention is not limited to the embodiments described above. Furthermore, the effects described in the embodiments are merely a list of the most preferred effects arising from the present invention, and the effects of the present invention are not limited to those described in the embodiments. The embodiments described above and the modified forms described later can be used in combination as appropriate, but a detailed explanation is omitted.

[0111] (Transformed form) (1) In the first embodiment, the wearable device 3 was described as smart glasses, but it is not limited to this. The wearable device may be other wearable devices such as a head-mounted display that covers the user's eyes. In the case of a head-mounted display, either a video see-through display or an optical see-through display may be used. Furthermore, if the wearable device 3 does not have an OS, for example, it may be connected to a mobile terminal and the same processing may be performed using the mobile terminal's OS. In that case, the wearable device is equipped with a communication interface unit for communicating with the mobile terminal. Then, the wearable device performs the processing related to shooting and display, and the mobile terminal performs the other processing. Furthermore, a mobile device or the like may be used instead of the wearable device 3. In that case, the user takes a picture of the exhibit with the camera on the mobile device and inputs their questions into the mobile device by voice. By doing so, just as with the wearable device 3, it is possible to obtain information about the exhibited artwork and answers to questions.

[0112] (2) In the first embodiment, an example was described in which each viewpoint image is formed from an image of one work, and each of the formed images is converted into vector information and stored in the work information storage unit 23, but the invention is not limited to this. Multiple images of a single work taken from multiple viewpoints or positions may be converted into vector information and stored in the work information storage unit. Furthermore, vector information for multiple images is not required; it is sufficient to store vector information for at least one image of a work.

[0113] (3) In the first embodiment, an example was described in which the artwork of the exhibited object in the captured image is identified using the captured image vector and the artwork image vector of the artwork information storage unit 23, but the embodiment is not limited to this. For example, the control unit (artwork prediction means) of the artwork explanation server may input a captured image and an artwork identification instruction text that instructs the AI ​​to identify the artwork based on the captured image, thereby obtaining artwork information output by the generating AI. In this case, it is desirable that the generating AI used be a multimodal generating AI because it takes images as input. A multimodal generating AI is a generating AI capable of image analysis. Furthermore, the instruction to identify the artwork might include something like, "The image data of the artwork being sent is one of the works housed in the XX Museum. Please identify this artwork and return its title." Furthermore, to account for cases where the identified work is not found in the work information storage unit, it is advisable to design the work identification instruction message to output multiple suspected work candidates.

[0114] Alternatively, for example, the artwork explanation server may be connected to an image search server (not shown) in a communicative manner, and the control unit (artwork prediction means) of the artwork explanation server may send a captured image to the image search server and predict the artwork of the exhibited object based on information about entities similar to the captured image received from the image search server.

[0115] (4) In the first embodiment, when registering artwork information in the artwork information storage unit, the system was described as acquiring related artworks based on the similarity between the artwork image vector and the artwork text vector, but the system is not limited to this. If related artworks are known in advance, that information can be registered.

[0116] (5) In the first embodiment, an example was described in which the storage unit 20 of the work commentary server 1 is equipped with a work information storage unit 23, but the invention is not limited to this. The work information storage unit may be stored in a server other than the work commentary server 1.

[0117] (6) In the first embodiment, a system using a work commentary server 1 and a mounting device 3 was described as an example, but the system is not limited thereto. The mounting device may be an all-in-one device that has the function of a work commentary server.

[0118] (7) In the second embodiment, an example was described in which a question about the artwork is received and an answer is output, but the embodiment is not limited to this. For example, when a question about how to operate the displayed cube-shaped object 243 (see Figure 20) is received, an instruction statement is generated to output the answer, and the generating AI is instructed to output the answer to the question about how to operate it.

[0119] (8) In each embodiment, the description was given using an example in which the artwork information storage unit stores information including an image of the artwork, but the invention is not limited to this. The image of the artwork may be stored separately in an artwork image storage unit. [Explanation of Symbols]

[0120] 1. Work Commentary Server 3. Mounting device 5. Information Registration Terminal 7. Generation AI Server 10, 30, 50, 210 Control Unit 11, 31 Image acquisition unit 12, 212 Question acquisition part 13, 52 Image vectorization section 14 Work Specification Department 15, 215 Text acquisition section 16,216 Answer generation processing unit 17. Answer truthfulness determination processing unit 18, 218 Answer Output Processing Unit 20, 40, 60, 220 memory section 21a, 221a Work Explanation Program 23, 223 Work information storage section 24, 224 Instruction sentence storage unit 32 Image transmission unit 33. Speech Acquisition Unit 34. Speech Voice Transmission Unit 35. Voice Receiver 36. Output section for the response audio. 45 Photography Department 46, 66 Display section 47, 247 Voice input section 48, 248 Audio output section 51 Work information acquisition department 53 Text Vectorization Unit 54 Information Registration Processing Unit 55 Related Works Acquisition Department 56 Information Update Processing Unit 100 Works Information Output System 201 Work Information Output Device 204 Touch Panel Display 206 Voice Recognition Devices 211 Image acquisition unit 219 Image output section 226 Generated AI storage unit Screens 240, 251 241 Speech display area 242, 252 Work display area P User

Claims

1. A work information storage unit that stores work information of the aforementioned work, including text about the work, An image acquisition means for acquiring an image specified by the user, A question acquisition means for acquiring the user's question regarding the image, A text acquisition means acquires the text of the artwork shown in the image acquired by the image acquisition means from the artwork information storage unit, The answer generation means inputs an answer generation instruction sentence to the generating AI that instructs the question acquisition means to generate an answer to the question acquired by the question acquisition means based on the text acquired by the text acquisition means, and acquires the answer output by the generating AI. A response output means that outputs the response obtained by the response generation means to the user, A device for outputting artwork information, equipped with the necessary components.

2. In the artwork information output device according to claim 1, The image acquisition means acquires a photographed image that includes at least a portion of the image of the exhibit obtained by photographing the exhibit, The question acquisition means acquires the user's questions regarding the exhibits, The image acquisition means includes a means for identifying the artwork of the exhibit from the captured image acquired, The text acquisition means is a work information output device that acquires the text of the work identified by the work identification means from the work information storage unit.

3. In the artwork information output device according to claim 1, A sentence decomposition means for decomposing the answer generated by the answer generation means into individual sentences, A truth value acquisition means obtains the truth value of each sentence output by the generating AI by inputting a truth value determination instruction sentence to the generating AI, which instructs the generating AI to determine the truth value of each sentence after decomposition by the sentence decomposition means based on the text acquired by the text acquisition means. Equipped with, The response output means is a work information output device that outputs the response generated by the response generation means to the user based on the truth value determination result of each sentence acquired by the truth value acquisition means.

4. In the artwork information output device according to claim 2, It is equipped with an image storage unit that stores images of the aforementioned works, The aforementioned artwork identification means is an artwork information output device that identifies the artwork from the photographed image by comparing the photographed image acquired by the image acquisition means with the image of the artwork in the artwork image storage unit.

5. In the artwork information output device according to claim 4, The aforementioned artwork information includes an artwork image vector obtained by converting the image of the artwork into vector information, The image acquisition means includes a vectorization means for converting the captured image acquired by the image acquisition means into vector information, The aforementioned artwork identification means is an artwork information output device that identifies the artwork of the exhibit based on the degree of similarity by comparing the captured image vector, which is the vector information converted by the captured image vectorization means, with the artwork image vector contained in the artwork information of the artwork information storage unit.

6. In the artwork information output device according to claim 5, The artwork information includes a plurality of images relating to the artwork and the artwork image vector of each of the plurality of images. The multiple images relating to the aforementioned work include at least one of the images of the aforementioned work taken from different angles and images of a portion of the aforementioned work. The aforementioned artwork identification means is an artwork information output device that compares the captured image vector converted by the captured image vectorization means with the artwork image vector contained in the artwork information of the artwork information storage unit, and identifies the artwork of the exhibit based on the similarity of each of the multiple images of the artwork.

7. In the artwork information output device according to claim 2, The system includes a means for predicting artworks that obtains artwork information output by a generating AI by inputting the aforementioned captured image and an instruction statement for identifying the artwork based on the captured image into the generating AI. The aforementioned artwork identification means is an artwork information output device that identifies the artwork among the exhibited items by searching the artwork information storage unit based on the artwork information acquired by the aforementioned artwork prediction means.

8. In the artwork information output device according to claim 2, A communication-enabled connection is established to the image search server, The system includes a means for transmitting the captured image to the image search server and inferring the artwork of the exhibit based on information about entities similar to the captured image received from the image search server. The artwork identification means is an artwork information output device that identifies the artwork in the exhibit by searching the artwork information storage unit based on the information of the entity corresponding to the artwork in the exhibit that has been inferred by the artwork inference means.

9. In the artwork information output device according to claim 4, A work information output device comprising a specific image output means for extracting and outputting the image of the work identified by the work identification means from the work image storage unit.

10. In the artwork information output device according to claim 9, The specified image output means outputs the images of one or more works identified by the work identification means, along with the identification result information. The system includes a selected work acquisition means that acquires a work selected by the user from the one or more works produced by the specified image output means, The text acquisition means is a work information output device that acquires the text of the work acquired by the selected work acquisition means from the work information storage unit.

11. In the artwork information output device according to any one of claims 1 to 10, The photography department, Voice input section, Audio output section, Equipped with, The image acquisition means acquires the image via the imaging unit, The question acquisition means acquires the question by converting the user's spoken voice acquired via the voice input unit into text. The aforementioned response output means is a work information output device that converts the response into audio and outputs it to the audio output unit.

12. In the artwork information output device according to any one of claims 1 to 10, Display unit and Voice input section, Audio output section, A work image storage unit that stores images of the aforementioned work, Artwork image output means for outputting images of at least one artwork stored in the artwork image storage unit to the display unit, Equipped with, The image acquisition means acquires the image from the title of the work obtained by converting the user's spoken voice acquired via the voice input unit into text. The question acquisition means acquires the question by converting the user's spoken voice acquired via the voice input unit into text. The aforementioned response output means is a work information output device that converts the response into audio and outputs it to the audio output unit.

13. In the artwork information output device according to claim 12, Equipped with an input section, The image acquisition means acquires the image by receiving a selection operation by the user for one of the images of at least one artwork displayed on the display unit via the input unit. The aforementioned response output means is a work information output device that outputs the response to the display unit.

14. Computers, An image acquisition means for acquiring an image specified by the user, A question acquisition means for acquiring the user's question regarding the image, A text acquisition means that acquires the text of the work shown in the image acquired by the image acquisition means from the work information storage unit which stores work information of the work including text about the work, The answer generation means inputs an answer generation instruction sentence to the generating AI that instructs the question acquisition means to generate an answer to the question acquired by the question acquisition means based on the text acquired by the text acquisition means, and acquires the answer output by the generating AI. A response output means that outputs the response obtained by the response generation means to the user, A program to make it work.

15. A work information output device according to any one of claims 1 to 10, A mounting device that is communicatively connected to the aforementioned artwork information output device, A work information output system equipped with, The aforementioned mounting device is A means for transmitting captured images taken by the shooting unit while the device is attached to the user's head to the artwork information output device, A question transmission means that transmits the user's spoken voice received by the voice input unit to the work information output device, A response receiving means that receives the response from the aforementioned artwork information output device and outputs it to the display unit, Equipped with, The question acquisition means is a work information output system that acquires the question by converting the spoken audio into text.

16. In the artwork information output system described in claim 15, The response output means outputs the response audio, which is the response converted into voice, to the wearable device. The aforementioned response receiving means is a work information output system that receives the response audio from the work information output device and outputs it to the audio output unit.

17. In the artwork information output system described in claim 15, The aforementioned mounting device is A question display means that outputs multiple questions stored in the memory unit to the display unit, The system includes a question receiving means that receives one question selected by the user from among the multiple questions output to the display unit, The aforementioned question transmission means transmits the question received by the question reception means to the aforementioned work information output device in place of the user's spoken voice, in a work information output system.

18. A means for acquiring artwork information to obtain an image and artwork attribute information of a single artwork, The aforementioned artwork information acquisition means converts the image of the aforementioned artwork into vector information; A text vectorization means that converts the work attribute information of the work acquired by the work information acquisition means into vector information, A registration means for registering an image of the aforementioned artwork in the artwork image storage unit, and registering the vector information of the image, the artwork attribute information of the aforementioned artwork, and the vector information of the artwork attribute information in the artwork information storage unit, Related work acquisition means for acquiring related works similar to the vector information of the image and the vector information of the work attribute information of the aforementioned work by referring to the work information storage unit based on each vector information, An update means that associates the information relating to the related work acquired by the related work acquisition means with the work information of the one work in the work information storage unit, A device for registering artwork information, equipped with the following features.

19. A work information output device according to any one of claims 1 to 10, A work information registration device for registering the work information to the work information storage unit, A work information output system equipped with, The aforementioned artwork information registration device is A means for acquiring artwork information to obtain an image and artwork attribute information of a single artwork, The aforementioned artwork information acquisition means converts the image of the aforementioned artwork into vector information; A text vectorization means that converts the work attribute information of the work acquired by the work information acquisition means into vector information, A registration means for registering an image of the aforementioned artwork in the artwork image storage unit, and registering the vector information of the image, the artwork attribute information of the aforementioned artwork, and the vector information of the artwork attribute information in the artwork information storage unit. Related work acquisition means for acquiring related works similar to the vector information of the image and the vector information of the work attribute information of the aforementioned work by referring to the work information storage unit based on each vector information, An update means that associates the information relating to the related work acquired by the related work acquisition means with the work information of the one work in the work information storage unit, A work information output system equipped with the following features.