Caption generation apparatus, caption generation method and program

The caption generation device addresses the challenge of generating role-specific captions for plant equipment images by using an image acquisition and characteristic information acquisition unit, enhancing communication efficiency and response to abnormalities.

JP2025098510APending Publication Date: 2025-07-02YOKOGAWA ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023214687
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-07-02

AI Technical Summary

Technical Problem

Existing systems lack the ability to generate captions for plant equipment images that are tailored to the specific roles and characteristics of the user viewing or capturing the images, leading to inefficiencies in communication and response to abnormalities.

Method used

A caption generation device that includes an image acquisition unit, word extraction unit, and characteristic information acquisition unit to generate captions based on the user's role and characteristics, using models to learn relationships between images, words, and captions, allowing for customized captions for different users.

Benefits of technology

Enables generation of captions that are relevant to the specific roles of users, improving communication efficiency and response to abnormalities in plant equipment by providing accurate and role-specific information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025098510000001_ABST
    Figure 2025098510000001_ABST
Patent Text Reader

Abstract

To provide a caption generation apparatus, a caption generation method, and a program.SOLUTION: A caption generation apparatus includes: an image acquisition unit which acquires an image captured in a plant; a word extraction unit which extracts a plurality of words representing characteristics of an object captured in the image; a property information acquisition unit which acquires property information indicating properties of a user; and a caption generation unit which generates a caption that describes the captured object according to the property information, based on the property information and the words. The property information may indicate at least one of the property of a person who captures an image, and the property of an administrator who receives a report using the generated caption.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a caption generation device, a caption generation method, and a program.

Background Art

[0002] Patent Document 1 describes that "when an abnormality (defect) occurs when an inexperienced worker or a field worker actually cuts a workpiece W using a cutting device 200, the video and vibration at that time are digitized and input as second information, and as a response, machining conditions described in a predetermined format including characters, numerical values, and symbols can be output as first information." (Paragraph 0119). Patent Document 2 describes that "obtain position information specified by GPS, display the current inspection location for plant staff, accumulate and analyze the image data captured by the installed camera, and determine whether there are signs of abnormality in the plant equipment from past equipment states and abnormal cases, etc., and at the same time, use a preset format / document to automatically create a periodic inspection report from the image data." (Paragraph 0053). Patent Document 3 describes that "a feature amount acquisition unit that acquires information about the pipe as a first feature amount from an image obtained by photographing the plant equipment of the work target and the pipes existing around the plant equipment with the camera, and a feature amount comparison unit that compares the first feature amount with a second feature amount about the pipe acquired from design data." (Claim 1). Patent Document 4 describes that "the abnormal mail creation function 102 is activated when the abnormality monitoring function 101 detects a plant abnormality, and creates a mail transmission text by formulating matters such as the date and time when the abnormality was detected, the name of the corresponding plant equipment, and the content of the abnormality that should be notified to the supervisor at an early stage." (Paragraph 0019). [Prior Art Documents] [Patent Documents] [Patent Document 1] Japanese Patent Application Laid-Open No. 2020-177547 [Patent Document 2] Patent No. 6099989 [Patent Document 3] Patent No. 6826509 [Patent Document 4] Japanese Patent Application Laid-Open No. 2003-51895

Summary of the Invention

[0003] In a first aspect of the present invention, a caption generation device is provided. The caption generation device includes an image acquisition unit that acquires an image captured within a plant, a word extraction unit that extracts a plurality of words representing the characteristics of the object imaged in the image, a characteristic information acquisition unit that acquires characteristic information indicating the characteristics of the user, and a caption generation unit that generates a caption for explaining the imaged object according to the characteristic information based on the characteristic information and the plurality of words.

[0004] In the above caption generation device, the characteristics of the user may include at least any one of the characteristics of the imager who captured the image and the characteristics of the viewer who views the report using the generated caption.

[0005] In any of the above caption generation devices, the characteristics of the user may include the type of the user's role.

[0006] In any of the above caption generation devices, the type of the role may include at least any one of the administrator of the plant, the monitor of the facilities within the plant, the maintenance staff of the facilities within the plant, and the recovery worker who deals with the abnormalities within the plant.

[0007] In any of the above caption generation devices, the type of the role may be related to a specific object within the plant. In any of the above caption generation devices, when an abnormality occurs in an object other than the specific object related to the type of the role of the user indicated by the characteristic information, the caption generation unit may generate a caption different from the caption generated when an abnormality occurs in the specific object indicated by the characteristic information.

[0008] Any of the above caption generation devices may further include a storage unit that stores the user ID of the user in association with the type of role. In any of the above caption generation devices, the characteristic information acquisition unit may extract the type of role associated with the user ID in response to obtaining the user ID.

[0009] Any of the above caption generation devices may further include a storage unit that stores the type of role in association with state information indicating the state of the target in the plant. In any of the above caption generation devices, the characteristic information acquisition unit acquires the characteristic information and extracts the state information associated with the type of role indicated by the characteristic information. The caption generation unit may generate a caption that describes the captured target according to the characteristic information and the state information based on the characteristic information, the plurality of words, and the state information.

[0010] Any of the above caption generation devices may further include a storage unit that stores the type of role in association with state information indicating the state of the target in the plant. In any of the above caption generation devices, the characteristic information acquisition unit acquires the characteristic information and may extract the state information associated with the type of role indicated by the characteristic information. In any of the above caption generation devices, for each of a plurality of predetermined types of roles, the caption generation unit generates a caption that describes the captured target according to the characteristic information and the state information corresponding to the type of role based on the characteristic information and the plurality of words, and selects and outputs at least one caption from the plurality of captions generated for the plurality of types of roles based on the state information extracted by the characteristic information acquisition unit.

[0011] In any of the above caption generation devices, the state information may include at least any one of the position information of the object, the deterioration information indicating the deterioration state of the object, the fluid information indicating the component or state of the fluid flowing inside the pipe that is the object, and the equipment information related to the equipment located on the upstream side or the downstream side of the pipe that is the object.

[0012] In any of the above caption generation devices, the word extraction unit may assign a likelihood according to the characteristic information to each of the plurality of words extracted from the image. In any of the above caption generation devices, the caption generation unit may generate the caption by preferentially using the words with higher likelihood among the plurality of words based on the characteristic information and the plurality of words.

[0013] In any of the above caption generation devices, the caption generation unit may assign a priority according to the likelihood of at least one word used for caption generation among the plurality of words to the generated caption.

[0014] In any of the above caption generation devices, when the likelihood of the combination of the plurality of words used for caption generation is lower than a predetermined threshold, the caption generation unit may output the caption including an instruction to re-capture the imaged object.

[0015] Any of the above caption generation devices may further include a storage unit that stores a plurality of sentences included in at least any one of the accident case collection, accident response manual, and maintenance history in the plant as example sentences. In any of the above caption generation devices, the caption generation unit may extract at least one of the example sentences by searching using at least any one of the plurality of words extracted, and generate the caption based on the characteristic information, the plurality of words, and the extracted example sentence.

[0016] In any of the above-caption generation devices, the caption generation unit may generate a plurality of captions and assign a similarity to the extracted example sentences.

[0017] In any of the above-caption generation devices, the caption may include a plurality of action options regarding the actions to be taken by the user, or instructions to the user.

[0018] In any of the above-caption generation devices, the caption may include at least one of an instruction to image the captured object from a different angle or a different shooting angle, and an instruction to image another object related to the captured object.

[0019] Any of the above-caption generation devices may further include a model storage unit that stores a caption generation model that has learned the relationship between the characteristic information, one or more words representing the characteristics of the object in the plant, and the caption describing the object in the plant. In any of the above-caption generation devices, the caption generation unit may generate the caption based on the newly input characteristic information and the plurality of words using the caption generation model.

[0020] Any of the above-caption generation devices may further include a word extraction model that extracts the plurality of words from the image using the result of the user's determination of the appropriateness of the generated caption, and a learning unit that learns the caption generation model that generates the caption from the characteristic information and the plurality of words. In any of the above-caption generation devices, the caption generation unit may generate the caption based on the newly input characteristic information and the plurality of words using the caption generation model.

[0021] Any of the above caption generation devices may further include a word extraction model that extracts the plurality of words from the image using the user's correction input for the generated caption, and a learning unit that learns a caption generation model that generates the caption from the characteristic information and the plurality of words. In any of the above caption generation devices, the caption generation unit may generate the caption based on the newly input characteristic information and the plurality of words using the caption generation model.

[0022] In a second aspect of the present invention, a caption generation method is provided. The caption generation method includes acquiring an image captured in a plant, extracting a plurality of words representing features of an object captured in the image, acquiring characteristic information indicating characteristics of a user, and generating a caption that describes the captured object according to the characteristic information based on the characteristic information and the plurality of words.

[0023] In a third aspect of the present invention, a program is provided. The program causes a computer to execute a procedure for acquiring an image captured in a plant, a procedure for extracting a plurality of words representing features of an object captured in the image, a procedure for acquiring characteristic information indicating characteristics of a user, and a procedure for generating a caption that describes the captured object according to the characteristic information based on the characteristic information and the plurality of words.

[0024] Note that the above summary of the invention does not enumerate all the features of the present invention. Also, sub-combinations of these feature groups can also be inventions.

Brief Description of the Drawings

[0025]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying Out the Invention

[0026] Hereinafter, the present invention will be described through embodiments of the invention. However, the following embodiments do not limit the invention claimed in the claims. Also, not all combinations of features described in the embodiments are essential for the solution means of the invention.

[0027] FIG. 1 is a schematic diagram of a caption generation system 5 including a caption generation device 100 according to the first embodiment. The caption generation system 5 generates a caption for explaining the imaged object 20 from an image 10 obtained by imaging the object 20 in a facility within a plant. More specifically, the caption generation system 5 generates a caption for explaining the imaged object 20 according to the characteristic information of the user of the caption generation system 5. The image 10 may be one or more still images or a moving image. Note that the caption for explaining the object 20 in the image 10 may refer to a descriptive text for explaining the state of the object 20 shown in the image 10.

[0028] Examples of plants include industrial plants such as those in the chemical industry, plants that manage and control wells and their surroundings in gas fields, oil fields, etc., plants that manage and control power generation from hydropower, thermal power, nuclear power, etc., plants that manage and control environmental power generation such as solar power and wind power, plants that manage and control water supply and sewerage, dams, etc.

[0029] In FIG. 1, as an example of the target 20 in the facilities within the plant, a pipe 21 and a tank 22 are shown. As shown in FIG. 1, the pipe 21 may be connected to the tank 22 on the upstream side, for example, and flow the liquid or gas stored in the tank 22 to the downstream side. In FIG. 1, illustration of the equipment on the downstream side of the pipe 21 is omitted.

[0030] The caption generation system 5 includes, as an example, an imaging device terminal 41 used by an imager 31, an administrator terminal 42 used by an administrator 32, a worker terminal 43 used by a recovery worker 33, and a caption generation device 100. The imaging device terminal 41, the administrator terminal 42, the worker terminal 43, and the caption generation device 100 can communicate with each other via a communication network 50. The communication network 50 may be a wired network, a wireless network, or may include both of these.

[0031] The imaging device terminal 41 captures an image 10 of the target 20. Instead of or in addition to this, an imaging device such as a surveillance camera installed in the plant, a camera provided in a traveling robot or drone that patrols the plant, etc. may capture the image 10 of the target 20. The imaging device terminal 41 may be a smartphone with a camera. The imaging device terminal 41 may be a tablet terminal. The imaging device terminal 41 may be a PC (Personal Computer). The imaging device terminal 41 may be a wearable terminal.

[0032] The imaging person 31 is an example of a user who uses the caption generation system 5, and carries and uses the imaging person terminal 41. The imaging person 31 uses the imaging person terminal 41 to capture an image 10 of the target 20 in the facilities within the plant. The imaging person 31 is, as an example, a maintenance worker for the facilities within the plant, and more specifically, a maintenance worker who maintains the tank 22. Note that there may be one or more imaging persons 31 within the plant.

[0033] The administrator terminal 42 displays a report using the caption generated by the caption generation system 5. The administrator terminal 42 transmits an instruction regarding the target 20 in the facilities within the plant to the worker terminal 43 via the communication network 50. The administrator terminal 42 may be a PC, and may be arranged at a location away from the target 20 in the facilities within the plant, such as in the control room within the plant. Note that the administrator terminal 42 may be a smartphone. The administrator terminal 42 may be a tablet terminal. The administrator terminal 42 may be a wearable terminal.

[0034] The administrator 32 is an example of a user who uses the caption generation system 5, and uses the administrator terminal 42. The administrator 32 browses a report using the caption generated by the caption generation system 5 via the administrator terminal 42. The administrator 32, as an example, browses the report in a situation where the target 20 described by the caption cannot be directly visually confirmed. The administrator 32 judges the situation of the target 20 in the facilities within the plant, and orders the recovery worker 33 or the like to perform recovery work or the like on the target 20. The administrator 32, as an example, judges the situation of the pipe 21 in the facilities within the plant, and orders the recovery worker 33 or the like for the pipe 21 to perform recovery work or the like on the pipe 21. Note that there may be one or more administrators 32 within the plant. Note that the administrator 32 is an example of a viewer who browses a report using the caption generated by the caption generation system 5.

[0035] The operator terminal 43 displays instructions regarding the target 20 in the facilities within the plant, reports using the captions generated by the caption generation system 5, etc. The operator terminal 43 may be a smartphone. The operator terminal 43 may be a tablet terminal. The operator terminal 43 may be a PC. The operator terminal 43 may be a wearable terminal.

[0036] The recovery operator 33 is an example of a user who uses the caption generation system 5 and uses the operator terminal 43. The recovery operator 33 browses instructions regarding the target 20 in the facilities within the plant and reports using the captions generated by the caption generation system 5 via the operator terminal 43. As an example, the recovery operator 33 browses the report in a situation where the target 20 explained by the caption cannot be directly visually confirmed. The recovery operator 33 deals with abnormalities within the plant. As an example, the recovery operator 33 deals with an abnormality in the pipe 21. As an example, the recovery operator 33 restores the pipe 21 in which the abnormality has occurred according to the instructions and reports from the administrator 32 displayed on the operator terminal 43. Note that there may be one or a plurality of recovery operators 33 within the plant. Note that the recovery operator 33 is an example of a viewer who browses a report using the caption generated by the caption generation system 5.

[0037] The caption generation device 100 receives, as an example, the image 10 of the target 20 in the facilities within the plant from the imaging terminal 41 via the communication network 50. The caption generation device 100 generates a caption that explains the target 20 in the plant according to the characteristic information of the user who uses the caption generation system 5 using the image 10 received from the imaging terminal 41. The caption generation device 100 generates, as an example, a caption that explains the target 20 in the plant according to the characteristic information of the imaging person 31, the administrator 32, the recovery operator 33, etc. In other words, the caption generation device 100 generates the caption required by the imaging person 31, etc.

[0038] The caption generation device 100 transmits the generated caption to the imaging device terminal 41 via the communication network 50. The caption generation device 100 may be arranged in a control room, an instrument room, etc. inside the plant, or may be arranged outside the plant. Some functions of the caption generation device 100 may be incorporated into the imaging device terminal 41.

[0039] FIG. 2 is a block diagram of the caption generation device 100 according to the first embodiment. In FIG. 2, the flow of data and the like is indicated by arrows. The caption generation device 100 includes an image acquisition unit 110, a word extraction unit 120, a characteristic information acquisition unit 130, a caption generation unit 140, a storage unit 150, and a report storage unit 155.

[0040] The image acquisition unit 110 acquires the image 10 captured inside the plant. As an example, the image acquisition unit 110 receives, from the imaging device terminal 41 via the communication network 50, the user ID of the imaging person 31 who uses the imaging device terminal 41 together with the image 10 of the target 20 inside the plant. The image acquisition unit 110 inputs the received image 10 together with the user ID of the imaging person 31 to the word extraction unit 120 and the characteristic information acquisition unit 130.

[0041] The word extraction unit 120 extracts a plurality of words representing the characteristics of the target 20 imaged in the image 10. The word extraction unit 120 may extract one or more words based on the newly input image 10 using a word extraction model that has learned the relationship between the image 10 and one or more words representing the characteristics of the target 20 imaged in the image 10. The word extraction unit 120 may, for example, read out the word extraction model stored in the storage unit 150. The word extraction unit 120 inputs the plurality of words extracted from the image 10, together with the user ID of the imaging person 31, etc., to the caption generation unit 140, for example, in vector representation.

[0042] The characteristic information acquisition unit 130 acquires characteristic information indicating the characteristics of the user. The characteristics of the user include, for example, the characteristics of the imaging person 31 who captures the image 10, and the characteristics of the viewer who views the report using the caption generated by the caption generation device 100. The characteristics of the user may additionally or alternatively include the type of role of the user.

[0043] The type of role may include at least any one of a plant administrator, a monitor of facilities in the plant, a maintenance worker of facilities in the plant, and a recovery worker who deals with abnormalities in the plant. The monitor may be a person who patrols the plant to monitor the facilities, or a person who monitors the facilities in the plant from a location away from the facilities in the plant, such as a plant monitoring room. The maintenance worker may be a person who periodically inspects, adjusts, or replaces parts so that the facilities in the plant can maintain a normal state. The recovery worker may be a person who repairs the location where an abnormality has occurred in the plant.

[0044] The type of role of the user may be related to a specific object 20 in the plant. Each type of role may be related to one or more objects 20. When the types of roles are different, the related one or more objects 20 may also be different. For example, when the work content of the imaging person 31 shown in FIG. 1 is the maintenance of a plurality of tanks 22, the imaging person 31 is related to the tanks 22 and not related to the piping 21. For example, when the administrator 32 shown in FIG. 1 belongs to a department that manages a plurality of pipings 21, the administrator 32 is related to the piping 21 and not related to the tanks 22. For example, when the recovery worker 33 shown in FIG. 1 specializes in dealing with abnormalities in a plurality of pipings 21, the recovery worker 33 is related to the piping 21 and not related to the tanks 22.

[0045] The characteristic information acquisition unit 130, as an example, may read out from the storage unit 150 characteristic information indicating the characteristics of the imaging person 31, which is associated with the user ID of the imaging person 31, in response to acquiring the user ID of the imaging person 31 together with the image 10. In this case, the characteristic information acquisition unit 130 may extract, from the storage unit 150, the type of role associated with the user ID of the imaging person 31, together with information on a specific target 20 related thereto, in response to acquiring the user ID of the imaging person 31 together with the image 10. Specifically, the characteristic information acquisition unit 130 may extract from the storage unit 150 information indicating that the imaging person 31 is a maintenance worker of the tank 22. In this case, the characteristic information acquisition unit 130 may input, as characteristic information indicating the characteristics of the imaging person 31, information indicating that the imaging person 31 is a maintenance worker of the tank 22 to the caption generation unit 140.

[0046] The characteristic information acquisition unit 130, as another example, may read out from the storage unit 150 characteristic information indicating the characteristics of viewers such as the administrator 32 and the recovery worker 33, which is associated with the user ID of the imaging person 31, in response to acquiring the user ID of the imaging person 31 together with the image 10. In this case, the user ID of the imaging person 31, the user IDs of viewers such as the administrator 32 and the recovery worker 33, and the characteristic information indicating the characteristics of the viewer may be associated and stored in the storage unit 150. In this case, the characteristic information acquisition unit 130 may extract, from the storage unit 150, the type of role of the viewer associated with the user ID of the imaging person 31, together with information on a specific target 20 related thereto, in response to acquiring the user ID of the imaging person 31 together with the image 10. Specifically, the characteristic information acquisition unit 130 may extract from the storage unit 150 information indicating that the viewer is the administrator 32 of the pipe 21, information indicating that the viewer is the recovery worker 33 of the pipe 21, etc. In this case, the characteristic information acquisition unit 130 may input, as characteristic information indicating the characteristics of the viewer, information indicating that the viewer is the administrator 32 of the pipe 21, information indicating that the viewer is the recovery worker 33 of the pipe 21, etc. to the caption generation unit 140.

[0047] The caption generation unit 140 generates a caption that describes the captured object 20 according to the characteristic information, based on the characteristic information and a plurality of words. The caption generation unit 140 may generate a caption based on the newly input characteristic information and a plurality of words, using a caption generation model that has learned the relationship between the characteristic information, one or more words representing the characteristics of the object 20 in the plant, and the caption that describes the object 20 in the plant. The caption generation unit 140 may, for example, read out the caption generation model stored in the storage unit 150. The caption generation unit 140 transmits the generated caption as text data to the imager terminal 41 via the communication network 50, based on the user ID of the imager 31.

[0048] The storage unit 150 stores the user ID of the user in association with the type of the user's role. The storage unit 150 may store, for example, a characteristic information table that associates the user ID of the user, the type of the user's role, and specific object 20 related to the type. The storage unit 150 may also store the above-mentioned word extraction model and caption generation model. Note that the storage unit 150 is an example of a model storage unit.

[0049] The report accumulation unit 155 receives a report from the imager terminal 41 via the communication network 50 and registers it in the memory. The report may be created by the imager 31 appropriately editing the caption using the imager terminal 41. The report accumulation unit 155 is accessed from the administrator terminal 42, the worker terminal 43, etc. via the communication network 50, and the report stored in the memory is read out.

[0050] FIG. 3 is an example of the characteristic information table stored in the storage unit 150. In the characteristic information table of FIG. 3, the user ID is stored as, for example, a four-digit number in the first column, the type of the role is stored in the second column, and the information of one or more objects 20 is stored in the third column. The item names are stored in the first row of the characteristic information table, and the user ID etc. of each user are stored in each row from the second row onward.

[0051] In the example of FIG. 3, the user ID of the imager 31 etc. is stored in the second row of the characteristic information table, the user ID of the administrator 32 etc. is stored in the third row, and the user ID of the recovery worker 33 etc. is stored in the fourth row.

[0052] FIG. 4 is a flowchart showing the flow of the caption generation method according to the first embodiment. As an example, the flow starts when the imager 31 transmits an image 10 of the target 20 in the plant captured by the imager terminal 41 to the caption generation device 100 via the communication network 50.

[0053] The caption generation device 100 acquires the image 10 captured in the plant (step S101). Specifically, the caption generation device 100 receives the user ID of the imager 31 together with the image 10 of the target 20 in the plant from the imager terminal 41 via the communication network 50.

[0054] The caption generation device 100 extracts a plurality of words representing the features of the target 20 imaged in the image 10 (step S102). Specifically, the caption generation device 100 extracts a plurality of words representing the features of the target 20 imaged in the image 10 using the above-described word extraction model.

[0055] The word extraction model learns the relationship between the image 10 and a plurality of image feature amounts, and also learns the relationship between the plurality of image feature amounts and a plurality of words. When a new image 10 is input, the word extraction model extracts a plurality of image feature amounts from the image 10 and outputs, as a vector representation, a plurality of words representing the features of the target 20 imaged in the image 10 from the plurality of image feature amounts. The word extraction model may be a deep learning model having a neural network structure including, for example, a CNN (Convolution Neural Network) and an FC layer (Fully Connected Layer). Instead of the CNN, the word extraction model may use the Self-Attention or Multi-Head Self-Attention of the Transformer.

[0056] The caption generation device 100 acquires characteristic information indicating the characteristics of the user (step S103). Specifically, the caption generation device 100 acquires, from the storage unit 150, the characteristic information indicating the characteristics of the imaging person 31 associated with the user ID of the imaging person 31. More specifically, the caption generation device 100 extracts, from the storage unit 150, the type of role and the information of the specific object 20 related to the type, which are associated with the user ID of the imaging person 31. More specifically, the caption generation device 100 extracts, from the storage unit 150, the information that the imaging person 31 is a maintenance worker of the tank 22.

[0057] The caption generation device 100 generates a caption for explaining the imaged object 20 according to the characteristic information based on the characteristic information and a plurality of words (step S104). Specifically, the caption generation device 100 uses a caption generation model to generate, as text data, a caption for explaining the imaged object 20 according to the characteristic information based on the characteristic information indicating the characteristics of the imaging person 31 and a plurality of words. The caption generation device 100 transmits the generated caption text data to the imaging person terminal 41 via the communication network 50, and thus the flow ends.

[0058] The caption generation model has learned the relationship between the characteristic information indicating the user's characteristics, the vector representation of words representing the features of the target 20, and the caption describing the target 20. When new characteristic information and the vector representations of a plurality of words representing the features of the target 20 captured in the image 10 are input, the caption generation model outputs the text data of the caption describing the target 20 captured in the image 10 from the characteristic information and the vector representations of the plurality of words. The caption generation model may be a deep learning model with a neural network structure including, for example, an RNN (Recurrent Neural Network). Instead of an RNN, the caption generation model may use an LSTM (Long Short-Term Memory), a Transformer, GPT-3 (Generative Pre-trained Transformer-3), GPT-2, a GPT-3-Clone, etc. The caption generation model may use a large number of example sentences as learning data and learn an RNN or the like so as to increase the likelihood of outputting each of the example sentences. The caption generation model may extract the words included in the example sentence, use the vectors of the extracted words as learning input data, use the example sentence as learning output data, and learn to increase the probability of outputting the learning output data for the learning input data.

[0059] Regarding steps S102 to S104, more specifically, the caption generation device 100 uses a word extraction model and a caption generation model to generate a caption that describes the imaged pipe 21 according to the information based on the plurality of words extracted from the image 10 and the information that the imager 31 is a maintenance worker of the tank 22.

[0060] For example, as shown in FIG. 1, when the image 10 shows an abnormality where a crack has occurred in the pipe 21 and white gas is jetting out from there, the caption generation device 100 may generate a caption such as "Do not approach the pipe 21." or a caption such as "Please report that a crack has occurred in the pipe 21 with a pipe ID of ○○ and white gas is jetting out from there." for the imager 31 who is a maintenance worker of the tank 22. The caption generation device 100 transmits the caption to the imager terminal 41 via the communication network 50.

[0061] Additionally or alternatively, when the caption generation device 100 acquires the user ID of the imager 31 together with the image 10, it may acquire characteristic information indicating the characteristics of the administrator 32 who is the viewer and is associated with the user ID of the imager 31. In this case, the caption generation device 100 uses the word extraction model and the caption generation model to generate a caption that explains the imaged pipe 21 according to the information based on a plurality of words extracted from the image 10 and the information that the viewer is the administrator 32 of the pipe 21. For example, the caption generation device 100 may generate a caption such as "A crack has occurred in the pipe 21 with a pipe ID of ○○ and white gas is jetting out from there. Please instruct the recovery worker 33 of the pipe 21 to recover the pipe 21." for the administrator 32 of the pipe 21 and transmit the caption to the imager terminal 41 via the communication network 50.

[0062] Additionally or alternatively, when the caption generation device 100 acquires the user ID of the imaging person 31 together with the image 10, it may acquire characteristic information indicating the characteristics of the recovery worker 33, who is the viewer, associated with the user ID of the imaging person 31. In this case, the caption generation device 100 uses the word extraction model and the caption generation model to generate a caption that describes the imaged pipe 21 according to the information based on the plurality of words extracted from the image 10 and the information that the viewer is the recovery worker 33 of the pipe 21. For example, the caption generation device 100 generates a caption such as "A crack has occurred in the pipe 21 with the pipe ID of ○○, and white gas is jetting out from there." for the recovery worker 33 of the pipe 21, and may transmit the caption to the imaging person terminal 41 via the communication network 50.

[0063] The imaging person 31 checks the caption displayed on the imaging person terminal 41 and transmits a report using the caption from the imaging person terminal 41 to the caption generation device 100 via the communication network 50. The imaging person 31 may appropriately edit the caption to create a report. When the caption generation device 100 receives a report from the imaging person terminal 41, it stores the report in the report storage unit 155. The report stored in the report storage unit 155 of the caption generation device 100 is viewed, for example, by the recovery worker 33 of the pipe 21, and the recovery worker 33 may repair the crack in the pipe 21 or adjust the valve of the pipe 21. The report stored in the report storage unit 155 of the caption generation device 100 is viewed, for example, by the administrator 32 of the pipe 21, and the administrator 32 may order the recovery worker 33 of the pipe 21 to perform a recovery operation or adjust the control parameters of the equipment in which the pipe 21 is installed.

[0064] According to the caption generation device 100 of the first embodiment described above, caption generation for explaining the target 20 is generated from the image 10 of the target 20 captured in the plant according to the characteristic information of the user. Thereby, the caption generation device 100 can generate captions required by each of the image capturer 31, the administrator 32, etc. from the image 10 of the target 20 captured in the plant.

[0065] In the caption generation device 100 according to the first embodiment, when an abnormality occurs in a target 20 other than the specific target 20 related to the type of role of the user indicated by the characteristic information acquired by the characteristic information acquisition unit 130, or when an abnormality occurs in the specific target 20 indicated by the characteristic information, a caption different from the caption generated in such a case may be generated.

[0066] For example, when the image 10 shows an abnormality such as a crack in the tank 22 and white gas is jetting out from there, the caption generation device 100 may generate a caption for the administrator 32 of the pipe 21, such as "A crack has occurred in the tank 22 with tank ID ○○, and white gas is jetting out from there. Please contact the department that manages the tank 22."

[0067] In the caption generation device 100 according to the first embodiment, when generating a caption for explaining the target 20 in the plant according to the characteristic information of the image capturer 31, the caption generated by the caption generation device 100 may include at least one of an instruction to image the captured target 20 from a different angle or a different picture angle, and an instruction to image another target 20 related to the captured target 20, in addition to the caption for explaining the target 20 in the plant. Note that the caption generation device 100 may generate both a caption for the image capturer 31 and a caption for viewers such as the administrator 32, and transmit the caption to the image capturer terminal 41.

[0068] In the caption generation device 100 according to the first embodiment, the word extraction unit 120 may assign a likelihood according to the characteristic information to each of the plurality of words extracted from the image 10. For example, the word extraction unit 120 acquires the user ID of the imaging person 31 together with the image 10 that has imaged the target 20, and extracts, from the storage unit 150, information indicating that the imaging person 31 is a maintenance worker of the tank 22 as characteristic information associated with the user ID of the imaging person 31. In this case, the word extraction unit 120 may assign a high likelihood to words having a high relevance to either the tank 22 or the maintenance worker of the tank 22, and assign a low likelihood to words having a low relevance to these. As a specific example, assume that the image 10 acquired by the word extraction unit 120 shows the tank 22, the pipe 21 connected to the tank 22, the valve attached to the pipe 21, other devices not directly or indirectly connected to the tank 22, and the wall around the tank 22. For example, when the reference range of the output value for each word is 0 to 1, the word extraction unit 120 may set the output value for the word "tank" to 0.9, the output value for the word "pipe" to 0.8, the output value for the word "valve" to 0.6, the output value for the word indicating other devices, and the output value for the word "wall" to 0.2 based on the characteristic information. In this case, the word extraction unit 120 may handle, for example, the probability of the word "tank" as 90%.

[0069] In this case, the caption generation unit 140 may generate a caption by preferentially using words with higher likelihoods among the plurality of words based on the characteristic information and the plurality of words. Thereby, the caption generation device 100 can further improve the accuracy of generating a caption required by the user.

[0070] The caption generation unit 140 may generate a plurality of captions. In this case, the caption generation unit 140 may further assign a priority to the generated captions according to the likelihood of at least one word used for caption generation among the plurality of words. The said priority may be the average value, total value, or integrated value of the likelihoods of the plurality of words. The caption generation unit 140 may transmit the caption with the highest priority among the generated plurality of captions to the imaging device terminal 41. Alternatively, the caption generation unit 140 may transmit the generated plurality of captions to the imaging device terminal 41 with priorities assigned to each of the generated plurality of captions. In this case, the imaging person 31 can judge the suitability of the plurality of captions while referring to the priorities.

[0071] As an example of the case where the caption generation unit 140 generates a plurality of captions, after attaching a negative weight to the words used when generating the first caption among the plurality of words extracted by the word extraction unit 120, the second caption may be generated so that the words used in the first caption are not included in the second caption. As another example of the case where the caption generation unit 140 generates a plurality of captions, after assigning arbitrary priorities to the plurality of words extracted by the word extraction unit 120, each time a word is used in caption generation, the priority of the word is reduced by a predetermined amount, so that the plurality of words extracted by the word extraction unit 120 may be used as a whole.

[0072] The caption generation unit 140 may also output a caption including an instruction to re-image the imaged object 20 when the likelihood of the combination of the plurality of words used for caption generation is lower than a predetermined threshold value. For example, the caption generation unit 140 may output a caption including an instruction to re-image the imaged object 20 when the average value, total value, integrated value, etc. of the likelihoods of the plurality of words used for caption generation are less than a predetermined threshold value stored in the storage unit 150.

[0073] In the caption generation device 100 according to the first embodiment, the storage unit 150 may store a plurality of sentences included in at least any one of the accident case collection, accident response manual, and maintenance history in the plant as example sentences. In this case, the caption generation unit 140 may extract at least one example sentence by searching among the plurality of example sentences stored in the storage unit 150 using at least any one of the plurality of words extracted by the word extraction unit 120.

[0074] The caption generation unit 140 may generate a caption based on the characteristic information acquired by the characteristic information acquisition unit 130, the plurality of words, and the extracted example sentence. Specifically, the caption generation unit 140 may search for an example sentence in which at least any one of the plurality of words is used among the plurality of example sentences stored in the storage unit 150, and generate a caption while referring to the extracted example sentence.

[0075] In this case, the caption generation unit 140 may generate a plurality of captions and assign a similarity to the extracted example sentence. Specifically, the caption generation unit 140 generates a plurality of captions based on the characteristic information and the plurality of words, and searches for an example sentence in which at least any one of the plurality of words is used among the plurality of example sentences stored in the storage unit 150, and may assign a similarity to each of the plurality of captions with the extracted example sentence. The caption generation unit 140 may, for example, convert each of the plurality of captions and the extracted example sentence into a feature vector indicating a combination of a plurality of words included therein, and calculate the inner product of the feature vectors as the similarity. The caption generation unit 140 may transmit the caption with the highest similarity among the generated plurality of captions to the imaging device terminal 41. Alternatively, the caption generation unit 140 may transmit to the imaging device terminal 41 with the similarity assigned to each of the generated plurality of captions. In this case, the imaging person 31 can determine the suitability of the plurality of captions while referring to the similarity.

[0076] The caption generated by the caption generation unit 140 while referring to the extracted example sentences may include, in addition to the caption that describes the target 20 in the plant, a plurality of action options regarding the actions to be taken by the user or instructions for the user. As a result of generating a caption by referring to the example sentences, the caption generation device 100 can present to the user action options or instructions incorporating abnormal handling methods that cannot be directly decoded from the image 10.

[0077] FIG. 5 is a flowchart showing an example of the detailed flow of the caption generation method according to the first embodiment. The operation flow shown in the flowchart of FIG. 5 may be a specific example of the caption generation method according to the first embodiment described with reference to FIG. 4. The caption generation device 100 according to the first embodiment described above may generate, as an example, a caption that describes the target 20 from the image 10 of the target 20 captured, according to the operation flow shown in the flowchart of FIG. 5.

[0078] The flow of FIG. 5 is started, as an example, when the imager 31 transmits the image 10 of the target 20 in the plant captured by the imager terminal 41 to the caption generation device 100 via the communication network 50.

[0079] The caption generation device 100 acquires the image 10 captured in the plant (step S201). In step S201, as shown in FIG. 6 that specifically describes a part of the flow of FIG. 5, the image acquisition unit 110 receives the image 10 of the target 20 in the plant from the imager terminal 41 via the communication network 50 and extracts Exif information from the image 10. The Exif information may include the user ID of the imager 31. The image acquisition unit 110 may output the image 10 to the word extraction unit 120 and output the Exif information to the characteristic information acquisition unit 130.

[0080] The caption generation device 100 crops the image 10 into a plurality of parts (step S202). For example, the caption generation device 100 creates a new image by discarding the peripheral area in the image 10 that does not contain the object 20, and divides it into patches of a fixed size. In step S202, as shown in FIG. 6, the word extraction unit 120 may create a new image by discarding the peripheral area in the image 10 that does not contain the pipe 21 and the tank 22, and divide the new image into nine patches by dividing it vertically and horizontally into three parts. Alternatively, the word extraction unit 120 may create a new image by discarding the peripheral area in the image 10 that does not contain the pipe 21 as a patch of a fixed size, and create a new image by discarding the peripheral area in the image 10 that does not contain the tank 22 as another patch of a fixed size. Note that the learning model of the word extraction unit 120 shown in FIG. 6 may also be referred to as a Vision Transformer and may be an example of the above-described word extraction model.

[0081] The caption generation device 100 converts the cropped image through a simple function, for example, a LINEAR function, into a numerical value A that can be used by the encoder (step S203). In step S203, as shown in FIG. 6, the word extraction unit 120 may put the nine patches into a linear projection layer, flatten each of them, and embed them (linearly) into a vector (or convert them into a one-dimensional array). The word extraction unit 120 may further add a CLS token to the beginning of the sequence of vectors output from the linear projection layer (extra learnable [class] embedding), and embed the position into each patch (each vector) (position embedding) to obtain the numerical value A. That is, the word extraction unit 120 may convert the nine patches divided from the image 10 into a sequence of 10-dimensional vectors that can be used by the encoder.

[0082] The caption generation device 100 causes the transformer encoder to convert a numerical value A into another value B (step S204) and unify the dimension and numerical width of the value B (step S205). In steps S204 to S205, as shown in FIG. 6, the word extraction unit 120 inputs a sequence of 10-dimensional vectors to the transformer encoder, and the transformer encoder converts this into another sequence of 10-dimensional vectors, and unifies the dimension and numerical width through, for example, a softmax function, and may output it as the output of the CLS token. Note that the transformer encoder may include a plurality of components such as Self-Attention described above. The output of the CLS token may be an aggregation of feature amounts necessary for classification from the entire image by Self-Attention.

[0083] The caption generation device 100 outputs image feature amounts (step S206). In step S206, as shown in FIG. 6, the word extraction unit 120 may input the output of the CLS token to a classification head (MLP: multi-layer perceptron) and output a plurality of image feature amounts from the classification head. The classification head may output after compressing the plurality of image feature amounts. The word extraction unit 120 may output a plurality of image feature amounts to the caption generation unit 140 as shown in FIG. 6, and as described with reference to FIGS. 1 to 4, extract word vectors of a plurality of words representing the features of the object 20 imaged in the image 10 from the plurality of image feature amounts and output them to the caption generation unit 140.

[0084] The caption generation device 100 acquires characteristic information (step S207), converts the image feature amount and the characteristic information into a numerical value C suitable for the decoder (step S208), and inputs the numerical value C into the decoder (step S209). In steps S207 to S209, as shown in FIG. 7, which specifically describes a part of the flow of FIG. 5, the caption generation unit 140 vectorizes the characteristic information input from the characteristic information acquisition unit 130 through, for example, a LINEAR function, and also vectorizes an empty text, and vectorizes these vectors together with a plurality of image feature amounts input from the word extraction unit 120. Thereby, the caption generation unit 140 converts each of the plurality of image feature amounts into the numerical value C. The caption generation unit 140 further inputs the numerical value C into the decoder. Note that the learning model of the decoder is a model using layer-by-layer adjustment, which may also be referred to as GPT-2 and may be an example of the above-described caption generation model.

[0085] The caption generation device 100 causes the decoder to convert the numerical value C into a numerical value D (step S210), and converts it into text (tokens) through a function that returns the numerical value D to text (step S211). In steps S210 to S211, as shown in FIG. 7, the decoder in the caption generation unit 140 may convert the numerical value C into the numerical value D. The decoder may weight each numerical value D based on the vectorized characteristic information. The decoder may compress the numerical value D. The decoder may convert the numerical value D into text (tokens) through a softmax function.

[0086] The caption generation device 100 outputs text (caption) (step S212), and this flow ends. In step S212, as shown in FIG. 7, for example, the decoder takes a sequence of consecutive texts X1 to X5 as input (source tokens), and uses a sequence of consecutive texts Y1 to Y5 that is the same sequence as the above sequence but with one token shifted to the right as target tokens. The source tokens and target tokens may be concatenated and processed with adjustment for each layer. For the source tokens and target tokens, their respective re - settable positions may be embedded together with their corresponding tokens. The position of the source tokens starts from zero, and for the target tokens, instead of increasing the position at the end of the source sentence, the position may be reset to zero again.

[0087] The decoder model shown in FIG. 7, as an example, includes an attention mechanism including a self - attention mechanism and a mixed - attention mechanism, an add - and - layer - normalization layer, a feed - forward (FNN) layer, and N layers stacked in this order with the add - and - layer - normalization layer. The decoder may put the processed result of N layers into a linearization layer, pass the output from the linearization layer through a softmax function, and output a caption consisting of a sequence of consecutive texts Y1 to Y5. The caption generation unit 140 may transmit the generated caption text data to the imaging device terminal 41 via the communication network 50.

[0088] FIG. 8 is another example of the characteristic information table stored in the storage unit 150. In the characteristic information table of FIG. 8, a fourth column is added to the characteristic information table illustrated in FIG. 3, and state information is stored in the fourth column.

[0089] In the caption generation device 100 according to the first embodiment, the storage unit 150 may store, in association with each other, the type of role and the state information indicating the state of the target 20 in the plant. As illustrated in FIG. 8, the storage unit 150 may store, in association with each other, the user ID, the type of role, a specific target 20 related to the type, and the state information.

[0090] The state information may include at least any one of the position information of the target 20, the degradation information indicating the degradation state of the target 20, the fluid information indicating the component or state of the fluid flowing inside the pipe that is the target 20, and the equipment information related to the equipment located on the upstream side or the downstream side of the pipe that is the target 20. The state information may include, for example, fixed information such as the position information of the pipe 21 or the information of the tank 22 upstream of the pipe 21, and may also include variable information indicating the type of gas or liquid flowing inside the pipe 21, the state such as the pressure and temperature of the gas or liquid, and the degree of degradation of the pipe 21. The degree of degradation of the pipe 21 may be an index of the degradation state according to the degree of rusting of the pipe 21 or the like. Such variable information may be stored in the storage unit 150 in advance, for example, by a user such as a maintenance worker periodically inspecting the target 20 or measuring the pressure value of the internal gas of the target 20.

[0091] In this case, the characteristic information acquisition unit 130 may acquire the characteristic information from the storage unit 150 and extract the state information associated with the type of role indicated by the characteristic information. The caption generation unit 140 may generate a caption that describes the imaged target 20 according to the characteristic information and the state information, based on the characteristic information, a plurality of words, and the state information.

[0092] In the example of FIG. 8, as information such as the user ID of the imaging person 31 stored in the second row of the characteristic information table, the imaging person 31 is a maintenance worker of the tank 22, and the tank 22 is located in area A, the degree of deterioration is 4, and it is shown that sulfuric acid gas is stored. As information such as the user ID of the administrator 32 stored in the third row of the characteristic information table, the administrator 32 is an administrator of the pipe 21, and the pipe 21 is located in area A, the degree of deterioration is 9, sulfuric acid gas is flowing inside, and it is shown that the tank 22 is connected to the upstream side. As information such as the user ID of the recovery worker 33 stored in the fourth row of the characteristic information table, the recovery worker 33 is a recovery worker of the pipe 21, and the pipe 21 is located in area A, the degree of deterioration is 9, sulfuric acid gas is flowing inside, and it is shown that the tank 22 is connected to the upstream side.

[0093] When the characteristic information table of FIG. 8 is stored in the storage unit 150, for example, as shown in FIG. 1, when the image 10 shows an abnormality that a crack has occurred in the pipe 21 and white gas is jetting out therefrom, the caption generation device 100 generates a caption for the administrator 32 of the pipe 21, "A crack has occurred in the pipe 21 with a pipe ID of ○○ in area A, and sulfuric acid gas is jetting out therefrom. Please order the recovery worker 33 of the pipe 21 to recover the pipe 21. There is a possibility that the crack in the pipe 21 has occurred due to the deterioration of the rust in the pipe 21. Please order to close the valve of the tank 22 on the upstream side of the pipe 21 before performing the recovery work." and transmit the caption to the imaging person terminal 41 via the communication network 50.

[0094] For example, when the image 10 shows an abnormality where a crack has occurred in the tank 22 and white gas is jetting out from there, the caption generation device 100 may generate a caption for the imager 31 who is the maintenance staff of the tank 22, such as "Please report that a crack has occurred in the tank 22 with a tank ID of ○○ in Area A and sulfuric acid gas is jetting out from there. Also, please report the results of the maintenance inspection of the tank 22. Please also report that there is a possibility that the crack in the tank 22 occurred due to factors other than rust." and transmit the caption to the imager terminal 41 via the communication network 50.

[0095] When the characteristic information acquisition unit 130 acquires the characteristic information and the state information, the caption generation unit 140 may generate a caption for explaining the imaged object 20 according to the characteristic information and the state information corresponding to the type of role, based on the characteristic information and a plurality of words for each of the plurality of predetermined role types. The caption generation unit 140 may further select and output at least one caption from the plurality of generated captions based on the state information extracted by the characteristic information acquisition unit 130.

[0096] For example, when there are an administrator 32 and a recovery worker 33 as viewers, the caption generation unit 140 may generate a caption for the administrator 32 and a caption for the recovery worker 33 respectively based on the characteristic information of the administrator 32, the characteristic information of the recovery worker 33, and a plurality of words. The caption generation unit 140 may further select the caption for the administrator 32 based on the state information corresponding to the administrator 32, or select the caption for the recovery worker 33 based on the state information corresponding to the recovery worker 33.

[0097] FIG. 9 is a block diagram of a caption generation device 200 according to the second embodiment. The caption generation device 200 according to the second embodiment includes a learning unit 260 in addition to the configuration included in the caption generation device 100 according to the first embodiment. Other configurations included in the caption generation device 200 according to the second embodiment are the same as those of the caption generation device 100 according to the first embodiment, and duplicate descriptions are omitted using the reference numerals of the respective configurations included in the caption generation device 100.

[0098] The learning unit 260 learns the above-described word extraction model and caption generation model using the result of the user's determination of the appropriateness of the caption generated by the caption generation unit 140. In addition to or instead of this, the learning unit 260 may learn the above-described word extraction model and caption generation model using the user's correction input for the caption generated by the caption generation unit 140. The learning unit 260 receives the above-described determination result or correction input by the user from the imaging device terminal 41 via the communication network 50.

[0099] The learning unit 260 updates the parameters of the model so as to reduce the error between the output of the model when each sample in the learning data including the above-described determination result is input to the model and the label. For example, when learning a word extraction model using a multi-layer neural network including CNN or the like, the learning unit 260 uses the error between the output value output by the neural network in response to input of each sample and the label, and uses a method such as backpropagation to adjust the weights between the neurons of the neural network and the biases of each neuron.

[0100] In this way, for the caption generation device 200 according to the second embodiment, information on whether the generated caption has been actually used by the user is fed back, and the word extraction model and the caption generation model are learned using the feedback content. For example, when the caption generation device 200 receives from the administrator terminal 42 the result that the administrator 32 determines that there is no problem as a result of viewing the caption, it may be determined that the generated caption has been actually used. The caption generation device 200 may also, for example, determine that the generated caption has been actually used when the caption generation unit 140 performs a search for example sentences using the generated caption, or when the generated caption is used in a report for the administrator 32. The caption generation device 200 according to such a second embodiment also has the same effects as the caption generation device 100 according to the first embodiment. According to the caption generation device 200 according to the second embodiment, the accuracy of generating captions required by the user can be further improved.

[0101] In the above-described plurality of embodiments, the characteristic information acquisition unit 130 may acquire a plurality of words instead of or in addition to the combination of the image 10 and the user ID of the imaging person 31. In response to acquiring the plurality of words, the words related to the target 20 included in the plurality of words may be specified, and the characteristic information associated with the target 20 may be read from the storage unit 150. For example, the characteristic information acquisition unit 130 may specify the name of the pipe 21 included in the plurality of words, and read the characteristic information of the administrator 32 and / or the recovery worker 33 associated with the pipe 21 from the storage unit 150. In this case, the characteristic information acquisition unit 130 may input to the caption generation unit 140 information indicating that the viewer is the administrator 32 of the pipe 21, information indicating that the viewer is the recovery worker 33 of the pipe 21, etc. as the characteristic information indicating the characteristics of the viewer.

[0102] Various embodiments of the present invention may be described with reference to flowcharts and block diagrams, where the blocks may represent (1) stages of a process in which an operation is performed or (2) sections of an apparatus having a role of performing an operation. Specific stages and sections may be implemented by a dedicated circuit, a programmable circuit supplied with computer-readable instructions stored on a computer-readable medium, and / or a processor supplied with computer-readable instructions stored on a computer-readable medium. The dedicated circuit may include digital and / or analog hardware circuits, and may include an integrated circuit (IC) and / or discrete circuits. The programmable circuit may include a reconfigurable hardware circuit including memory elements such as logical AND, logical OR, logical XOR, logical NAND, logical NOR, and other logical operations, flip-flops, registers, a field programmable gate array (FPGA), a programmable logic array (PLA), etc.

[0103] The computer-readable medium may include any tangible device capable of storing instructions executable by an appropriate device, and as a result, a computer-readable medium having instructions stored therein will comprise a product including instructions executable to create means for performing the operations specified in the flowchart or block diagram. Examples of the computer-readable medium may include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, etc. More specific examples of the computer-readable medium may include a floppy (registered trademark) disk, a diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an electrically erasable programmable read-only memory (EEPROM), a static random access memory (SRAM), a compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a Blu-ray (registered trademark) disc, a memory stick, an integrated circuit card, etc.

[0104] Computer-readable instructions may include source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk®, JAVA®, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0105] Computer-readable instructions may be provided locally or via a wide area network (WAN) such as a local area network (LAN), the Internet, etc., to a processor or programmable circuit of a programmable data processing apparatus such as a general-purpose computer, a special-purpose computer, or other computer, and may be executed to create means for performing the operations specified in a flowchart or block diagram. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.

[0106] FIG. 10 shows an example of a computer 2200 in which multiple aspects of the present invention may be embodied in whole or in part. Programs installed on the computer 2200 can cause the computer 2200 to function as an operation associated with the apparatus according to an embodiment of the present invention or as one or more sections of the apparatus, or can cause the operation or the one or more sections to be executed, and / or can cause the computer 2200 to execute a process according to an embodiment of the present invention or a stage of the process. Such a program may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.

[0107] The computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphic controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes an input / output unit such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.

[0108] The CPU 2212 operates according to programs stored in the ROM 2230 and the RAM 2214, thereby controlling each unit. The graphic controller 2216 acquires image data generated by the CPU 2212 in a frame buffer or the like provided in the RAM 2214 or in itself, and causes the image data to be displayed on the display device 2218.

[0109] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads a program or data from the DVD-ROM 2201 and provides the program or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to the IC card.

[0110] The ROM 2230 stores therein a boot program or the like executed by the computer 2200 upon activation and / or a program dependent on the hardware of the computer 2200. The input / output chip 2240 may also be connected to the input / output controller 2220 via various input / output units such as a parallel port, a serial port, a keyboard port, a mouse port, etc.

[0111] The program is provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The program is read from the computer-readable medium, installed in a hard disk drive 2224, a RAM 2214, or a ROM 2230 which are also examples of computer-readable media, and executed by the CPU 2212. The information processing described in these programs is read by the computer 2200, resulting in the cooperation between the programs and the various types of hardware resources described above. The apparatus or method may be configured by realizing the operation or processing of information according to the use of the computer 2200.

[0112] For example, when communication is executed between the computer 2200 and an external device, the CPU 2212 may execute a communication program loaded in the RAM 2214 and instruct the communication interface 2222 to perform communication processing based on the processing described in the communication program. The communication interface 2222 reads the transmission data stored in a transmission buffer processing area provided in a recording medium such as the RAM 2214, the hard disk drive 2224, the DVD-ROM 2201, or an IC card under the control of the CPU 2212, transmits the read transmission data to the network, or writes the received data received from the network to a reception buffer processing area or the like provided on the recording medium.

[0113] Further, the CPU 2212 may cause all or necessary parts of files or databases stored in external recording media such as a hard disk drive 2224, a DVD-ROM drive 2226 (DVD-ROM 2201), and an IC card to be read into the RAM 2214, and may execute various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording media.

[0114] Various types of information such as various types of programs, data, tables, and databases may be stored in the recording media and may undergo information processing. The CPU 2212 may perform various types of processing on the data read from the RAM 2214, including various types of operations, information processing, condition judgment, conditional branch, unconditional branch, information search / replacement, etc., described throughout this disclosure and specified by the instruction sequence of the program, and write back the results to the RAM 2214. Further, the CPU 2212 may search for information in files, databases, etc. within the recording media. For example, when a plurality of entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored in the recording media, the CPU 2212 searches for an entry that matches the condition where the attribute value of the first attribute is specified from among the plurality of entries, reads the attribute value of the second attribute stored in the entry, and thereby may obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.

[0115] The programs or software modules described above may be stored on a computer-readable medium on or near the computer 2200. Also, a recording medium such as a hard disk or RAM provided within a server system connected to a dedicated communication network or the Internet can be used as a computer-readable medium, thereby providing the program to the computer 2200 via the network.

[0116] As described above, the present invention has been described using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments. It is obvious to those skilled in the art that various changes or improvements can be made to the above embodiments. It is clear from the description of the claims that forms with such changes or improvements can also be included in the technical scope of the present invention.

[0117] For example, the control system may be a computer housed in a single housing. That is, the controller may be realized by executing a program on a processor of the computer, and each input / output device may be implemented as an I / O device of the computer. Further, the controller may be implemented as a virtual machine executed by one or more processors. In such a configuration, the control system does not include a network that is a general-purpose or dedicated network, and the controller and the input / output devices can be connected by a chipset such as a memory controller hub and an I / O controller hub that connect between the processor and the I / O devices.

[0118] It should be noted that the execution order of each process such as operations, procedures, steps, and stages in the devices, systems, programs, and methods shown in the claims, the specification, and the drawings is not explicitly indicated as "before" or "preceding" etc., and can be realized in any order unless the output of the previous process is used in the subsequent process. Regarding the operation flows in the claims, the specification, and the drawings, even if "first," "next," etc. are used for convenience of explanation, it does not mean that it is essential to implement in this order.

Explanation of Reference Numerals

[0119] 5 Caption generation system 10 Image 20 Target 21 Pipe 22 Tank 31 Imager 32 Administrator 33 Recovery worker 41 Imager terminal 42 Manager terminal 43 Operator terminal 100 Caption generation device 110 Image acquisition unit 120 Word extraction unit 130 Characteristic information acquisition unit 140 Caption generation unit 150 Memory unit 155 Report storage unit 200 Caption generation device 260 Learning unit 2200 Computer 2201 DVD-ROM 2210 Host controller 2212 CPU 2214 RAM 2216 Graphics controller 2218 Display device 2220 Input / output controller 2222 Communication interface 2224 Hard disk drive 2226 DVD-ROM drive 2230 ROM 2240 Input / output chip 2242 Keyboard

Claims

1. An image acquisition unit that acquires an image captured within a plant; A word extraction unit that extracts a plurality of words representing features of an object imaged in the image; A characteristic information acquisition unit that acquires characteristic information indicating characteristics of a user; A caption generation unit that generates a caption for explaining the imaged object according to the characteristic information based on the characteristic information and the plurality of words; A caption generation device comprising the above.

2. The characteristics of the user include at least one of the characteristics of the imager who captured the image and the characteristics of the viewer who views a report using the generated caption. The caption generation device according to Claim 1.

3. The characteristics of the user include the type of the user's role. The caption generation device according to Claim 1.

4. The type of the role includes at least one of an administrator of the plant, a monitor of equipment within the plant, a maintenance worker of equipment within the plant, and a recovery worker who deals with abnormalities within the plant. The caption generation device according to Claim 3.

5. The type of the role is related to a specific object within the plant. When an abnormality has occurred in an object other than the specific object related to the type of the role of the user indicated in the characteristic information, the caption generation unit generates a caption different from the caption generated when an abnormality has occurred in the specific object indicated in the characteristic information. The caption generation device according to Claim 4.

6. The caption generation device further comprises a storage unit that stores the user ID of the user in association with the type of the role. The characteristic information acquisition unit extracts the type of the role associated with the user ID in response to acquiring the user ID. The caption generation device according to Claim 3.

7. The caption generation device further comprises a storage unit that stores the type of the role in association with state information indicating the state of an object within the plant. The characteristic information acquisition unit acquires the characteristic information and extracts the state information associated with the type of the role indicated in the characteristic information. The caption generation unit generates the caption for explaining the imaged object according to the characteristic information and the state information based on the characteristic information, the plurality of words, and the state information. The caption generation device according to Claim 3.

8. It further includes a storage unit that stores by associating the type of the role with state information indicating the state of the target within the plant. The characteristic information acquisition unit acquires the characteristic information and extracts the state information associated with the type of the role indicated by the characteristic information. The caption generation unit For each of a plurality of predetermined types of the role, based on the characteristic information and the plurality of words, generates a caption that describes the captured target according to the characteristic information and the state information corresponding to the type of the role. Based on the state information extracted by the characteristic information acquisition unit, at least one caption is selected from the plurality of captions generated for the plurality of types of the role and output. The caption generation device according to claim 3.

9. The state information includes at least any one of position information of the target, deterioration information indicating a deterioration state of the target, fluid information indicating a component or state of a fluid flowing inside a pipe that is the target, and equipment information related to equipment located upstream or downstream of the pipe that is the target. The caption generation device according to claim 7 or 8.

10. The word extraction unit assigns a likelihood corresponding to the characteristic information to each of the plurality of words extracted from the image. The caption generation unit generates the caption by preferentially using words with higher likelihoods among the plurality of words based on the characteristic information and the plurality of words. The caption generation device according to claim 1.

11. The caption generation unit assigns a priority corresponding to the likelihood of at least one word used for caption generation among the plurality of words to the generated caption. The caption generation device according to claim 10.

12. When the likelihood of a combination of a plurality of words used for caption generation is lower than a predetermined threshold, the caption generation unit outputs a caption including an instruction to re-capture the captured target. The caption generation device according to claim 10.

13. It further includes a storage unit that stores a plurality of sentences included in at least any one of an accident case collection, an accident response manual, and a maintenance history within the plant as example sentences. The caption generation unit extracts at least one of the examples by searching using at least any one of the plurality of words extracted from among the plurality of examples stored in the storage unit, and generates the caption based on the characteristic information, the plurality of words, and the extracted example. The caption generation device according to claim 1.

14. The caption generation unit generates a plurality of captions and assigns a similarity to the extracted example. The caption generation device according to claim 13.

15. The caption includes a plurality of action options regarding actions to be taken by the user or an instruction to the user. The caption generation device according to claim 13.

16. The caption includes at least any one of an instruction to image the captured object from a different angle or a different view angle and an instruction to image another object related to the captured object. The caption generation device according to claim 1.

17. The caption generation device further includes a model storage unit that stores a caption generation model that has learned the relationship between the characteristic information, one or more words representing the characteristics of the object in the plant, and the caption explaining the object in the plant. The caption generation unit generates the caption based on the newly input characteristic information and the plurality of words using the caption generation model. The caption generation device according to claim 1.

18. The caption generation device further includes a word extraction model that extracts the plurality of words from the image using the result of the user's determination of the appropriateness of the generated caption, and a learning unit that learns the caption generation model that generates the caption from the characteristic information and the plurality of words. The caption generation unit generates the caption based on the newly input characteristic information and the plurality of words using the caption generation model. The caption generation device according to claim 1.

19. The caption generation device further includes a word extraction model that extracts the plurality of words from the image using the user's correction input for the generated caption, and a learning unit that learns the caption generation model that generates the caption from the characteristic information and the plurality of words. The caption generation unit generates the caption based on the newly input characteristic information and the plurality of words using the caption generation model. The caption generation device according to claim 1.

20. Obtaining an image captured in a plant; Extracting a plurality of words representing features of an object imaged in the image; Obtaining characteristic information indicating characteristics of a user; Generating a caption for explaining the imaged object according to the characteristic information based on the characteristic information and the plurality of words A caption generation method comprising:

21. On a computer, A procedure for obtaining an image captured in a plant; A procedure for extracting a plurality of words representing features of an object imaged in the image; A procedure for obtaining characteristic information indicating characteristics of a user; A procedure for generating a caption for explaining the imaged object according to the characteristic information based on the characteristic information and the plurality of words A program for causing the computer to execute the above procedures.

Citation Information

Patent Citations

  • Facility management system and method

    JP2015125745A

  • Display system, wearable device and monitoring control device

    JP2019219917A