Caption generation apparatus, caption generation method and program

The caption generation device improves industrial plant monitoring by generating accurate captions from plant images using image and state information, addressing inefficiencies in existing systems.

JP2025098517APending Publication Date: 2025-07-02YOKOGAWA ELECTRIC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023214705
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-07-02

AI Technical Summary

Technical Problem

Existing systems lack the ability to accurately generate captions that describe the state of objects within industrial plants using image analysis and state information, leading to inefficiencies in maintenance and monitoring.

Method used

A caption generation device that includes an image acquisition unit, word extraction unit, and state information acquisition unit to generate captions based on the state information and extracted words from plant images, utilizing models to improve accuracy and relevance.

Benefits of technology

Enhances the accuracy of caption generation, allowing for more precise descriptions of plant object states, facilitating better maintenance and monitoring through improved image analysis and state information integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025098517000001_ABST
    Figure 2025098517000001_ABST
Patent Text Reader

Abstract

To provide a caption generation apparatus, a caption generation method, and a program for explaining a caption that describes an object captured, from an image obtained by imaging the inside of a plant.SOLUTION: A caption generation apparatus 100 includes: an image acquisition unit 110 which acquires an image captured in a plant; a word extraction unit 120 which extracts a plurality of words representing characteristics of an object captured in the image; a state information acquisition unit 130 which acquires state information indicating the state of the captured object; and a caption generation unit 140 which generates a caption that describes the captured object according to the state information, based on the state information and the words. The state information includes at least one of position information of the object, degradation information indicating degradation state of the object, fluid information indicating components or the state of a fluid flowing inside a target piping, and device information related to a device located upstream or downstream of the target piping.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a caption generation device, a caption generation method, and a program.

Background Art

[0002] Patent Document 1 describes that "the video function abnormality determination means 16 has a function of inputting the result of the abnormality determination by the abnormality determination means 12 and determining from the result, the data indicating the state of the plant input from the input means 11, and the video of the plant situation of the ITV camera 19 that there is an abnormality in the function elements obtained by subdividing the plant function" (paragraph 0024). Patent Document 2 describes that "when the alarm condition is satisfied, the display 5 displays an alarm display screen using the alarm function. On the alarm display screen, the date and time when the abnormality occurred and a comment explaining the content of the abnormality are displayed." (paragraph 0039). Patent Document 3 describes that "obtain position information specified by GPS, display the current inspection location for plant staff, accumulate and analyze the image data captured by the installed camera, and determine whether there are signs of abnormality in the plant equipment from the past equipment state and abnormality cases, etc., and at the same time, use a preset format / document to automatically create a report of the regular inspection from the image data" (paragraph 0053). Patent Document 4 describes that "a feature amount acquisition unit that acquires information about the pipe as a first feature amount from an image obtained by photographing the plant equipment of the work target and the pipes existing around the plant equipment with the camera, and a feature amount comparison unit that compares the first feature amount with a second feature amount about the pipe acquired from design data" (Claim 1). Patent Document 5 describes that "the abnormal mail creation function 102 is activated when the abnormality monitoring function 101 detects a plant abnormality, and formulates matters to be communicated to the supervisor at an early stage, such as the abnormality, the detected date and time, the name of the corresponding plant equipment, and the content of the abnormality, to create a mail transmission text" (paragraph 0019). [Prior Art Documents] [Patent Documents] [Patent Document 1] Japanese Patent Application Laid-Open No. 2000-76570 [Patent Document 2] Japanese Patent No. 6419399 [Patent Document 3] Japanese Patent No. 6099989 [Patent Document 4] Japanese Patent No. 6826509 [Patent Document 5] Japanese Patent Application Laid-Open No. 2003-51895

Summary of the Invention

[0003] In a first aspect of the present invention, a caption generation device is provided. The caption generation device includes an image acquisition unit that acquires an image captured in a plant, a word extraction unit that extracts a plurality of words representing the features of the object imaged in the image, a state information acquisition unit that acquires state information indicating the state of the imaged object, and a caption generation unit that generates a caption for explaining the imaged object according to the state information based on the state information and the plurality of words.

[0004] In the caption generation device described above, the state information may include at least any one of the position information of the object, the deterioration information indicating the deterioration state of the object, the fluid information indicating the component or state of the fluid flowing inside the pipe that is the object, and the equipment information related to the equipment located upstream or downstream of the pipe that is the object.

[0005] In any of the caption generation devices described above, the state information acquisition unit may acquire the state information by analyzing the image.

[0006] In any of the above caption generation devices, the state information may include the position information of the target. Any of the above caption generation devices may further include a storage unit that stores by associating the past image of the target in the plant with the position information. Any of the above caption generation devices may further include an image extraction unit that extracts the past image having position information common to the position information included in the new state information in response to acquiring the new state information of the new image. In any of the above caption generation devices, the caption generation unit may generate the caption using both the new image and the past image corresponding to the new image.

[0007] In any of the above caption generation devices, the word extraction unit may assign a likelihood according to the state information to each of the plurality of words extracted from the image. In any of the above caption generation devices, the caption generation unit may generate the caption by preferentially using the words having a higher likelihood among the plurality of words based on the state information and the plurality of words.

[0008] In any of the above caption generation devices, the caption generation unit may assign a priority according to the likelihood of at least one word used for caption generation among the plurality of words to the generated caption.

[0009] In any of the above caption generation devices, when the likelihood of the combination of the plurality of words used for caption generation is lower than a predetermined threshold, the caption generation unit may output the caption including an instruction to re-capture the imaged target.

[0010] Any of the above caption generation devices may further include a storage unit that stores a plurality of sentences included in at least any one of the accident case collection, accident response manual, and maintenance history in the plant as example sentences. In any of the above caption generation devices, the caption generation unit may extract at least one of the example sentences by searching using at least any one of the plurality of words extracted from among the plurality of example sentences stored in the storage unit, and generate the caption based on the state information, the plurality of words, and the extracted example sentence.

[0011] In any of the above caption generation devices, the caption generation unit may generate a plurality of captions and assign a similarity degree to the extracted example sentences.

[0012] In any of the above caption generation devices, the caption may include a plurality of action options regarding actions to be taken by the user or instructions for the user.

[0013] In any of the above caption generation devices, the caption may include at least any one of an instruction to image the captured object from a different angle or a different shooting angle, and an instruction to image another object related to the captured object.

[0014] Any of the above caption generation devices may further include a model storage unit that stores a caption generation model that has learned the relationship between the state information, one or more words representing the characteristics of the object in the plant, and the caption describing the object in the plant. In any of the above caption generation devices, the caption generation unit may generate the caption based on the newly input state information and the plurality of words using the caption generation model.

[0015] Any of the above caption generation devices may further include a learning unit that learns a word extraction model that extracts the plurality of words from the image using the result of the user's determination of the suitability of the generated caption, and a caption generation model that generates the caption from the state information and the plurality of words. In any of the above caption generation devices, the caption generation unit may generate the caption based on the newly input state information and the plurality of words using the caption generation model.

[0016] Any of the above caption generation devices may further include a learning unit that learns a word extraction model that extracts the plurality of words from the image using the user's correction input for the generated caption, and a caption generation model that generates the caption from the state information and the plurality of words. In any of the above caption generation devices, the caption generation unit may generate the caption based on the newly input state information and the plurality of words using the caption generation model.

[0017] In a second aspect of the present invention, a caption generation method is provided. The caption generation method includes obtaining an image captured in a plant, extracting a plurality of words representing features of an object captured in the image, obtaining state information indicating a state of the captured object, and generating a caption that describes the captured object according to the state information based on the state information and the plurality of words.

[0018] In a third aspect of the present invention, a program is provided. The program causes a computer to execute a procedure for obtaining an image captured in a plant, a procedure for extracting a plurality of words representing features of an object captured in the image, a procedure for obtaining state information indicating a state of the captured object, and a procedure for generating a caption that describes the captured object according to the state information based on the state information and the plurality of words.

[0019] Note that the above summary of the invention does not enumerate all the features of the present invention. Also, sub - combinations of these feature groups can also be inventions.

Brief Description of Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0021] Hereinafter, the present invention will be described through embodiments of the invention. However, the following embodiments do not limit the invention according to the claims. Also, not all combinations of features described in the embodiments are essential for the solution means of the invention.

[0022] FIG. 1 is a schematic diagram of a caption generation system 5 including a caption generation device 100 according to the first embodiment. The caption generation system 5 generates a caption that describes the imaged object 20 from an image 10 obtained by imaging the object 20 in a facility within a plant. More specifically, the caption generation system 5 generates a caption that describes the imaged object 20 according to state information indicating the state of the imaged object 20. The image 10 may be one or more still images or a moving image. Note that the caption for describing the object 20 in the image 10 may refer to a descriptive text that describes the state of the object 20 shown in the image 10. Here, the caption may refer to a sentence separated by a period, or may refer to a clause more finely separated than a sentence.

[0023] Examples of the plant include industrial plants such as chemical plants, plants that manage and control wells and their surroundings in gas fields and oil fields, plants that manage and control power generation such as hydraulic power, thermal power, and nuclear power, plants that manage and control environmental power generation such as solar power and wind power, and plants that manage and control water supply and sewerage, dams, and the like.

[0024] FIG. 1 shows a pipe 21 and a tank 22 as an example of the object 20 in a facility within a plant. As shown in FIG. 1, the pipe 21 may be connected to the tank 22 on the upstream side, for example, and may flow the liquid or gas stored in the tank 22 to the downstream side. Note that in FIG. 1, illustration of the equipment on the downstream side of the pipe 21 is omitted.

[0025] The caption generation system 5 includes, as an example, an imaging device terminal 41 used by an imager 31, an administrator terminal 42 used by an administrator 32, a worker terminal 43 used by a recovery worker 33, and a caption generation device 100. The imaging device terminal 41, the administrator terminal 42, the worker terminal 43, and the caption generation device 100 can communicate with each other via a communication network 50. The communication network 50 may be a wired network, a wireless network, or may include both of these.

[0026] The imaging device 41 captures the image 10 of the object 20. Instead of or in addition to this, imaging devices such as surveillance cameras installed in the plant, cameras equipped on mobile robots or drones that patrol the plant, etc. may capture the image 10 of the object 20. The imaging device 41 may be a smartphone with a camera. The imaging device 41 may be a tablet terminal. The imaging device 41 may be a PC (Personal Computer). The imaging device 41 may be a wearable terminal.

[0027] The imager 31 is an example of a user who uses the caption generation system 5 and carries and uses the imaging device 41. The imager 31 uses the imaging device 41 to capture the image 10 of the object 20 in the facilities within the plant. The imager 31 is, as an example, a maintenance worker for the equipment in the plant, and more specifically, a maintenance worker who maintains the tank 22. Note that there may be one or multiple imagers 31 in the plant.

[0028] The administrator terminal 42 displays reports using the captions generated by the caption generation system 5. The administrator terminal 42 transmits instructions regarding the object 20 in the equipment within the plant to the worker terminal 43 via the communication network 50. The administrator terminal 42 may be a PC, and may be arranged at a location away from the object 20 in the facilities within the plant, such as in the control room within the plant. Note that the administrator terminal 42 may be a smartphone. The administrator terminal 42 may be a tablet terminal. The administrator terminal 42 may be a wearable terminal.

[0029] The administrator 32 is an example of a user who uses the caption generation system 5 and uses the administrator terminal 42. The administrator 32 browses reports using the captions generated by the caption generation system 5 via the administrator terminal 42. As an example, the administrator 32 browses the report in a situation where the object 20 described by the caption cannot be directly visually confirmed. The administrator 32 judges the situation of the object 20 in the facilities within the plant and orders the recovery worker 33 or the like to perform recovery work on the object 20. As an example, the administrator 32 judges the situation of the pipe 21 in the facilities within the plant and orders the recovery worker 33 or the like for the pipe 21 to perform recovery work on the pipe 21. Note that there may be one or a plurality of administrators 32 in the plant. Note that the administrator 32 is an example of a viewer who browses reports using the captions generated by the caption generation system 5.

[0030] The worker terminal 43 displays instructions regarding the object 20 in the facilities within the plant, reports using the captions generated by the caption generation system 5, and the like. The worker terminal 43 may be a smartphone. The worker terminal 43 may be a tablet terminal. The worker terminal 43 may be a PC. The worker terminal 43 may be a wearable terminal.

[0031] The recovery worker 33 is an example of a user who uses the caption generation system 5 and uses the worker terminal 43. The recovery worker 33 browses, via the worker terminal 43, instructions regarding the target 20 in the facilities within the plant, reports using the captions generated by the caption generation system 5, etc. The recovery worker 33, as an example, browses the report in a situation where the target 20 described by the caption cannot be directly visually confirmed. The recovery worker 33 addresses abnormalities within the plant. The recovery worker 33, as an example, addresses an abnormality in the pipe 21. The recovery worker 33, as an example, restores the pipe 21 in which an abnormality has occurred in accordance with instructions and reports from the administrator 32 displayed on the worker terminal 43. Note that there may be one or a plurality of recovery workers 33 within the plant. Note that the recovery worker 33 is an example of a viewer who browses a report using the captions generated by the caption generation system 5.

[0032] The caption generation device 100, as an example, receives, via the communication network 50, the image 10 of the target 20 in the facilities within the plant from the imaging device terminal 41. The caption generation device 100 generates a caption that describes the target 20 in the plant according to the state information indicating the state of the imaged target 20 using the image 10 received from the imaging device terminal 41. The caption generation device 100 generates a caption with higher accuracy, that is, outputs a description text that more accurately describes the state of the target 20, compared to the case of generating a caption that describes the target 20 without using the state information from the image 10 of the target 20.

[0033] The caption generation device 100 transmits the generated caption to the imaging device terminal 41 via the communication network 50. The caption generation device 100 may be arranged in a control room, an instrument room, etc. within the plant, or may be arranged outside the plant. Some functions of the caption generation device 100 may be incorporated in the imaging device terminal 41.

[0034] FIG. 2 is a block diagram of the caption generation device 100 according to the first embodiment. In FIG. 2, the flow of data and the like is indicated by arrows. The caption generation device 100 includes an image acquisition unit 110, a word extraction unit 120, a state information acquisition unit 130, an image extraction unit 135, a caption generation unit 140, a storage unit 150, and a report storage unit 155.

[0035] The image acquisition unit 110 acquires the image 10 captured within the plant. As an example, the image acquisition unit 110 receives, from the imaging terminal 41 via the communication network 50, the user ID of the imager 31 who uses the imaging terminal 41, together with the image 10 of the target 20 within the plant. As an example, the image acquisition unit 110 receives, from the imaging terminal 41 via the communication network 50, the user ID of the imager 31 and the imaging information indicating the imaging position and imaging direction when the imaging terminal 41 captures the image 10 of the target 20, together with the image 10 of the target 20 captured within the plant. The image acquisition unit 110 inputs the received image 10 to the word extraction unit 120 and inputs the image 10 together with the user ID of the imager 31 and the like to the state information acquisition unit 130.

[0036] The word extraction unit 120 extracts a plurality of words representing the features of the target 20 imaged in the image 10. The word extraction unit 120 may extract one or more words based on the newly input image 10 using a word extraction model that has learned the relationship between the image 10 and one or more words representing the features of the target 20 imaged in the image 10. The word extraction unit 120 may, for example, read out the word extraction model stored in the storage unit 150. The word extraction unit 120 inputs the plurality of words extracted from the image 10 to the caption generation unit 140, for example, as a vector representation.

[0037] The state information acquisition unit 130 acquires state information indicating the state of the target 20 imaged in the image 10. As an example, the state information acquisition unit 130 acquires state information by analyzing the image 10. Specifically, the state information acquisition unit 130 extracts the target 20 from the image 10 by analyzing the image 10, and determines which facility in the plant the extracted target 20 corresponds to by comparing the features of the extracted target 20 with the features of a plurality of targets 20 stored in the storage unit 150. The state information acquisition unit 130 may further access a distributed control system that controls the plant, which is external to the caption generation device 100, and receive state information regarding the facility corresponding to the target 20 from the distributed control system. The facility corresponding to the target 20 may include the facility itself of the target 20, or may include facilities located upstream or downstream of the target 20. Note that the state information acquisition unit 130 may extract the target 20 from the image 10 by analyzing the image 10, determine which facility in the plant the target 20 corresponds to, and read out the state information regarding the facility corresponding to the target 20, which is stored in the storage unit 150.

[0038] As another example, the state information acquisition unit 130 may acquire state information by acquiring imaging information from the imaging terminal 41. Specifically, the state information acquisition unit 130 may determine which facility in the plant the target 20 imaged in the image 10 corresponds to by comparing the imaging information from the imaging terminal 41 with a plurality of pieces of imaging information stored in the storage unit 150. The state information acquisition unit 130 may further access the distributed control system and receive state information regarding the facility corresponding to the target 20 from the distributed control system. Note that the imaging information may additionally indicate, for example, the degree of zooming when the imaging terminal 41 captured the image 10 and the distance to the target 20. The state information acquisition unit 130 inputs the acquired state information to the image extraction unit 135, and inputs the state information, for example, in vector representation, together with the user ID of the imager 31, etc., to the caption generation unit 140.

[0039] The state information may include at least any one of the position information of the target 20, the deterioration information indicating the deterioration state of the target 20, the fluid information indicating the component or state of the fluid flowing inside the pipe that is the target 20, and the equipment information related to the equipment located upstream or downstream of the pipe that is the target 20. The state information may include, for example, fixed information such as the position information of the pipe 21 or the information of the tank 22 upstream of the pipe 21, and may also include variable information indicating the type of gas or liquid flowing inside the pipe 21, the state such as the pressure and temperature of the gas or liquid, and the degree of deterioration of the pipe 21. The degree of deterioration of the pipe 21 may be an index of the deterioration state according to the degree of rusting of the pipe 21 or the like. As an example, the fixed state information may be stored in the storage unit 150. For example, as the fixed state information, the design data of the plant in which a plurality of targets 20 and the equipment located upstream or downstream of each target 20 are shown may be stored in the storage unit 150.

[0040] When the state information includes the position information of the target 20, the image extraction unit 135 extracts, from the storage unit 150, a past image having position information common to the position information included in the new state information in response to obtaining the new state information of the new image 10. The image extraction unit 135 may extract one or a plurality of past images corresponding to the image 10 from the storage unit 150. The image extraction unit 135 inputs the past image extracted from the storage unit 150 to the caption generation unit 140 together with the image 10. Note that when the state information does not include the position information of the target 20, the image extraction unit 135 may not input anything to the caption generation unit 140.

[0041] The caption generation unit 140 generates a caption that describes the captured object 20 according to the state information, based on the state information and a plurality of words. The caption generation unit 140 may generate a caption based on newly input state information and a plurality of words, using a caption generation model that has learned the relationship between the state information, one or more words representing the characteristics of the object 20 in the plant, and the caption that describes the object 20 in the plant. The caption generation unit 140 may, for example, read out the caption generation model stored in the storage unit 150. As an example, the caption generation unit 140 may generate a caption using both a new image 10 input from the image extraction unit 135 and a past image corresponding to the new image 10. The caption generation unit 140 transmits the generated caption as text data to the imaging device terminal 41 via the communication network 50, based on the user ID of the imager 31.

[0042] The storage unit 150 stores the past images of the object 20 in the plant in association with the location information. The storage unit 150 may also store the above-mentioned fixed state information, word extraction model, and caption generation model. Note that the storage unit 150 is an example of a model storage unit.

[0043] The report storage unit 155 receives a report from the imaging device terminal 41 via the communication network 50 and registers it in the memory. The report may be created by the imager 31 appropriately editing the caption using the imaging device terminal 41. The report storage unit 155 is accessed via the communication network 50 from, for example, the administrator terminal 42 or the worker terminal 43, and the report stored in the memory is read out.

[0044] FIG. 3 is a flowchart showing the flow of the caption generation method according to the first embodiment. As an example, the flow starts when the imager 31 transmits an image 10 of the object 20 in the plant captured by the imaging device terminal 41 to the caption generation device 100 via the communication network 50.

[0045] The caption generation device 100 acquires the image 10 captured in the plant (step S101). Specifically, the caption generation device 100 receives the user ID of the imager 31 from the imager terminal 41 via the communication network 50, together with the image 10 of the target 20 in the plant. Alternatively, the caption generation device 100 may receive the user ID of the imager 31 and the position information of the imager terminal 41 from the imager terminal 41 via the communication network 50, together with the image 10.

[0046] The caption generation device 100 extracts a plurality of words representing the features of the target 20 imaged in the image 10 (step S102). Specifically, the caption generation device 100 uses the above-mentioned word extraction model to extract a plurality of words representing the features of the target 20 imaged in the image 10.

[0047] The word extraction model learns the relationship between the image 10 and a plurality of image feature amounts, and also learns the relationship between the plurality of image feature amounts and a plurality of words. When a new image 10 is input, the word extraction model extracts a plurality of image feature amounts from the image 10, and outputs, as a vector representation, a plurality of words representing the features of the target 20 imaged in the image 10 from the plurality of image feature amounts. The word extraction model may be a deep learning model having a neural network structure including, for example, a CNN (Convolution Neural Network) and an FC layer (Fully Connected Layer). Instead of the CNN, the word extraction model may use the Self-Attention or Multi-Head Self-Attention of Transformer.

[0048] The caption generation device 100 acquires state information indicating the state of the imaged object 20 (step S103). Specifically, the caption generation device 100 acquires the state information by analyzing the image 10. More specifically, the caption generation device 100 extracts the object 20 from the image 10 by analyzing the image 10, and determines which facility in the plant the extracted object 20 corresponds to by comparing the features of the extracted object 20 with the features of a plurality of objects 20 stored in the storage unit 150. For example, as a result of analyzing the image 10, the caption generation device 100 identifies features such as the shape and model number of the object 20 imaged in the image 10, compares them with the features of a plurality of objects 20 stored in the storage unit 150, and extracts information on the facility in the plant that is associated with the object 20 whose features match. For example, as a result of analyzing the image 10, the caption generation device 100 acquires the identification information of the object 20 located on the surface or around the object 20 imaged in the image 10, compares it with the identification information of a plurality of objects 20 stored in the storage unit 150, and extracts information on the facility in the plant that is associated with the object 20 whose identification information matches. The caption generation device 100 further accesses the distributed control system and receives state information regarding the facility corresponding to the object 20 from the distributed control system.

[0049] Additionally or alternatively, the caption generation device 100 acquires the state information by receiving the above-described imaging information together with the image 10 from the imaging terminal 41. More specifically, the caption generation device 100 determines which facility in the plant the object 20 imaged in the image 10 corresponds to by comparing the imaging information received from the imaging terminal 41 with a plurality of pieces of imaging information stored in the storage unit 150. The caption generation device 100 further accesses the distributed control system and receives state information regarding the facility corresponding to the object 20 from the distributed control system.

[0050] The caption generation device 100 generates a caption that describes the captured object 20 according to the state information, based on the state information and a plurality of words (step S104). Specifically, the caption generation device 100 uses a caption generation model to generate, as text data, a caption that describes the captured object 20 according to the state information, based on the state information indicating the state of the object 20 and a plurality of words. The caption generation device 100 transmits the generated text data of the caption to the imaging device terminal 41 via the communication network 50, and thus this flow ends.

[0051] The caption generation model has learned the relationship between the state information indicating the state of the object 20, the vector representation of the words representing the features of the object 20, and the caption that describes the object 20. When a new vector representation of the state information of the object 20 captured in the image 10 and vector representations of a plurality of words representing the features of the object 20 are input, the caption generation model outputs text data of a caption that describes the object 20 captured in the image 10 from the vector representation of the state information and the vector representations of the plurality of words. The caption generation model may be a deep learning model having a neural network structure including, for example, an RNN (Recurrent Neural Network). Instead of an RNN, the caption generation model may use an LSTM (Long Short-Term Memory), a Transformer, GPT-3 (Generative Pre-trained Transformer-3), GPT-2, a GPT-3-Clone, or the like. The caption generation model may use a large number of example sentences as learning data and learn an RNN or the like so as to increase the likelihood of outputting each of the example sentences. The caption generation model may extract the words included in the example sentences, use the vectors of the extracted words as learning input data, use the example sentences as learning output data, and learn so as to increase the probability of outputting the learning output data for the learning input data.

[0052] Regarding steps S102 to S104, more specifically, the caption generation device 100 uses a word extraction model and a caption generation model to generate a caption that describes the imaged pipe 21 according to the state information, based on a plurality of words extracted from the image 10 and the state information of the pipe 21 imaged in the image 10.

[0053] For example, as shown in FIG. 1, when the image 10 shows an abnormality where a crack has occurred in the pipe 21 and white gas is jetting out therefrom, the caption generation device 100 that has acquired the image 10 may generate a caption such as "A crack has occurred in the pipe with pipe ID ○○ in Area A, and sulfuric acid gas is jetting out therefrom" according to the position information, deterioration information, and fluid information. For example, the caption generation device 100 may generate a caption such as "There is a possibility that a crack has occurred in the pipe due to the deterioration of the rust on the pipe" according to the deterioration information. For example, the caption generation device 100 may generate a caption such as "Since a crack has occurred in the pipe and gas is jetting out, please close the valve of the tank on the upstream side of the pipe" according to the deterioration information and equipment information. For example, the caption generation device 100 may generate a caption such as "White smoke is coming out of the pipe that has been in use for one year" according to the maintenance information of the pipe 21, which is an example of the deterioration information. For example, the caption generation device 100 may generate a caption such as "The pressure in the pipe is ○○ kPa, and a liquid of 〇〇〇 is leaking and dripping from the joint" according to the fluid information including the measured value of the pressure of the fluid flowing inside the pipe 21. The caption generation device 100 transmits the caption to the imager terminal 41 via the communication network 50.

[0054] The imaging person 31 checks the caption displayed on the imaging person terminal 41, and transmits a report using the caption from the imaging person terminal 41 to the caption generation device 100 via the communication network 50. The imaging person 31 may appropriately edit the caption to create a report. When the caption generation device 100 receives a report from the imaging person terminal 41, it stores the report in the report storage unit 155. The report stored in the report storage unit 155 of the caption generation device 100 is viewed by, for example, the pipe repair worker 33 of the pipe 21, and the pipe repair worker 33 may plug the crack in the pipe 21 or adjust the valve of the tank 22 on the upstream side of the pipe 21. The report stored in the report storage unit 155 of the caption generation device 100 is viewed by, for example, the administrator 32 of the pipe 21, and the administrator 32 may order the pipe repair worker 33 of the pipe 21 to perform a repair operation or adjust the control parameters of the equipment in which the pipe 21 is installed.

[0055] Note that in steps S103 to S104, when the state information includes the position information of the target 20, the caption generation device 100 extracts one or more past images having the position information common to the position information from the storage unit 150, and may generate a caption using both the image 10 and the one or more past images. For example, the caption generation device 100 may compare one or more past images with the image 10 according to the date and time when each image was captured, and identify changes in the target 20 captured in the image 10, such as the progress of rust occurring on the surface of the target 20. In this case, the caption generation device 100 may generate a caption corresponding to the change and the state information based on the state information, the change of the target 20 specified as a result of the comparison, and a plurality of words.

[0056] According to the caption generation device 100 of the first embodiment described above, caption generation for explaining the target 20 is generated from the image 10 of the target 20 captured in the plant according to the state information indicating the state of the target 20. Thereby, compared with the case where the caption generation device 100 generates a caption for explaining the target 20 without using the state information from the image 10 of the captured target 20, the caption generation device 100 can generate a caption with high accuracy, that is, output a description text that more accurately explains the state of the target 20 and the like.

[0057] In the caption generation device 100 according to the first embodiment, the word extraction unit 120 may be input with the state information acquired by the state information acquisition unit 130, and a likelihood corresponding to the state information may be assigned to each of the plurality of words extracted from the image 10. For example, the word extraction unit 120 may be input with the degree of deterioration, which is an index of the deterioration state according to the degree of rusting of the target 20, as the state information of the target 20 captured in the image 10 together with the image 10 of the captured target 20. In this case, the word extraction unit 120 may assign a high likelihood to words having a high relevance to the deterioration state of the target 20 and a low likelihood to words having a low relevance to the deterioration state of the target 20.

[0058] More specifically, the word extraction unit 120 may assign a likelihood corresponding to the state information to each of the plurality of words extracted from the image 10 using a word extraction model. The word extraction model inputs, for example, the pixel values of each pixel of the image to each input node in the input stage, and outputs a scalar value for each word from each output node in the output stage. The scalar value may be an output value with a reference range of 0 to 1. In this case, the word vector is a vector having the scalar value of 0 to 1 for each word. When the output value of a word is 0, it means that the image 10 does not represent that word, and when the output value of a word is 1, it means that the image 10 represents that word. The word extraction model is, for example, trained to output a value closer to 1 as the likelihood that the image 10 represents that word is higher. The word extraction model is further trained by, for example, using a certain image 10 and the state information obtained from the image 10 by the state information acquisition unit 130 as input teacher data, and using the plurality of words extracted by decomposing the caption actually generated from the image 10 and the state information by the caption generation unit 140 as output teacher data, and repeating the training. In this case, the word extraction model outputs a value closer to 1 as the likelihood that the image 10 represents that word is higher, and outputs a value closer to 1 as the relevance between the state information and the word is higher. The word extraction unit 120 treats, for example, the scalar value of each word output from the word extraction model as a likelihood for convenience and assigns it to each word.

[0059] As a specific example, assume that in the image 10 acquired by the word extraction unit 120, a tank 22, a pipe 21 connected to the tank 22, a crack occurring in the pipe 21, gas generated from the crack, a valve attached to the pipe 21, other devices not directly or indirectly connected to the tank 22 and the pipe 21, and the surrounding walls of the tank 22 and the pipe 21 are shown. For example, when the reference range of the output value for each word is 0 to 1, based on the degree of deterioration, the output value for the word "tank" may be 0.9, the output value for the word "pipe" may be 0.9, the output value for the word "crack" may be 0.8, the output value for the word "gas" may be 0.5, the output value for the word "valve" may be 0.4, the output value for the word indicating other devices and the word "wall" may be 0.2. In this case, the word extraction unit 120 may handle the probability of the word "crack" as 80% using the output value from the word extraction model.

[0060] In this case, the caption generation unit 140 may generate a caption by preferentially using the word with a higher likelihood among the plurality of words based on the state information and the plurality of words. Thereby, the caption generation device 100 can generate a caption with even higher accuracy.

[0061] The caption generation unit 140 may generate a plurality of captions. In this case, the caption generation unit 140 may further assign a priority corresponding to the likelihood of at least one word used for caption generation among the plurality of words to the generated captions. The priority may be an average value, a total value, or an integrated value of the likelihoods of the plurality of words. The caption generation unit 140 may transmit the caption with the highest priority among the generated plurality of captions to the imaging device terminal 41. Alternatively, the caption generation unit 140 may transmit the generated plurality of captions to the imaging device terminal 41 with priorities assigned to each of the generated plurality of captions. In this case, the imaging device 31 can determine the suitability of the plurality of captions while referring to the priorities.

[0062] As an example when the caption generation unit 140 generates a plurality of captions, after attaching a negative weight to the words used when generating the first caption among the plurality of words extracted by the word extraction unit 120, by generating the second caption, the words used in the first caption may not be included in the second caption. As another example when the caption generation unit 140 generates a plurality of captions, after assigning an arbitrary priority to the plurality of words extracted by the word extraction unit 120, each time a word is used in caption generation, the priority of the word is decreased by a predetermined amount, so that the plurality of words extracted by the word extraction unit 120 may be used as a whole.

[0063] The caption generation unit 140 may also output a caption including an instruction to re-capture the imaged object 20 when the likelihood of the combination of the plurality of words used for caption generation is lower than a predetermined threshold. For example, the caption generation unit 140 may output a caption including an instruction to re-capture the imaged object 20 when the average value, total value, integrated value, etc. of the likelihoods of the plurality of words used for caption generation are less than a predetermined threshold stored in the storage unit 150.

[0064] In the caption generation device 100 according to the first embodiment, the storage unit 150 may store a plurality of sentences included in at least any one of the accident case collection, accident response manual, and maintenance history in the plant as example sentences. In this case, the caption generation unit 140 may extract at least one example sentence by searching using at least any one of the plurality of words extracted by the word extraction unit 120 from among the plurality of example sentences stored in the storage unit 150.

[0065] The caption generation unit 140 may generate a caption based on the state information acquired by the state information acquisition unit 130, the plurality of words, and the extracted example sentence. Specifically, the caption generation unit 140 may search for an example sentence in which at least any one of the plurality of words is used from among the plurality of example sentences stored in the storage unit 150, and generate a caption while referring to the extracted example sentence.

[0066] In this case, the caption generation unit 140 may generate a plurality of captions and assign a similarity to the extracted example sentence. Specifically, the caption generation unit 140 generates a plurality of captions based on the state information and the plurality of words, and searches for an example sentence in which at least any one of the plurality of words is used from among the plurality of example sentences stored in the storage unit 150, and may assign a similarity to each of the plurality of captions with the extracted example sentence. The caption generation unit 140 may, for example, convert each of the plurality of captions and the extracted example sentence into a feature vector indicating a combination of a plurality of words included therein, and calculate the inner product of the feature vectors as the similarity. The caption generation unit 140 may transmit the caption with the highest similarity among the generated plurality of captions to the imaging device terminal 41. Alternatively, the caption generation unit 140 may transmit to the imaging device terminal 41 in a state where the similarity is assigned to each of the generated plurality of captions. In this case, the imager 31 can determine the suitability of the plurality of captions while referring to the similarity.

[0067] The caption generated by the caption generation unit 140 while referring to the extracted example sentence may include, in addition to the caption explaining the target 20 in the plant, a plurality of action options regarding the actions to be taken by the user, or instructions to the user. As a result of generating a caption with reference to the example sentence, the caption generation device 100 can present to the user an action option or an instruction incorporating an abnormality handling method that cannot be directly decoded from the image 10.

[0068] The caption generated by the caption generation device 100 according to the first embodiment may include at least one of an instruction to image the captured target 20 from another angle or another field of view, and an instruction to image another target 20 related to the captured target 20, in addition to the caption that describes the target 20 in the plant.

[0069] FIG. 4 is a flowchart showing an example of a detailed flow of the caption generation method according to the first embodiment. The operation flow shown in the flowchart of FIG. 4 may be a specific example of the caption generation method according to the first embodiment described with reference to FIG. 3. The caption generation device 100 according to the first embodiment described above may generate, as an example, a caption that describes the target 20 from the image 10 of the captured target 20 according to the operation flow shown in the flowchart of FIG. 4.

[0070] The flow of FIG. 4 is started, as an example, when the imager 31 transmits the image 10 of the target 20 in the plant captured by the imager terminal 41 to the caption generation device 100 via the communication network 50.

[0071] The caption generation device 100 acquires the image 10 captured in the plant (step S201). In step S201, as shown in FIG. 5 which specifically describes a part of the flow of FIG. 4, the image acquisition unit 110 receives the image 10 of the target 20 in the plant from the imager terminal 41 via the communication network 50, and extracts Exif information from the image 10. The Exif information may include the user ID of the imager 31. The image acquisition unit 110 may output the image 10 to the word extraction unit 120 and output the Exif information to the state information acquisition unit 130 together with the image 10.

[0072] The caption generation device 100 crops the image 10 into a plurality of parts (step S202). For example, the caption generation device 100 creates a new image by discarding the peripheral area in the image 10 that does not contain the object 20, and divides it into patches of a fixed size. In step S202, as shown in FIG. 5, the word extraction unit 120 may create a new image by discarding the peripheral area in the image 10 that does not contain the pipe 21 and the tank 22, and divide the new image into nine patches by dividing it vertically and horizontally into three parts. Alternatively, the word extraction unit 120 may create a new image by discarding the peripheral area in the image 10 that does not contain the pipe 21 as a single patch of a fixed size, and create a new image by discarding the peripheral area in the image 10 that does not contain the tank 22 as another patch of a fixed size. Note that the learning model of the word extraction unit 120 shown in FIG. 5 may also be referred to as a Vision Transformer and may be an example of the word extraction model described above.

[0073] The caption generation device 100 converts the cropped image into a numerical value A that can be used by the encoder through a simple function, for example, a LINEAR function (step S203). In step S203, as shown in FIG. 5, the word extraction unit 120 may put nine patches into a linear projection layer, flatten each of them, and embed them linearly (or convert them into a one-dimensional array) into vectors. The word extraction unit 120 may further add a CLS token to the beginning of the sequence of vectors output from the linear projection layer (extra learnable [class] embedding), and embed positions into each patch (each vector) to obtain the numerical value A. That is, the word extraction unit 120 may convert the nine patches divided from the image 10 into a sequence of 10-dimensional vectors that can be used by the encoder.

[0074] The caption generation device 100 causes the transformer encoder to convert a numerical value A into another value B (step S204), and unifies the dimension and numerical width of the value B (step S205). In steps S204 to S205, as shown in FIG. 5, the word extraction unit 120 inputs a sequence of 10-dimensional vectors to the transformer encoder, and the transformer encoder converts this into another sequence of 10-dimensional vectors, and unifies the dimension and numerical width through, for example, a softmax function, and may output it as the output of the CLS token. Note that the transformer encoder may include a plurality of components such as Self-Attention described above. The output of the CLS token may be an aggregation of feature amounts necessary for classification from the entire image by Self-Attention.

[0075] The caption generation device 100 outputs image feature amounts (step S206). In step S206, as shown in FIG. 5, the word extraction unit 120 may input the output of the CLS token to a classification head (MLP: multi-layer perceptron) and output a plurality of image feature amounts from the classification head. The classification head may output after compressing the plurality of image feature amounts. The word extraction unit 120 may output a plurality of image feature amounts to the caption generation unit 140 as shown in FIG. 5, and as described with reference to FIGS. 1 to 3, extract word vectors of a plurality of words representing the features of the object 20 imaged in the image 10 from the plurality of image feature amounts, and output them to the caption generation unit 140.

[0076] The caption generation device 100 acquires state information (step S207), converts the image feature amount and the state information into a numerical value C suitable for the decoder (step S208), and inputs the numerical value C into the decoder (step S209). In steps S207 to S209, as shown in FIG. 6, which specifically describes a part of the flow of FIG. 4, the caption generation unit 140 vectorizes the state information input from the state information acquisition unit 130 and the empty text through, for example, a LINEAR function, and vectorizes these vectors together with a plurality of image feature amounts input from the word extraction unit 120. Thereby, the caption generation unit 140 converts each of the plurality of image feature amounts into the numerical value C. The caption generation unit 140 further inputs the numerical value C into the decoder. Note that the learning model of the decoder is a model that uses layer-by-layer adjustment, may also be referred to as GPT-2, and may be an example of the caption generation model described above.

[0077] The caption generation device 100 causes the decoder to convert the numerical value C into a numerical value D (step S210), and converts it into text (tokens) through a function that returns the numerical value D to text (step S211). In steps S210 to S211, as shown in FIG. 6, the decoder in the caption generation unit 140 may convert the numerical value C into the numerical value D. The decoder may weight each numerical value D based on the vectorized state information. The decoder may compress the numerical value D. The decoder may convert the numerical value D into text (tokens) through a softmax function.

[0078] The caption generation device 100 outputs text (caption) (step S212), and the flow ends. In step S212, as shown in FIG. 6, the decoder takes, for example, a sequence of consecutive texts X1 to X5 as an input (source token), and uses a sequence of consecutive texts Y1 to Y5 that is the same as the sequence but shifted one token to the right as a target token. The source token and the target token may be concatenated and processed with adjustment for each layer. Resettable positions of the source token and the target token may be embedded together with their respective corresponding tokens. The position of the source token starts from zero, and for the target token, instead of increasing the position at the end of the source sentence, the position may be reset from zero again.

[0079] The decoder model shown in FIG. 6 includes, as an example, an attention mechanism including a self-attention mechanism and a mixed attention mechanism, an add and layer normalization (Add&Layer Normalization) layer, a feed-forward neural network (FNN) layer, and N layers stacked in this order with an add and layer normalization layer. The decoder may put the processed result of N layers into a linearization layer, pass the output from the linearization layer through a softmax function, and output a caption consisting of a sequence of consecutive texts Y1 to Y5. The caption generation unit 140 may transmit the generated caption text data to the imaging device terminal 41 via the communication network 50.

[0080] FIG. 7 is a block diagram of the caption generation device 200 according to the second embodiment. The caption generation device 200 according to the second embodiment includes a learning unit 260 in addition to the configuration included in the caption generation device 100 according to the first embodiment. Other configurations included in the caption generation device 200 according to the second embodiment are the same as those of the caption generation device 100 according to the first embodiment, and redundant descriptions are omitted using the reference numerals of the respective configurations included in the caption generation device 100.

[0081] The learning unit 260 learns the above-described word extraction model and caption generation model using the result of the user's determination of the appropriateness of the caption generated by the caption generation unit 140. In addition to or instead of this, the learning unit 260 may learn the above-described word extraction model and caption generation model using the user's correction input for the caption generated by the caption generation unit 140. The learning unit 260 receives the above-described determination result or correction input by the user from the imaging device terminal 41 via the communication network 50.

[0082] The learning unit 260 updates the parameters of the model so as to reduce the error between the output of the model when each sample in the training data including the result of the above-described determination is input to the model and the label. For example, when learning a word extraction model using a multi-layer neural network including a CNN or the like, the learning unit 260 uses the error between the output value output by the neural network in response to inputting each sample and the label, and by a method such as backpropagation, adjusts the weights between the neurons of the neural network and the biases of each neuron.

[0083] As described above, in the caption generation device 200 according to the second embodiment, information on whether the generated caption has been actually used by the user is fed back, and the word extraction model and the caption generation model are learned using the feedback content. For example, when the caption generation device 200 receives, from the administrator terminal 42, the result determined by the administrator 32 that there is no problem as a result of viewing the caption, the caption generation device 200 may determine that the generated caption has been actually used. The caption generation device 200 may also determine that the generated caption has been actually used, for example, when the caption generation unit 140 performs a search for example sentences using the generated caption, or when the generated caption is used in a report for the administrator 32. The caption generation device 200 according to such a second embodiment also has the same effects as the caption generation device 100 according to the first embodiment. According to the caption generation device 200 according to the second embodiment, captions with even higher accuracy can be generated.

[0084] Various embodiments of the present invention may be described with reference to flowcharts and block diagrams, where the blocks may represent (1) stages of a process in which operations are performed or (2) sections of a device having the role of performing the operations. The specific stages and sections may be implemented by a dedicated circuit, a programmable circuit supplied with computer-readable instructions stored on a computer-readable medium, and / or a processor supplied with computer-readable instructions stored on a computer-readable medium. The dedicated circuit may include digital and / or analog hardware circuits, and may include integrated circuits (ICs) and / or discrete circuits. The programmable circuit may include a reconfigurable hardware circuit including memory elements such as logical AND, logical OR, logical XOR, logical NAND, logical NOR, and other logical operations, flip-flops, registers, field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), etc.

[0085] A computer-readable medium may include any tangible device capable of storing instructions executable by a suitable device. As a result, a computer-readable medium having instructions stored therein will comprise a product that includes instructions that can be executed to create means for performing the operations specified in a flowchart or block diagram. Examples of computer-readable media may include electronic memory media, magnetic memory media, optical memory media, electromagnetic memory media, semiconductor memory media, and the like. More specific examples of computer-readable media may include floppy (registered trademark) disks, diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), Blu-ray (registered trademark) disc, memory stick, integrated circuit card, and the like.

[0086] Computer-readable instructions may include any combination of one or more programming languages, including source code or object code written in any combination of assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk (registered trademark), JAVA (registered trademark), C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0087] Computer-readable instructions may be provided to a processor or programmable circuitry of a programmable data processing apparatus such as a general-purpose computer, a special-purpose computer, or other computers, either locally or via a wide area network (WAN) such as a local area network (LAN), the Internet, etc., and may execute the computer-readable instructions to create means for performing the operations specified in the flowchart or block diagram. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.

[0088] FIG. 8 shows an example of a computer 2200 in which multiple aspects of the present invention may be embodied, in whole or in part. Programs installed on the computer 2200 can cause the computer 2200 to function as an operation associated with the apparatus according to an embodiment of the present invention or as one or more sections of the apparatus, or can cause the operation or the one or more sections to be executed, and / or can cause the computer 2200 to execute a process according to an embodiment of the present invention or a stage of the process. Such a program may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.

[0089] The computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphic controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes input / output units such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.

[0090] The CPU 2212 operates according to programs stored in the ROM 2230 and the RAM 2214, thereby controlling each unit. The graphic controller 2216 acquires image data generated by the CPU 2212 in a frame buffer or the like provided in the RAM 2214 or in itself, and causes the image data to be displayed on the display device 2218.

[0091] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads a program or data from the DVD-ROM 2201 and provides the program or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to the IC card.

[0092] The ROM 2230 stores therein a boot program or the like executed by the computer 2200 at activation and / or a program dependent on the hardware of the computer 2200. The input / output chip 2240 may also be connected to the input / output controller 2220 via various input / output units such as a parallel port, a serial port, a keyboard port, a mouse port, etc.

[0093] The program is provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The program is read from the computer-readable medium, installed in the hard disk drive 2224, the RAM 2214, or the ROM 2230, which are also examples of computer-readable media, and executed by the CPU 2212. The information processing described in these programs is read by the computer 2200, bringing about cooperation between the programs and the various types of hardware resources described above. The apparatus or method may be configured by realizing the operation or processing of information according to the use of the computer 2200.

[0094] For example, when communication is executed between the computer 2200 and an external device, the CPU 2212 may execute a communication program loaded in the RAM 2214 and instruct the communication interface 2222 to perform communication processing based on the processing described in the communication program. The communication interface 2222 reads the transmission data stored in the transmission buffer processing area provided in a recording medium such as the RAM 2214, the hard disk drive 2224, the DVD-ROM 2201, or the IC card under the control of the CPU 2212, transmits the read transmission data to the network, or writes the received data received from the network to the reception buffer processing area or the like provided on the recording medium.

[0095] Further, the CPU 2212 may cause all or a necessary part of a file or database stored in an external recording medium such as a hard disk drive 2224, a DVD-ROM drive 2226 (DVD-ROM 2201), an IC card, etc. to be read into the RAM 2214, and execute various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording medium.

[0096] Various types of information such as various types of programs, data, tables, and databases may be stored in the recording medium and may be subjected to information processing. The CPU 2212 may execute various types of processing on the data read from the RAM 2214, including various types of operations, information processing, conditional judgment, conditional branch, unconditional branch, information search / replacement, etc. described throughout this disclosure and specified by the instruction sequence of the program, and write back the result to the RAM 2214. Also, the CPU 2212 may search for information in files, databases, etc. in the recording medium. For example, when a plurality of entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored in the recording medium, the CPU 2212 searches for an entry that matches the condition where the attribute value of the first attribute is specified from among the plurality of entries, reads the attribute value of the second attribute stored in the entry, and thereby obtains the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.

[0097] The programs or software modules described above may be stored in a computer-readable medium on or near the computer 2200. Also, a recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can be used as a computer-readable medium, thereby providing the program to the computer 2200 via the network.

[0098] As described above, the present invention has been explained using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments. It is obvious to those skilled in the art that various changes or improvements can be made to the above embodiments. It is clear from the description of the claims that forms with such changes or improvements can also be included in the technical scope of the present invention.

[0099] For example, the control system may be a computer housed in a single housing. That is, the controller may be realized by executing a program on a processor of the computer, and each input / output device may be implemented as an I / O device of the computer. Further, the controller may be implemented as a virtual machine executed by one or more processors. In such a configuration, the control system does not include a network that is a general-purpose or dedicated network, and the controller and the input / output devices can be connected by a chipset such as a memory controller hub and an I / O controller hub that connect between the processor and the I / O devices.

[0100] It should be noted that the execution order of each process such as operations, procedures, steps, and stages in the devices, systems, programs, and methods shown in the claims, the specification, and the drawings is not explicitly indicated as "earlier" or "preceding" etc., and can be realized in any order unless the output of the previous process is used in the subsequent process. Regarding the operation flows in the claims, the specification, and the drawings, even if "first," "next," etc. are used for convenience of explanation, it does not mean that it is essential to implement in this order.

Explanation of Reference Numerals

[0101] 5 Caption generation system 10 Image 20 Target 21 Pipe 22 Tank 31 Imager 32 Administrator 33 Recovery worker 41 Imager terminal 42 Manager terminal 43 Operator terminal 100 Caption generation device 110 Image acquisition unit 120 Word extraction unit 130 Status information acquisition unit 135 Image extraction unit 140 Caption generation unit 150 Memory unit 155 Report storage unit 200 Caption generation device 260 Learning unit 2200 Computer 2201 DVD-ROM 2210 Host controller 2212 CPU 2214 RAM 2216 Graphics controller 2218 Display device 2220 Input / output controller 2222 Communication interface 2224 Hard disk drive 2226 DVD-ROM drive 2230 ROM 2240 Input / output chip 2242 Keyboard

Claims

1. An image acquisition unit that acquires an image captured within a plant; A word extraction unit that extracts a plurality of words representing features of an object imaged in the image; A state information acquisition unit that acquires state information indicating the state of the imaged object; A caption generation unit that generates a caption for explaining the imaged object according to the state information based on the state information and the plurality of words; A caption generation device comprising the above.

2. The state information includes at least any one of position information of the object, deterioration information indicating the deterioration state of the object, fluid information indicating the components or state of a fluid flowing inside a pipe that is the object, and equipment information related to equipment located upstream or downstream of the pipe that is the object; The caption generation device according to Claim 1.

3. The state information acquisition unit acquires the state information by analyzing the image; The caption generation device according to Claim 2.

4. The state information includes the position information of the object; A storage unit that stores by associating past images of the object in the plant with the position information; An image extraction unit that extracts a past image having position information common to the position information included in the new state information in response to acquiring new state information of the new image; further comprising; The caption generation unit generates the caption using both the new image and the past image corresponding to the new image; The caption generation device according to Claim 1.

5. The word extraction unit assigns a likelihood corresponding to the state information to each of the plurality of words extracted from the image; The caption generation unit generates the caption by preferentially using words with higher likelihoods among the plurality of words based on the state information and the plurality of words; The caption generation device according to Claim 1.

6. The caption generation unit assigns a priority corresponding to the likelihood of at least one word used for caption generation among the plurality of words to the generated caption; The caption generation device according to Claim 5.

7. When the likelihood of a combination of a plurality of words used for caption generation is lower than a predetermined threshold, the caption generation unit outputs the caption including an instruction to re-capture the imaged object; The caption generation device according to Claim 5.

8. The apparatus further includes a storage unit that stores a plurality of sentences included in at least any one of an accident case collection, an accident response manual, and a maintenance history in the plant as example sentences. The caption generation unit extracts at least one of the example sentences by searching from among the plurality of example sentences stored in the storage unit using at least any one of the plurality of words that have been extracted, and generates the caption based on the state information, the plurality of words, and the extracted example sentence. The caption generation apparatus according to claim 1.

9. The caption generation unit generates a plurality of captions and assigns a similarity to the extracted example sentences. The caption generation apparatus according to claim 8.

10. The caption includes a plurality of action options regarding actions to be taken by the user or instructions for the user. The caption generation apparatus according to claim 8.

11. The caption includes at least any one of an instruction to image the captured object from a different angle or a different shooting angle and an instruction to image another object related to the captured object. The caption generation apparatus according to claim 1.

12. The apparatus further includes a model storage unit that stores a caption generation model that has learned the relationship between the state information, one or more words representing the characteristics of the object in the plant, and the caption that describes the object in the plant. The caption generation unit generates the caption based on the newly input state information and the plurality of words using the caption generation model. The caption generation apparatus according to claim 1.

13. The apparatus further includes a learning unit that learns a word extraction model that extracts the plurality of words from the image using the result of the user's determination of the appropriateness of the generated caption, and a caption generation model that generates the caption from the state information and the plurality of words. The caption generation unit generates the caption based on the newly input state information and the plurality of words using the caption generation model. The caption generation apparatus according to claim 1.

14. The apparatus further includes a learning unit that learns a word extraction model that extracts the plurality of words from the image using the user's correction input for the generated caption, and a caption generation model that generates the caption from the state information and the plurality of words. The caption generation unit generates the caption based on the newly input state information and the plurality of words using the caption generation model. The caption generation device according to claim 1.

15. Obtaining an image captured within a plant; Extracting a plurality of words representing features of an object imaged in the image; Obtaining state information indicating the state of the imaged object; Generating a caption that describes the imaged object according to the state information based on the state information and the plurality of words A caption generation method comprising:

16. On a computer, A procedure for obtaining an image captured within a plant; A procedure for extracting a plurality of words representing features of an object imaged in the image; A procedure for obtaining state information indicating the state of the imaged object; A procedure for generating a caption that describes the imaged object according to the state information based on the state information and the plurality of words A program for causing execution.

Citation Information

Patent Citations

  • Optical display monitoring device, optical display monitoring system and optical display monitoring program

    JP2017049678A

  • Oil spill monitoring system and oil spill monitoring method

    JP2022112341A

  • Control system

    JP2023049535A