Caption generation device, caption generation method, and program
The caption generation device addresses the challenge of generating accurate captions for plant monitoring by using image analysis and machine learning to extract relevant information from plant images, resulting in improved monitoring and reporting efficiency.
Patent Information
- Application Number
- PCT/JP2024/038557
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-10-29
- Publication Date
- 2025-06-26
AI Technical Summary
Existing technologies lack an efficient method for generating accurate captions that describe the state of objects within a plant, particularly in industrial settings where real-time monitoring and reporting are crucial.
A caption generation device that acquires images from a plant, extracts relevant words, and generates captions based on state information such as position, deterioration, fluid state, and device information, using a combination of image analysis and machine learning models.
The solution enables the generation of accurate and informative captions that describe the state of objects within a plant, improving monitoring and reporting efficiency by providing detailed and relevant information in real-time.
Smart Images

Figure JP2024038557_26062025_PF_FP_ABST
Abstract
Description
Caption generation device, caption generation method, and program
[0001] The present invention relates to a caption generating device, a caption generating method, and a program.
[0002] Patent Document 1 states, "The video function abnormality determination means 16 has the function of inputting the abnormality determination result of the abnormality determination means 12, and determining whether there is an abnormality in a functional element that subdivides the plant's functions based on the result, data indicating the plant status input from the input means 11, and the image of the plant status from the ITV camera 19" (paragraph 0024). Patent Document 2 states, "When an alarm condition is met, the display device 5 uses the alarm function to display an alarm display screen. The alarm display screen displays the date and time of the abnormality and a comment explaining the abnormality" (paragraph 0039). Patent Document 3 states, "The system obtains location information identified by GPS, displays the current inspection location to plant personnel, and accumulates and analyzes image data captured by a built-in camera...to determine whether there are signs of an abnormality in the plant equipment based on past equipment status and abnormality cases, and simultaneously automatically creates a periodic inspection report from the image data using a preset format and text" (paragraph 0053). Patent Document 4 states, "The system includes a feature acquisition unit that acquires information about the piping as a first feature from an image obtained by photographing the plant equipment and piping around the plant equipment with the camera, and a feature comparison unit that compares the first feature with a second feature about the piping acquired from design data" (Claim 1). Patent Document 5 states, "The anomaly email creation function 102 is activated when the anomaly monitoring function 101 detects a plant anomaly, and creates an email message by documenting information that should be notified to a monitor as soon as possible, such as the date and time the anomaly was detected, the name of the plant equipment, and the details of the anomaly" (Paragraph 0019). [Prior Art Literature] [Patent Documents] [Patent Document 1] JP 2000-76570 A [Patent Document 2] JP 6419399 A [Patent Document 3] JP 6099989 A [Patent Document 4] JP 6826509 A [Patent Document 5] JP 2003-51895 A General disclosure
[0003] A first aspect of the present invention provides a caption generation device including an image acquisition unit that acquires an image captured within a plant, a word extraction unit that extracts a plurality of words that represent characteristics of an object captured in the image, a state information acquisition unit that acquires state information that indicates a state of the captured object, and a caption generation unit that generates a caption that describes the captured object in accordance with the state information, based on the state information and the plurality of words.
[0004] In the above-mentioned caption generation device, the status information may include at least any of the following: location information of the target, deterioration information indicating the deterioration state of the target, fluid information indicating the components or state of the fluid flowing inside the target pipe, and equipment information related to equipment located upstream or downstream of the target pipe.
[0005] In any of the above caption generation devices, the state information acquisition unit may acquire the state information by analyzing the image.
[0006] In any of the above caption generation devices, the status information may include position information of the object. Any of the above caption generation devices may further include a storage unit that stores past images of objects in the plant in association with position information. Any of the above caption generation devices may further include an image extraction unit that, in response to acquiring new status information for the new image, extracts the past image having position information common to the position information included in the new status information. In any of the above caption generation devices, the caption generation unit may generate the caption using both the new image and the past image corresponding to the new image.
[0007] In any of the above caption generation devices, the word extraction unit may assign a likelihood according to the state information to each of the plurality of words extracted from the image. In any of the above caption generation devices, the caption generation unit may generate the caption by preferentially using a word of the plurality of words having a higher likelihood based on the state information and the plurality of words.
[0008] In any of the above caption generation devices, the caption generation unit may assign a priority to the generated caption based on the likelihood of at least one word used to generate the caption from among the plurality of words.
[0009] In any of the above caption generation devices, the caption generation unit may output the caption including an instruction to re-image the captured object if the likelihood of the combination of multiple words used to generate the caption is lower than a predetermined threshold.
[0010] Any of the above caption generation devices may further include a storage unit that stores, as example sentences, sentences included in at least one of a collection of accident case studies in the plant, an accident response manual, and a maintenance history. In any of the above caption generation devices, the caption generation unit may extract at least one example sentence from the example sentences stored in the storage unit by searching using at least one of the extracted words, and generate the caption based on the status information, the words, and the extracted example sentence.
[0011] In any of the above caption generation devices, the caption generation unit may generate a plurality of captions and assign a degree of similarity to the extracted example sentence.
[0012] In any of the above caption generating devices, the caption may include a plurality of action options regarding actions that the user should take, or instructions to the user.
[0013] In any of the above caption generation devices, the caption may include at least one of instructions to image the imaged object from a different angle or a different angle of view, and instructions to image another object related to the imaged object.
[0014] Any of the above caption generation devices may further include a model storage unit that stores a caption generation model that has learned the relationship between the state information, one or more words that represent characteristics of objects within the plant, and captions that describe the objects within the plant. In any of the above caption generation devices, the caption generation unit may use the caption generation model to generate the caption based on newly input state information and the plurality of words.
[0015] Any of the above caption generation devices may further include a learning unit that uses a result of a user's judgment on the appropriateness of the generated caption to train a word extraction model that extracts the plurality of words from the image and a caption generation model that generates the caption from the state information and the plurality of words. In any of the above caption generation devices, the caption generation unit may use the caption generation model to generate the caption based on newly input state information and the plurality of words.
[0016] Any of the above caption generation devices may further include a learning unit that uses a user's correction input for the generated caption to train a word extraction model that extracts the plurality of words from the image and a caption generation model that generates the caption from the state information and the plurality of words. In any of the above caption generation devices, the caption generation unit may use the caption generation model to generate the caption based on newly input state information and the plurality of words.
[0017] In a second aspect of the present invention, there is provided a caption generation method, comprising: acquiring an image captured within a plant; extracting a plurality of words that represent characteristics of an object captured in the image; acquiring status information that indicates a status of the captured object; and generating a caption that describes the captured object in accordance with the status information based on the status information and the plurality of words.
[0018] In a third aspect of the present invention, there is provided a program that causes a computer to execute the steps of acquiring an image captured within a plant, extracting a plurality of words that represent characteristics of an object captured in the image, acquiring status information that indicates a status of the captured object, and generating a caption that describes the captured object in accordance with the status information based on the status information and the plurality of words.
[0019] The above summary of the invention does not list all of the features of the present invention, and subcombinations of these features may also be inventions.
[0020] FIG. 1 is a schematic overview diagram of a caption generation system 5 including a caption generation device 100 according to a first embodiment. FIG. 2 is a block diagram of the caption generation device 100 according to the first embodiment. FIG. 3 is a flowchart showing the flow of a caption generation method according to the first embodiment. FIG. 4 is a flowchart showing an example of a detailed flow of a caption generation method according to the first embodiment. FIG. 5 is a diagram for specifically explaining a part of the flow of FIG. 4. FIG. 6 is a block diagram of a caption generation device 200 according to a second embodiment. FIG. 7 shows an example of a computer 2200 in which multiple aspects of the present invention may be embodied in whole or in part.
[0021] The present invention will be described below through embodiments of the invention, but the following embodiments do not limit the scope of the invention as claimed. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention.
[0022] FIG. 1 is a schematic diagram of a caption generation system 5 including a caption generation device 100 according to a first embodiment. The caption generation system 5 generates a caption describing an image of an object 20 from an image 10 obtained by capturing the object 20 in a facility within a plant. More specifically, the caption generation system 5 generates a caption describing the captured object 20 in accordance with status information indicating the status of the captured object 20. The image 10 may be one or more still images, or may be a video. Note that a caption describing the object 20 in the image 10 may refer to an explanatory sentence describing the status of the object 20 shown in the image 10. The caption referred to here may refer to a sentence separated by periods, or may refer to a phrase separated more finely than a sentence.
[0023] Plants include industrial plants such as chemical plants, as well as plants that manage and control wellheads and surrounding areas of gas and oil fields, plants that manage and control hydroelectric, thermal, and nuclear power generation, plants that manage and control environmental power generation such as solar and wind power, and plants that manage and control water supply and sewage systems and dams.
[0024] 1 shows a pipe 21 and a tank 22 as an example of an object 20 in a facility within a plant. As shown in Fig. 1, the pipe 21 may be connected, for example, to the tank 22 on its upstream side, and liquid or gas stored in the tank 22 may flow downstream. Note that equipment on the downstream side of the pipe 21 is not shown in Fig. 1.
[0025] As an example, the caption generation system 5 includes an image capturer terminal 41 used by an image capturer 31, an administrator terminal 42 used by an administrator 32, a worker terminal 43 used by a recovery worker 33, and a caption generation device 100. The image capturer terminal 41, the administrator terminal 42, the worker terminal 43, and the caption generation device 100 can communicate with each other via a communication network 50. The communication network 50 may be a wired network, a wireless network, or may include both.
[0026] The imager terminal 41 captures the image 10 of the target 20. Alternatively or additionally, the image 10 of the target 20 may be captured by an imaging device such as a surveillance camera installed in the plant or a camera provided on a mobile robot or drone that patrols the plant. The imager terminal 41 may be a smartphone with a camera. The imager terminal 41 may be a tablet terminal. The imager terminal 41 may be a PC (Personal Computer). The imager terminal 41 may be a wearable terminal.
[0027] The photographer 31 is an example of a user who uses the caption generation system 5, and carries and uses a photographer terminal 41. The photographer 31 uses the photographer terminal 41 to capture an image 10 of an object 20 in a facility within a plant. As an example, the photographer 31 is a maintenance worker for equipment within the plant, and more specifically, a maintenance worker who maintains a tank 22. Note that there may be one or more photographers 31 within the plant.
[0028] The manager terminal 42 displays a report using captions generated by the caption generation system 5. The manager terminal 42 transmits instructions regarding the target 20 in the equipment within the plant to the worker terminal 43 via the communication network 50. The manager terminal 42 may be a PC, or may be located in a location away from the target 20 in the facility within the plant, such as a control room within the plant. The manager terminal 42 may be a smartphone. The manager terminal 42 may be a tablet terminal. The manager terminal 42 may be a wearable terminal.
[0029] The manager 32 is an example of a user who uses the caption generation system 5, and uses the manager terminal 42. The manager 32 uses the manager terminal 42 to view a report using captions generated by the caption generation system 5. As an example, the manager 32 views the report in a situation where the object 20 described in the caption cannot be directly visually confirmed. The manager 32 determines the status of the object 20 in a facility within the plant, and orders a recovery worker 33 or the like to perform recovery work on the object 20. As an example, the manager 32 determines the status of a pipe 21 in a facility within the plant, and orders a recovery worker 33 or the like to perform recovery work on the pipe 21. It should be noted that there may be one or more managers 32 in the plant. It should be noted that the manager 32 is an example of a viewer who views a report using captions generated by the caption generation system 5.
[0030] The worker terminal 43 displays commands related to the target 20 in the equipment within the plant, reports using captions generated by the caption generation system 5, etc. The worker terminal 43 may be a smartphone. The worker terminal 43 may be a tablet terminal. The worker terminal 43 may be a PC. The worker terminal 43 may be a wearable terminal.
[0031] The recovery worker 33 is an example of a user who uses the caption generation system 5 and uses the worker terminal 43. The recovery worker 33 uses the worker terminal 43 to view instructions regarding the object 20 in the equipment within the plant and reports using captions generated by the caption generation system 5. As an example, the recovery worker 33 views the report in a situation where the object 20 described in the caption cannot be directly visually confirmed. The recovery worker 33 deals with an abnormality within the plant. As an example, the recovery worker 33 deals with an abnormality in the piping 21. As an example, the recovery worker 33 restores the piping 21 where the abnormality has occurred in accordance with instructions and reports from the manager 32 displayed on the worker terminal 43. Note that there may be one or more recovery workers 33 within the plant. Note that the recovery worker 33 is an example of a viewer who views a report using captions generated by the caption generation system 5.
[0032] As an example, the caption generation device 100 receives an image 10 of an object 20 captured in a facility within a plant from the photographer terminal 41 via the communication network 50. The caption generation device 100 uses the image 10 received from the photographer terminal 41 to generate a caption that describes the object 20 within the plant in accordance with status information that indicates the status of the captured object 20. The caption generation device 100 generates a caption with higher accuracy than when a caption that describes the object 20 is generated from the image 10 of the object 20 without using status information, i.e., outputs an explanatory text that more accurately describes the status of the object 20.
[0033] The caption generation device 100 transmits the generated captions to the cameraman terminal 41 via the communication network 50. The caption generation device 100 may be located in a control room, an instrument room, or the like within the plant, or may be located outside the plant. Some of the functions of the caption generation device 100 may be incorporated into the cameraman terminal 41.
[0034] 2 is a block diagram of the caption generation device 100 according to the first embodiment. In FIG. 2, the flow of data and the like is indicated by arrows. The caption generation device 100 includes an image acquisition unit 110, a word extraction unit 120, a state information acquisition unit 130, an image extraction unit 135, a caption generation unit 140, a storage unit 150, and a report accumulation unit 155.
[0035] The image acquisition unit 110 acquires images 10 captured within a plant. As an example, the image acquisition unit 110 receives, from the image capturer terminal 41 via the communication network 50, images 10 captured of targets 20 within the plant, as well as the user ID of the photographer 31 using the image capturer terminal 41. As an example, the image acquisition unit 110 receives, from the image capturer terminal 41 via the communication network 50, images 10 captured of targets 20 within the plant, as well as the user ID of the photographer 31 and imaging information indicating the imaging position and imaging direction when the image capturer terminal 41 captured the images 10 of the targets 20. The image acquisition unit 110 inputs the received images 10 to the word extraction unit 120, and also inputs the images 10 together with the user ID of the photographer 31, etc., to the status information acquisition unit 130.
[0036] The word extraction unit 120 extracts a plurality of words that represent characteristics of the object 20 captured in the image 10. The word extraction unit 120 may extract one or more words based on a newly input image 10 using a word extraction model that has learned the relationship between the image 10 and one or more words that represent characteristics of the object 20 captured in the image 10. The word extraction unit 120 may read out a word extraction model stored in the storage unit 150, for example. The word extraction unit 120 inputs the multiple words extracted from the image 10 to the caption generation unit 140, for example, as vector expressions.
[0037] The status information acquisition unit 130 acquires status information indicating the status of the object 20 captured in the image 10. For example, the status information acquisition unit 130 acquires the status information by analyzing the image 10. Specifically, the status information acquisition unit 130 may analyze the image 10 to extract the object 20 from the image 10, and determine which piece of equipment in the plant the object 20 corresponds to by comparing the features of the extracted object 20 with the features of multiple objects 20 stored in the storage unit 150. The status information acquisition unit 130 may further access a distributed control system that controls the plant and is external to the caption generation device 100, and receive status information about the equipment corresponding to the object 20 from the distributed control system. The equipment corresponding to the object 20 may include the equipment itself, or may include equipment located upstream or downstream of the object 20. In addition, the status information acquisition unit 130 may extract the object 20 from the image 10 by analyzing the image 10, determine which equipment in the plant the object 20 corresponds to, and read out the status information regarding the equipment corresponding to the object 20 stored in the memory unit 150.
[0038] As another example, the status information acquisition unit 130 may acquire status information by acquiring imaging information from the photographer terminal 41. Specifically, the status information acquisition unit 130 may determine to which facility in the plant the target 20 captured in the image 10 corresponds by comparing the imaging information from the photographer terminal 41 with multiple pieces of imaging information stored in the storage unit 150. The status information acquisition unit 130 may further access a distributed control system and receive status information regarding the facility corresponding to the target 20 from the distributed control system. Note that the imaging information may additionally indicate the zoom degree when the photographer terminal 41 captured the image 10 and the distance to the target 20. The status information acquisition unit 130 inputs the acquired status information to the image extraction unit 135 and also inputs the status information, for example as a vector expression, to the caption generation unit 140 together with the user ID of the photographer 31, etc.
[0039] The status information may include at least any of the following: location information of the object 20, deterioration information indicating the deterioration state of the object 20, fluid information indicating the composition or state of a fluid flowing inside the piping that is the object 20, and equipment information related to equipment located upstream or downstream of the piping that is the object 20. The status information may include, for example, fixed information such as location information of the piping 21 and information about the tank 22 located upstream of the piping 21, or may include variable information indicating the type of gas or liquid flowing inside the piping 21, the state of the gas or liquid, such as its pressure and temperature, and the deterioration level of the piping 21. The deterioration level of the piping 21 may be an index of the deterioration state corresponding to the degree of rust on the piping 21, for example. As an example, the fixed status information may be stored in the storage unit 150. For example, the fixed status information may include plant design data indicating multiple objects 20 and equipment located upstream or downstream of each object 20.
[0040] When the status information includes the position information of the target 20, the image extraction unit 135, in response to acquiring new status information of the new image 10, extracts from the storage unit 150 past images having position information common to the position information included in the new status information. The image extraction unit 135 may extract one or more past images corresponding to the image 10 from the storage unit 150. The image extraction unit 135 inputs the past images extracted from the storage unit 150 together with the image 10 to the caption generation unit 140. Note that when the status information does not include the position information of the target 20, the image extraction unit 135 does not need to input anything to the caption generation unit 140.
[0041] The caption generation unit 140 generates a caption that describes the captured object 20 in accordance with the state information based on the state information and a plurality of words. The caption generation unit 140 may generate a caption based on newly input state information and a plurality of words using a caption generation model that has learned the relationship between the state information, one or more words that represent characteristics of the object 20 in the plant, and captions that describe the object 20 in the plant. The caption generation unit 140 may read out a caption generation model stored in the storage unit 150, for example. For example, the caption generation unit 140 may generate a caption using both a new image 10 input from the image extraction unit 135 and a past image corresponding to the new image 10. The caption generation unit 140 transmits the generated caption as text data to the image capturer terminal 41 via the communication network 50 based on the user ID of the image capturer 31.
[0042] The storage unit 150 stores past images of the target 20 in the plant in association with position information. The storage unit 150 may also store the above-mentioned fixed state information, word extraction models, and caption generation models. The storage unit 150 is an example of a model storage unit.
[0043] The report storage unit 155 receives reports from the photographer terminal 41 via the communication network 50 and registers them in memory. The report may be one that the photographer 31 has created by appropriately editing captions using the photographer terminal 41. The report storage unit 155 is accessed via the communication network 50 from the manager terminal 42, worker terminal 43, etc., and the reports stored in the memory are read out.
[0044] 3 is a flowchart showing the flow of the caption generation method according to the first embodiment. As an example, the flow starts when a photographer 31 captures an image 10 of an object 20 in a plant using a photographer terminal 41 and transmits the image 10 to the caption generation device 100 via a communication network 50.
[0045] The caption generation device 100 acquires an image 10 captured within a plant (step S101). Specifically, the caption generation device 100 receives the image 10, which is a captured image of an object 20 within the plant, along with the user ID of the photographer 31 from the photographer terminal 41 via the communication network 50. Alternatively, the caption generation device 100 may receive the user ID of the photographer 31 and location information of the photographer terminal 41 along with the image 10 from the photographer terminal 41 via the communication network 50.
[0046] The caption generation device 100 extracts a plurality of words that represent characteristics of the object 20 captured in the image 10 (step S102). Specifically, the caption generation device 100 extracts a plurality of words that represent characteristics of the object 20 captured in the image 10 using the above-described word extraction model.
[0047] The word extraction model learns the relationship between an image 10 and multiple image features, and also learns the relationship between the multiple image features and multiple words. When a new image 10 is input, the word extraction model extracts multiple image features from the image 10 and outputs multiple words representing the characteristics of the object 20 captured in the image 10 as a vector representation from the multiple image features. The word extraction model may be, for example, a deep learning model with a neural network structure including a CNN (Convolution Neural Network) and an FC layer (Fully Connected Layer). Instead of a CNN, the word extraction model may use a Transformer Self-Attention or Multi-Head Self-Attention.
[0048] The caption generation device 100 acquires status information indicating the status of the captured object 20 (step S103). Specifically, the caption generation device 100 acquires the status information by analyzing the image 10. More specifically, the caption generation device 100 analyzes the image 10 to extract the object 20 from the image 10, and determines which piece of equipment in the plant the object 20 corresponds to by comparing the features of the extracted object 20 with the features of multiple objects 20 stored in the storage unit 150. For example, the caption generation device 100 may identify features such as the shape and model number of the object 20 captured in the image 10 as a result of analyzing the image 10, compare them with the features of multiple objects 20 stored in the storage unit 150, and extract information about the equipment in the plant that is stored in association with the object 20 with matching features. For example, the caption generation device 100 may acquire, as a result of analyzing the image 10, identification information of the object 20 located on or around the surface of the object 20 captured in the image 10, compare it with the identification information of multiple objects 20 stored in the storage unit 150, and extract information about equipment in a plant that is stored in association with the object 20 with matching identification information. The caption generation device 100 further accesses a distributed control system and receives, from the distributed control system, status information about the equipment corresponding to the object 20.
[0049] Additionally or alternatively, the caption generation device 100 acquires status information by receiving the above-mentioned imaging information together with the image 10 from the photographer terminal 41. More specifically, the caption generation device 100 determines to which piece of equipment in the plant the object 20 captured in the image 10 corresponds by comparing the imaging information received from the photographer terminal 41 with multiple pieces of imaging information stored in the storage unit 150. The caption generation device 100 further accesses a distributed control system and receives status information regarding the equipment corresponding to the object 20 from the distributed control system.
[0050] The caption generation device 100 generates a caption that describes the captured object 20 in accordance with the state information based on the state information and a plurality of words (step S104). Specifically, the caption generation device 100 uses a caption generation model to generate, as text data, a caption that describes the captured object 20 in accordance with the state information based on the state information indicating the state of the object 20 and a plurality of words. The caption generation device 100 transmits the generated text data of the caption to the photographer terminal 41 via the communication network 50, thereby completing the flow.
[0051] The caption generation model has learned the relationship between state information indicating the state of the object 20, vector representations of words representing the characteristics of the object 20, and captions describing the object 20. When a vector representation of state information of the object 20 captured in the image 10 and a vector representation of multiple words representing the characteristics of the object 20 are newly input, the caption generation model outputs text data of a caption describing the object 20 captured in the image 10 from the vector representation of the state information and the vector representations of the multiple words. The caption generation model may be, for example, a deep learning model with a neural network structure including an RNN (Recurrent Neural Network). Instead of an RNN, the caption generation model may use a Long Short-Term Memory (LSTM), a Transformer, a Generative Pre-trained Transformer-3 (GPT-3), GPT-2, a GPT-3-Clone, or the like. The caption generation model may use a large number of example sentences as training data, and train an RNN or the like for each example sentence so as to increase the likelihood of outputting that example sentence. The caption generation model may extract words contained in example sentences, use vectors of the extracted words as training input data, and use example sentences as training output data, and train so as to increase the probability of outputting training output data for the training input data.
[0052] More specifically, regarding steps S102 to S104, the caption generation device 100 uses a word extraction model and a caption generation model to generate a caption that describes the imaged pipe 21 in accordance with the status information, based on multiple words extracted from the image 10 and the status information of the pipe 21 imaged in the image 10.
[0053] For example, as shown in FIG. 1 , when image 10 shows an abnormality in which a crack has appeared in pipe 21 and white gas is escaping from it, caption generation device 100 that has acquired image 10 may generate a caption such as, "A crack has appeared in the pipe with pipe ID XX in area A, and sulfuric acid gas is escaping from it" based on location information, deterioration information, and fluid information. For example, caption generation device 100 may generate a caption such as, "The crack may have occurred in the pipe due to worsening rust in the pipe" based on deterioration information. For example, caption generation device 100 may generate a caption such as, "A crack has appeared in the pipe and gas is escaping. Please close the valve of the tank upstream of the pipe" based on deterioration information and equipment information. For example, caption generation device 100 may generate a caption such as, "White smoke is escaping from a pipe that has been in use for one year" based on maintenance information for pipe 21, which is an example of deterioration information. For example, the caption generation device 100 may generate a caption such as "The pressure inside the pipe is XX kPa, and XX liquid is leaking and dripping from the joint" in accordance with fluid information including a measurement value of the pressure of the fluid flowing inside the pipe 21. The caption generation device 100 transmits the caption to the photographer terminal 41 via the communication network 50.
[0054] The photographer 31 checks the caption displayed on the photographer terminal 41 and transmits a report using the caption from the photographer terminal 41 to the caption generation device 100 via the communication network 50. The photographer 31 may edit the caption as appropriate to create a report. Upon receiving the report from the photographer terminal 41, the caption generation device 100 stores the report in the report storage unit 155. The report stored in the report storage unit 155 of the caption generation device 100 may be viewed, for example, by a restoration worker 33 of the piping 21, and the restoration worker 33 may seal a crack in the piping 21 or adjust a valve of a tank 22 upstream of the piping 21. The report stored in the report storage unit 155 of the caption generation device 100 may be viewed, for example, by a manager 32 of the piping 21, and the manager 32 may order the restoration worker 33 of the piping 21 to perform restoration work or adjust control parameters of the equipment to which the piping 21 is installed.
[0055] In steps S103 and S104, if the status information includes location information of the object 20, the caption generation device 100 may extract from the storage unit 150 one or more past images having location information common to the location information, and generate a caption using both the image 10 and the one or more past images. For example, the caption generation device 100 may identify a change in the object 20 captured in the image 10, such as the progress of rust on the surface of the object 20, by comparing the one or more past images with the image 10 according to the date and time at which each image was captured. In this case, the caption generation device 100 may generate a caption according to the change and the status information, based on the status information, the change in the object 20 identified as a result of the comparison, and a plurality of words.
[0056] According to the caption generation device 100 of the first embodiment described above, a caption describing the object 20 is generated from the image 10 of the object 20 captured within a plant, in accordance with status information indicating the status of the object 20. As a result, the caption generation device 100 generates a caption with higher accuracy than when a caption describing the object 20 is generated from the image 10 of the object 20 without using status information, i.e., it is possible to output an explanatory sentence that more accurately describes the status of the object 20.
[0057] In the caption generation device 100 according to the first embodiment, the word extraction unit 120 may receive the state information acquired by the state information acquisition unit 130, and may assign a likelihood corresponding to the state information to each of a plurality of words extracted from the image 10. For example, the word extraction unit 120 may receive the image 10 of the object 20, along with a deterioration degree that is an index of the deterioration state corresponding to the degree of rust of the object 20, as state information of the object 20 captured in the image 10. In this case, the word extraction unit 120 may assign a high likelihood to words that are highly related to the deterioration state of the object 20, and a low likelihood to words that are less related to the deterioration state of the object 20.
[0058] More specifically, the word extraction unit 120 may use a word extraction model to assign likelihoods to each of a plurality of words extracted from the image 10 according to the state information. The word extraction model, for example, inputs pixel values of each pixel of the image to each input node of the input stage and outputs a scalar value for each word from each output node of the output stage. The scalar value may be an output value with a reference range of 0 to 1. In this case, the word vector is a vector having the scalar value of 0 to 1 for each word. When the output value of a word is 0, it means that the image 10 does not represent that word, and when the output value of a word is 1, it means that the image 10 represents that word. The word extraction model has been trained, for example, to output a value closer to 1 the more likely the image 10 is to represent that word. The word extraction model has further been repeatedly trained using, for example, a certain image 10 and state information acquired from the image 10 by the state information acquisition unit 130 as input training data, and multiple words extracted by decomposing a caption actually generated from the image 10 and the state information by the caption generation unit 140 as output training data. In this case, the word extraction model outputs a value closer to 1 the more likely the image 10 represents the word, and the closer the relevance between the state information and the word, the more likely it outputs a value closer to 1. For convenience, the word extraction unit 120 treats the scalar value of each word output from the word extraction model as a likelihood and assigns it to each word.
[0059] As a specific example, assume that the image 10 acquired by the word extraction unit 120 shows a tank 22, a pipe 21 connected to the tank 22, a crack in the pipe 21, gas emanating from the crack, a valve attached to the pipe 21, other equipment that is not directly or indirectly connected to the tank 22 or the pipe 21, and a wall surrounding the tank 22 and the pipe 21. For example, when the reference range of the output value for each word is 0 to 1, the word extraction model may set the output value for the word "tank" to 0.9, the output value for the word "piping," the output value for the word "crack," the output value for the word "gas," the output value for the word "valve," and the output value for the words referring to other equipment and "wall" to 0.2 based on the degree of deterioration. In this case, the word extraction unit 120 may treat the probability of the word "crack" as 80%, for example, using the output values from the word extraction model.
[0060] In this case, the caption generation unit 140 may generate a caption by giving priority to a word with a higher likelihood among the multiple words based on the state information and the multiple words, thereby enabling the caption generation device 100 to generate captions with even higher accuracy.
[0061] The caption generation unit 140 may generate multiple captions. In this case, the caption generation unit 140 may further assign a priority to the generated caption based on the likelihood of at least one word used to generate the caption among the multiple words. The priority may be an average value, a total value, or an integrated value of the likelihoods of the multiple words. The caption generation unit 140 may transmit the caption with the highest priority among the multiple generated captions to the image capturer terminal 41. Alternatively, the caption generation unit 140 may assign a priority to each of the multiple generated captions and transmit the multiple generated captions to the image capturer terminal 41. In this case, the image capturer 31 can determine the appropriateness of the multiple captions by referring to the priority.
[0062] As an example of generating multiple captions, the caption generation unit 140 may assign a negative weight to a word used in generating a first caption among the multiple words extracted by the word extraction unit 120, and then generate a second caption so that the word used in the first caption is not included in the second caption. As another example of generating multiple captions, the caption generation unit 140 may assign an arbitrary priority to the multiple words extracted by the word extraction unit 120, and then use the multiple words extracted by the word extraction unit 120 as a whole by decreasing the priority of the word by a predetermined amount each time the word is used in caption generation.
[0063] Furthermore, when the likelihood of a combination of multiple words used to generate the caption is lower than a predetermined threshold, the caption generation unit 140 may output a caption including an instruction to re-image the imaged object 20. For example, when the average value, total value, integrated value, or the like of the likelihood of multiple words used to generate the caption is lower than a predetermined threshold stored in the storage unit 150, the caption generation unit 140 may output a caption including an instruction to re-image the imaged object 20.
[0064] In the caption generation device 100 according to the first embodiment, the storage unit 150 may store, as example sentences, sentences contained in at least one of a collection of accident case studies in a plant, an accident response manual, and a maintenance history. In this case, the caption generation unit 140 may extract at least one example sentence from the example sentences stored in the storage unit 150 by searching using at least one of the words extracted by the word extraction unit 120.
[0065] The caption generation unit 140 may generate a caption based on the state information acquired by the state information acquisition unit 130, the plurality of words, and the extracted example sentences. Specifically, the caption generation unit 140 may search for example sentences in which at least any of the plurality of words is used from the plurality of example sentences stored in the storage unit 150, and generate a caption while referring to the extracted example sentences.
[0066] In this case, the caption generation unit 140 may generate multiple captions and assign a similarity to the extracted example sentence. Specifically, the caption generation unit 140 may generate multiple captions based on the state information and multiple words, search for example sentences containing at least one of the multiple words among the multiple example sentences stored in the storage unit 150, and assign a similarity to each of the multiple captions with the extracted example sentence. For example, the caption generation unit 140 may convert each of the multiple captions and the extracted example sentence into a feature vector indicating a combination of multiple words contained therein, and calculate the dot product of the feature vectors to determine the similarity. The caption generation unit 140 may transmit the caption with the highest similarity among the multiple generated captions to the image capturer terminal 41. Alternatively, the caption generation unit 140 may transmit the multiple generated captions to the image capturer terminal 41 with the similarity assigned to each caption. In this case, the photographer 31 can judge the suitability of a plurality of captions by referring to the similarity.
[0067] The caption generated by the caption generation unit 140 with reference to the extracted example sentences may include, in addition to a caption describing the object 20 in the plant, multiple action options regarding actions that the user should take or instructions to the user. As a result of generating a caption with reference to the example sentences, the caption generation device 100 can present to the user action options and instructions that incorporate methods for dealing with abnormalities that cannot be directly deciphered from the image 10.
[0068] The captions generated by the caption generation device 100 according to the first embodiment may include, in addition to captions describing the objects 20 within the plant, at least one of instructions to capture the captured object 20 at a different angle or a different angle of view, and instructions to capture other objects 20 related to the captured object 20.
[0069] Fig. 4 is a flowchart showing an example of a detailed flow of the caption generation method according to the first embodiment. The operational flow shown in the flowchart of Fig. 4 may be a specific example of the caption generation method according to the first embodiment described using Fig. 3. As an example, the caption generation device 100 according to the first embodiment described above may generate a caption describing the target 20 from an image 10 capturing the target 20, in accordance with the operational flow shown in the flowchart of Fig. 4.
[0070] As an example, the flow of Figure 4 begins when a photographer 31 captures an image 10 of an object 20 in a plant using a photographer terminal 41 and sends it to the caption generation device 100 via a communication network 50.
[0071] The caption generation device 100 acquires an image 10 captured within a plant (step S201). In step S201, as shown in FIG. 5 , which specifically describes a portion of the flow in FIG. 4 , the image acquisition unit 110 receives an image 10 of an object 20 within a plant from the photographer terminal 41 via the communication network 50, and extracts Exif information from the image 10. The Exif information may include the user ID of the photographer 31. The image acquisition unit 110 may output the image 10 to the word extraction unit 120, and may output the Exif information together with the image 10 to the state information acquisition unit 130.
[0072] The caption generation device 100 crops the image 10 into multiple images (step S202). For example, the caption generation device 100 discards peripheral areas of the image 10 that do not include the object 20, creating a new image and dividing the image into fixed-size patches. In step S202, as shown in FIG. 5 , the word extraction unit 120 may create a new image from the image 10 by discarding peripheral areas that do not include the pipe 21 and the tank 22, and divide the new image vertically and horizontally into three to create nine patches. Alternatively, the word extraction unit 120 may create one fixed-size patch from the image 10 by discarding peripheral areas that do not include the pipe 21, and another fixed-size patch from the image 10 by discarding peripheral areas that do not include the tank 22. The learning model of the word extraction unit 120 shown in FIG. 5 may also be referred to as a Vision Transformer and may be an example of the above-mentioned word extraction model.
[0073] The caption generation device 100 applies a simple function, such as the LINEAR function, to the cropped image to convert it into a numerical value A that can be used by the encoder (step S203). In step S203, as shown in FIG. 5, the word extraction unit 120 may put the nine patches into a linear projection layer, flatten each, and linearly embed them into vectors (convert them into a one-dimensional array). The word extraction unit 120 may further add a CLS token to the beginning of the sequence of vectors coming out of the linear projection layer (extra learnable [class] embedding) and embed a position into each patch (each vector) (position embedding) to obtain the numerical value A. In other words, the word extraction unit 120 may convert the nine patches divided from the image 10 into a 10-dimensional sequence of vectors that can be used by the encoder.
[0074] In the caption generation device 100, the Transformer Encoder converts a numerical value A into another value B (step S204) and standardizes the dimensions and numerical width of value B (step S205). In steps S204 and S205, as shown in FIG. 5, the word extraction unit 120 inputs a 10-dimensional vector sequence into the Transformer Encoder, which then converts it into another 10-dimensional vector sequence, standardizes the dimensions and numerical width using, for example, a softmax function, and outputs the result as a CLS token. Note that the Transformer Encoder may include multiple components, such as the above-mentioned self-attention. The output of the CLS token may be a collection of features required for classification from the entire image using self-attention.
[0075] The caption generation device 100 outputs image features (step S206). In step S206, as shown in FIG. 5, the word extraction unit 120 may input the output of the CLS tokens into a classification head (MLP: Multilayer Perceptron) and output multiple image features from the classification head. The classification head may compress the multiple image features before outputting them. As shown in FIG. 5, the word extraction unit 120 may output multiple image features to the caption generation unit 140, or, as described with reference to FIGS. 1 to 3, may extract word vectors of multiple words representing characteristics of the object 20 captured in the image 10 from the multiple image features and output them to the caption generation unit 140.
[0076] The caption generation device 100 acquires state information (step S207), converts the image features and the state information into a numerical value C suitable for the decoder (step S208), and inputs the numerical value C into the decoder (step S209). In steps S207 to S209, as shown in FIG. 6 , which specifically describes a portion of the flow of FIG. 4 , the caption generation unit 140 vectorizes the state information input from the state information acquisition unit 130, for example, through the LINEAR function, vectorizes empty text, and vectorizes these vectors together with the multiple image features input from the word extraction unit 120. As a result, the caption generation unit 140 converts each of the multiple image features into a numerical value C. The caption generation unit 140 further inputs the numerical value C into the decoder. The decoder's learning model is a model that uses layer-by-layer training, may also be referred to as GPT-2, and may be an example of the caption generation model described above.
[0077] In the caption generation device 100, the decoder converts the numerical value C into a numerical value D (step S210), and then converts the numerical value D into text (tokens) through a function that converts the numerical value D back into text (step S211). In steps S210 to S211, as shown in FIG. 6 , the decoder in the caption generation unit 140 may convert the numerical value C into a numerical value D. The decoder may weight each numerical value D based on vectorized state information. The decoder may compress the numerical value D. The decoder may convert the numerical value D into text (tokens) through a softmax function.
[0078] The caption generation device 100 outputs the text (caption) (step S212), and the flow ends. In step S212, as shown in FIG. 1 ~X 5 A sequence of continuous text is taken as input (source token), and a sequence Y is obtained by shifting one token to the right. 1 ~Y 5The sequence of continuous text may be treated as a target token, and the source token and target token may be concatenated and adjusted layer by layer for processing. The source token and target token may have their respective resettable positions embedded in them along with their respective corresponding tokens. The position of the source token starts at zero, and for the target token, instead of incrementing the position at the end of the source sentence, the position may be reset again to zero.
[0079] The decoder model shown in FIG. 6 includes, as an example, an N-layer stack of layers, in which an attention mechanism including a self-attention mechanism and a mixed attention mechanism, an add & layer normalization layer, a feedforward (FNN) layer, and an add & layer normalization layer are arranged in this order. The decoder inputs the processed data from the N layers into a linearization layer, and passes the data output from the linearization layer through a softmax function to obtain Y 1 ~Y 5 The caption generator 140 may transmit the generated caption text data to the photographer terminal 41 via the communication network 50.
[0080] 7 is a block diagram of a caption generation device 200 according to the second embodiment. The caption generation device 200 according to the second embodiment includes a learning unit 260 in addition to the components included in the caption generation device 100 according to the first embodiment. Other components included in the caption generation device 200 according to the second embodiment are similar to those of the caption generation device 100 according to the first embodiment, and the reference numbers of the components included in the caption generation device 100 will be used to omit redundant explanations.
[0081] The learning unit 260 learns the above-mentioned word extraction model and caption generation model using the result of a user's judgment on the appropriateness of the caption generated by the caption generation unit 140. Additionally or alternatively, the learning unit 260 may learn the above-mentioned word extraction model and caption generation model using a user's correction input for the caption generated by the caption generation unit 140. The learning unit 260 receives the above-mentioned user's judgment result and correction input from the image capturer terminal 41 via the communication network 50.
[0082] The learning unit 260 updates the model parameters so as to reduce the error between the model output and the label when each sample in the learning data including the results of the above-mentioned determination is input to the model. For example, when training a word extraction model using a multilayer neural network including a CNN or the like, the learning unit 260 adjusts the weights between each neuron in the neural network and the bias of each neuron by a method such as backpropagation using the error between the label and the output value output by the neural network in response to the input of each sample.
[0083] In this way, the caption generation device 200 according to the second embodiment receives feedback on whether the generated caption was actually used by the user, and uses the feedback to train the word extraction model and the caption generation model. The caption generation device 200 may determine that the generated caption was actually used, for example, when it receives from the administrator terminal 42 a result indicating that the administrator 32 has viewed the caption and determined that there is no problem with it. The caption generation device 200 may also determine that the generated caption was actually used, for example, when the caption generation unit 140 uses the generated caption to search for example sentences or uses the generated caption in a report to the administrator 32. The caption generation device 200 according to the second embodiment achieves the same effects as the caption generation device 100 according to the first embodiment. The caption generation device 200 according to the second embodiment can generate captions with even higher accuracy.
[0084] Various embodiments of the present invention may be described with reference to flowcharts and block diagrams, where the blocks may represent (1) stages of a process in which operations are performed or (2) sections of apparatus responsible for performing the operations. Particular stages and sections may be implemented by dedicated circuitry, programmable circuitry provided with computer-readable instructions stored on a computer-readable medium, and / or a processor provided with computer-readable instructions stored on a computer-readable medium. Dedicated circuitry may include digital and / or analog hardware circuitry, and may include integrated circuits (ICs) and / or discrete circuits. Programmable circuitry may include reconfigurable hardware circuitry including logical AND, OR, XOR, NAND, NOR, and other logic operations, flip-flops, registers, memory elements such as field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), and the like.
[0085] A computer-readable medium may include any tangible device capable of storing instructions that are executed by an appropriate device, such that the computer-readable medium having instructions stored thereon comprises an article of manufacture containing instructions that can be executed to create means for performing the operations specified in the flowcharts or block diagrams. Examples of computer-readable media may include electronic, magnetic, optical, electromagnetic, and semiconductor storage media. More specific examples of computer-readable media may include floppy disks, diskettes, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), electrically erasable programmable read-only memories (EEPROMs), static random access memories (SRAMs), compact disc read-only memories (CD-ROMs), digital versatile discs (DVDs), Blu-ray discs, memory sticks, integrated circuit cards, and the like.
[0086] The computer readable instructions may include either assembler instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, JAVA®, C++, etc., and conventional procedural programming languages such as the “C” programming language or similar programming languages.
[0087] The computer-readable instructions may be provided to a processor or programmable circuitry of a programmable data processing apparatus, such as a general-purpose computer, special-purpose computer, or other computer, either locally or over a local area network (LAN), a wide area network (WAN) such as the Internet, etc., which executes the computer-readable instructions to create means for performing the operations specified in the flowcharts or block diagrams. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.
[0088] 8 illustrates an example of a computer 2200 in which aspects of the present invention may be embodied, in whole or in part. Programs installed on the computer 2200 may cause the computer 2200 to function as or perform operations associated with an apparatus or one or more sections of the apparatus according to embodiments of the present invention, and / or to perform a process or steps of a process according to embodiments of the present invention. Such programs may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks in the flowcharts and block diagrams described herein.
[0089] A computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphics controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes input / output units such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.
[0090] The CPU 2212 operates according to programs stored in the ROM 2230 and RAM 2214, thereby controlling each unit. The graphics controller 2216 acquires image data generated by the CPU 2212 into a frame buffer or the like provided in the RAM 2214 or into the graphics controller 2216 itself, and causes the image data to be displayed on the display device 2218.
[0091] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads programs or data from the DVD-ROM 2201 and provides the programs or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to an IC card.
[0092] ROM 2230 stores therein a boot program or the like that is executed by computer 2200 upon activation, and / or programs that depend on the hardware of computer 2200. I / O chip 2240 may also connect various I / O units to I / O controller 2220 via parallel ports, serial ports, keyboard ports, mouse ports, etc.
[0093] The programs are provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The programs are read from the computer-readable medium, installed in the hard disk drive 2224, RAM 2214, or ROM 2230, which are also examples of computer-readable media, and executed by the CPU 2212. Information processing described in these programs is read by the computer 2200, and brings about cooperation between the programs and the various types of hardware resources described above. An apparatus or method may be configured by implementing information manipulation or processing in accordance with the use of the computer 2200.
[0094] For example, when communication is performed between computer 2200 and an external device, CPU 2212 may execute a communication program loaded in RAM 2214 and instruct communication interface 2222 to perform communication processing based on the processing described in the communication program. Under the control of CPU 2212, communication interface 2222 reads transmission data stored in a transmission buffer processing area provided in RAM 2214, hard disk drive 2224, DVD-ROM 2201, or a recording medium such as an IC card, and transmits the read transmission data to the network, or writes received data received from the network to a reception buffer processing area or the like provided on the recording medium.
[0095] Furthermore, the CPU 2212 may cause all or a necessary portion of a file or database stored on an external recording medium such as the hard disk drive 2224, the DVD-ROM drive 2226 (DVD-ROM 2201), an IC card, etc. to be read into the RAM 2214, and may perform various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording medium.
[0096] Various types of information, such as various types of programs, data, tables, and databases, may be stored on the recording medium and may undergo information processing. The CPU 2212 may perform various types of processing on data read from the RAM 2214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., as described throughout this disclosure and specified by the instruction sequences of the programs, and write the results back to the RAM 2214. The CPU 2212 may also search for information in a file, database, etc. on the recording medium. For example, if multiple entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored on the recording medium, the CPU 2212 may search for an entry that matches a condition specified by the attribute value of the first attribute from among the multiple entries, read the attribute value of the second attribute stored in the entry, and thereby obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.
[0097] The above-described programs or software modules may be stored in a computer-readable medium on or near the computer 2200. A recording medium such as a hard disk or RAM provided in a server system connected to a dedicated communication network or the Internet can also be used as a computer-readable medium, thereby providing the programs to the computer 2200 via the network.
[0098] Although the present invention has been described above using embodiments, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various modifications and improvements can be made to the above embodiments. It is clear from the claims that such modifications and improvements can also be included within the technical scope of the present invention.
[0099] For example, the control system may be a computer housed in a single housing. That is, the controller may be realized by executing a program on the computer's processor, and each input / output device may be implemented as an I / O device of the computer. The controller may also be implemented as a virtual machine executed by one or more processors. In such a configuration, the control system does not include a network, whether a general-purpose or dedicated network, and the controller and the input / output devices may be connected by a chipset, such as a memory controller hub and an I / O controller hub, that connect the processor and the I / O devices.
[0100] It should be noted that the order of execution of each process, such as operations, procedures, steps, and stages, in the devices, systems, programs, and methods shown in the claims, specifications, and drawings is not specifically stated as "before," "prior to," etc., and that the processes can be performed in any order unless the output of a previous process is used in a subsequent process. Even if the operational flow in the claims, specifications, and drawings is described using "first," "next," etc. for convenience, this does not mean that the processes must be performed in this order.
[0101] 5 Caption generation system 10 Image 20 Object 21 Pipe 22 Tank 31 Photographer 32 Manager 33 Recovery worker 41 Photographer terminal 42 Manager terminal 43 Worker terminal 100 Caption generation device 110 Image acquisition unit 120 Word extraction unit 130 Status information acquisition unit 135 Image extraction unit 140 Caption generation unit 150 Memory unit 155 Report accumulation unit 200 Caption generation device 260 Learning unit 2200 Computer 2201 DVD-ROM 2210 Host controller 2212 CPU 2214 RAM 2216 Graphics controller 2218 Display device 2220 Input / output controller 2222 Communication interface 2224 Hard disk drive 2226 DVD-ROM drive 2230 ROM 2240 Input / output chip 2242 Keyboard
Claims
1. A caption generation device comprising: an image acquisition unit that acquires images captured within a plant; a word extraction unit that extracts a plurality of words that express characteristics of an object captured in the image; a status information acquisition unit that acquires status information that indicates the status of the imaged object; and a caption generation unit that generates a caption that describes the imaged object in accordance with the status information based on the status information and the plurality of words.
2. A caption generating device as described in claim 1, wherein the status information includes at least any of the following: location information of the object, deterioration information indicating the deterioration state of the object, fluid information indicating the components or state of the fluid flowing inside the target piping, and equipment information related to equipment located upstream or downstream of the target piping.
3. The caption generation device according to claim 2, wherein the state information acquisition unit acquires the state information by analyzing the image.
4. The caption generation device of claim 1, further comprising: a memory unit which stores past images of the objects in the plant in association with the location information, the status information including location information of the objects; and an image extraction unit which extracts the past images having location information in common with the location information included in the new status information in response to obtaining new status information of the new image; and wherein the caption generation unit generates the caption using both the new image and the past image corresponding to the new image.
5. The caption generation device described in claim 1, wherein the word extraction unit assigns a likelihood corresponding to the state information to each of the multiple words extracted from the image, and the caption generation unit generates the caption based on the state information and the multiple words, giving priority to words among the multiple words having a higher likelihood.
6. The caption generating device according to claim 5, wherein the caption generating unit assigns a priority to the generated caption according to the likelihood of at least one word used in generating the caption among the plurality of words.
7. The caption generation device of claim 5, wherein the caption generation unit outputs the caption including an instruction to re-image the captured object when the likelihood of a combination of multiple words used to generate the caption is lower than a predetermined threshold.
8. A caption generation device as described in claim 1, further comprising a memory unit that stores a plurality of example sentences contained in at least one of a collection of accident cases within the plant, an accident response manual, and a maintenance history, wherein the caption generation unit extracts at least one example sentence from the plurality of example sentences stored in the memory unit by searching using at least one of the extracted plurality of words, and generates the caption based on the status information, the plurality of words, and the extracted example sentence.
9. The caption generating device according to claim 8, wherein the caption generating unit generates a plurality of the captions and assigns a degree of similarity to the extracted example sentence.
10. The caption generating device of claim 8, wherein the caption includes a plurality of action options or instructions to the user regarding an action to be taken by the user.
11. The caption generating device of claim 1, wherein the caption includes at least one of instructions to image the imaged object from a different angle or different angle of view, and instructions to image another object related to the imaged object.
12. A caption generation device as described in claim 1, further comprising a model memory unit that stores a caption generation model that has learned the relationship between the status information, one or more words that represent characteristics of objects within the plant, and captions that describe the objects within the plant, wherein the caption generation unit uses the caption generation model to generate the caption based on the newly input status information and the multiple words.
13. A caption generation device as described in claim 1, further comprising a learning unit that uses the results of a user's judgment on the appropriateness of the generated caption to train a word extraction model that extracts the multiple words from the image, and a caption generation model that generates the caption from the state information and the multiple words, wherein the caption generation unit uses the caption generation model to generate the caption based on newly input state information and the multiple words.
14. A caption generation device as described in claim 1, further comprising a learning unit that uses user correction input for the generated caption to train a word extraction model that extracts the multiple words from the image, and a caption generation model that generates the caption from the state information and the multiple words, wherein the caption generation unit uses the caption generation model to generate the caption based on the newly input state information and the multiple words.
15. A caption generation method comprising: acquiring an image taken within a plant; extracting a plurality of words that represent characteristics of an object captured in the image; acquiring status information that indicates a status of the imaged object; and generating a caption that describes the imaged object in accordance with the status information based on the status information and the plurality of words.
16. A program for causing a computer to execute the following steps: acquiring an image taken within a plant; extracting a plurality of words that express characteristics of an object captured in the image; acquiring status information that indicates the status of the captured object; and generating a caption that describes the captured object in accordance with the status information based on the status information and the plurality of words.
Citation Information
Patent Citations
Optical display monitoring device, optical display monitoring system and optical display monitoring program
JP2017049678A
Oil spill monitoring system and oil spill monitoring method
JP2022112341A
Control system
JP2023049535A