Caption generation apparatus, caption generation method and program

The caption generation device enhances plant maintenance by accurately describing equipment through image analysis, addressing the ambiguity in existing systems and improving maintenance efficiency.

JP2025098527APending Publication Date: 2025-07-02YOKOGAWA ELECTRIC CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023214720
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-07-02

AI Technical Summary

Technical Problem

Existing systems lack an efficient method to generate descriptive captions for objects within plant environments, leading to ambiguous and unclear descriptions that hinder effective maintenance and monitoring.

Method used

A caption generation device that includes an image acquisition unit, feature extraction unit, region specification unit, grouping unit, and word extraction unit to generate detailed captions for objects within plant images, utilizing machine learning models to improve accuracy and clarity.

Benefits of technology

Facilitates clear and actionable descriptions of plant equipment, enhancing maintenance efficiency by providing precise captions that aid in identifying issues and guiding maintenance actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025098527000001_ABST
    Figure 2025098527000001_ABST
Patent Text Reader

Abstract

To provide a caption generation apparatus, a caption generation method, and a program for generating a caption that describes an object captured, from an image obtained by imaging the inside of a plant.SOLUTION: A caption generation apparatus 100 includes: an image acquisition unit 110 which acquires an image captured in a plant; a feature extraction unit 112 which extracts a plurality of features from the image; a region specifying unit 114 which specifies image regions corresponding to the features, respectively; a grouping unit 116 which groups the features based on the image regions; and a word extraction unit 120 which extracts a plurality of words representing characteristics of an object captured in the image; and a caption generation unit 140 which generates, for each of the groups, based on at least one word in the group, a first caption that describes the object in an image range corresponding to the group.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a caption generation device, a caption generation method, and a program.

Background Art

[0002] Patent Document 1 describes that "location information specified by GPS is obtained, the current inspection location is displayed to plant staff, and image data captured by an installed camera is... accumulated and analyzed to determine whether there are signs of abnormality in plant equipment from past equipment states and abnormal cases, etc., and at the same time, using a preset format and text, a report of regular inspection is automatically created from the image data" (paragraph 0053). Patent Document 2 describes that "a feature amount acquisition unit that acquires information about the pipe as a first feature amount from an image obtained by photographing the plant equipment to be worked on and the pipes existing around the plant equipment with the camera, and a feature amount comparison unit that compares the first feature amount with a second feature amount about the pipe obtained from design data" (claim 1). Patent Document 3 describes that "the abnormal mail creation function 102 is activated when the abnormality monitoring function 101 detects a plant abnormality, and creates a mail transmission text by formulating matters to be communicated to the supervisor at an early stage, such as the date and time when the abnormality was detected, the name of the corresponding plant equipment, and the content of the abnormality" (paragraph 0019). [Prior Art Documents] [Patent Documents] [Patent Document 1] Patent No. 6099989 [Patent Document 2] Patent No. 6826509 [Patent Document 3] Japanese Unexamined Patent Application Publication No. 2003 - 51895

Summary of the Invention

[0003] In a first aspect of the present invention, a caption generation device is provided. The caption generation device includes an image acquisition unit that acquires an image captured within a plant, a feature extraction unit that extracts a plurality of features from the image, a region identification unit that identifies an image region corresponding to each of the plurality of features, a grouping unit that groups the plurality of features based on the image region, a word extraction unit that extracts a plurality of words representing the features of the object imaged in the image, and a caption generation unit that generates a first caption for explaining the features of the object within the image range corresponding to the group, based on at least one of the words within the group for each of the plurality of groups.

[0004] In the caption generation device described above, the caption generation unit may generate a second caption for explaining the object within the image by combining the first captions for each of the image ranges.

[0005] In any of the caption generation devices described above, the caption generation unit may generate a plurality of the second captions by combining the first captions of a plurality of the image ranges adjacent to each other.

[0006] Any of the caption generation devices described above may further include a model storage unit that stores a caption generation model that learns the relationship between at least one of the words representing the features of the object and the first caption for explaining the features of the object, and learns the relationship between the combination of the plurality of the first captions and the second caption by annotation of one or more relationships between the words representing the features of the object within one of the image ranges and the words representing the features of the object within another of the image ranges. In any of the caption generation devices described above, the caption generation unit may use the caption generation model to generate the first caption based on the at least one word newly input, and generate the second caption from the plurality of the first captions newly generated.

[0007] Any of the above caption generation devices may further include a learning unit that learns a feature extraction model that extracts the plurality of features from the image, a word extraction model that extracts the plurality of words from the image, and a caption generation model that generates the first caption from the plurality of words and generates the second caption from the plurality of generated first captions, using the result of the user's determination of the suitability of the generated second caption. In any of the above caption generation devices, the caption generation unit may generate the first caption based on the at least one newly input word using the caption generation model, and generate the second caption based on the plurality of newly generated first captions.

[0008] Any of the above caption generation devices may further include a learning unit that learns a feature extraction model that extracts the plurality of features from the image, a word extraction model that extracts the plurality of words from the image, and a caption generation model that generates the first caption from the plurality of words and generates the second caption from the plurality of generated first captions, using the user's correction input for the generated second caption. In any of the above caption generation devices, the caption generation unit may generate the first caption based on the at least one newly input word using the caption generation model, and generate the second caption based on the plurality of newly generated first captions.

[0009] In any of the above caption generation devices, the caption generation unit may generate the second caption including both the word representing the feature of the fluid-related object within one image range corresponding to one of the first captions and the word representing the liquid or gas within another image range adjacent to the one image range and corresponding to another of the first captions, from one or more relationships between them.

[0010] Any of the above caption generation devices may further include a storage unit that stores a plurality of sentences included in at least any one of the accident case collection, accident response manual, and maintenance history in the plant as example sentences. In any of the above caption generation devices, the caption generation unit extracts at least one of the example sentences by searching using at least any one of the plurality of words extracted from among the plurality of example sentences stored in the storage unit, generates the first caption based on the at least one word and the extracted example sentence, and may generate the second caption based on the newly generated plurality of first captions and the extracted example sentence.

[0011] In any of the above caption generation devices, the caption generation unit may generate a plurality of the second captions and assign a similarity to the extracted example sentences.

[0012] In any of the above caption generation devices, the second caption may include a plurality of action options regarding actions to be taken by the user or instructions to the user.

[0013] In any of the above caption generation devices, the second caption may include at least any one of an instruction to image the imaged object from another angle or another image angle and an instruction to image another object related to the imaged object.

[0014] In a second aspect of the present invention, a caption generation method is provided. The caption generation method includes obtaining an image captured in a plant, extracting a plurality of features from the image, identifying an image region corresponding to each of the plurality of features, grouping the plurality of features based on the image region, extracting a plurality of words representing the features of the object imaged in the image, and for each of the plurality of groups, generating a first caption that describes the object within the image range corresponding to the group based on at least one of the words within the group.

[0015] In a third aspect of the present invention, a program is provided. The program causes a computer to execute procedures including: acquiring an image captured within a plant; extracting a plurality of features from the image; specifying an image region corresponding to each of the plurality of features; grouping the plurality of features based on the image regions; extracting a plurality of words representing features of an object imaged in the image; and for each of the plurality of groups, generating a first caption for explaining the object within an image range corresponding to the group based on at least one of the words within the group.

[0016] Note that the above summary of the invention does not enumerate all the features of the present invention. Also, sub-combinations of these feature groups can also be inventions.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Embodiments for Carrying Out the Invention

[0018] Hereinafter, the present invention will be described through embodiments of the invention. However, the following embodiments do not limit the invention claimed in the claims. Also, not all combinations of features described in the embodiments are essential for the solution means of the invention.

[0019] FIG. 1 is a schematic overview of a caption generation system 5 including a caption generation device 100 according to the first embodiment. The caption generation system 5 generates a caption that describes the object 20 imaged in the image 10 from the image 10 of the object 20 in the facility within the plant. More specifically, the caption generation system 5 generates a caption that describes the object 20 imaged in the image 10 for each image range obtained by dividing the image 10 into a plurality of parts. The caption generation system 5 further combines the captions for each such image range to generate a caption that describes the object 20 within the image 10. The image 10 may be one or more still images or a moving image. Note that the caption that describes the object 20 within the image 10 may refer to a descriptive text that explains the state of the object 20 shown in the image 10.

[0020] Examples of plants include industrial plants such as chemical plants, plants that manage and control wells and their surroundings in gas fields, oil fields, etc., plants that manage and control power generation such as hydraulic, thermal, and nuclear power, plants that manage and control environmental power generation such as solar power and wind power, plants that manage and control water supply and sewerage, dams, etc.

[0021] FIG. 1 shows, as an example of the object 20 in the facilities within the plant, a pipe 21 and a tank 22. As shown in FIG. 1, the pipe 21 may be connected to the tank 22 on the upstream side, for example, and may flow the liquid or gas stored in the tank 22 to the downstream side. Note that in FIG. 1, illustration of the equipment on the downstream side of the pipe 21 is omitted.

[0022] The caption generation system 5 includes, as an example, an imaging device terminal 41 used by an imager 31, an administrator terminal 42 used by an administrator 32, a worker terminal 43 used by a recovery worker 33, and a caption generation device 100. The imager terminal 41, the administrator terminal 42, the worker terminal 43, and the caption generation device 100 can communicate with each other via a communication network 50. The communication network 50 may be a wired network, a wireless network, or may include both of these.

[0023] The imager terminal 41 captures an image 10 of the object 20. Instead of or in addition to this, an imaging device such as a monitoring camera installed in the plant, a traveling robot or a drone that patrols within the plant, etc. may capture the image 10 of the object 20. The imager terminal 41 may be a smartphone with a camera. The imager terminal 41 may be a tablet terminal. The imager terminal 41 may be a PC (Personal Computer). The imager terminal 41 may be a wearable terminal.

[0024] The imager 31 is an example of a user who uses the caption generation system 5 and carries and uses the imager terminal 41. The imager 31 uses the imager terminal 41 to capture an image 10 of the object 20 in the facilities within the plant. The imager 31 is, as an example, a maintenance worker of the equipment within the plant, and more specifically, a maintenance worker who maintains the tank 22. Note that there may be one or a plurality of imagers 31 within the plant.

[0025] The administrator terminal 42 displays a report using the caption generated by the caption generation system 5. The administrator terminal 42 transmits an instruction regarding the target 20 in the facilities within the plant to the worker terminal 43 via the communication network 50. The administrator terminal 42 may be a PC and may be arranged at a location away from the target 20 in the facilities within the plant, such as in the control room within the plant. Note that the administrator terminal 42 may be a smartphone. The administrator terminal 42 may be a tablet terminal. The administrator terminal 42 may be a wearable terminal.

[0026] The administrator 32 is an example of a user who utilizes the caption generation system 5 and uses the administrator terminal 42. The administrator 32 browses a report using the caption generated by the caption generation system 5 via the administrator terminal 42. As an example, the administrator 32 browses the report in a situation where the target 20 described by the caption cannot be directly visually confirmed. The administrator 32 judges the situation of the target 20 in the facilities within the plant and orders the recovery worker 33 or the like to perform recovery work on the target 20. As an example, the administrator 32 judges the situation of the pipe 21 in the facilities within the plant and orders the recovery worker 33 or the like to perform recovery work on the pipe 21. Note that there may be one or multiple administrators 32 within the plant. Note that the administrator 32 is an example of a viewer who browses a report using the caption generated by the caption generation system 5.

[0027] The worker terminal 43 displays an instruction regarding the target 20 in the facilities within the plant, a report using the caption generated by the caption generation system 5, and the like. The worker terminal 43 may be a smartphone. The worker terminal 43 may be a tablet terminal. The worker terminal 43 may be a PC. The worker terminal 43 may be a wearable terminal.

[0028] The restoration worker 33 is an example of a user who uses the caption generation system 5 and uses the worker terminal 43. The restoration worker 33 browses, via the worker terminal 43, commands regarding the target 20 in the facilities within the plant, reports using the captions generated by the caption generation system 5, etc. The restoration worker 33, as an example, browses the report in a situation where the target 20 described by the caption cannot be directly visually confirmed. The restoration worker 33 deals with abnormalities within the plant. The restoration worker 33, as an example, deals with an abnormality in the pipe 21. The restoration worker 33, as an example, restores the pipe 21 in which the abnormality has occurred according to the commands and reports from the administrator 32 displayed on the worker terminal 43. Note that there may be one or a plurality of restoration workers 33 within the plant. Note that the restoration worker 33 is an example of a viewer who browses a report using the caption generated by the caption generation system 5.

[0029] The caption generation device 100, as an example, receives the image 10 obtained by imaging the target 20 in the facilities within the plant from the imager terminal 41 via the communication network 50. The caption generation device 100 generates captions that explain the target 20 imaged in the image 10 received from the imager terminal 41 for each image range obtained by dividing the image 10 into a plurality of parts. The caption generation device 100 further generates a caption that explains the target 20 in the image 10 by combining the captions for each such image range.

[0030] The caption generation device 100 transmits the generated captions to the imager terminal 41 via the communication network 50. The caption generation device 100 may be arranged in a control room, an instrument room, etc. within the plant, or may be arranged outside the plant. Some of the functions of the caption generation device 100 may be incorporated into the imager terminal 41.

[0031] Figure 2 is an example of an image 10 of an object 20 captured within a plant. In the image 10, as an example, a pipe 21 and a tank 22, which are examples of the object 20, are captured. More specifically, in the image 10, a straight pipe 11, a curved pipe 12, and a curved pipe 13 are captured, and the straight pipe 11 etc. that are connected to each other constitute the pipe 21. Further, in the image 10, a valve 14 provided on the pipe 21 is captured. Further, in the image 10, a tank body 15, a tank ladder 16, and a tank support 17 are captured, and the tank body 15 etc. that are connected to each other constitute the tank 22. Further, in the image 10, a gas 18 and a wall 19 are captured. The straight pipe 11, the curved pipe 12, the curved pipe 13, the valve 14, the tank body 15, the tank ladder 16, the tank support 17, the gas 18, and the wall 19 captured in the image 10 are examples of a plurality of features included in the image 10. The features included in the image 10 may refer to the objects captured in the image 10.

[0032] Figure 3 is a block diagram of a caption generation device 100 according to the first embodiment. In Figure 3, the flow of data etc. is indicated by arrows. The caption generation device 100 includes an image acquisition unit 110, a feature extraction unit 112, a region specification unit 114, a grouping unit 116, a word extraction unit 120, a caption generation unit 140, a storage unit 150, and a report accumulation unit 155.

[0033] The image acquisition unit 110 acquires the image 10 captured within the plant. The image acquisition unit 110, as an example, receives, from an imaging device terminal 41 via a communication network 50, the user ID of an imager 31 who uses the imaging device terminal 41, together with the image 10 of the object 20 within the plant. The image acquisition unit 110 inputs the received image 10 to the feature extraction unit 112. The image acquisition unit 110 also inputs the received image 10 to the word extraction unit 120 together with the user ID of the imager 31.

[0034] The feature extraction unit 112 extracts a plurality of features from the image 10. As an example, the feature extraction unit 112 extracts a plurality of features such as the linear pipe 11 from the image 10 shown in FIG. 2. The feature extraction unit 112 may extract a plurality of features based on the newly input image 10 using a feature extraction model that extracts the plurality of features from the image 10. The feature extraction unit 112 may, for example, read out the feature extraction model stored in the storage unit 150. The feature extraction unit 112 inputs the plurality of features extracted from the image 10 to the region identification unit 114, for example, in vector representation.

[0035] The region identification unit 114 identifies the image regions corresponding to each of the plurality of features extracted by the feature extraction unit 112 in the image 10. In other words, the region identification unit 114 identifies the image region for each feature in the image 10. Further in other words, the region identification unit 114 divides the image 10 into a plurality of image regions based on the plurality of features extracted from the image 10. As an example, the region identification unit 114 identifies the region surrounding each feature identified in the image 10 as the image region of each feature. The region identification unit 114 may identify, for example, a region such as a circle, an ellipse, or a rectangle surrounding each feature identified in the image 10 as the image region of each feature. The region identification unit 114 may identify the region surrounded by, for example, the contour of each feature, which is identified by identifying each feature pixel by pixel by segmentation such as semantic segmentation, instance segmentation, or panoptic segmentation, as the image region of each feature. The region identification unit 114 inputs the plurality of image regions identified in the image 10 to the grouping unit 116, for example, in vector representation.

[0036] The grouping unit 116 groups a plurality of features extracted by the feature extraction unit 112 based on the image regions specified by the region specifying unit 114. As an example, when two or more image regions in the image 10 satisfy a predetermined relationship, the grouping unit 116 may regard the plurality of features included in the two or more image regions as one group. In the image 10, the grouping unit 116 inputs a plurality of image ranges corresponding to a plurality of groups, each of which includes one or more features, to the caption generation unit 140 as, for example, a vector representation.

[0037] The word extraction unit 120 extracts a plurality of words representing the features of the object 20 imaged in the image 10. The word extraction unit 120 may extract one or more words based on the newly input image 10 by using a word extraction model that has learned the relationship between the image 10 and one or more words representing the features of the object 20 imaged in the image 10. For example, the word extraction unit 120 may read out the word extraction model stored in the storage unit 150. The word extraction unit 120 inputs the plurality of words extracted from the image 10 to the caption generation unit 140 as, for example, a vector representation, together with the user ID of the imager 31 or the like.

[0038] For each of the plurality of groups grouped by the grouping unit 116, the caption generation unit 140 generates a first caption that describes the features of the object 20 within the image range corresponding to the group based on at least one word within the group among the plurality of words extracted by the word extraction unit 120. The caption generation unit 140 may generate the first caption based on at least one newly input word by using a caption generation model that has learned the relationship between at least one word representing the features of the object 20 and the first caption that describes the features of the object 20. For example, the caption generation unit 140 may read out the caption generation model stored in the storage unit 150.

[0039] The caption generation unit 140 combines the first captions for each of the above-described image ranges to generate a second caption that describes the object 20 in the image 10. Specifically, the caption generation unit 140 may generate a plurality of second captions by combining the plurality of first captions of a plurality of adjacent image ranges. The caption generation unit 140 uses a caption generation model that has learned the relationship between the combination of the plurality of first captions and the second caption based on one or more relationships between the words representing the features of the object 20 in one image range and the words representing the features of the object 20 in another image range, and generates a second caption from the newly generated plurality of first captions. The caption generation unit 140 transmits the second caption as text data to the imaging device terminal 41 via the communication network 50 based on the user ID of the imager 31.

[0040] The storage unit 150 stores the above-described caption generation model. The storage unit 150 may also store the above-described feature extraction model and word extraction model. Note that the storage unit 150 is an example of a model storage unit.

[0041] The report storage unit 155 receives a report from the imaging device terminal 41 via the communication network 50 and registers it in the memory. The report may be created by the imager 31 appropriately editing the second caption using the imaging device terminal 41. The report storage unit 155 is accessed from the administrator terminal 42, the worker terminal 43, etc. via the communication network 50, and the report stored in the memory is read out.

[0042] FIG. 4 is a flowchart showing the flow of the caption generation method according to the first embodiment. As an example, the flow is started when the imager 31 transmits an image 10 of the object 20 in the plant captured by the imaging device terminal 41 to the caption generation device 100 via the communication network 50.

[0043] The caption generation device 100 acquires the image 10 captured in the plant (step S101). Specifically, the caption generation device 100 receives the user ID of the imager 31 together with the image 10 obtained by imaging the target 20 in the plant from the imager terminal 41 via the communication network 50.

[0044] The caption generation device 100 extracts a plurality of features from the image 10 (step S102). Specifically, the caption generation device 100 extracts a plurality of features from the image 10 using the above-described feature extraction model.

[0045] The feature extraction model learns the relationship between the image 10 and a plurality of image feature amounts, and when a new image 10 is input, extracts a plurality of image feature amounts from the image 10 and outputs each image feature amount as a vector representation. The feature extraction model may be, for example, a deep learning model having a neural network structure including a CNN (Convolution Neural Network) and an FC layer (Fully Connected Layer, fully connected layer).

[0046] The caption generation device 100 identifies the image regions corresponding to each of the plurality of features (step S103). Specifically, the caption generation device 100 identifies circular or elliptical image regions surrounding each feature identified in the image 10.

[0047] The caption generation device 100 groups the plurality of features based on the image regions (step S104). Specifically, when two or more image regions in the image 10 satisfy a predetermined relationship, the caption generation device 100 groups the plurality of features included in the two or more image regions into one group.

[0048] More specifically, the grouping unit 116 of the caption generation device 100 defines the position on the rectangular image 10 with XY coordinates, and when at least one of the ranges of the X coordinates and Y coordinates of a plurality of circular or elliptical image regions coincides with a predetermined threshold or more, a plurality of features within the image range obtained by combining the plurality of image regions may be regarded as one group. The grouping unit 116 has a size in which at least one of the widths of the X coordinates and Y coordinates of a specific circular or elliptical image region is equal to or greater than a predetermined threshold, and when there are other circular or elliptical image regions adjacent to the image region in the coordinate axis direction satisfying the condition, a plurality of features included in the specific image region and the adjacent image regions may be regarded as one group.

[0049] The caption generation device 100 extracts a plurality of words representing the features of the object 20 imaged in the image 10 (step S105). Specifically, the caption generation device 100 uses a word extraction model to extract a plurality of words representing the features of the object 20 imaged in the image 10.

[0050] The word extraction model learns the relationship between the image 10 and a plurality of image feature amounts, and also learns the relationship between the plurality of image feature amounts and a plurality of words. When a new image 10 is input, the word extraction model extracts a plurality of image feature amounts from the image 10, and outputs, as a vector representation, a plurality of words representing the features of the object 20 imaged in the image 10 from the plurality of image feature amounts. The word extraction model may be, for example, a deep learning model having a neural network structure including a CNN and an FC layer. Instead of the CNN, the word extraction model may use Transformer's Self-Attention or Multi-Head Self-Attention.

[0051] The word extraction model may assign a likelihood to each of the plurality of words extracted from the image 10. For example, when the reference range of the output value for each word is 0 to 1, when an image 10 is input and the output value for a certain word is 0.8, the probability of the word may be treated as 80%.

[0052] The caption generation device 100 generates a first caption (step S106) that describes the features of the object 20 within the image range corresponding to a group, based on at least one word within the group among the plurality of extracted words for each of the plurality of groups. Specifically, the caption generation device 100 uses a caption generation model to generate, as text data, a first caption that describes the features of the object 20 within the image range corresponding to each group, based on at least one word within each group. The caption generation device 100 may generate a first caption that preferentially uses words with higher likelihood among the plurality of words using the caption generation model.

[0053] The caption generation device 100 combines the first captions for each image range to generate a second caption (step S107) that describes the object 20 within the image 10. Specifically, the caption generation device 100 uses a caption generation model to generate, as text data, a plurality of second captions by combining the plurality of first captions of a plurality of adjacent image ranges.

[0054] As a more specific example, the caption generation unit 140 may generate a second caption that includes both a word representing the features of the fluid-related object 20 within one image range corresponding to one first caption and a word representing a liquid or gas within another image range adjacent to the one image range corresponding to another first caption, from one or more relationships between them. The fluid-related object 20 refers to a member through which a fluid flows inside, a member that stores a fluid, etc., such as a pipe 21 or a tank 22. The caption generation device 100 transmits the generated text data of the second caption to the imaging device terminal 41 via the communication network 50, and thus the flow ends.

[0055] The caption generation model has learned the relationship between the vector representation of words representing the features of the target 20 and the first caption explaining the features of the target 20. When at least one vector representation of a word representing the features of the target 20 captured in the image 10 is newly input, the caption generation model outputs, from the at least one vector representation of a word, the text data of the first caption explaining the features of the target 20 within the image range corresponding to the group containing the at least one word.

[0056] The caption generation model learns the relationship between a plurality of first captions and the second caption, and when a plurality of newly generated first captions are input, outputs the second caption explaining the target 20 within the image 10 as text data.

[0057] The caption generation model may be a deep learning model with a neural network structure including, for example, an RNN (Recurrent Neural Network). Instead of the RNN, the caption generation model may use LSTM (Long Short-Term Memory), Transformer, GPT-3 (Generative Pre-trained Transformer-3), GPT-2, GPT-3-Clone, etc.

[0058] The caption generation model may learn the relationship between a combination of multiple first captions and a second caption based on annotations of one or more relationships between words representing the features of the object 20 within one image range and words representing the features of the object 20 within another image range. The annotation may mean, for example, as a relationship between the word linear pipe representing the features of the pipe 21 within one image range and the words gas and ejection representing the features related to the pipe 21 within another image range, that the relationship "the pipe is ejecting gas" is inappropriate and the relationship "gas is ejecting from the pipe" is appropriate, and providing this as teacher data. The caption generation model may use a large number of example sentences as learning data and, for each example sentence, learn an RNN or the like to increase the likelihood of outputting that example sentence. The caption generation model may extract the words included in the example sentence, use the vectors of the extracted words as learning input data, and use the example sentence as learning output data, and learn to increase the probability of outputting the learning output data for the learning input data.

[0059] The imaging person 31 checks the second caption displayed on the imaging person terminal 41 and transmits a report using the second caption from the imaging person terminal 41 to the caption generation device 100 via the communication network 50. The imaging person 31 may appropriately edit the second caption to create a report. When the caption generation device 100 receives a report from the imaging person terminal 41, it stores the report in the report storage unit 155. The report stored in the report storage unit 155 of the caption generation device 100 is viewed, for example, by the pipe repair worker 33 of the pipe 21, and the pipe repair worker 33 may repair the crack of the pipe 21 or adjust the valve of the pipe 21. The report stored in the report storage unit 155 of the caption generation device 100 is viewed, for example, by the administrator 32 of the pipe 21, and the administrator 32 may order the pipe repair worker 33 of the pipe 21 to perform repair work or adjust the control parameters of the equipment where the pipe 21 is installed.

[0060] FIG. 5 is a diagram for explaining a method of specifying a plurality of image regions in an example of an image 10 of an object 20 captured in a plant. In FIG. 5, image regions 51 to 59 of circles or ellipses surrounding a plurality of features in the image 10 are indicated by broken lines.

[0061] As shown in FIG. 5, the caption generation device 100 may extract each feature such as the linear pipe 11 on the image 10. Further, as shown in FIG. 5, the caption generation device 100 may specify image regions 51 to 59 of circles or ellipses surrounding the linear pipe 11 or the like on the image 10 as image regions corresponding to each of the plurality of features. In the example shown in FIG. 5, the elliptical image region 51 corresponds to the linear pipe 11, the elliptical image region 52 corresponds to the curved pipe 12, the elliptical image region 53 corresponds to the curved pipe 13, and the elliptical image region 54 corresponds to the valve 14. The elliptical image region 55 corresponds to the tank body 15, the elliptical image region 56 corresponds to the tank ladder 16, and the elliptical image region 57 corresponds to the tank support 17. The elliptical image region 58 corresponds to the gas 18, and the elliptical image region 59 corresponds to the wall 19.

[0062] FIG. 6 is a diagram for explaining a method of grouping a plurality of features based on image regions in an example of an image 10 of an object 20 captured in a plant. In FIG. 6, overlapping with FIG. 5, polygonal image ranges 61 to 65 surrounding one or more of the circular or elliptical image regions 51 to 59 are indicated by solid lines.

[0063] As shown in FIG. 6, when the caption generation device 100 determines that the elliptical image regions 51 to 54 surrounding the linear pipe 11, the curved pipe 12, the curved pipe 13, and the valve 14 on the image 10 satisfy a predetermined relationship, the caption generation device 100 may set the combination of the linear pipe 11, the curved pipe 12, the curved pipe 13, and the valve 14 as a first group. As shown in FIG. 6, the caption generation device 100 may specify a polygonal image range 61 surrounding the elliptical image regions 51 to 54 corresponding to the combination included in the first group as an image range corresponding to the first group. In this case, within the image range 61, the feature of the pipe 21 is included as a feature of the object 20.

[0064] Here, as a plurality of words representing the features of the object 20 imaged in the image 10, the caption generation device 100 may extract words such as, for example, "straight pipe / curved pipe / valve / crack / tank body / tank ladder / tank support / gas / jet". In this case, the caption generation device 100 may identify the words "straight pipe / curved pipe / valve / crack" as at least one word within the first group among the plurality of words extracted from the image 10.

[0065] In this case, for the first group, the caption generation device 100 may generate one or a plurality of first captions that describe the features of the pipe 21 within the image range 61, such as, for example, "a curved pipe is connected to the left end of the straight pipe", "a valve is provided in the curved pipe", "a curved pipe is connected to the right end of the straight pipe", "a crack has occurred in the straight pipe".

[0066] Similarly, as shown in FIG. 6, the caption generation device 100 may determine that the elliptical image regions 55 to 56 surrounding the tank body 15 and the tank ladder 16 on the image 10 satisfy a predetermined relationship, and may set the combination of the tank body 15 and the tank ladder 16 as the second group. As shown in FIG. 6, the caption generation device 100 may identify a polygonal image range 62 surrounding the elliptical image regions 55 to 56 corresponding to the combination included in the second group as the image range corresponding to the second group. In this case, the features of the tank 22 are included as the features of the object 20 within the image range 62.

[0067] The caption generation device 100 may identify, as at least one word in the second group, the words "tank body / tank ladder" among the plurality of words "linear pipe / curved pipe / valve / crack / tank body / tank ladder / tank support / gas / jet" extracted from the image 10. In this case, for the second group, the caption generation device 100 may generate one or more first captions that describe the features of the tank 22 within the image range 62, such as "a tank ladder is provided on the left side of the tank body", based on these words.

[0068] Similarly, as shown in FIG. 6, the caption generation device 100 may determine that there is no other image area that satisfies a predetermined relationship with the elliptical image area 57 surrounding the tank support 17 on the image 10, and may use only the tank support 17 as the third group. The caption generation device 100 may specify, as the image range corresponding to the third group, a polygonal image range 63 that surrounds the elliptical image area 57 corresponding to the tank support 17 included in the third group. In this case, the features of the tank 22 are included as the features of the object 20 within the image range 63.

[0069] The caption generation device 100 may identify, as at least one word in the third group, the words "tank body / tank support" among the plurality of words "linear pipe / curved pipe / valve / crack / tank body / tank ladder / tank support / gas / jet" extracted from the image 10. In this case, for the third group, the caption generation device 100 may generate one or more first captions that describe the features of the tank 22 within the image range 63, such as "a tank body is provided on the tank support", based on these words.

[0070] Similarly, as shown in FIG. 6, the caption generation device 100 may determine that there is no other image area that satisfies a predetermined relationship with the elliptical image area 58 surrounding the gas 18 on the image 10, and may use only the gas 18 as the fourth group. As shown in FIG. 6, the caption generation device 100 may specify a polygonal image range 64 that surrounds the elliptical image area 58 corresponding to the gas 18 included in the fourth group, as the image range corresponding to the fourth group. In this case, within the image range 64, as features of the object 20, features of the pipe 21 or the tank 22 that is the gas ejection source are included.

[0071] The caption generation device 100 may specify the words "gas / ejection" as at least one of the plurality of words "linear pipe / curved pipe / valve / crack / tank body / tank ladder / tank support / gas / ejection" extracted from the image 10 as being within the fourth group. In this case, for the fourth group, the caption generation device 100 may generate one or more first captions that describe features related to the pipe 21 or the tank 22 within the image range 64, such as "gas is ejecting", based on these words.

[0072] Similarly, as shown in FIG. 6, the caption generation device 100 may determine that there is no other image area that satisfies a predetermined relationship with the elliptical image area 59 surrounding the wall 19 on the image 10, and may use only the wall 19 as the fifth group. As shown in FIG. 6, the caption generation device 100 may specify a polygonal image range 65 that surrounds the elliptical image area 59 corresponding to the wall 19 included in the fifth group, as the image range corresponding to the fifth group. In this case, no features of the object 20 are included within the image range 65. Therefore, the caption generation device 100 may determine that there are no words representing the features of the object 20 extracted from the image 10 within the fifth group, and may not generate a first caption for the fifth group.

[0073] According to the caption generation device 100 according to the above first embodiment, a first caption is generated that describes the target 20 in the plant imaged in the image 10 for each image range obtained by dividing the image 10 into a plurality of parts. More specifically, as in the example described with reference to FIGS. 5 and 6, the caption generation device 100 may generate a first caption that describes the pipe 21 and the tank 22 imaged in the image 10 for each of the four image ranges 61 to 64.

[0074] As a comparative example with the caption generation device 100 according to the first embodiment, assume a device that generates one or more captions for the target 20 in the plant imaged in the image 10 without dividing the image 10 into a plurality of parts, but as a whole image 10. According to the device of the comparative example, for example, based on a plurality of words such as "linear pipe / curved pipe / valve / crack / tank body / tank ladder / tank support / gas / jet" extracted from the image 10, captions such as "gas is connected to the tank ladder", "a valve is jetting from the curved pipe", "a crack has occurred in the gas", etc., which are inappropriate and ambiguous for describing the target 20 imaged in the image 10, may be generated.

[0075] On the other hand, according to the caption generation device 100 according to the first embodiment, for the target 20 imaged in the image 10, a first caption is generated that describes the characteristics of the target 20 for each image range obtained by dividing the image 10 into a plurality of parts, so that it is possible to facilitate the judgment of the reader of the caption.

[0076] The caption generation device 100 according to the first embodiment may further generate a second caption that describes the target 20 in the image 10 by combining the captions for each such image range. More specifically, the caption generation device 100 may generate a plurality of second captions by combining a plurality of first captions for a plurality of adjacent image ranges.

[0077] In the case of the example described with reference to FIGS. 5 and 6, the caption generation device 100 may combine the first captions of the image ranges 61 and 64 adjacent to each other to generate a second caption such as "A crack has occurred in the straight pipe and gas is jetting out". The caption generation device 100 may combine the first captions of the image ranges 61 and 62 adjacent to each other to generate a second caption such as "Above the straight pipe, there is a tank body provided with a ladder for the tank". The caption generation device 100 may combine the first captions of the image ranges 62 and 64 adjacent to each other to generate a second caption such as "A ladder for the tank is provided on the left side of the tank body, and gas is jetting out on the left side of the ladder for the tank". The caption generation device 100 may combine the first captions of the image ranges 61 and 63 adjacent to each other to generate a second caption such as "The tank body is provided on the tank support, and there is a curved pipe in front of the tank body". Thereby, the caption generation device 100 can further facilitate the judgment of the caption reader.

[0078] In the caption generation device 100 according to the first embodiment described above, when generating a plurality of captions, the caption generation unit 140 may assign a negative weight to the words used when generating the first caption among the plurality of words extracted by the word extraction unit 120, and then generate the second caption so that the words used in the first caption are not included in the second caption. As another example when generating a plurality of captions, the caption generation unit 140 may assign arbitrary priorities to the plurality of words extracted by the word extraction unit 120, and then reduce the priority of the word by a predetermined amount each time the word is used in caption generation, so as to use the plurality of words extracted by the word extraction unit 120 as a whole.

[0079] In the caption generation device 100 according to the first embodiment, the storage unit 150 may store a plurality of sentences included in at least any one of an accident case collection, an accident response manual, and a maintenance history in the plant as example sentences. In this case, the caption generation unit 140 may extract at least one example sentence by searching among the plurality of example sentences stored in the storage unit 150 using at least any one of the plurality of words extracted by the word extraction unit 120.

[0080] For each of the plurality of groups, the caption generation unit 140 may generate a first caption based on at least one word within the group among the plurality of words extracted by the word extraction unit 120 and the extracted example sentence. Specifically, the caption generation unit 140 may search for an example sentence in which at least any one of the at least one word is used among the plurality of example sentences stored in the storage unit 150, and generate a first caption while referring to the extracted example sentence.

[0081] The caption generation unit 140 may generate a second caption based on the newly generated plurality of first captions and the extracted example sentences. Specifically, the caption generation unit 140 may search for an example sentence in which at least any one of the plurality of first captions is used among the plurality of example sentences stored in the storage unit 150, and generate a second caption while referring to the extracted example sentence.

[0082] The caption generation unit 140 may generate a plurality of second captions and assign a similarity to the extracted example sentences. Specifically, the caption generation unit 140 generates a plurality of second captions, searches for example sentences in which at least any one of the plurality of second captions is used from among the plurality of example sentences stored in the storage unit 150, and assigns a similarity to each of the plurality of second captions with the extracted example sentences. The caption generation unit 140 may, for example, convert each of the plurality of second captions and the extracted example sentences into feature vectors indicating combinations of a plurality of words included therein, and calculate the inner product of the feature vectors as the similarity. The caption generation unit 140 may transmit the second caption with the highest similarity among the plurality of generated second captions to the imaging device terminal 41. Alternatively, the caption generation unit 140 may transmit to the imaging device terminal 41 with the similarity assigned to each of the plurality of generated second captions. In this case, the imaging person 31 can determine the appropriateness of the plurality of second captions while referring to the similarity.

[0083] The second caption generated by the caption generation unit 140 while referring to the extracted example sentences may include, in addition to the caption explaining the target 20 in the plant, a plurality of action options regarding actions to be taken by the user, or instructions to the user. As a result of generating the second caption with reference to the example sentences, the caption generation device 100 can present to the user action options or instructions incorporating abnormal handling methods that cannot be directly interpreted from the image 10.

[0084] The second caption generated by the caption generation unit 140 while referring to the extracted example sentences may include at least one of an instruction to image the imaged target 20 from another angle or another shooting angle, and an instruction to image another target 20 related to the imaged target 20, in addition to the caption explaining the target 20 in the plant.

[0085] FIG. 7 is a flowchart showing an example of a detailed flow of the caption generation method according to the first embodiment. The operation flow shown in the flowchart of FIG. 7 may be a specific example of the caption generation method according to the first embodiment described with reference to FIG. 4. The caption generation apparatus 100 according to the first embodiment described above may generate, as an example, a caption for explaining the object 20 from the image 10 obtained by imaging the object 20 according to the operation flow shown in the flowchart of FIG. 7.

[0086] The flow of FIG. 7 is started, as an example, when the image 10 of the object 20 in the plant captured by the imager 31 using the imager terminal 41 is transmitted to the caption generation apparatus 100 via the communication network 50.

[0087] The caption generation apparatus 100 acquires the image 10 captured in the plant (step S201). In step S201, as shown in FIG. 8 which specifically describes a part of the flow of FIG. 7, the image acquisition unit 110 may receive the image 10 of the object 20 in the plant from the imager terminal 41 via the communication network 50 and output it to an integrated component including the feature extraction unit 112, the region specification unit 114, the grouping unit 116, and the word extraction unit 120.

[0088] The caption generation device 100 crops the image 10 into a plurality of parts (step S202). For example, the caption generation device 100 creates a new image by discarding the peripheral area in the image 10 that does not contain the target 20, and divides it into patches of a fixed size. In step S202, as shown in FIG. 8, the above-described integrated component may create a new image by discarding the peripheral area in the image 10 that does not contain the pipe 21 and the tank 22, and divide the new image into nine patches by dividing it vertically and horizontally into three parts. Alternatively, the integrated component may create a new image by discarding the peripheral area in the image 10 that does not contain the pipe 21 as a patch of a fixed size, and create a new image by discarding the peripheral area in the image 10 that does not contain the tank 22 as another patch of a fixed size. Note that the learning model of the integrated component shown in FIG. 8 may also be referred to as a Vision Transformer and may be an example of the above-described word extraction model. The learning model of the integrated component may also include a model (not shown) referred to as U-Net. U-Net is one of the FCNs (Fully Convolution Network, fully convolutional network), and after extracting features with an encoder, it may return to the original resolution with a decoder, and in that case, information of the original high-resolution image may also be used with a skip connection.

[0089] The caption generation device 100 converts the cropped image through a simple function, for example, a LINEAR function, into a numerical value A that can be used by the encoder (step S203). In step S203, as shown in FIG. 8, an integrated component may put nine patches into a linear projection layer, flatten each of them, and embed them (linearly) into a vector (or convert them into a one-dimensional array). The integrated component may further add a CLS token to the beginning of the sequence of vectors output from the linear projection layer (extra learnable [class] embedding), embed positions into each patch (each vector), and use it as the numerical value A. That is, the integrated component may convert the nine patches divided from the image 10 into a sequence of 10-dimensional vectors that can be used by the encoder.

[0090] The caption generation device 100 causes the transformer encoder to convert the numerical value A into another value B and value C (step S204), and unifies the dimensions and numerical widths of the value B and value C (step S205). From step S204 to S205, as shown in FIG. 8, an integrated component puts a sequence of 10-dimensional vectors into two transformer encoders, and each transformer encoder converts this into another sequence of 10-dimensional vectors, and unifies the dimensions and numerical widths, for example, through a softmax function, and outputs it as the output of the CLS token. Each transformer encoder may include a plurality of components such as Self-Attention described above. The output of the CLS token may be a feature quantity necessary for classification aggregated from the entire image by Self-Attention.

[0091] The caption generation device 100 outputs image feature amounts and region information (step S206). In step S206, as shown in FIG. 8, an integrated component may input the output of the CLS token into a classification head (MLP: multi-layer perceptron) and output a plurality of image feature amounts and region information from the classification head. The classification head may output the plurality of image feature amounts and region information after compressing them. As shown in FIG. 8, the integrated component may output the plurality of image feature amounts and region information to the caption generation unit 140. As described with reference to FIGS. 1 to 3, a plurality of word vectors of words representing the features of the object 20 imaged in the image 10 may be extracted from the plurality of image feature amounts and output to the caption generation unit 140 together with the region information.

[0092] The caption generation device 100 converts the image feature amounts and region information into numerical values D and E suitable for the decoder (step S207) and inputs the numerical values D and E into the decoder (step S208). From step S207 to S208, as shown in FIG. 9, which specifically describes a part of the flow of FIG. 7, the caption generation unit 140 vectorizes the region information input from the integrated component through, for example, a LINEAR function, and also vectorizes an empty text, and vectorizes these vectors together with the plurality of image feature amounts input from the integrated component. Thereby, the caption generation unit 140 converts the plurality of image feature amounts and region information into numerical values D and E. The caption generation unit 140 further inputs the numerical values D and E into the decoder. Note that the learning model of the decoder is a model using layer-by-layer adjustment, may also be referred to as GPT-2, and may be an example of the caption generation model described above.

[0093] The caption generation device 100 converts the numerical values D and E into a numerical value F by a decoder (step S209), and converts it into text (token) through a function that returns the numerical value F to text (step S210). From step S209 to S210, as shown in FIG. 9, the decoder in the caption generation unit 140 may convert the numerical values D and E into a numerical value F and weight each numerical value F. The decoder may compress the numerical value F. The decoder may convert the numerical value F into text (token) through a softmax function.

[0094] The caption generation device 100 outputs text (caption) (step S211), and the flow ends. In step S211, as shown in FIG. 9, for example, the decoder uses a sequence of consecutive texts X1 to X5 as an input (source token), and a sequence of consecutive texts Y1 to Y5 that is the same sequence as the source token but shifted one token to the right as a target token, concatenates the source token and the target token, and may process them by adjusting for each layer. For the source token and the target token, a reconfigurable position for each of the source token and the target token may be embedded together with each corresponding token. The position of the source token starts from zero, and for the target token, instead of incrementing the position at the end of the source sentence, the position may be reset to zero again.

[0095] The decoder model shown in FIG. 9 includes, as an example, an attention mechanism including a self-attention mechanism and a mixed attention mechanism, an add and layer normalization (Add&Layer Normalization) layer, a feed-forward neural network (FNN) layer, and N layers stacked in this order of add and layer normalization layers. The decoder may put the processed data of N layers into a linearization layer, pass the output from the linearization layer through a softmax function, and output a caption consisting of a sequence of continuous texts of Y1 to Y5. The caption generation unit 140 may transmit the generated text data of the caption to the imaging device terminal 41 via the communication network 50.

[0096] FIG. 10 is a block diagram of a caption generation device 200 according to the second embodiment. The caption generation device 200 according to the second embodiment includes a learning unit 260 in addition to the configuration included in the caption generation device 100 according to the first embodiment. Other configurations included in the caption generation device 200 according to the second embodiment are the same as those of the caption generation device 100 according to the first embodiment, and redundant descriptions are omitted using the reference numerals of the respective configurations included in the caption generation device 100.

[0097] The learning unit 260 learns the above-described feature extraction model, word extraction model, and caption generation model using the result of the user's determination of the appropriateness of the second caption generated by the caption generation unit 140. In addition to or instead of this, the learning unit 260 may learn the above-described feature extraction model, word extraction model, and caption generation model using the user's correction input for the second caption generated by the caption generation unit 140. The learning unit 260 receives the above-described determination result or correction input by the user from the imaging device terminal 41 via the communication network 50.

[0098] The learning unit 260 updates the parameters of the model so as to reduce the error between the output of the model when each sample in the learning data including the result of the above determination is input to the model and the label. For example, when learning a feature extraction model or a word extraction model using a multi-layer neural network including a CNN or the like, the learning unit 260 uses the error between the output value output by the neural network in response to inputting each sample and the label, and adjusts the weights between the neurons of the multi-layer neural network and the biases of each neuron by a method such as backpropagation.

[0099] As described above, in the caption generation device 200 according to the second embodiment, information on whether the generated second caption is actually used by the user is fed back, and the feature extraction model, the word extraction model, and the caption generation model are learned using the feedback content. The caption generation device 200 may determine that the generated second caption is actually used, for example, when receiving from the administrator terminal 42 the result determined by the administrator 32 that there is no problem as a result of viewing the second caption. The caption generation device 200 may also determine that the generated second caption is actually used, for example, when the caption generation unit 140 performs a search for example sentences using the generated second caption or when the generated second caption is used in a report for the administrator 32. The caption generation device 200 according to such a second embodiment also has the same effect as the caption generation device 100 according to the first embodiment. According to the caption generation device 200 according to the second embodiment, the accuracy of generating a caption required by the user can also be improved.

[0100] Various embodiments of the present invention may be described with reference to flowcharts and block diagrams, where the blocks may represent (1) stages of a process in which an operation is performed or (2) sections of an apparatus having a role of performing an operation. Specific stages and sections may be implemented by a dedicated circuit, a programmable circuit supplied with computer-readable instructions stored on a computer-readable medium, and / or a processor supplied with computer-readable instructions stored on a computer-readable medium. The dedicated circuit may include digital and / or analog hardware circuits, and may include integrated circuits (ICs) and / or discrete circuits. The programmable circuit may include a reconfigurable hardware circuit including memory elements such as logical AND, logical OR, logical XOR, logical NAND, logical NOR, and other logical operations, flip-flops, registers, field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), etc.

[0101] A computer-readable medium may include any tangible device capable of storing instructions executable by an appropriate device, and as a result, a computer-readable medium having instructions stored therein will comprise a product including instructions executable to create means for performing the operations specified in the flowchart or block diagram. Examples of computer-readable media may include electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, etc. More specific examples of computer-readable media may include floppy (registered trademark) disks, diskettes, hard disks, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM or flash memory), electrically erasable programmable read only memory (EEPROM), static random access memory (SRAM), compact disc read only memory (CD-ROM), digital versatile disc (DVD), Blu-ray (registered trademark) disc, memory stick, integrated circuit card, etc.

[0102] Computer-readable instructions may include source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or object-oriented programming languages such as Smalltalk®, JAVA®, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages.

[0103] Computer-readable instructions may be provided locally or via a wide area network (WAN) such as a local area network (LAN), the Internet, etc. to a processor or programmable circuit of a programmable data processing apparatus such as a general-purpose computer, a special-purpose computer, or other computer, and the computer-readable instructions may be executed to create means for performing the operations specified in a flowchart or block diagram. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.

[0104] FIG. 11 shows an example of a computer 2200 in which multiple aspects of the present invention may be embodied, in whole or in part. Programs installed on the computer 2200 can cause the computer 2200 to function as an operation associated with the apparatus according to an embodiment of the present invention or as one or more sections of the apparatus, or can cause the operation or the one or more sections to be executed, and / or can cause the computer 2200 to execute a process according to an embodiment of the present invention or a stage of the process. Such a program may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.

[0105] The computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphic controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes an input / output unit such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.

[0106] The CPU 2212 operates according to programs stored in the ROM 2230 and the RAM 2214, thereby controlling each unit. The graphic controller 2216 acquires image data generated by the CPU 2212 in a frame buffer provided in the RAM 2214 or the like or in itself, and causes the image data to be displayed on the display device 2218.

[0107] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads a program or data from the DVD-ROM 2201 and provides the program or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to the IC card.

[0108] ROM 2230 stores therein a boot program or the like executed by the computer 2200 upon activation, and / or a program dependent on the hardware of the computer 2200. The input / output chip 2240 may also be connected to the input / output controller 2220 via various input / output units through a parallel port, a serial port, a keyboard port, a mouse port, or the like.

[0109] The program is provided by a computer-readable medium such as a DVD-ROM 2201 or an IC card. The program is read from the computer-readable medium, installed in the hard disk drive 2224, the RAM 2214, or the ROM 2230, which is also an example of a computer-readable medium, and executed by the CPU 2212. The information processing described in these programs is read by the computer 2200, resulting in the cooperation between the programs and the various types of hardware resources described above. The apparatus or method may be configured by realizing the operation or processing of information according to the use of the computer 2200.

[0110] For example, when communication is executed between the computer 2200 and an external device, the CPU 2212 may execute a communication program loaded in the RAM 2214 and instruct the communication interface 2222 to perform communication processing based on the processing described in the communication program. The communication interface 2222 reads the transmission data stored in the transmission buffer processing area provided in a recording medium such as the RAM 2214, the hard disk drive 2224, the DVD-ROM 2201, or the IC card under the control of the CPU 2212, transmits the read transmission data to the network, or writes the received data received from the network to the reception buffer processing area or the like provided on the recording medium.

[0111] Further, the CPU 2212 may cause all or necessary parts of files or databases stored in external recording media such as a hard disk drive 2224, a DVD-ROM drive 2226 (DVD-ROM 2201), and an IC card to be read into the RAM 2214, and may execute various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording media.

[0112] Various types of information such as various types of programs, data, tables, and databases may be stored in the recording media and may undergo information processing. The CPU 2212 may execute various types of processing on the data read from the RAM 2214, including various types of operations, information processing, conditional judgment, conditional branching, unconditional branching, information search / replacement, etc., described throughout this disclosure and specified by the instruction sequence of the program, and write back the results to the RAM 2214. Further, the CPU 2212 may search for information in files, databases, etc. within the recording media. For example, when a plurality of entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored in the recording media, the CPU 2212 searches for an entry that matches the condition where the attribute value of the first attribute is specified from among the plurality of entries, reads the attribute value of the second attribute stored in the entry, and thereby may obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.

[0113] The programs or software modules described above may be stored in a computer-readable medium on or near the computer 2200. Also, a recording medium such as a hard disk or RAM provided within a server system connected to a dedicated communication network or the Internet can be used as a computer-readable medium, thereby providing the program to the computer 2200 via the network.

[0114] As described above, the present invention has been described using embodiments. However, the technical scope of the present invention is not limited to the scope described in the above embodiments. It will be apparent to those skilled in the art that various changes or improvements can be made to the above embodiments. It is clear from the description of the claims that forms with such changes or improvements can also be included in the technical scope of the present invention.

[0115] For example, the control system may be a computer housed in a single housing. That is, the controller may be realized by program execution on a processor of the computer, and each input / output device may be implemented as an I / O device of the computer. Further, the controller may be implemented as a virtual machine executed by one or more processors. In such a configuration, the control system does not include a network that is a general-purpose or dedicated network, and the controller and the input / output devices can be connected by a chipset such as a memory controller hub and an I / O controller hub that connect between the processor and the I / O devices.

[0116] It should be noted that the execution order of each process such as operations, procedures, steps, and stages in the devices, systems, programs, and methods shown in the claims, the specification, and the drawings is not explicitly indicated as "before" or "preceding" etc. in particular, and can be realized in any order unless the output of the previous process is used in the subsequent process. Regarding the operation flows in the claims, the specification, and the drawings, even if "first," "next," etc. are used for convenience of explanation, it does not mean that it is essential to implement in this order.

Explanation of Reference Numerals

[0117] 5 Caption generation system 10 Image 11 Straight pipe 12 Curved pipe 13 Curved pipe 15 Tank body 16 Tank ladder 17 Tank support stand 18 Gas 19 Wall 20 Target 21 Pipe 22 Tank 31 Cameraman 32 Administrator 33 Restoration worker 41 Cameraman's terminal 42 Administrator's terminal 43 Worker's terminal 51, 52, 53, 55, 56, 57, 58, 59 Image area 61, 62, 63, 64, 65 Image range 100 Caption generation device 110 Image acquisition unit 112 Feature extraction unit 114 Region identification unit 116 Grouping unit 120 Word extraction unit 140 Caption generation unit 150 Memory unit 155 Report storage unit 200 Caption generation device 260 Learning unit 2200 Computer 2201 DVD-ROM 2210 Host controller 2212 CPU 2214 RAM 2216 Graphics controller 2218 Display device 2220 Input / output controller 2222 Communication interface 2224 Hard disk drive 2226 DVD-ROM drive 2230 ROM 2240 Input / output chip 2242 Keyboard

Claims

1. An image acquisition unit that acquires an image captured within a plant; A feature extraction unit that extracts a plurality of features from the image; A region identification unit that identifies an image region corresponding to each of the plurality of features; A grouping unit that groups the plurality of features based on the image regions; A word extraction unit that extracts a plurality of words representing the features of the object imaged in the image; A caption generation unit that generates a first caption for explaining the features of the object within the image range corresponding to the group, based on at least one of the words within the group, for each of the plurality of groups; A caption generation device comprising the above.

2. The caption generation unit generates a second caption for explaining the object within the image by combining the first captions for each of the image ranges. The caption generation device according to Claim 1.

3. The caption generation unit generates a plurality of the second captions by combining the first captions for a plurality of the image ranges adjacent to each other. The caption generation device according to Claim 2.

4. The caption generation device further comprises a model storage unit that stores a caption generation model that learns the relationship between at least one of the words representing the features of the object and the first caption for explaining the features of the object, and learns the relationship between the combination of the plurality of the first captions and the second caption by the annotation of one or more relationships between the words representing the features of the object within one of the image ranges and the words representing the features of the object within another of the image ranges. The caption generation unit uses the caption generation model to generate the first caption based on the newly input at least one word, and generates the second caption from the newly generated plurality of the first captions. The caption generation device according to Claim 2 or 3.

5. The caption generation device further comprises a learning unit that learns a feature extraction model that extracts the plurality of features from the image, a word extraction model that extracts the plurality of words from the image, and a caption generation model that generates the first caption from the plurality of words and generates the second caption from the generated plurality of the first captions, using the result of the user's determination of the suitability of the generated second caption. ​ The caption generation unit generates the first caption based on the at least one newly input word using the caption generation model, and generates the second caption based on the plurality of newly generated first captions. The caption generation device according to claim 2 or 3.

6. The feature extraction model that extracts the plurality of features from the image, the word extraction model that extracts the plurality of words from the image, and the learning unit that learns the caption generation model that generates the first caption from the plurality of words and generates the second caption from the plurality of generated first captions, using the user's correction input for the generated second caption. The caption generation unit generates the first caption based on the at least one newly input word using the caption generation model, and generates the second caption based on the plurality of newly generated first captions. The caption generation device according to claim 2 or 3.

7. The caption generation unit generates the second caption including both the word representing the feature of the fluid-related object within one image range corresponding to one of the first captions and the word representing the liquid or gas within another image range adjacent to the one image range corresponding to another of the first captions, from one or more relationships between them. The caption generation device according to claim 2 or 3.

8. The apparatus further includes a storage unit that stores a plurality of sentences included in at least any one of the accident case collection, accident response manual, and maintenance history in the plant as example sentences. The caption generation unit extracts at least one example sentence by searching using at least any one of the plurality of extracted words from among the plurality of example sentences stored in the storage unit, generates the first caption based on the at least one word and the extracted example sentence, and generates the second caption based on the plurality of newly generated first captions and the extracted example sentence. The caption generation device according to claim 2 or 3.

9. The caption generation unit generates a plurality of the second captions and assigns a similarity to the extracted example sentence. The caption generation device according to claim 8.

10. The second caption includes a plurality of action options regarding actions to be taken by the user, or instructions for the user. The caption generation device according to claim 8.

11. The second caption includes at least one of an instruction to image the captured object from a different angle or a different shooting angle, and an instruction to image another object related to the captured object. The caption generation device according to claim 8.

12. Obtaining an image captured within a plant; Extracting a plurality of features from the image; Identifying an image region corresponding to each of the plurality of features; Grouping the plurality of features based on the image regions; Extracting a plurality of words representing features of the object imaged in the image; Generating, for each of the plurality of groups, a first caption that describes the features of the object within the image range corresponding to the group based on at least one of the words within the group. A caption generation method comprising the steps of:

13. Causing a computer to perform a procedure for obtaining an image captured within a plant; perform a procedure for extracting a plurality of features from the image; perform a procedure for identifying an image region corresponding to each of the plurality of features; perform a procedure for grouping the plurality of features based on the image regions; perform a procedure for extracting a plurality of words representing features of the object imaged in the image; perform a procedure for generating, for each of the plurality of groups, a first caption that describes the features of the object within the image range corresponding to the group based on at least one of the words within the group. A program for causing the computer to execute the above procedures.

Citation Information

Patent Citations

  • Image processing apparatus, image processing method, and program

    JP2021071808A

  • Photographing method and terminal device

    JP2021520163A

  • Oil spill monitoring system and oil spill monitoring method

    JP2022112341A

  • Estimation model generation program, estimation model generation method and information processing device

    JP2023043341A

  • Automatically evaluating caption quality of rich media using context learning

    US20210064879A1