Caption generation apparatus, caption generation method and program
The caption generation device improves caption accuracy by prioritizing less frequent words in plant environments, effectively highlighting abnormal conditions through a deep learning model, reducing irrelevant captions and enhancing clarity in abnormality notification.
Patent Information
- Application Number
- JP2023214731
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-07-02
- Estimated Expiration
- 2043-12-20
AI Technical Summary
Existing caption generation systems for plant environments struggle to accurately describe abnormal conditions due to the prevalence of common words masking critical abnormality indicators, leading to a high likelihood of irrelevant captions being generated.
A caption generation device that assigns higher weights to less frequent words extracted from past images of the same object, using a deep learning model to prioritize these words in generating captions, thereby emphasizing abnormality indicators.
Enhances the accuracy of captions in identifying abnormal conditions by preferentially using less frequent words associated with abnormalities, reducing the likelihood of irrelevant captions and improving the clarity of abnormality notification.
Smart Images

Figure 2025098533000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a caption generation device, a caption generation method, and a program.
Background Art
[0002] Patent Document 1 describes that "when the keyword extracted from the input data is a specific keyword not registered in the dictionary DB231 and the input data includes an image or voice, the category corresponding to the specific keyword is estimated based on the image or voice included in the input data and the case DB232." (paragraph 0049), "When the category is estimated by the category estimation unit 223, the DB update unit 222 registers the input data in the case DB232 in association with the category estimated by the category estimation unit 223. That is, the DB update unit 222 performs semantic annotation on the input data and then accumulates the input data." (paragraph 0051). Patent Document 2 describes that "obtain position information specified by GPS, display the current inspection location for plant staff, accumulate and analyze image data captured by an installed camera, and determine whether there are signs of abnormality in plant equipment from past equipment states and abnormal cases, etc., and at the same time, use a preset format / document to automatically create a report for regular inspection from the image data." (paragraph 0053). Patent Document 3 describes that "a feature amount acquisition unit that acquires information about the pipe as a first feature amount from an image obtained by photographing the plant equipment to be worked on and the pipes existing around the plant equipment with the camera, and a feature amount comparison unit that compares the first feature amount with a second feature amount about the pipe acquired from design data" (claim 1). Patent Document 4 describes that "the abnormal mail creation function 102 is activated when the abnormal monitoring function 101 detects a plant abnormality, and creates a mail transmission text by formulating matters to be communicated to the supervisor at an early stage, such as the date and time when the abnormality was detected, the name of the corresponding plant equipment, and the content of the abnormality." (paragraph 0019). [Prior Art Documents] [Patent Documents] [Patent Document 1] Japanese Patent Application Laid-Open No. 2020-102101 [Patent Document 2] Japanese Patent No. 6099989 [Patent Document 3] Japanese Patent No. 6826509 [Patent Document 4] Japanese Patent Application Laid-Open No. 2003-51895
Summary of the Invention
[0003] In a first aspect of the present invention, a caption generation device is provided. The caption generation device includes an image acquisition unit that acquires an image captured in a plant, a word extraction unit that extracts a plurality of words representing features of an object imaged in the image, a weighting unit that assigns a greater weight to words with a lower frequency extracted from past images of the same object as the imaged object among the plurality of words, and a caption generation unit that generates a caption explaining the imaged object based on the plurality of words and the respective weights of the plurality of words.
[0004] In the above caption generation device, the caption generation unit may generate the caption by preferentially using words with greater weights among the plurality of words.
[0005] In any of the above caption generation devices, the caption generation unit may assign a priority corresponding to the weight of at least one word used for caption generation among the plurality of words to the generated caption.
[0006] In any of the above caption generation devices, the word extraction unit may assign a likelihood to each of the plurality of words extracted from the image. In any of the above caption generation devices, the caption generation unit may generate the caption based on the plurality of words, the respective likelihoods and weights of the plurality of words.
[0007] In any of the above caption generation devices, the caption generation unit may generate the caption by preferentially using words among the plurality of words that have a higher likelihood and a larger weight.
[0008] In any of the above caption generation devices, the caption generation unit may generate the caption by preferentially using words among the plurality of words that have a larger integrated value of the likelihood and the weight.
[0009] In any of the above caption generation devices, the caption generation unit may generate the caption using words among the plurality of words that have a likelihood greater than a predetermined likelihood threshold and a weight greater than a predetermined weight threshold.
[0010] In any of the above caption generation devices, the caption generation unit may assign a priority according to the likelihood and the weight of at least one word used for caption generation among the plurality of words to the generated caption.
[0011] In any of the above caption generation devices, when the likelihood of a combination of a plurality of words used for caption generation is lower than a predetermined second likelihood threshold, the caption generation unit may output the caption including an instruction to re-capture the captured object.
[0012] Any of the above caption generation devices may further include a storage unit that stores a plurality of sentences included in at least any one of the accident case collection, accident response manual, and maintenance history in the plant as example sentences. In any of the above caption generation devices, the caption generation unit extracts at least one example sentence by searching using at least any one of the plurality of words extracted from among the plurality of example sentences stored in the storage unit, and generates the caption based on the plurality of words, the respective weights of the plurality of words, and the extracted example sentence.
[0013] In any of the above caption generation devices, the caption generation unit may generate a plurality of captions and assign a similarity to the extracted example sentences.
[0014] In any of the above caption generation devices, the caption may include a plurality of action options regarding actions to be taken by the user, or instructions to the user.
[0015] In any of the above caption generation devices, the caption may include at least one of an instruction to image the captured object from a different angle or a different shooting angle, and an instruction to image another object related to the captured object.
[0016] Any of the above caption generation devices may further include a model storage unit that stores a caption generation model that has learned the relationship between each of the weights of one or more words representing the characteristics of the object in the plant and the caption that describes the object in the plant. In any of the above caption generation devices, the caption generation unit may use the caption generation model to generate the caption based on the newly input plurality of words and each of the weights of the plurality of words.
[0017] Any of the above caption generation devices may further include a learning unit that learns a word extraction model that extracts the plurality of words from the image using the result of the user's determination of the appropriateness of the generated caption, and a caption generation model that generates the caption from each of the weights of the plurality of words. In any of the above caption generation devices, the caption generation unit may use the caption generation model to generate the caption based on the newly input plurality of words and each of the weights of the plurality of words.
[0018] Any of the above caption generation devices may further include a learning unit that learns a word extraction model that extracts the plurality of words from the image using a user's correction input for the generated caption, and a caption generation model that generates the caption from the weights of the respective ones of the plurality of words. In any of the above caption generation devices, the caption generation unit may generate the caption based on the newly input plurality of words and the weights of the respective ones of the plurality of words using the caption generation model.
[0019] In a second aspect of the present invention, a caption generation method is provided. The caption generation method includes acquiring an image captured within a plant, extracting a plurality of words representing features of an object imaged in the image, assigning a greater weight to a word having a lower frequency of extraction from past images that captured the same object as the imaged object among the plurality of words, and generating a caption that describes the imaged object based on the plurality of words and the weights of the respective ones of the plurality of words.
[0020] In a third aspect of the present invention, a program is provided. The program causes a computer to execute a procedure for acquiring an image captured within a plant, a procedure for extracting a plurality of words representing features of an object imaged in the image, a procedure for assigning a greater weight to a word having a lower frequency of extraction from past images that captured the same object as the imaged object among the plurality of words, and a procedure for generating a caption that describes the imaged object based on the plurality of words and the weights of the respective ones of the plurality of words.
[0021] Note that the above summary of the invention does not list all the features of the present invention. Also, sub-combinations of these feature groups can also be inventions.
Brief Description of the Drawings
[0022]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0023] Hereinafter, the present invention will be described through embodiments of the invention. However, the following embodiments do not limit the invention according to the claims. Also, not all combinations of features described in the embodiments are essential for the solution means of the invention.
[0024] FIG. 1 is a schematic overview of a caption generation system 5 including a caption generation apparatus 100 according to the first embodiment. The caption generation system 5 generates a caption that describes the object 20 imaged in the image 10 from the image 10 of the object 20 in the facility within the plant. More specifically, the caption generation system 5 assigns weights to a plurality of words extracted from the image 10 according to the frequencies extracted from past images of the object 20 imaged in the image 10, and generates a caption that describes the object 20 imaged in the image 10 from the plurality of weighted words. The image 10 may be one or more still images or a moving image. Note that the caption that describes the object 20 in the image 10 may refer to a descriptive text that describes the state of the object 20 shown in the image 10. The caption referred to here may refer to a sentence delimited by a period, or may refer to a clause delimited more finely than a sentence.
[0025] Examples of the plant include industrial plants such as chemical plants, plants that manage and control wells and their surroundings in gas fields, oil fields, etc., plants that manage and control power generation from hydropower, thermal power, nuclear power, etc., plants that manage and control environmental power generation from solar power, wind power, etc., plants that manage and control water supply and sewerage, dams, etc.
[0026] FIG. 1 shows a pipe 21 and a tank 22 as an example of the object 20 in the facility within the plant. As shown in FIG. 1, for example, the upstream side of the pipe 21 is connected to the tank 22, and the liquid or gas stored in the tank 22 may flow downstream. Note that in FIG. 1, the illustration of the equipment on the downstream side of the pipe 21 is omitted.
[0027] The caption generation system 5 includes, as an example, an imaging device terminal 41 used by an imager 31, an administrator terminal 42 used by an administrator 32, a worker terminal 43 used by a recovery worker 33, and a caption generation device 100. The imaging device terminal 41, the administrator terminal 42, the worker terminal 43, and the caption generation device 100 can communicate with each other via a communication network 50. The communication network 50 may be a wired network, a wireless network, or may include both of them.
[0028] The imaging device terminal 41 captures an image 10 of the target 20. Instead of or in addition to this, an imaging device such as a surveillance camera installed in the plant, a traveling robot or a drone equipped with a camera that patrols the plant, etc. may capture the image 10 of the target 20. The imaging device terminal 41 may be a smartphone with a camera. The imaging device terminal 41 may be a tablet terminal. The imaging device terminal 41 may be a PC (Personal Computer). The imaging device terminal 41 may be a wearable terminal.
[0029] The imager 31 is an example of a user who uses the caption generation system 5 and carries and uses the imaging device terminal 41. The imager 31 uses the imaging device terminal 41 to capture an image 10 of the target 20 in the facilities in the plant. The imager 31 is, as an example, a maintenance worker for the equipment in the plant, and more specifically, a maintenance worker who maintains the tank 22. Note that there may be one or more imagers 31 in the plant.
[0030] The administrator terminal 42 displays reports using the captions generated by the caption generation system 5. The administrator terminal 42 transmits instructions regarding the target 20 in the facilities within the plant to the worker terminal 43 via the communication network 50. The administrator terminal 42 may be a PC and may be arranged at a location away from the target 20 in the facilities within the plant, such as in a control room within the plant. Note that the administrator terminal 42 may be a smartphone. The administrator terminal 42 may be a tablet terminal. The administrator terminal 42 may be a wearable terminal.
[0031] The administrator 32 is an example of a user who uses the caption generation system 5 and uses the administrator terminal 42. The administrator 32 browses reports using the captions generated by the caption generation system 5 via the administrator terminal 42. As an example, the administrator 32 browses the report in a situation where the target 20 described by the caption cannot be directly visually confirmed. The administrator 32 judges the situation of the target 20 in the facilities within the plant and orders the recovery worker 33 or the like to perform recovery work on the target 20. As an example, the administrator 32 judges the situation of the pipe 21 in the facilities within the plant and orders the recovery worker 33 or the like to perform recovery work on the pipe 21. Note that there may be one or more administrators 32 in the plant. Note that the administrator 32 is an example of a viewer who browses reports using the captions generated by the caption generation system 5.
[0032] The worker terminal 43 displays instructions regarding the target 20 in the facilities within the plant, reports using the captions generated by the caption generation system 5, and the like. The worker terminal 43 may be a smartphone. The worker terminal 43 may be a tablet terminal. The worker terminal 43 may be a PC. The worker terminal 43 may be a wearable terminal.
[0033] The recovery worker 33 is an example of a user who uses the caption generation system 5 and uses the worker terminal 43. The recovery worker 33 browses, via the worker terminal 43, instructions regarding the target 20 in the facilities within the plant, reports using the captions generated by the caption generation system 5, etc. The recovery worker 33, as an example, browses the report in a situation where the target 20 described by the caption cannot be directly visually confirmed. The recovery worker 33 deals with abnormalities within the plant. The recovery worker 33, as an example, deals with an abnormality in the pipe 21. The recovery worker 33, as an example, restores the pipe 21 in which the abnormality has occurred according to the instructions and reports from the administrator 32 displayed on the worker terminal 43. Note that there may be one or a plurality of recovery workers 33 within the plant. Note that the recovery worker 33 is an example of a viewer who browses a report using the caption generated by the caption generation system 5.
[0034] The caption generation device 100, as an example, receives the image 10 obtained by imaging the target 20 in the facilities within the plant from the imager terminal 41 via the communication network 50. The caption generation device 100 assigns weights according to the frequencies extracted from the past images of the target 20 imaged in the image 10 to a plurality of words extracted from the image 10, and generates a caption that describes the target 20 imaged in the image 10 from the plurality of weighted words. The caption generation device 100 generates a caption that describes the target 20 with higher accuracy when an abnormality occurs in the target 20 compared to the case where a caption that describes the target 20 is generated without assigning the weight to the plurality of words extracted from the image 10 of the target 20. Note that the frequency of an abnormality occurring in the target 20 is, for example, about once every few years.
[0035] The caption generation device 100 transmits the generated caption to the imager terminal 41 via the communication network 50. The caption generation device 100 may be arranged in a control room, an instrument room, etc. within the plant, or may be arranged outside the plant. Some functions of the caption generation device 100 may be incorporated into the imager terminal 41.
[0036] Figure 2 is a block diagram of the caption generation device 100 according to the first embodiment. In Figure 2, the flow of data and the like is indicated by arrows. The caption generation device 100 includes an image acquisition unit 110, a word extraction unit 120, a weighting unit 130, a caption generation unit 140, a storage unit 150, and a report storage unit 155.
[0037] The image acquisition unit 110 acquires the image 10 captured within the plant. As an example, the image acquisition unit 110 receives, from the imaging device terminal 41 via the communication network 50, the user ID of the imager 31 who uses the imaging device terminal 41, together with the image 10 of the target 20 within the plant. The image acquisition unit 110 may receive, from the imaging device terminal 41 via the communication network 50, the user ID of the imager 31 and the imaging information indicating the imaging position and imaging direction when the imaging device terminal 41 captured the image 10 of the target 20, together with the image 10 of the target 20 captured within the plant. The image acquisition unit 110 inputs the received image 10 to the word extraction unit 120 together with the user ID of the imager 31 and the like.
[0038] The word extraction unit 120 extracts a plurality of words representing the features of the target 20 imaged in the image 10. The word extraction unit 120 may extract one or more words based on the newly input image 10 using a word extraction model that has learned the relationship between the image 10 and one or more words representing the features of the target 20 imaged in the image 10. The word extraction unit 120 may, for example, read out the word extraction model stored in the storage unit 150. The word extraction unit 120 inputs the plurality of words extracted from the image 10, for example, in vector representation, to the weighting unit 130 together with the user ID of the imager 31 and the like.
[0039] The weighting unit 130 assigns a greater weight to words with a lower frequency of extraction from past images that captured the same object as the object 20 imaged in the image 10 among the plurality of words extracted by the word extraction unit 120. The weighting unit 130 may, for example, compare the image 10 with past images included in each of the plurality of past reports stored in the report storage unit 155, and select one or more past images that captured the same object 20 as the image 10. In this case, the weighting unit 130 may extract a plurality of words associated with each of the selected one or more past images from the reports for each of the selected one or more past images. Alternatively, the weighting unit 130 may, for example, compare the image 10 with the plurality of past images stored in the storage unit 150, and select one or more past images that captured the same object 20 as the image 10. In this case, the weighting unit 130 may extract a plurality of words associated with each of the selected one or more past images from the storage unit 150.
[0040] The weighting unit 130 may specify the frequency with which each word was extracted from one or more past images by comparing the vector representation of each word extracted from the image 10 with the vector representations of the plurality of words stored in the storage unit 150 associated with one or more past images. The weighting unit 130 inputs the plurality of words each weighted, for example, as a vector representation, together with the user ID of the imager 31 to the caption generation unit 140.
[0041] The caption generation unit 140 generates a caption that describes the imaged object 20 based on a plurality of words and respective weights of the plurality of words. The caption generation unit 140 may use a caption generation model that has learned the relationship between the respective weights of one or more words representing the features of the object 20 in the plant and the caption that describes the object 20 in the plant, and generate a caption based on the newly input plurality of words and the respective weights of the plurality of words. The caption generation unit 140 may read, for example, the caption generation model stored in the storage unit 150. The caption generation unit 140 transmits the generated caption as text data to the imager terminal 41 via the communication network 50 based on the user ID of the imager 31.
[0042] The storage unit 150 may store a plurality of past images, each of which is an image of an arbitrary object 20. In this case, the storage unit 150 may store, in association with each past image, imaging information indicating the imaging position and imaging direction when each past image was captured. The storage unit 150 may store the above-described word extraction model and caption generation model. Note that the storage unit 150 is an example of a model storage unit.
[0043] The report storage unit 155 receives reports and images from the imager terminal 41 via the communication network 50, and registers the reports and images in association in the memory. The report may be created by the imager 31 appropriately editing the caption using the imager terminal 41. The image corresponds to the above-described past image. The report storage unit 155 may be accessed by the weighting unit 130 and read the reports and past images stored in the memory. The report storage unit 155 may also be accessed from the administrator terminal 42, the worker terminal 43, etc. via the communication network 50, and the reports, etc. may be read.
[0044] FIG. 3 is a flowchart showing the flow of the caption generation method according to the first embodiment. As an example, this flow starts when an image 10 of a target 20 in a plant captured by an imager 31 is transmitted to a caption generation device 100 via a communication network 50 by an imager terminal 41.
[0045] The caption generation device 100 acquires the image 10 captured in the plant (step S101). Specifically, as an example, an image acquisition unit 110 of the caption generation device 100 receives, from the imager terminal 41 via the communication network 50, the image 10 of the target 20 in the plant, the user ID of the imager 31, and the imaging information of the image 10. In addition to the imaging position and imaging direction, the imaging information may also indicate at least one of the zooming degree when the imager terminal 41 captures the image 10 and the distance to the target 20.
[0046] The caption generation device 100 extracts a plurality of words representing the features of the target 20 imaged in the image 10 (step S102). Specifically, the caption generation device 100 extracts, as word vectors, a plurality of words representing the features of the target 20 imaged in the image 10 using the above-described word extraction model.
[0047] The word extraction model learns the relationship between the image 10 and a plurality of image feature amounts, and also learns the relationship between the plurality of image feature amounts and the word vectors of a plurality of words. When a new image 10 is input, the word extraction model extracts a plurality of image feature amounts from the image 10 and outputs, from the plurality of image feature amounts, word vectors of a plurality of words representing the features of the target 20 imaged in the image 10. The word extraction model may be a deep learning model having a neural network structure including, for example, a CNN (Convolution Neural Network) and an FC layer (Fully Connected Layer). Instead of the CNN, the word extraction model may use the Self-Attention or Multi-Head Self-Attention of a Transformer.
[0048] The caption generation device 100 assigns a greater weight to words extracted from the plurality of words extracted from the image 10 and having a lower frequency of extraction from past images that captured the same object as the object 20 imaged in the image 10 (step S103). Specifically, the weighting unit 130 of the caption generation device 100 first compares the image 10 with the past images included in each of the plurality of past reports stored in the report storage unit 155, and selects one or more past images that captured the same object 20 as the image 10. As an example, the weighting unit 130 uses the imaging information of the image 10 to select one or more past images stored in the report storage unit 155. More specifically, the weighting unit 130 selects the imaging information of one or more past images from the imaging information of the plurality of past images by collating the imaging information of the image 10 with the imaging information of each past image stored in the report storage unit 155.
[0049] The weighting unit 130 may select the imaging information of the past image when the imaging position of the past image is within a range of a predetermined distance threshold from the imaging position of the image 10 and the imaging direction of the past image is within a range of a predetermined angle threshold from the imaging direction of the image 10. In this case, the weighting unit 130 may read out the distance threshold and the angle threshold stored in the storage unit 150.
[0050] When the imaging information also includes at least one of the zooming degree when the imaging device 41 captures the image 10 and the distance to the target 20, the weighting unit 130, in addition to the imaging position of the past image being within the range of a predetermined distance threshold from the imaging position of the image 10 and the imaging direction of the past image being within the range of a predetermined angle threshold from the imaging direction of the image 10, the zooming degree of the past image is within the range of a predetermined degree threshold from the zooming degree of the image 10, and the distance to the target in the past image is within the range of a predetermined second distance threshold from the distance to the target 20 in the image 10. When at least one of these conditions is satisfied, the imaging information of the past image may be selected. In this case, the weighting unit 130 may read out the degree threshold and the second distance threshold stored in the storage unit 150.
[0051] As another example, the weighting unit 130 selects one or more past images stored in the storage unit 150 using the image 10 itself. Specifically, the weighting unit 130 selects one or more past images from among the plurality of past images by performing pattern matching between the image 10 and each of the plurality of past images stored in the storage unit 150. More specifically, the weighting unit 130 may perform pattern matching between the image 10 and each of the plurality of past images stored in the storage unit 150, and select one or more past images having a degree of coincidence equal to or greater than a predetermined degree of coincidence threshold with the image 10. In this case, the weighting unit 130 may read out the degree of coincidence threshold stored in the storage unit 150.
[0052] Next, the weighting unit 130 identifies the frequency with which each word extracted from the image 10 has been extracted from one or more past images by comparing the word vector of each word extracted from the image 10 with the word vectors of a plurality of words stored in the storage unit 150 associated with one or more past images. The weighting unit 130 assigns a larger weight to the word vector of each word extracted from the image 10 as the frequency is lower. As an example, the weighting unit 130 may use a value with a reference range of 0.2 to 1 as the weight. For example, when the weighting unit 130 reads out the image IDs of five past images corresponding to the image 10 from the storage unit 150 and reads out five reports corresponding to the five image IDs from the report storage unit 155, if the word "wall" is used in any of the five reports, a weight of 0.2 may be assigned to the word "wall". In this case, if the word "gas" is used in only one of the five reports, a weight of 0.8 may be assigned to the word "gas". As an example, the weighting unit 130 may use the weight assigned to the word vector of each word as a weight vector and output the integrated value of the word vector of each word and the weight vector. Hereinafter, the integrated value will be referred to as a priority vector.
[0053] The caption generation device 100 generates a caption (step S104) for explaining the imaged object 20 based on a plurality of words and the respective weights of the plurality of words. The caption generation unit 140 may generate a caption by preferentially using the words with larger weights among the plurality of words. Specifically, the caption generation unit 140 of the caption generation device 100 uses the caption generation model stored in the storage unit 150 to generate, as text data, a caption for explaining the object 20 imaged in the image 10 based on the plurality of words extracted from the image 10 and the respective weights of the plurality of words. The caption generation unit 140 transmits the generated caption text data to the imaging device terminal 41 via the communication network 50, and thus the flow ends.
[0054] The imaging person 31 checks the caption displayed on the imaging person terminal 41, and transmits a report using the caption from the imaging person terminal 41 to the caption generation device 100 via the communication network 50. The imaging person 31 may appropriately edit the caption to create a report. When the caption generation device 100 receives a report from the imaging person terminal 41, it stores the report in the report storage unit 155. The report stored in the report storage unit 155 of the caption generation device 100 is viewed by, for example, the repair worker 33 of the pipe 21, and the repair worker 33 may block the crack in the pipe 21 or adjust the valve of the tank 22 on the upstream side of the pipe 21. The report stored in the report storage unit 155 of the caption generation device 100 is viewed by, for example, the administrator 32 of the pipe 21, and the administrator 32 may order the repair worker 33 of the pipe 21 to perform a repair work or adjust the control parameters of the equipment in which the pipe 21 is provided.
[0055] Regarding step S104, the caption generation model stored in the storage unit 150 has learned the relationship between the weights of each of one or more words representing the characteristics of the target 20 in the plant and the caption explaining the target 20 in the plant. More specifically, as an example, the caption generation model has learned the relationship between a plurality of priority vectors and a permutation of a plurality of word vectors. Regarding step S104, more specifically, the caption generation unit 140 inputs a plurality of priority vectors to the caption generation model, arranges the text data of a plurality of words extracted from the image 10 according to the permutation of the plurality of word vectors output from the caption generation model, and generates the text data of the caption. The permutation of the plurality of word vectors output from the caption generation model may be configured using, for example, those with a larger magnitude of the priority vector proportional to the magnitude of the weight vector. As an example, the permutation may be configured by adopting them in order from the one with a larger priority vector, or may be configured by adopting those with a magnitude of the priority vector equal to or greater than a predetermined threshold.
[0056] The caption generation model may have learned the relationship between a plurality of word vectors, a plurality of weight vectors each attached to the plurality of word vectors, and a permutation of the plurality of word vectors, instead of the relationship between a plurality of priority vectors and a permutation of the plurality of word vectors. In this case, the caption generation unit 140 inputs a plurality of word vectors and a plurality of weight vectors each attached to the plurality of word vectors to the caption generation model, and arranges the text data of the plurality of words extracted from the image 10 according to the permutation of the plurality of word vectors output from the caption generation model, and may generate the text data of the caption.
[0057] The caption generation model may be a deep learning model having a neural network structure including, for example, an RNN (Recurrent Neural Network). Instead of the RNN, the caption generation model may use an LSTM (Long Short-Term Memory), a Transformer, GPT-3 (Generative Pre-trained Transformer-3), GPT-2, a GPT-3-Clone, or the like. The caption generation model may use a large number of example sentences as learning data, and for each example sentence, learn an RNN or the like so as to increase the likelihood of outputting that example sentence. The caption generation model may extract words included in the example sentence, use the vectors of the extracted words as learning input data, use the example sentence as learning output data, and learn so as to increase the probability of outputting the learning output data for the learning input data.
[0058] As a comparative example with the caption generation device 100 according to the first embodiment, assume a device that generates a caption for explaining the target 20 in the plant imaged in the image 10 without using weights such that words with lower frequencies extracted from past images capturing the same target as the target 20 have larger values. For example, as shown in FIG. 1, assume a case where an abnormality is shown in the image 10, such as a crack occurring in a straight portion of the pipe 21 and gas jetting out from there. According to the device of the comparative example, for example, a plurality of words such as "straight pipe / curved pipe / valve / crack / tank body / tank ladder / tank support / gas / jet" may be extracted from the image 10.
[0059] Among these plurality of words, the word groups "straight pipe / curved pipe / valve / crack / tank body / tank ladder / tank support" are likely to be used even if there is no abnormality in the pipe 21. On the other hand, the word groups "crack / gas / jet" are likely to be used when there is an abnormality in the pipe 21. That is, the former word groups have a high frequency extracted from one or more past images capturing the pipe 21, while the latter word groups have a lower frequency extracted from one or more past images capturing the pipe 21.
[0060] However, according to the device of the comparative example, since weights corresponding to such frequencies are not used, based on the above-mentioned plurality of words, captions that do not explain the abnormality occurring in the pipe 21, such as "a curved pipe is connected to a straight pipe", "a valve is provided in the curved pipe", "a tank ladder is provided beside the tank body", that is, captions unrelated to the abnormality, may be generated with the same priority as the captions explaining the abnormality. When the device of the comparative example generates a plurality of captions and provides them to the user, even if a caption explaining the abnormality occurring in the pipe 21, such as "gas is jetting out from the straight pipe", may be generated, there is a high possibility that a large number of captions unrelated to the abnormality will be generated, and as a result, the caption explaining the abnormality is likely to be buried among a large number of captions unrelated to the abnormality.
[0061] On the other hand, according to the caption generation device 100 of the first embodiment, a plurality of words representing the characteristics of the target 20 are extracted from the image 10 of the captured target 20, and a greater weight is assigned to words with a lower frequency extracted from past images of the same target as the target 20 captured in the image 10. According to the caption generation device 100, a caption explaining the target 20 is generated based on the plurality of words and the respective weights of the plurality of words. According to such a caption generation device 100, compared with the device of the comparative example, the possibility of generating a caption for explaining the target 20 in an abnormal state from the image 10 in which the target 20 in an abnormal state is captured can be increased. For example, according to the caption generation device 100, when the above-mentioned plurality of words are extracted from the image 10 of the captured pipe 21, compared with the device of the comparative example, a caption explaining the pipe 21 in an abnormal state such as "a crack has occurred in the straight pipe and gas is jetting out" can be generated. In other words, according to the caption generation device 100, compared with the device of the comparative example, a caption for explaining the target 20 in an abnormal state can be generated with high accuracy from the image 10 in which the target 20 in an abnormal state is captured. Therefore, according to the caption generation device 100, when an abnormality occurs in the target 20 while the state of the target 20 in the facilities in the plant changes over time, the abnormality of the target 20 can be accurately notified to the user.
[0062] The caption generation device 100 may, for example, generate a caption by preferentially using words with a greater weight among the plurality of words, that is, words that are likely to be necessary for explaining the abnormality occurring in the target 20. Thereby, the caption generation device 100 can preferentially generate a caption explaining the abnormality rather than a caption unrelated to the abnormality, and can generate a caption with higher accuracy compared to the device of the comparative example.
[0063] In the caption generation device 100 according to the above first embodiment, when the caption generation unit 140 generates a caption by preferentially using words with larger weights among a plurality of words, the caption generation unit 140 may generate a plurality of captions. In this case, the caption generation unit 140 may assign a priority corresponding to the weight of at least one word used for caption generation among the plurality of words to the generated captions. The priority may be, for example, an average value, a total value, a cumulative value, etc. of the weights of the plurality of words. That is, the higher the average value, the total value, the cumulative value, etc. of the weights of the plurality of words, the higher the priority may be. The caption generation unit 140 may transmit the caption with the highest priority among the generated plurality of captions to the imaging device terminal 41. Alternatively, the caption generation unit 140 may transmit the generated plurality of captions to the imaging device terminal 41 with priorities assigned to each of the generated plurality of captions. The caption generation unit 140 may not transmit to the imaging device terminal 41 captions among the generated plurality of captions whose priorities are equal to or lower than a predetermined threshold value stored in the storage unit 150. When the caption generation device 100 transmits a plurality of captions to the imaging device terminal 41, the imaging device 31 can determine the suitability of the plurality of captions while referring to the priorities.
[0064] As an example of the case where the caption generation unit 140 generates a plurality of captions, after assigning arbitrary priorities to the plurality of words extracted by the word extraction unit 120, each time a word is used in caption generation, the priority of the word is decreased by a predetermined amount, so that the plurality of words extracted by the word extraction unit 120 may be used as a whole. In this case, a lower limit may be defined for the priority, and when the priority of a certain word reaches the lower limit, the caption generation unit 140 may set the priority of the word to be the same as the lower limit. Note that, instead of decreasing the priority of the word by a predetermined amount, the caption generation unit 140 may multiply the priority of the word by a coefficient greater than 0 and less than 1.
[0065] In the caption generation device 100 according to the above first embodiment, the word extraction unit 120 may assign a likelihood to each of a plurality of words extracted from the image 10 that represent the features of the object 20 imaged in the image 10. The likelihood may indicate the possibility that the word represents the features of the object 20 imaged in the image 10. The word extraction unit 120 may assign a high likelihood to a word that is highly likely to be a word representing the features of the object 20, and a low likelihood to a word that is less likely to be a word representing the features of the object 20.
[0066] More specifically, the word extraction unit 120 may assign a likelihood to each of a plurality of words extracted from the image 10 using a word extraction model. The word extraction model, for example, inputs the pixel values of each pixel of the image to each input node in the input stage, and outputs a scalar value for each word from each output node in the output stage. The scalar value may be an output value with a reference range of 0 to 1. In this case, the word vector is a vector having the above scalar value of 0 to 1 for each word. When the output value of the word is 0, it means that the image 10 does not represent that word, and when the output value of the word is 1, it means that the image 10 represents that word. The word extraction model is, for example, trained to output a value closer to 1 as the possibility that the image 10 represents that word is higher. The word extraction model is further, for example, repeatedly trained with a certain image 10 as input teacher data and a plurality of words extracted by decomposing the caption actually generated from the image 10 by the caption generation unit 140 as output teacher data. In this case, the word extraction model outputs a value closer to 1 as the possibility that the image 10 represents that word is higher, and also outputs a value closer to 1 as the possibility that the word represents the features of the object 20 imaged in the image 10 is higher. The word extraction unit 120, for example, treats the scalar value of each word output from the word extraction model as a likelihood for the sake of convenience and assigns it to each word.
[0067] As a specific example, assume that in the image 10 acquired by the word extraction unit 120, a tank 22, a pipe 21 connected to the tank 22, a crack occurring in the pipe 21, gas generated from the crack, a valve attached to the pipe 21, other devices not directly or indirectly connected to the tank 22 and the pipe 21, and a wall around the tank 22 and the pipe 21 are shown. For example, when the reference range of the output value for each word is 0 to 1 in the word extraction model, the output value for the word "tank" may be 0.9, the output value for the word "pipe" may be 0.9, the output value for the word "crack" may be 0.4, the output value for the word "gas" may be 0.3, the output value for the word "valve" may be 0.3, and the output values for the word referring to other devices and the word "wall" may be 0.2. In this case, the word extraction unit 120 may handle the probability of the word "tank" as 90% using the output value from the word extraction model.
[0068] In this case, the caption generation unit 140 may generate a caption based on a plurality of words, and the likelihood and weight of each of the plurality of words. More specifically, the caption generation unit 140 may generate a caption by preferentially using words with higher likelihood and larger weight among the plurality of words. Alternatively or in addition, the caption generation unit 140 may generate a caption by preferentially using words with a larger integrated value of likelihood and weight among the plurality of words. Alternatively or in addition, the caption generation unit 140 may generate a caption using words among the plurality of words whose likelihood is greater than a predetermined likelihood threshold and whose weight is greater than a predetermined weight threshold. Thereby, the caption generation device 100 can generate a caption with even higher accuracy. Note that the caption generation unit 140 may read the above-described likelihood threshold and weight threshold from the storage unit 150.
[0069] In this case, the caption generation model has already learned the relationship between the vector representation of the words representing the features of the target 20, the likelihood and weight assigned to the words, and the caption that describes the target 20. When a plurality of vector representations of words representing the features of the target 20 and the likelihood and weight assigned to each word are newly input, the caption generation model outputs text data of a caption that describes the target 20 imaged in the image 10 from the plurality of vector representations of words and the likelihood and weight of each word.
[0070] The caption generation unit 140 may generate a plurality of captions, for example, when generating a caption by preferentially using words with higher likelihood and larger weight among a plurality of words based on the plurality of words and the likelihood and weight of each of the plurality of words. In this case, the caption generation unit 140 may further assign a priority corresponding to the likelihood and weight of at least one word used for caption generation among the plurality of words to the generated caption. The priority may be, for example, an average value, a total value, a product value, etc. of the integrated values of the likelihood and weight assigned to each of the plurality of words. That is, the larger the average value, the total value, the product value, etc. of the integrated values of the likelihood and weight assigned to each of the plurality of words, the larger the priority may be. The caption generation unit 140 may transmit the caption with the highest priority among the generated plurality of captions to the imaging device terminal 41. Alternatively, the caption generation unit 140 may transmit the generated plurality of captions to the imaging device terminal 41 with priorities assigned to each of the generated plurality of captions. The caption generation unit 140 may not transmit to the imaging device terminal 41 captions among the generated plurality of captions whose priorities are below a predetermined threshold stored in the storage unit 150. When the caption generation device 100 transmits a plurality of captions to the imaging device terminal 41, the imaging person 31 can determine the suitability of the plurality of captions while referring to the priorities.
[0071] The caption generation unit 140 may also output a caption including an instruction to re-capture the imaged object 20 when the likelihood of a combination of a plurality of words used for caption generation is lower than a predetermined second likelihood threshold. For example, when the average value, total value, integrated value, etc. of the likelihoods of a plurality of words used for caption generation is less than a predetermined second likelihood threshold stored in the storage unit 150, the caption generation unit 140 may output a caption including an instruction to re-capture the imaged object 20.
[0072] In the caption generation device 100 according to the first embodiment, the storage unit 150 may store a plurality of sentences included in at least any one of an accident case collection, an accident response manual, and a maintenance history in the plant as example sentences. In this case, the caption generation unit 140 may extract at least one example sentence by searching among the plurality of example sentences stored in the storage unit 150 using at least any one of the plurality of words extracted by the word extraction unit 120.
[0073] The caption generation unit 140 may generate a caption based on the plurality of words, the respective weights of the plurality of words, and the extracted example sentence. Alternatively, the caption generation unit 140 may generate a caption based on the plurality of words, the respective likelihoods and weights of the plurality of words, and the extracted example sentence. Specifically, the caption generation unit 140 may search for an example sentence in which at least any one of the plurality of words is used among the plurality of example sentences stored in the storage unit 150, and generate a caption based on the weights of each word while referring to the extracted example sentence.
[0074] In this case, the caption generation unit 140 may generate a plurality of captions and assign a similarity to the extracted example sentences. Specifically, the caption generation unit 140 generates a plurality of captions based on a plurality of words and respective weights of the plurality of words, and searches for an example sentence in which at least any one of the plurality of words is used from among the plurality of example sentences stored in the storage unit 150, and may assign a similarity to each of the plurality of captions with the extracted example sentence. The caption generation unit 140 may, for example, convert each of the plurality of captions and the extracted example sentence into a feature vector indicating a combination of a plurality of words included therein, and calculate the inner product between the feature vectors as the similarity. The caption generation unit 140 may transmit the caption having the highest similarity among the generated plurality of captions to the imaging device terminal 41. Alternatively, the caption generation unit 140 may transmit to the imaging device terminal 41 with the similarity assigned to each of the generated plurality of captions. In this case, the imager 31 can determine the suitability of the plurality of captions while referring to the similarity.
[0075] The captions generated by the caption generation unit 140 while referring to the extracted example sentences may include, in addition to the captions explaining the target 20 in the plant, a plurality of action options regarding actions to be taken by the user or instructions to the user. As a result of generating captions with reference to example sentences, the caption generation device 100 can present to the user action options or instructions incorporating abnormal handling methods that cannot be directly decoded from the image 10.
[0076] The captions generated by the caption generation device 100 according to the first embodiment may include at least one of an instruction to image the imaged target 20 from a different angle or a different shooting angle and an instruction to image another target 20 related to the imaged target 20, in addition to the captions explaining the target 20 in the plant.
[0077] FIG. 4 is a flowchart showing an example of a detailed flow of the caption generation method according to the first embodiment. The operation flow shown in the flowchart of FIG. 4 may be a specific example of the caption generation method according to the first embodiment described with reference to FIG. 3. The caption generation apparatus 100 according to the first embodiment described above may generate a caption for explaining the object 20 from the image 10 obtained by imaging the object 20 according to the operation flow shown in the flowchart of FIG. 4, as an example.
[0078] The flow of FIG. 4 starts, as an example, when the image 10 of the object 20 in the plant captured by the imager 31 using the imager terminal 41 is transmitted to the caption generation apparatus 100 via the communication network 50.
[0079] The caption generation apparatus 100 acquires the image 10 captured in the plant (step S201). In step S201, as shown in FIG. 5, which specifically describes a part of the flow of FIG. 4, the image acquisition unit 110 receives the image 10 of the object 20 in the plant from the imager terminal 41 via the communication network 50, and extracts Exif information from the image 10. The Exif information may include the user ID of the imager 31 in addition to the imaging information described above. The image acquisition unit 110 may output the image 10 to the word extraction unit 120 and output the Exif information to the weighting unit 130.
[0080] The caption generation device 100 crops the image 10 into a plurality of parts (step S202). For example, the caption generation device 100 creates a new image by discarding the peripheral area in the image 10 that does not include the target 20, and divides it into patches of a fixed size. In step S202, as shown in FIG. 5, the word extraction unit 120 may create a new image by discarding the peripheral area in the image 10 that does not include the pipe 21 and the tank 22, and divide the new image into nine patches by dividing it vertically and horizontally into three parts. Alternatively, the word extraction unit 120 may create a new image by discarding the peripheral area in the image 10 that does not include the pipe 21 as a patch of a fixed size, and create a new image by discarding the peripheral area in the image 10 that does not include the tank 22 as another patch of a fixed size. Note that the learning model of the word extraction unit 120 shown in FIG. 5 may also be referred to as a Vision Transformer and may be an example of the above-described word extraction model.
[0081] The caption generation device 100 converts the cropped image into a numerical value A that can be used by the encoder through a simple function, for example, a LINEAR function (step S203). In step S203, as shown in FIG. 5, the word extraction unit 120 may put nine patches into a linear projection layer, flatten each of them, and embed them (linearly) into vectors (or convert them into one-dimensional arrays). The word extraction unit 120 may further add a CLS token to the head of the sequence of vectors output from the linear projection layer (extra learnable [class] embedding), and embed positions into each patch (each vector) to obtain the numerical value A. That is, the word extraction unit 120 may convert the nine patches divided from the image 10 into a sequence of 10-dimensional vectors that can be used by the encoder.
[0082] The caption generation device 100 causes the transformer encoder to convert a numerical value A into another value B (step S204) and unify the dimension and numerical width of the value B (step S205). In steps S204 to S205, as shown in FIG. 5, the word extraction unit 120 inputs a sequence of 10-dimensional vectors to the transformer encoder, and the transformer encoder converts this into another sequence of 10-dimensional vectors, and unifies the dimension and numerical width, for example, through a softmax function, and may output it as the output of the CLS token. Note that the transformer encoder may include a plurality of components such as Self-Attention described above. The output of the CLS token may be an aggregation of feature amounts necessary for classification from the entire image by Self-Attention.
[0083] The caption generation device 100 outputs image feature amounts (step S206). In step S206, as shown in FIG. 5, the word extraction unit 120 may input the output of the CLS token to a classification head (MLP: multi-layer perceptron) and output a plurality of image feature amounts from the classification head. The classification head may output after compressing a plurality of image feature amounts. As shown in FIG. 5, the word extraction unit 120 may output a plurality of image feature amounts to the weighting unit 130, and as described with reference to FIGS. 1 to 3, extract word vectors of a plurality of words representing the features of the object 20 imaged in the image 10 from the plurality of image feature amounts and output them to the weighting unit 130.
[0084] The caption generation device 100 converts the image feature amount into a numerical value C suitable for the decoder (step S207), and inputs the numerical value C into the decoder (step S208). From step S207 to S208, as shown in FIG. 6 which specifically describes a part of the flow in FIG. 4, an integrated component including the weighting unit 130 and the caption generation unit 140 vectorizes the Exif information input from the image acquisition unit 110 and vectorizes the empty text through, for example, a LINEAR function, and vectorizes these vectors together with a plurality of image feature amounts input from the word extraction unit 120. Thereby, the integrated component converts each of the plurality of image feature amounts into the numerical value C. The integrated component further inputs the numerical value C into the decoder. Note that the learning model of the decoder is a model using layer-by-layer adjustment, which can also be referred to as GPT-2 and may be an example of the above-described caption generation model.
[0085] The caption generation device 100 causes the decoder to convert the numerical value C into a numerical value D (step S209), and converts it into text (tokens) through a function that returns the numerical value D to text (step S210). From step S209 to S210, as shown in FIG. 6, the decoder in the above-described integrated component may convert the numerical value C into the numerical value D and weight each numerical value D. The decoder may weight each numerical value D in the same manner as the weighting unit 130 described with reference to FIGS. 1 to 3 based on the vectorized Exif information. The decoder may compress the numerical value D. The decoder may convert the numerical value D into text (tokens) through a softmax function.
[0086] The caption generation device 100 outputs text (caption) (step S211), and this flow ends. In step S211, as shown in FIG. 6, for example, the decoder takes a sequence of continuous text such as X1 to X5 as input (source tokens), and uses a sequence of continuous text such as Y1 to Y5 that is the same as the sequence but with one token shifted to the right as target tokens, concatenates the source tokens and the target tokens, and may process them by adjusting layer by layer. For the source tokens and the target tokens, their respective resettable positions may be embedded together with their corresponding tokens. The position of the source tokens starts from zero, and for the target tokens, instead of increasing the position at the end of the source sentence, the position may be reset to zero again.
[0087] The decoder model shown in FIG. 6, as an example, includes an attention mechanism including a self-attention mechanism and a mixed attention mechanism, an add and layer normalization (Add&Layer Normalization) layer, a feed-forward neural network (FNN) layer, and N layers stacked in this order with the add and layer normalization layers. The decoder may put the processed result of N layers into a linearization layer, pass the output from the linearization layer through a softmax function, and output a caption consisting of a sequence of continuous text such as Y1 to Y5. The integrated component may transmit the generated caption text data to the imaging device terminal 41 via the communication network 50.
[0088] FIG. 7 is a block diagram of the caption generation device 200 according to the second embodiment. The caption generation device 200 according to the second embodiment includes a learning unit 260 in addition to the configuration included in the caption generation device 100 according to the first embodiment. Other configurations included in the caption generation device 200 according to the second embodiment are the same as those of the caption generation device 100 according to the first embodiment, and redundant descriptions are omitted using the reference numbers of each configuration included in the caption generation device 100.
[0089] The learning unit 260 learns the above-described word extraction model and caption generation model using the result of the user's determination of the appropriateness of the caption generated by the caption generation unit 140. In addition to or instead of this, the learning unit 260 may learn the above-described word extraction model and caption generation model using the user's correction input for the caption generated by the caption generation unit 140. The learning unit 260 receives the above-described determination result and correction input by the user from the imaging device terminal 41 via the communication network 50.
[0090] The learning unit 260 updates the parameters of the model so as to reduce the error between the output of the model when each sample in the training data including the result of the above-described determination is input to the model and the label. For example, when learning a word extraction model using a multi-layer neural network including a CNN or the like, the learning unit 260 uses the error between the output value output by the neural network in response to inputting each sample and the label, and by a method such as backpropagation, adjusts the weights between the neurons of the neural network and the biases of each neuron.
[0091] As described above, in the caption generation device 200 according to the second embodiment, information on whether the generated caption has been actually used by the user is fed back, and the word extraction model and the caption generation model are learned using the feedback content. For example, when the caption generation device 200 receives, from the administrator terminal 42, the result determined by the administrator 32 as being problem-free as a result of viewing the caption, it may be determined that the generated caption has been actually used. The caption generation device 200 may also determine that the generated caption has been actually used, for example, when the caption generation unit 140 performs a search for example sentences using the generated caption, or when the generated caption is used in a report for the administrator 32. The caption generation device 200 according to such a second embodiment also has the same effects as the caption generation device 100 according to the first embodiment. According to the caption generation device 200 according to the second embodiment, captions with even higher accuracy can be generated.
[0092] Various embodiments of the present invention may be described with reference to flowcharts and block diagrams, where the blocks may represent (1) stages of a process in which an operation is performed or (2) sections of a device having a role of performing an operation. Specific stages and sections may be implemented by a dedicated circuit, a programmable circuit supplied with computer-readable instructions stored on a computer-readable medium, and / or a processor supplied with computer-readable instructions stored on a computer-readable medium. The dedicated circuit may include digital and / or analog hardware circuits, and may include integrated circuits (ICs) and / or discrete circuits. The programmable circuit may include a reconfigurable hardware circuit including memory elements such as logical AND, logical OR, logical XOR, logical NAND, logical NOR, and other logical operations, flip-flops, registers, field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), etc.
[0093] A computer-readable medium may include any tangible device capable of storing instructions executable by an appropriate device, and as a result, a computer-readable medium having instructions stored therein will comprise a product including instructions that can be executed to create means for performing the operations specified in a flowchart or block diagram. Examples of computer-readable media may include electronic memory media, magnetic memory media, optical memory media, electromagnetic memory media, semiconductor memory media, and the like. More specific examples of computer-readable media may include floppy (registered trademark) disks, diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), electrically erasable programmable read-only memory (EEPROM), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disc (DVD), Blu-ray (registered trademark) disc, memory stick, integrated circuit card, and the like.
[0094] Computer-readable instructions may include any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source code or object code written in an object-oriented programming language such as Smalltalk (registered trademark), JAVA (registered trademark), C++, and a conventional procedural programming language such as the "C" programming language or a similar programming language.
[0095] Computer-readable instructions may be provided to a processor or programmable circuitry of a programmable data processing apparatus such as a general-purpose computer, a special-purpose computer, or other computers, either locally or via a wide area network (WAN) such as a local area network (LAN), the Internet, etc., and may execute the computer-readable instructions to create means for performing the operations specified in a flowchart or block diagram. Examples of processors include computer processors, processing units, microprocessors, digital signal processors, controllers, microcontrollers, etc.
[0096] FIG. 8 shows an example of a computer 2200 in which multiple aspects of the present invention may be embodied, in whole or in part. Programs installed on the computer 2200 can cause the computer 2200 to function as an operation associated with the apparatus according to an embodiment of the present invention or as one or more sections of the apparatus, or can cause the computer 2200 to execute the operation or the one or more sections, and / or can cause the computer 2200 to execute a process according to an embodiment of the present invention or a stage of the process. Such a program may be executed by the CPU 2212 to cause the computer 2200 to perform specific operations associated with some or all of the blocks of the flowcharts and block diagrams described herein.
[0097] The computer 2200 according to this embodiment includes a CPU 2212, a RAM 2214, a graphic controller 2216, and a display device 2218, which are interconnected by a host controller 2210. The computer 2200 also includes an input / output unit such as a communication interface 2222, a hard disk drive 2224, a DVD-ROM drive 2226, and an IC card drive, which are connected to the host controller 2210 via an input / output controller 2220. The computer also includes legacy input / output units such as a ROM 2230 and a keyboard 2242, which are connected to the input / output controller 2220 via an input / output chip 2240.
[0098] The CPU 2212 operates according to programs stored in the ROM 2230 and the RAM 2214, thereby controlling each unit. The graphic controller 2216 acquires image data generated by the CPU 2212 in a frame buffer or the like provided in the RAM 2214 or in itself, and causes the image data to be displayed on the display device 2218.
[0099] The communication interface 2222 communicates with other electronic devices via a network. The hard disk drive 2224 stores programs and data used by the CPU 2212 in the computer 2200. The DVD-ROM drive 2226 reads a program or data from the DVD-ROM 2201 and provides the program or data to the hard disk drive 2224 via the RAM 2214. The IC card drive reads programs and data from an IC card and / or writes programs and data to the IC card.
[0100] ROM 2230 stores therein a boot program or the like executed by computer 2200 upon activation, and / or a program dependent on the hardware of computer 2200. Input / output chip 2240 may also be connected to input / output controller 2220 via various input / output units through a parallel port, a serial port, a keyboard port, a mouse port, etc.
[0101] The program is provided by a computer-readable medium such as DVD-ROM 2201 or an IC card. The program is read from the computer-readable medium, installed in hard disk drive 2224, RAM 2214, or ROM 2230, which is also an example of a computer-readable medium, and executed by CPU 2212. The information processing described in these programs is read by computer 2200, resulting in cooperation between the programs and the various types of hardware resources described above. The apparatus or method may be configured by realizing the operation or processing of information according to the use of computer 2200.
[0102] For example, when communication is executed between computer 2200 and an external device, CPU 2212 may execute a communication program loaded in RAM 2214 and instruct communication interface 2222 to perform communication processing based on the processing described in the communication program. Communication interface 2222 reads transmission data stored in a transmission buffer processing area provided in a recording medium such as RAM 2214, hard disk drive 2224, DVD-ROM 2201, or an IC card under the control of CPU 2212, transmits the read transmission data to the network, or writes the received data received from the network to a reception buffer processing area or the like provided on the recording medium.
[0103] Further, the CPU 2212 may cause all or necessary parts of files or databases stored in external recording media such as a hard disk drive 2224, a DVD-ROM drive 2226 (DVD-ROM 2201), an IC card, etc. to be read into the RAM 2214, and may execute various types of processing on the data on the RAM 2214. The CPU 2212 then writes back the processed data to the external recording media.
[0104] Various types of information such as various types of programs, data, tables, and databases may be stored in the recording media and may undergo information processing. The CPU 2212 may perform various types of processing on the data read from the RAM 2214, including various types of operations, information processing, conditional judgments, conditional branches, unconditional branches, information search / replacement, etc. described throughout this disclosure and specified by the program instruction sequence, and write back the results to the RAM 2214. Also, the CPU 2212 may search for information in files, databases, etc. within the recording media. For example, when a plurality of entries each having an attribute value of a first attribute associated with an attribute value of a second attribute are stored in the recording media, the CPU 2212 searches for an entry that matches the condition where the attribute value of the first attribute is specified from among the plurality of entries, reads the attribute value of the second attribute stored within the entry, and thereby may obtain the attribute value of the second attribute associated with the first attribute that satisfies a predetermined condition.
[0105] The programs or software modules described above may be stored in a computer-readable medium on or near the computer 2200. Also, a recording medium such as a hard disk or RAM provided within a server system connected to a dedicated communication network or the Internet can be used as a computer-readable medium, thereby providing the program to the computer 2200 via the network.
[0106] As described above, the present invention has been described using embodiments, but the technical scope of the present invention is not limited to the scope described in the above embodiments. It is obvious to those skilled in the art that various changes or improvements can be made to the above embodiments. It is clear from the description of the claims that forms with such changes or improvements can also be included in the technical scope of the present invention.
[0107] For example, the control system may be a computer housed in a single housing. That is, the controller may be realized by executing a program on a processor of the computer, and each input / output device may be implemented as an I / O device of the computer. Further, the controller may be implemented as a virtual machine executed by one or more processors. In such a configuration, the control system does not include a network that is a general-purpose or dedicated network, and the controller and the input / output devices can be connected by a chipset such as a memory controller hub and an I / O controller hub that connect between the processor and the I / O devices.
[0108] It should be noted that the execution order of each process such as operations, procedures, steps, and stages in the devices, systems, programs, and methods shown in the claims, the specification, and the drawings is not explicitly indicated as "before" or "preceding" etc., and can be realized in any order unless the output of the previous process is used in the subsequent process. Regarding the operation flows in the claims, the specification, and the drawings, even if "first," "next," etc. are used for convenience of explanation, it does not mean that it is essential to implement in this order.
Explanation of Reference Numerals
[0109] 5 Caption generation system 10 Image 20 Target 21 Pipe 22 Tank 31 Imager 32 Administrator 33 Recovery worker 41 Imager terminal 42 Manager terminal 43 Operator terminal 100 Caption generation device 110 Image acquisition unit 120 Word extraction unit 130 Weighting unit 140 Caption generation unit 150 Storage unit 155 Report storage unit 200 Caption generation device 260 Learning unit 2200 Computer 2201 DVD-ROM 2210 Host controller 2212 CPU 2214 RAM 2216 Graphics controller 2218 Display device 2220 Input / output controller 2222 Communication interface 2224 Hard disk drive 2226 DVD-ROM drive 2230 ROM 2240 Input / output chip 2242 Keyboard
Claims
1. An image acquisition unit that acquires an image captured within a plant; A word extraction unit that extracts a plurality of words representing features of an object imaged in the image; A weighting unit that assigns a greater weight to words among the plurality of words that have a lower frequency of extraction from past images of an object identical to the imaged object; A caption generation unit that generates a caption for explaining the imaged object based on the plurality of words and the respective weights of the plurality of words. A caption generation device comprising the above.
2. The caption generation unit generates the caption by preferentially using words among the plurality of words that have a greater weight. The caption generation device according to Claim 1.
3. The caption generation unit assigns a priority corresponding to the weight of at least one word used for caption generation among the plurality of words to the generated caption. The caption generation device according to Claim 2.
4. The word extraction unit assigns a likelihood to each of the plurality of words extracted from the image. The caption generation unit generates the caption based on the plurality of words, the respective likelihoods of the plurality of words, and the weights. The caption generation device according to Claim 1.
5. The caption generation unit generates the caption by preferentially using words among the plurality of words that have a higher likelihood and a greater weight. The caption generation device according to Claim 4.
6. The caption generation unit generates the caption by preferentially using words among the plurality of words that have a greater integrated value of the likelihood and the weight. The caption generation device according to Claim 4.
7. The caption generation unit generates the caption using words among the plurality of words that have a likelihood greater than a predetermined likelihood threshold and a weight greater than a predetermined weight threshold. The caption generation device according to Claim 4.
8. The caption generation unit assigns a priority corresponding to the likelihood and the weight of at least one word used for caption generation among the plurality of words to the generated caption. The caption generation device according to any one of Claims 5 to 7.
9. When the likelihood of a combination of a plurality of words used for caption generation is lower than a predetermined second likelihood threshold, the caption generation unit outputs the caption including an instruction to re-capture the imaged object. The caption generation device according to any one of claims 5 to 7.
10. The apparatus further includes a storage unit that stores a plurality of sentences included in at least any one of an accident case collection, an accident response manual, and a maintenance history in the plant as example sentences. The caption generation unit extracts at least one example sentence by searching using at least any one of the plurality of extracted words from among the plurality of example sentences stored in the storage unit, and generates the caption based on the plurality of words, the respective weights of the plurality of words, and the extracted example sentence. The caption generation device according to claim 1.
11. The caption generation unit generates a plurality of captions and assigns a similarity to the extracted example sentence. The caption generation device according to claim 10.
12. The caption includes a plurality of action options regarding actions to be taken by the user or an instruction to the user. The caption generation device according to claim 10.
13. The caption includes at least any one of an instruction to image the imaged object at a different angle or a different shooting angle and an instruction to image another object related to the imaged object. The caption generation device according to claim 1.
14. The apparatus further includes a model storage unit that stores a caption generation model that has learned the relationship between the respective weights of one or a plurality of words representing the characteristics of the object in the plant and a caption for explaining the object in the plant. The caption generation unit generates the caption based on the newly input plurality of words and the respective weights of the plurality of words using the caption generation model. The caption generation device according to claim 1.
15. The apparatus further includes a learning unit that learns a word extraction model that extracts the plurality of words from the image and a caption generation model that generates the caption from the respective weights of the plurality of words using the result of the user's determination of the appropriateness of the generated caption. The caption generation unit generates the caption based on the newly input plurality of words and the respective weights of the plurality of words using the caption generation model. The caption generation device according to claim 1.
16. The apparatus further includes a word extraction model that extracts the plurality of words from the image using a user's correction input for the generated caption, and a learning unit that learns the caption generation model from the respective weights of the plurality of words. The caption generation unit generates the caption based on the newly input plurality of words and the respective weights of the plurality of words using the caption generation model. The caption generation device according to claim 1.
17. Obtaining an image captured within a plant; Extracting a plurality of words representing features of an object imaged in the image; Assigning a greater weight to a word having a lower frequency of extraction from past images that captured the same object as the imaged object among the plurality of words; Generating a caption that describes the imaged object based on the plurality of words and the respective weights of the plurality of words A caption generation method comprising:
18. A computer, A procedure for obtaining an image captured within a plant; A procedure for extracting a plurality of words representing features of an object imaged in the image; A procedure for assigning a greater weight to a word having a lower frequency of extraction from past images that captured the same object as the imaged object among the plurality of words; A procedure for generating a caption that describes the imaged object based on the plurality of words and the respective weights of the plurality of words A program for causing the computer to execute the procedures.
Citation Information
Patent Citations
Inspection support processing server and inspection support system
JP2014139724A
Denoising autoencoder image captioning
US20220012534A1
Automatic concrete dam defect image description generation method based on graph attention network
US20230401390A1