Image processing apparatus, method, and program

The image processing apparatus efficiently generates accurate image description texts by integrating event name and semantic role information with image data within a single neural network, addressing the cost and error issues of existing multi-step methods.

JP7687527B2Active Publication Date: 2025-06-03NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024522764
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-24
Publication Date
2025-06-03
Estimated Expiration
2042-05-24

AI Technical Summary

Technical Problem

Existing image description text generation methods, such as those described in Non-Patent Document 2, are costly in terms of learning and inference due to the need for multiple neural network models and steps, which can lead to errors in sentence generation, particularly in the ordering of semantic roles.

Method used

An image processing apparatus and method that integrates image data, event name information, and semantic role information to generate image description texts efficiently by using a single neural network that estimates regions, order, and words simultaneously, while ensuring accurate positioning and semantic roles in the generated text.

Benefits of technology

The proposed solution enables the generation of accurate and relevant image description texts at a lower cost compared to existing methods, reducing the likelihood of errors in word order and ensuring that generated sentences are natural and contextually appropriate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007687527000003
    Figure 0007687527000003
  • Figure 0007687527000004
    Figure 0007687527000004
  • Figure 0007687527000005
    Figure 0007687527000005
Patent Text Reader

Abstract

An image processing device according to one embodiment includes an input unit and an output unit. The input unit receives input of image data, name information that gives names for situations depicted by the image data, semantic role information that gives the respective semantic roles of the words in captions that describe the situations depicted by the image data, correct answer information for captions that describe the situations depicted by the image data, and correct answer information for the respective semantic roles of the words in captions that describe the situations. The output unit outputs captions that describe situations depicted by image data on the basis of image data, name information, semantic role information, correct answer information for captions, and correct answer information for semantic roles that have been input via the input unit and also outputs respective positions within a display area for the image data for the words of the captions that describe the situations and information that indicates the respective semantic roles of the words of the outputted captions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an image processing apparatus, method, and program.

Background Art

[0002] There is an image description text generation technology as a technology for describing in words the situation depicted in image data (which may also be simply referred to as an image). This technology is expected to be used to reduce the cost of manual clerical work, such as automatic recording of work photographed in a factory or automatic entry of electronic medical records at a medical site, and is widely studied.

[0003] For example, in Non-Patent Document 1, a set of a large number of images and sentences describing the situation depicted in the images is used as learning data, the images are input into an image description text generation model, and the model is learned to estimate the sentences describing the situation depicted in the images. This image description text generation model can generate simple sentences using common words that are often included in the data set, but cannot control the content mentioned in the sentences according to the purpose.

[0004] The technology for generating controllable image description texts is a technology that gives a control signal together with image data as an input to an image description text generator and generates an image description text while controlling the content mentioned, and has recently begun to be studied.

[0005] In addition, there is an object region specification type image description text generation technology that gives a partial region of an image to be mentioned as a control signal. Since the object region of the control signal in this technology is automatically selected from a plurality of object regions detected by an object detection technology that specifies the position and name of an object in the display region of the image, there may be included a region that has no direct relation to an event indicating the situation of the object to be mentioned shown by the image. At this time, there is a possibility that an unnatural sentence may be generated by using words that are not related to the event of interest. For example, while the event of interest is blood pressure measurement, when the area of a chair is included as control information, the phrase "with the chair", which has no direct relation to the event, will be included in the generated sentence.

[0006] Non-Patent Document 2 discloses a technique for generating an image description sentence in which an event name of interest and semantic role information related to the event are given as control signals. The semantic role may also be referred to as a thematic role. The event name in this technique is, for example, “test”, which is a name indicating the activity to be mentioned in the image. The semantic role information is, for example, “subject”, “object” or “location”, which are elements necessary when explaining the event name indicating the activity of interest in a sentence. There is little possibility that a word not related to the event to be mentioned is included in the sentence generated using this technique.

Prior Art Documents

Non-Patent Documents

[0007]

Non-Patent Document 1

Non-Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0008] However, the image description text generation method described in Non-Patent Document 2 is divided into a plurality of steps, and a neural network model is required in each step. Therefore, the problem is that the cost related to learning and inference is large. The neural network models in the above-mentioned plurality of steps include a model for estimating regions in an image for each semantic role given as a control signal, a model for estimating the order of semantic roles, and a model for estimating words for each semantic role arranged in order. During learning and inference, since it is necessary to learn and infer each model in order, it is assumed that it takes time such as adjusting the parameters of the model, and the cost increases.

[0009] In addition, since Non-Patent Document 2 also performs inference in the subsequent step using the estimation result of the previous step, an error in the inference in the previous step cannot be corrected in the subsequent step. If an error occurs during the inference of the order of semantic roles, there is a possibility that a sentence with an unnatural word order will be generated.

[0010] The present invention has been made paying attention to the above circumstances, and an object thereof is to provide an image processing apparatus, method, and program capable of appropriately generating an image description text for explaining a situation depicted in image data.

Means for Solving the Problems

[0011] An image processing apparatus according to an aspect of the present invention includes: an input unit that receives inputs of image data, name information indicating a name of a situation depicted in the image data, semantic role information indicating a semantic role of each word in a description text that describes the situation depicted in the image data, correct answer information of the description text that describes the situation depicted in the image data, and correct answer information of the semantic role of each word in the description text that describes the situation; and an output unit that outputs a description text that describes the situation depicted in the image data based on the image data, name information, semantic role information, correct answer information of the description text, and correct answer information of the semantic role input by the input unit, and outputs information indicating positions related to each word of the description text that describes the situation in a display area of the image data and a semantic role of each word of the output description text.

[0012] An image processing method according to an aspect of the present invention is a method performed by an image processing apparatus, the method including: the image processing apparatus receiving inputs of image data, name information indicating a name of a situation depicted in the image data, semantic role information indicating a semantic role of each word in a description text that describes the situation depicted in the image data, correct answer information of the description text that describes the situation depicted in the image data, and correct answer information of the semantic role of each word in the description text that describes the situation; and the image processing apparatus outputting a description text that describes the situation depicted in the image data based on the input image data, name information, semantic role information, correct answer information of the description text, and correct answer information of the semantic role, and outputting information indicating positions related to each word of the description text that describes the situation in a display area of the image data and a semantic role of each word of the output description text.

Advantages of the Invention

[0013] According to the present invention, an image description text that describes a situation depicted in image data can be appropriately generated.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

[0015] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. <Configuration> First, the configuration of the image processing apparatus according to an embodiment of the present invention will be described. This image processing apparatus may also be referred to as an image description text generation apparatus or an image description apparatus. FIG. 1 is a diagram showing an application example of the image processing apparatus 100 according to an embodiment of the present invention. The image processing apparatus 100 is configured by a computer including a CPU (Central Processing Unit), a RAM (Random Access Memory), and a ROM (Read Only Memory) storing a program for executing an image description text processing routine described later, and is functionally configured as follows.

[0016] As shown in FIG. 1, the image processing apparatus 100 according to the present embodiment includes a fusion information creation unit 1, a storage unit 2, an image description unit 3, a parameter update unit 4, an output unit 5, and a decoder fusion information creation unit.

[0017] The fusion information creation unit 1 receives the input of the image feature amount x, the event information y, and the semantic role information z from the storage unit 2, and creates the fusion information w in which these image feature amount x, event information y, and semantic role information z are fused.

[0018] The image feature amount x can be anything as long as it is a tensor extracted from a certain image data I. For example, it is a tensor output by inputting an image to the VGG network (Visual Geometry Group network) in Non-Patent Document 2.

[0019] The event information y is not particularly limited as long as it is a vector indicating the event name, which is the name of the situation depicted in the display area of a certain image data I. For example, the event information y is a vector with a length corresponding to the number of event types, and only the value of the index corresponding to the event shown in the image is "1", and the values of the other indexes are "0". For example, a vector indicating the event in a certain image may be created manually from among the types of events given in advance, or the event shown in a certain image may be recognized by an event recognition model capable of recognizing the given event types in advance, and a vector indicating the event in the above image may be created using this recognition result.

[0020] The semantic role information z is not particularly limited as long as it is a vector indicating the information necessary for explaining the content of the event shown by a certain image data I, that is, the semantic role information indicating the semantic role of each word in the sentence explaining the situation depicted in the image data. For example, the semantic role information z is a vector with a length corresponding to the number of types of semantic roles, and only the value of the index corresponding to the semantic role necessary for explaining the content of the event shown in the image is "1", and the values of the other indexes are "0". For example, from among the types of semantic roles given in advance, a vector indicating the semantic role in a certain image may be created manually, or a vector indicating the semantic role of the above image may be created using the result of parsing each word in a sentence explaining the content of the same event as the event shown in a certain image by a language parser that classifies the words into the types of semantic roles given in advance.

[0021] The fused information w is not particularly limited as long as it is a tensor created from the image feature amount x, the event information y, and the semantic role information z. Here, for example, assume that the size of the image feature amount x is represented by the horizontal width w, the vertical height h, and the number of channels c, the event information y is a vector having a length l_y, and the semantic role information z is a vector having a length l_z. In this case, the fused information w is a tensor in which the vectors of the event information y and the semantic role information z are replicated according to the sizes of the horizontal width w and the vertical height h of the image feature amount x. The size of this tensor when the tensor of the image feature x amount, the tensor formed by replicating the vector of the event information y, and the tensor formed by replicating the vector of the semantic role information z are superimposed in the channel direction is a tensor having a horizontal width w, a vertical height h, and the number of channels (c + l_y + l_z).

[0022] The storage unit 2 stores one or more sets of the neural network of the image description text generator A (sometimes referred to as a screen description model), the image feature amount x, the event information y, the semantic role information z, the correct sentence C, the correct position B, and the correct semantic role sequence S. FIG. 2 is a diagram showing an example of the configuration of the image description text generator. In this FIG. 2, the image feature amount x, the event information y, the semantic role information z, the fused information w created by the fused information creation unit 1, and the concept of the configuration of the image description text generator A are shown.

[0023] As shown in FIG. 2, the neural network of the image caption generator A is composed of an encoder neural network, a decoder neural network, a caption estimation neural network, a position estimation neural network, and a semantic role estimation neural network.

[0024] The neural network of this image caption generator A is not particularly limited as long as it is a neural network that inputs the fusion information w into the encoder neural network to obtain the encoder output feature amount e, then inputs the encoder output feature amount e and the decoder fusion information v into the decoder neural network to output the common feature amount h, inputs the common feature amount h into the caption estimation neural network to output the caption estimation result c, inputs the common feature amount h into the position estimation neural network to output the position estimation result b, and inputs the common feature amount h into the semantic role estimation neural network to output the semantic role estimation result s.

[0025] The encoder neural network is not particularly limited as long as it is a neural network that inputs the fusion information w and outputs the encoder output feature amount e. The encoder output feature amount e is not particularly limited as long as it is a tensor indicating the feature amount extracted from the event information y, the semantic role information z, and the image feature amount x, and is, for example, a tensor of size 100x512.

[0026] The decoder neural network is not particularly limited as long as it is a neural network that inputs the encoder output feature amount e and the decoder fusion information v and outputs the common feature amount h. This common feature amount h is not particularly limited as long as it is a tensor composed of the feature vectors indicating each word of the output sentence, and is, for example, a tensor of size output sentence length × l_h.

[0027] The caption estimation neural network is not particularly limited as long as it is a network that inputs the common feature amount h and outputs the caption estimation result c. This description text estimation result c is not particularly limited as long as it is a tensor indicating the word sequence of the output sentence that is the image description text. For example, it is a tensor with a size of "output sentence length l_h × vocabulary size D", and each element of this tensor is the occurrence probability of each word in the output sentence.

[0028] The position estimation neural network is not particularly limited as long as it is a network that takes the common feature quantity h as input and outputs the position estimation result b. This position estimation result b is not particularly limited as long as it is a tensor indicating the position in the display area of the image corresponding to each word of the output sentence that describes the situation depicted in the image data. For example, it is a tensor with a size of "output sentence length l_h × 4", and each element of this tensor is, for example, the values of the coordinates x, y, the horizontal width w, and the vertical height h with respect to the upper left of the area corresponding to each word of the output sentence in the image data.

[0029] The semantic role estimation neural network is not particularly limited as long as it is a network that takes the common feature quantity h as input and outputs the semantic role estimation result s. This semantic role estimation result s is not particularly limited as long as it is a tensor indicating the semantic role of each word in the above output sentence. For example, it is a tensor with a size of "output sentence length l_h × number of types of semantic roles p", and each element of this tensor is the occurrence probability regarding the semantic role of each word in the output sentence.

[0030] The correct sentence C stored in the storage unit 2 as described above is not particularly limited as long as it is a tensor indicating the correct information of the word sequence of the output sentence that describes the situation depicted in a certain image data I. For example, it is a tensor with a size of "output sentence length l_h × vocabulary size D", and each element is a tensor in which only the value of the index corresponding to each word in the output sentence is "1" and the values of the other indices are "0".

[0031] The correct position B stored in the storage unit 2 as described above is not particularly limited as long as it is a tensor indicating the position in the display area of the image data corresponding to each word of the output sentence, which is a sentence explaining the situation depicted in a certain image data I. For example, it is a tensor with a size of "output sentence length l_h × 4", and each element of this tensor is, for example, the values of the coordinates x, y, the horizontal width w, and the vertical height h with respect to the upper left of the area corresponding to each word of the output sentence in the image data.

[0032] The correct semantic role sequence S stored in the storage unit 2 as described above is not particularly limited as long as it is a tensor indicating the semantic role of each word of the output sentence, which is a sentence explaining the situation depicted in a certain image data I. For example, it is a tensor with a size of "output sentence length l_h × number of types of semantic roles", and each element is a tensor where only the index value corresponding to the semantic role of each word of the output sentence is "1" and the values of the other indexes are "0".

[0033] During the learning process, the decoder fusion information creation unit 6 reads and accepts the correct sentence C and the correct semantic role sequence S from the storage unit 2, and creates the decoder fusion information v based on these accepted results.

[0034] During the inference process, the decoder fusion information creation unit 6 accepts the partially estimated sentence estimation result c´ and the partial semantic role estimation result s´ from the image description unit 3, and creates the decoder fusion information v based on these accepted results.

[0035] The partially estimated sentence estimation result c´ is not particularly limited as long as it is a tensor indicating the word sequence output so far. For example, when outputting up to the third word from the beginning of the sentence, it is a tensor with a size of "3 × vocabulary size D", and each element of the tensor is the probability of each word.

[0036] The partial semantic role prediction result s´ is not particularly limited as long as it is a tensor indicating the semantic role of each word for the word sequence output up to a certain point. For example, when the third word from the beginning of the sentence is output, it is a tensor with a size of "3 × number of semantic role types", and each element of the tensor is the probability regarding the semantic role of each word.

[0037] The decoder fusion information v is not particularly limited as long as it is a tensor created from the correct sentence C and the correct semantic role sequence S as described above during the learning process. The decoder fusion information creation unit 6, for example, takes a tensor where the size of the correct sentence C is "output sentence length I_h × vocabulary size D", and converts it into a tensor with a size of "output sentence length I_h × 512" through a neural network that converts it from the "vocabulary size D dimension" to 512 dimensions, and designates this as the language feature tensor.

[0038] And the decoder fusion information creation unit 6, for example, takes a tensor where the size of the correct semantic role sequence S is "output sentence length l_h × number of semantic role types", and converts it into a tensor with a size of "output sentence length I_h × 512" through a neural network that converts it from the "number of semantic role types dimension" to 512 dimensions, and designates this as the semantic role feature tensor.

[0039] Then, the decoder fusion information creation unit 6 designates a tensor with a size of "output sentence length I_h × 512" created using the positional encoder proposed in Non-Patent Document 2 as the position information tensor, and designates a tensor that is the sum of each element of the language feature tensor, the semantic role feature tensor, and the position information tensor, and has a size of "output sentence length I_h × 512" as the decoder fusion information v.

[0040] Also, during the inference process, the decoder fusion information v is not particularly limited as long as it is a tensor created from the partial hypothesis sentence prediction result c´ and the partial semantic role prediction result s´ as described above. The decoder fusion information creation unit 6, for example, sets a tensor with a size of "3 × vocabulary number D" as the partial description text estimation result c´, and converts this tensor with a size of "3 × 512" that has been converted by a neural network that converts it from the "vocabulary number D dimension" to 512 dimensions as the language feature tensor.

[0041] Then, the decoder fusion information creation unit 6, for example, sets a tensor with a size of "3 × the number of semantic role types" as the partial semantic role estimation result s´, and converts this tensor with a size of "3 × 512" that has been converted by a neural network that converts it from the "semantic role type dimension" to 512 dimensions as the semantic role feature tensor.

[0042] Then, the decoder fusion information creation unit 6 sets a tensor with a size of "3 × 512" created using, for example, the positional encoder proposed in Non-Patent Document 2 as the position information tensor, and sets a tensor with a size of "3 × 512" that is the sum of each element of the language feature tensor, the semantic role feature tensor, and the position information tensor as the decoder fusion information v.

[0043] The image description unit 3 receives the fusion information w from the fusion information creation unit 1, receives the decoder fusion information v from the decoder fusion information creation unit 6, receives the image description text generator A from the storage unit 2, and inputs the fusion information w and the decoder fusion information v into the neural network of the image description text generator A. During the learning process, the image description unit 3 outputs the description text estimation result c, the position estimation result b, and the semantic role estimation result s output from this neural network, respectively. Also, during the inference process, the image description unit 3 outputs the partial description text estimation result c´, the partial position estimation result b´, and the partial semantic role estimation result s´ output from the neural network of the image description text generator A, respectively, and performs a sentence generation end determination process described later.

[0044] The partial position estimation result b´ is not particularly limited as long as it is a tensor indicating the positions in the image corresponding to each word for the word sequence output up to a certain point. For example, when the output includes the third word from the beginning of the sentence, it is a tensor with a size of "3×4", and each element of the tensor is the upper left coordinates "x, y", the horizontal width w, and the vertical height h of the region corresponding to each word.

[0045] The above-described sentence generation end determination process is not particularly limited as long as it is a process for determining whether sentence generation has ended. The image description unit 3 indicates, for example, the end of the sentence. <eos>When [output] is output, it is determined that sentence generation ends, and when other words are output, it is determined that sentence generation continues.

[0046] The parameter update unit 4 receives the description text estimation result c, the position estimation result b, and the semantic role estimation result s from the image description unit 3, and receives the image description text generator A, the correct sentence C, the correct position B, and the correct semantic role sequence S from the storage unit 2, and updates the parameters of each neural network of the image description text generator A (which may also be referred to as the parameters of the image description text generator A) so as to satisfy the following three constraints.

[0047] The first constraint is to update the parameters of the image description text generator A so that the content of the description text estimation result c and the correct sentence C approaches or becomes the same, and there is no particular limitation as long as it is a learning method set to satisfy this constraint. For example, the parameter update unit 4 calculates the cross-entropy loss between the description text estimation result c and the correct sentence C as in the following formula (1), and updates the parameters of the description text estimation neural network of the image description text generator A so that this error becomes small, for example, below a certain value, or becomes zero.

[0048]

Equation

[0049] Here, k in formula (1) is the index of the description text estimation result c and the correct sentence C, and y k is the value in the description text estimation result c output from the neural network of the image description text generator A, and t k is the value in the correct sentence C. t k is a value in which only the value of the index that becomes the correct class is "1" and the values of other indexes are "0".

[0050] The second constraint is to update the parameters of the image caption generator A so that the position estimation result b approaches or becomes the same as the correct position B, and there is no particular limitation as long as it is a learning method set to satisfy this constraint. For example, when the position estimation result b is {x b , y b , w b , h b} and the correct position B is {x B , y B , w B , h B}, the L1 distance between {x b , y b , w b , h b} and {x B , y B , w B , h B} is calculated, and the parameters of the position estimation neural network of the image caption detector A are updated so that this distance becomes smaller than a certain value, for example, or zero. The L1 distance is expressed as follows. |x b -x B |+|y b -y B |+|w b -w B |+|h b -h B |

[0051] The third constraint is to update the parameters of the image caption generator A so that the content of the semantic role estimation result s approaches or becomes the same as the correct semantic role sequence S, and there is no particular limitation as long as it is a learning method set to satisfy this constraint. The parameter update unit 4 calculates, for example, the cross-entropy error between the semantic role estimation result s and the correct semantic role sequence S as in the following formula (2), and updates the parameters of the semantic role neural network of the image caption generator A so that this error becomes smaller than a certain value, for example, or zero.

[0052]

Equation

[0053] Here, m in Expression (2) is an index of the semantic role estimation result s and the correct semantic role sequence S, and y m is a value in the semantic role estimation result s output from the neural network of the image caption generator A, and t m is a value in the correct semantic role sequence S. t m is a value in which only the value of the index that becomes the correct class is "1" and the values of other indexes are "0".

[0054] The output unit 5 receives the caption estimation result c, the position estimation result b, and the semantic role estimation result s from the image caption unit 3, and outputs these estimation results. These estimation results may be only the output sentence c' obtained by converting the caption estimation result c into a word sequence. In addition to this output sentence c', the position output information b' based on the position estimation result b may be further output as an estimation result, or the output semantic role s' in which each word of the output sentence is estimated from the semantic role estimation result s may be further output as an estimation result.

[0055] The above output sentence c' is not particularly limited as long as it is a word sequence obtained based on the caption estimation result c. For example, when the caption estimation result c is a tensor with a size of "output sentence length I_h × vocabulary size D" and each element of this tensor is the appearance probability of each word of the output sentence, the word sequence obtained by searching for the sentence with the maximum probability with a beam width of "5" from the beginning of the output sentence by beam search may be used, or the word sequence with the maximum probability obtained by calculating the appearance probability for all possible word sequences by grid search may also be used.

[0056] The above position output information b' is not particularly limited as long as it is data based on the position estimation result b. For example, it may be a visualization image in which a rectangle is superimposed on the position indicated by the position estimation result b on the display area of the image data, or it may be a file in which the position estimation result b is output as text data.

[0057] The above output semantic role s´ is not particularly limited as long as it is data based on the semantic role estimation result s. For example, a rectangle may be superimposed on the position indicated by the position estimation result b on the display area of the image data, and a visualization image in which the index of the maximum value of the semantic role estimation result s is superimposed near this rectangle, for example, in the upper left, may be used, or a file in which the index of the maximum value of the semantic role estimation result s is output as text data may be used.

[0058] <Operation by the image processing apparatus> Next, the operation of the image processing apparatus 100 according to the present embodiment will be described. The image processing apparatus 100 executes a learning processing routine and an inference processing routine described below, respectively. <<Learning processing routine>> First, the learning processing routine will be described. FIG. 3 is a flowchart showing an example of the learning processing routine executed by the image processing apparatus. In this learning processing routine, first, an input of an image feature amount x, event information y, and semantic role information z is received, and fusion information w obtained by fusing these pieces of information is input to the neural network of the image caption generator A. Then, a caption estimation result c, a position estimation result b, and a semantic role estimation result s are output from the neural network of the image caption generator A.

[0059] Then, the correct caption C, the correct position B, and the correct semantic role sequence S are received from the storage unit 2, and (1) the content of the output caption estimation result c approaches or becomes the same as the correct caption C, and (2) the output position estimation result b approaches or becomes the same as the correct position B, and (3) the content of the output semantic role estimation result s approaches or becomes the same as the correct semantic role sequence S. The parameters of the various neural networks of the image caption generator A are updated so that the above three constraints are satisfied.

[0060] First, in step S101, the fusion information creation unit 1 receives the input of the image feature amount x, the event information y, and the semantic role information z from the storage unit 2, creates the fusion information w formed by fusing these image feature amount x, event information y, and semantic role information z, and outputs the created fusion information w to the image description unit 3.

[0061] In step S102, the decoder fusion information creation unit 6 receives the correct sentence C and the correct semantic role sequence S from the storage unit 2, creates the decoder fusion information v based on the received results, and transmits the decoder fusion information v to the image description unit 3.

[0062] In step S103, the image description unit 3 receives the fusion information w output in step S101, the decoder fusion information v received in step S102, and the image description text generator A stored in the storage unit 2 respectively, and inputs the received fusion information w and decoder fusion information v into the neural network of the image description text estimator A. The image description unit 3 outputs the description text estimation result c, the position estimation result b, and the semantic role estimation result s from the neural network of this image description text estimator A respectively, and outputs these description text estimation result c, position estimation result b, and semantic role estimation result s to the parameter update unit 4.

[0063] In step S104, the parameter update unit 4 receives the description text estimation result c, the position estimation result b, and the semantic role estimation result s output in step S103, receives the image description text generator A, the correct sentence C, the correct position B, and the correct semantic role sequence S stored in the storage unit 2, and calculates the error between the description text estimation result c and the correct sentence C (sometimes referred to as the image description text loss), the error between the position estimation result b and the correct position B (sometimes referred to as the position estimation loss), and the error between the semantic role estimation result s and the correct semantic role sequence S (sometimes referred to as the semantic role estimation loss).

[0064] Then, in step S105, the parameter update unit 4 updates the parameters (parameters of the image description text model) of various neural networks of the image description text generator A so that the following three constraints are satisfied: (1) the content of the description text estimation result c approaches or becomes the same as the correct text C, (2) the position estimation result b approaches or becomes the same as the correct position B, and (3) the content of the semantic role estimation result s approaches or becomes the same as the correct semantic role sequence S. The parameter update unit 4 stores the image description text generator A with updated parameters in the storage unit 2.

[0065] <<Inference Processing Routine>> Next, the inference processing routine will be described. FIG. 4 is a flowchart showing an example of the inference processing routine executed by the image processing apparatus. In this inference processing routine, first, an input of the image feature amount x, the event information y, and the semantic role information z is received, and the fusion information w formed by fusing these is input to the neural network of the image description text generator A. Then, the description text estimation result c, the position estimation result b, and the semantic role estimation result s are output from the neural network of the image description text generator A.

[0066] Then, an output sentence c´ formed by converting the description text estimation result c into a word sequence, position output information b´ formed by visualizing the position estimation result b, and output semantic role s´ in which each word of the output sentence is estimated from the semantic role estimation result s are output respectively.

[0067] First, in step S201, the fusion information creation unit 1 receives an input of the image feature amount x, the event information y, and the semantic role information z from the storage unit 2, creates the fusion information w formed by fusing these image feature amount x, event information y, and semantic role information z, and outputs this fusion information w to the image description unit 3.

[0068] In step S202, the decoder fusion information creation unit 6 receives the partial description text estimation result c' and the partial semantic role estimation result s' from the image description unit 3, creates decoder fusion information v based on these received results, and outputs this decoder fusion information v to the image description unit 3.

[0069] In step S203, the image description unit 3 receives the fusion information w output in step S201, the decoder fusion information v output in step S202, and the image description text generator A stored in the storage unit 2 respectively, and inputs these received fusion information w and decoder fusion information v into the neural network of the image description text generator A.

[0070] The image description unit 3 outputs the partial description text estimation result c', the partial position estimation result b', and the partial semantic role estimation result s' from the neural network of this image description text estimator A, and performs sentence generation end determination processing on the above partial description text estimation result c'.

[0071] In step S204, when the image description unit 3 determines to continue sentence generation, it outputs the above partial description text estimation result c', partial position estimation result b', and partial semantic role estimation result s' to the image description unit 3.

[0072] On the other hand, in step S204, when the image description unit 3 determines to end sentence generation, it outputs the above partial description text estimation result c', partial position estimation result b', and partial semantic role estimation result s' to the output unit 5 as the description text estimation result c, position estimation result b, and semantic role estimation result s respectively.

[0073] In step S205, based on the description text estimation result c, position estimation result b, and semantic role estimation result s output in step S203, the output unit 5 outputs an output sentence c' obtained by converting the description text estimation result c into a word sequence, position output information b' obtained by visualizing the position estimation result b, and an output semantic role s' obtained by estimating each word of the output sentence from the semantic role estimation result s respectively.

[0074] According to an embodiment of the present invention, by using an image data, an event name as a control signal, and semantic role information as inputs, and using a neural network, an effect can be obtained that an image description text in which the event name and the semantic role information are given as control signals can be generated at low cost and with high accuracy.

[0075] Also, according to an embodiment of the present invention, by inputting an image data, an event name as a control signal, and semantic role information, and using a method disclosed in Non-Patent Document 2 above, it is possible to learn a neural network that simultaneously estimates a region estimation of a semantic role, an order estimation of a semantic role, and a word estimation from a semantic role in a single neural network, and an effect can also be obtained.

[0076] FIG. 5 is a block diagram showing an example of a hardware configuration of an image processing apparatus according to an embodiment of the present invention. In the example shown in FIG. 5, the image processing apparatus 100 according to the above embodiment is constituted by, for example, a server computer or a personal computer, and has a hardware processor 111A such as a CPU. Then, a program memory 111B, a data memory 112, an input / output interface 113, and a communication interface 114 are connected to the hardware processor 111A via a bus 115.

[0077] The communication interface 114 includes, for example, one or more wireless communication interface units, and enables transmission and reception of information with a communication network NW. As the wireless interface, for example, an interface adopting a low-power wireless data communication standard such as a wireless LAN (Local Area Network) is used.

[0078] Connected to the input / output interface 113 are an input device 200 and an output device 300 that are attached to the image processing apparatus 100 and used by users or the like. The input / output interface 113 captures operation data input by users or the like through an input device 200 such as a keyboard, touch panel, touchpad, mouse, etc., and performs processing to output and display the output data to an output device 300 including a display device using liquid crystal or organic EL (Electro Luminescence), etc. Note that as the input device 200 and the output device 300, devices built into the image processing apparatus 100 may be used, or input devices and output devices of other information terminals that can communicate with the image processing apparatus 100 via the network NW may also be used.

[0079] The program memory 111B is used in combination as a non-temporary tangible storage medium, for example, a non-volatile memory such as an HDD (Hard Disk Drive) or SSD (Solid State Drive) that can be written to and read from at any time, and a non-volatile memory such as a ROM, and stores programs necessary for executing various control processes and the like according to one embodiment.

[0080] The data memory 112 is used in combination as a tangible storage medium, for example, the above non-volatile memory and a volatile memory such as a RAM, and is used to store various data acquired and created during the process of performing various processes.

[0081] The image processing apparatus 100 according to one embodiment of the present invention can be configured as a data processing apparatus having, as a software-based processing functional unit, each unit shown in FIG. 1, that is, a fusion information creation unit 1, an image description unit 3, a parameter update unit 4, an output unit 5, and a decoder fusion information creation unit 6.

[0082] Each information storage unit and storage unit 2 used as a working memory or the like by each part of the image processing apparatus 100 can be configured by using the data memory 112 shown in FIG. 5. However, these configured storage areas are not essential configurations within the image processing apparatus 100. For example, they may be areas provided in an external storage medium such as a USB (Universal Serial Bus) memory or a storage device such as a database server arranged in the cloud.

[0083] The processing functional units in each of the above-described fusion information creation unit 1, image description unit 3, parameter update unit 4, output unit 5, and decoder fusion information creation unit 6 can all be realized by causing the hardware processor 111A to read and execute the program stored in the program memory 111B. Note that some or all of these processing functional units may be realized in other various forms including integrated circuits such as an application specific integrated circuit (ASIC) or a field-programmable gate array (FPGA).

[0084] In addition, the methods described in each embodiment can be implemented as a program (software means) to be executed by a computer, and can be stored in a recording medium such as a magnetic disk (e.g., a floppy (registered trademark) disk, hard disk, etc.), an optical disc (e.g., CD-ROM, DVD, MO, etc.), or a semiconductor memory (e.g., ROM, RAM, flash memory, etc.), and can also be transmitted and distributed via a communication medium. Note that the program stored on the medium side includes a setting program for configuring software means (including not only the execution program but also a table and data structure) to be executed by a computer in the computer. The computer that realizes this apparatus reads the program recorded on the recording medium, and in some cases, constructs software means using the setting program, and executes the above-described processing by controlling the operation by this software means. Note that the recording medium referred to in this specification includes not only a medium for distribution but also storage media such as a magnetic disk and a semiconductor memory provided inside a computer or in a device connected via a network.

[0085] Note that the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the gist thereof at the implementation stage. Also, the respective embodiments may be implemented in appropriate combinations, and in such cases, the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combinations selected from a plurality of disclosed constituent elements. For example, even if some constituent elements are deleted from all the constituent elements shown in the embodiments, and the problem can be solved and the effects can be obtained, the configuration from which these constituent elements are deleted can be extracted as an invention.

Explanation of Reference Numerals

[0086] 100... Image processing apparatus 1... Fusion information creation unit 2... Storage unit 3... Image description unit 4... Parameter update unit 5... Output unit 6…Decoder fusion information creation unit< / eos>

Claims

1. An input unit that receives an input of image data, name information indicating the name of the situation depicted in the image data, semantic role information indicating the semantic role of each word in the description text explaining the situation depicted in the image data, correct answer information of the description text explaining the situation depicted in the image data, and correct answer information of the semantic role of each word in the description text explaining the situation; An output unit that outputs a description text explaining the situation depicted in the image data based on the image data, name information, semantic role information, correct answer information of the description text, and correct answer information of the semantic role input by the input unit, and outputs information indicating the position related to each word of the description text explaining the situation in the display area of the image data and the semantic role of each word of the output description text; An image processing apparatus comprising the above.

2. The output unit: Inputs the image data, name information, semantic role information, correct answer information of the description text, and correct answer information of the semantic role input by the input unit into a neural network, and based on the result of this input, outputs a description text explaining the situation depicted in the image data, and outputs information indicating the position related to each word of the description text explaining the situation in the display area of the image data and the semantic role of each word of the output description text, The image processing apparatus according to Claim 1.

3. A first fusion information creation unit that creates first fusion information formed by fusing the image data, name information, and semantic role information input by the input unit; A second fusion information creation unit that creates second fusion information formed by fusing the correct answer information of the description text explaining the situation depicted in the image data and the correct answer information of the semantic role of each word in the description text explaining the situation, and further comprising: The output unit: Inputs the first fusion information created by the first fusion information creation unit and the second fusion information created by the second fusion information creation unit into a neural network, and based on the result of these inputs, outputs feature information indicating the features of each word of the description text explaining the situation; Based on the feature information, outputs an estimation result of the description text explaining the situation, an estimation result of the position related to each word of the description text explaining the situation in the display area of the image data, and an estimation result of the semantic role of each word in the sentence explaining the situation, respectively; The image processing apparatus according to Claim 1.

4. The correct information of the explanatory text describing the situation depicted in the image data, the correct information of the positions related to each word of the explanatory text describing the situation in the display area of the image data, and the correct information of the semantic roles of each word of the explanatory text describing the situation, and the explanatory text describing the situation depicted in the image data output using the neural network, the positions related to each word of the explanatory text describing the situation in the image data output using the neural network, and the information indicating the semantic role of each word of the explanatory text output using the neural network, based on update the parameters of the neural network so that the explanatory text describing the situation depicted in the image data output using the neural network approaches the correct information of the explanatory text; update the parameters of the neural network so that the positions related to each word of the explanatory text describing the situation in the image data output using the neural network approach the correct information of the positions; update the parameters of the neural network so that the information indicating the semantic role of each word of the explanatory text output using the neural network approaches the correct information of the semantic role; further comprising an update unit; The image processing apparatus according to claim 2.

5. A method performed by an image processing apparatus, comprising: the image processing apparatus receiving an input of the image data, name information indicating the name of the situation depicted in the image data, semantic role information indicating the semantic role of each word of the explanatory text describing the situation depicted in the image data, correct information of the explanatory text describing the situation depicted in the image data, and correct information of the semantic role of each word of the explanatory text describing the situation; the image processing apparatus outputting an explanatory text describing the situation depicted in the image data based on the input image data, name information, semantic role information, the correct information of the explanatory text, and the correct information of the semantic role, and outputting the positions related to each word of the explanatory text describing the situation in the display area of the image data and information indicating the semantic role of each word of the output explanatory text; An image processing method comprising the above.

6. The outputting Input the input image data, name information, semantic role information, correct information of the description text, and correct information of the semantic role of the image data into a neural network, and based on the result of this input, output a description text that describes the situation depicted in the image data, and output information indicating the position related to each word of the description text that describes the situation in the display area of the image data and the semantic role of each word of the output description text. The image processing method according to claim 5.

7. Create first fusion information formed by fusing the input image data, name information, and semantic role information. Further include creating second fusion information formed by fusing the correct information of the description text that describes the situation depicted in the image data and the correct information of the semantic role of each word of the description text that describes the situation. The outputting includes Input the created first and second fusion information into a neural network, and based on the results of these inputs, output feature information indicating the features of each word of the description text that describes the situation. Based on the feature information, output the estimated result of the description text that describes the situation, the estimated result of the position related to each word of the description text that describes the situation in the display area of the image data, and the estimated result of the semantic role of each word of the text that describes the situation, respectively. The image processing method according to claim 5.

8. An image processing program that causes a processor to function as each part of the image processing apparatus according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Data structure and image processing device

    JP2021036388A

  • Method and system for visio-linguistic understanding using contextual language model reasoners

    US20220019734A1