Object detection device, object detection method, learning device, learning method, and program

The object detection device addresses the challenges of over-detection and omission by integrating image and language features, enhancing the accuracy of important object detection in images and improving situational understanding.

JP7694350B2Active Publication Date: 2025-06-18NIPPON TELEGRAPH & TELEPHONE CORP
5 Cites 0 Cited by

Patent Information

Application Number
JP2021186377
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-11-16
Publication Date
2025-06-18
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

Conventional object detection technologies often suffer from over-detection of non-important objects and failure to detect important objects in images, which hampers accurate understanding of the situation depicted.

Method used

An object detection device that integrates image features and language information from captions to enhance object detection accuracy. This device includes an acquisition unit for images and captions, feature generation units for image and language information, a fusion unit to combine these features, and an object detection model trained using machine learning to provide precise object position and class estimation.

Benefits of technology

The proposed solution enables more accurate detection of important objects in images, reducing over-detection and omission errors, thereby improving the understanding of the situation depicted in the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007694350000004
    Figure 0007694350000004
  • Figure 0007694350000005
    Figure 0007694350000005
  • Figure 0007694350000006
    Figure 0007694350000006
Patent Text Reader

Abstract

To provide a technology which enables detecting with higher accuracy from an image an object important for understanding a situation in the image.SOLUTION: An object detection device comprises: an acquisition unit which acquires an image and a caption for explaining the image; an image feature generation unit which generates an image feature indicating a feature of the image; a linguistic information generation unit which generates linguistic information indicating the feature of the caption; a fused information generation unit which generates fused information including the image feature and the linguistic information; an object detection unit which inputs the fused information to an object detection model generated in advance by machine learning, and obtains object position estimated results indicating areas where the object in the image exists and object class estimated results indicating probabilities that the object belongs to respective classes, which are outputted from the object detection model; and an output unit which outputs object detection results based on the object position estimated results and the object class estimated results.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for detecting an object from an image.

Background Art

[0002] In order for a machine to understand the current situation captured by a camera, an object detection technique for specifying the position and class of an object existing in an image is utilized. For example, there are cases where the object detection technique is utilized to determine whether a person has entered a dangerous area by detecting a person and a machine in a factory, and cases where the object detection technique is utilized to detect a person's fall from a platform by detecting a person and a railway track on a platform of a station.

[0003] In recent years, object detection techniques using neural networks have been widely studied. For example, Non-Patent Document 1 discloses a technique for estimating a candidate region of an object from an image and estimating the position and class of the object from feature amounts extracted from the candidate region of the object.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In conventional technologies such as the technology disclosed in Non-Patent Document 1, situations frequently occur where objects that need to be recognized to understand the situation cannot be accurately detected. For example, as shown in FIG. 9, there may be over-detection of background people who are not important for understanding that a person is playing a tennis match. Also, as shown in FIG. 10, there may be a failure to detect a smartphone that is important for understanding that a person is looking at a smartphone. Thus, in conventional technologies, over-detection of objects that are not important for understanding the situation or failure to detect important objects often occurs.

[0006] An object of the present invention is to provide a technology that enables more accurate detection of objects important for understanding the situation in an image from the image.

Means for Solving the Problem

[0007] An object detection device according to an aspect of the present invention includes an acquisition unit that acquires an image and a caption that describes the image, an image feature generation unit that generates an image feature indicating a feature of the image, a language information generation unit that generates language information indicating a feature of the caption, a fusion information generation unit that generates fusion information including the image feature and the language information, an object detection unit that inputs the fusion information into an object detection model generated in advance by machine learning and obtains an object position estimation result indicating a region where an object exists in the image and an object class estimation result indicating a probability that the object belongs to each class output from the object detection model, and an output unit that outputs an object detection result based on the object position estimation result and the object class estimation result.

Effects of the Invention

[0008] According to the present invention, a technology is provided that enables more accurate detection of objects important for understanding the situation in an image from the image.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0011] <First Embodiment> [Configuration] FIG. 1 schematically shows an object detection apparatus 100 according to the first embodiment of the present invention. As shown in FIG. 1, the object detection apparatus 100 includes an image processing unit 110 and an object detection model generation unit 120. The object detection model generation unit 120 generates an object detection model for detecting an object from an image by a machine learning method. The object detection model is implemented by a neural network architecture. The image processing unit 110 uses the object detection model previously generated by the object detection model generation unit 120 to detect an object from the image input to the object detection apparatus 100.

[0012] The image processing unit 110 includes an acquisition unit 111, an image feature generation unit 112, a language information generation unit 113, a fusion information generation unit 114, an object detection unit 115, an output unit 116, and a model storage unit 117.

[0013] The model storage unit 117 stores an object detection model generated by the object detection model generation unit 120. The object detection model is a neural network configured to receive, as input, information (fusion information described later) generated from an image and a caption that describes the image, and output an object position estimation result and an object class estimation result.

[0014] A caption is a sentence that describes the situation (context) in the image. For example, for an image showing a situation where a woman is looking at a smartphone, the caption may be the sentence “The woman is watching the phone.” or “The person grabs the smart phone.” The caption is created manually. The caption may be a sentence described in a language other than English, such as Japanese.

[0015] The object position estimation result indicates the region where each object in the image exists. For example, when the coordinates of the top-leftmost point of the region where the object exists are (x′, y′), the width of the region is w′, and the height of the region is h′, the object position estimation result includes a combination of four values x′, y′, w′, and h′ for each object. The object class estimation result indicates the probability (likelihood) that each object in the image belongs to an individual class. In this embodiment, the object class estimation result is a tensor. Here, a tensor refers to data represented by a multi-dimensional array such as a vector or a matrix. For example, when three classes such as "phone", "woman", and "person" are assumed and 10 objects are detected from the image, the object class estimation result is a tensor of size 10×3. Each element of the tensor is the probability that a certain object belongs to a certain class. For example, the object class estimation result includes the probability that the class of the first object is "phone", the probability that the class of the first object is "woman", the probability that the class of the first object is "person", the probability that the class of the second object is "phone", the probability that the class of the second object is "woman", the probability that the class of the second object is "person", ···, the probability that the class of the tenth object is "phone", the probability that the class of the tenth object is "woman", and the probability that the class of the tenth object is "person".

[0016] The acquisition unit 111 acquires an image to be processed and a caption that describes the image to be processed. Hereinafter, the image to be processed may also be referred to as an input image.

[0017] The image feature generation unit 112 generates image features from the input image acquired by the acquisition unit 111. The image features can indicate the features of the input image. The image features can be any information as long as it includes values generated from the input image. For example, the image features include a plurality of numerical values that depend on the pixel values included in the input image. In this embodiment, the image features are a tensor.

[0018] For example, the image feature generation unit 112 applies c filters to the input image to generate c feature maps. Here, c is called the number of channels. c can be an integer greater than or equal to 2. Each feature map has w×h pixels. In other words, each feature map is an image with a size of w×h. w is the width of the feature map (the number of pixels in the X-axis direction), and h is the height of the feature map (the number of pixels in the Y-axis direction). Typically, the size of the feature map is smaller than the size of the input image. The image features include c feature maps. In this case, the image features are a tensor with a size of w×h×c. In other words, the image features are a tensor that includes w×h×c numerical values as elements. As the filter, for example, the convolutional filter included in the VGG (Visual Geometry Group) network proposed in Karen Simonyan and Andrew Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition”, in Proc. of ICLR2015 can be used.

[0019] Note that the image feature generation unit 112 may perform pooling (e.g., average pooling or max pooling) on each of the c feature maps, thereby obtaining c feature maps with reduced sizes as image features.

[0020] The language information generation unit 113 generates language information from the caption acquired by the acquisition unit 111. The language information may indicate the features of the caption. The language information may be any information as long as it includes values generated from the caption. For example, the language information includes a plurality of numerical values that depend on the words included in the caption. In this embodiment, the language information is a vector. Here, the vector refers to data represented by a one-dimensional array.

[0021] For example, the language information generation unit 113 divides the caption into a plurality of words by morphological analysis of the caption, converts each word into a word vector using a vector space model, and obtains a vector obtained by averaging or concatenating the word vectors as language information. The language information generation unit 113 may extract words of a predetermined part of speech from the words obtained by morphological analysis of the caption. The predetermined part of speech includes at least one part of speech. In one example, the predetermined part of speech may be a noun. When one word is extracted from the caption, the language information generation unit 113 obtains the word vector obtained by converting that word as language information. As the vector space model, for example, a model proposed in Tomas Mikorov, Kai Chen, G.s. Corrado, Jeffrey Dean, “Efficient Estimation of Word Representations in Vector Space”, In Proc of workshop at ICLR2013 and also referred to as word2vec can be used.

[0022] The fusion information generation unit 114 fuses the image features generated by the image feature generation unit 112 and the language information generated by the language information generation unit 113, and generates fusion information including the image features and the language information. In the present embodiment, the fusion information is a tensor. For example, the fusion information generation unit 114 replicates the vector that is the language information w×h times, generates a tensor obtained by combining the w×h replications, and superimposes the generated tensor and the tensor of the image features in the channel direction to obtain the fusion information. The fusion information is a tensor of size w×h×(c + v). Here, v is the length (number of elements) of the vector that is the language information.

[0023] The object detection unit 115 uses the object detection model stored in the model storage unit 117 to obtain an object position estimation result and an object class estimation result regarding the input image based on the fusion information generated by the fusion information generation unit 114. Specifically, the object detection unit 115 inputs the fusion information into the object detection model, and obtains the object position estimation result and the object class estimation result output from the object detection model as the object position estimation result and the object class estimation result regarding the input image. The object position estimation result regarding the input image indicates the region where each object exists in the input image, and the object class estimation result regarding the input image indicates the probability that each object in the input image belongs to each of a plurality of classes.

[0024] The output unit 116 generates an object detection result from the object position estimation result and the object class estimation result regarding the input image obtained by the object detection unit 115, and outputs the object detection result. The output unit 116 may display the object detection result on a display device. Alternatively, the output unit 116 may transmit the object detection result to another device such as a server. The object detection result may be any information as long as it is information based on the object position estimation result and the object class estimation result regarding the input image. In one example, the object detection result may be a visualization image in which a rectangular frame surrounding each object on the input image and the class name of each object are superimposed. The position where the rectangular frame is arranged is determined based on the object position estimation result. The class name is arranged near (for example, at the upper left) the rectangular frame. The class name corresponds to the index at which the probability regarding each object is the maximum value in the object class estimation result. In another example, the object detection result may be text data including the object position estimation result and the class name.

[0025] The object detection model generation unit 120 includes a selection unit 121, a language information generation unit 122, a fusion information generation unit 123, an object detection unit 124, an update unit 125, an output unit 126, a learning data storage unit 127, and a model storage unit 128.

[0026] The learning data storage unit 127 stores a learning dataset which is learning data used to train the object detection model. The learning dataset includes a plurality of data (samples) for each of a plurality of reference images which are learning images. The data for each reference image includes an image feature generated from the reference image, a caption explaining the reference image, ground truth object position information indicating the position of each object existing in the reference image, and ground truth object class information indicating the class of each object existing in the reference image. The objects existing in the reference image refer to the objects to be detected from the reference image. Specifically, the objects existing in the reference image refer to the objects important for understanding the situation in the reference image.

[0027] The image feature is generated by performing the same processing as described in relation to the image feature generation unit 112 on the reference image. The image feature is a tensor including w×h×c pixel values included in c feature maps with a size of w×h as elements. The caption is created manually. The ground truth object position information may be any information as long as it indicates the position of each object existing in the reference image. For example, when the top-left coordinates of the region where the object exists are (x′, y′), the width of the region is w′, and the height of the region is h′, the ground truth object position information includes a combination of four values x′, y′, w′, h′ for each object. The ground truth object class information may be any information as long as it indicates the class of each object existing in the image. For example, when three classes “phone”, “woman”, and “person” are assumed and 10 objects exist in the image, the object class estimation result is a tensor with a size of 10×3. In the tensor, only the element of the index corresponding to the ground truth class of the object is 1, and the other elements are 0. For example, when the ground truth class of the first object is “phone” and the ground truth class of the second object is “woman”, for the first object, the element corresponding to the class “phone” is 1, the element corresponding to the class “woman” is 0, and the element corresponding to the class “person” is 0, and for the second object, the element corresponding to the class “phone” is 0, the element corresponding to the class “woman” is 1, and the element corresponding to the class “person” is 0.

[0028] The model storage unit 128 stores an object detection model. As described above, the object detection model is a neural network configured to receive, as input, fusion information including image features indicating features of an image and language information indicating features of a caption that describes the image, and output an object position estimation result and an object class estimation result. The object detection model includes a plurality of parameters.

[0029] The learning process repeats a learning routine including selecting data regarding at least one reference image from a data set, and updating the object detection model using the selected data. The selection unit 121 selects, for example randomly, data to be used in each learning routine from the data set. When numbers are associated with samples (data) included in the data set, the selection unit 121 may select samples to be used in each learning routine from the data set in order of the numbers. The data selected by the selection unit 121 is referred to as target data. The selection unit 121 passes the caption included in the target data to the language information generation unit 122, passes the image features included in the target data to the fusion information generation unit 123, and passes the correct object position information and the correct object class information included in the target data to the update unit 125.

[0030] The language information generation unit 122, the fusion information generation unit 123, and the object detection unit 124 each perform the same processes as those described in relation to the language information generation unit 113, the fusion information generation unit 114, and the object detection unit 115. Therefore, detailed descriptions of the language information generation unit 122, the fusion information generation unit 123, and the object detection unit 124 are omitted.

[0031] The language information generation unit 122 generates language information from the caption included in the target data. The language information is a vector including, as elements, a plurality of numerical values depending on words included in the caption. For example, the language information generation unit 122 uses a predetermined vector space model such as word2vec to convert each word included in the caption into a word vector, and obtains, as the language information, a vector obtained by averaging or concatenating the word vectors.

[0032] The fusion information generation unit 123 fuses the image features included in the target data and the language information generated by the language information generation unit 122 to generate fusion information. The fusion information is a tensor formed by fusing the tensor that is the image feature and the vector that is the language information. For example, the fusion information generation unit 114 duplicates the vector that is the language information w×h times, generates a tensor obtained by combining the w×h duplicates, and obtains the fusion information by superimposing the generated tensor and the tensor of the image features in the channel direction.

[0033] The object detection unit 124 receives the object detection model from the model storage unit 128, inputs the fusion information generated by the fusion information generation unit 123 into the object detection model, and obtains the object position estimation result and the object class estimation result output from the object detection model.

[0034] The update unit 125 updates the object detection model based on the object position estimation result and the object class estimation result obtained by the object detection unit 124, and the correct object position information and the correct object class information included in the target data. Specifically, the update unit 125 updates the parameters constituting the object detection model so as to satisfy the following two constraints.

[0035] The first constraint is that the object position estimation result approaches or matches the correct object position information. The learning method may be any learning method as long as it is set to satisfy the first constraint. For example, the first constraint is that the L1 distance (Manhattan distance) between the object position estimation result and the correct object position information becomes small. Assuming the object position estimation result is (x b , y b , w b , h b ) and the correct object position information is (x B , y B , w B , h B ), the L1 distance d between the object position estimation result and the correct object position information is as shown in the following formula (1).

Equation

[0036] The second constraint is that the object class estimation result approaches or matches the correct object class information. The learning method may be any learning method as long as it is set to satisfy the second constraint. For example, the second constraint is that the cross-entropy error between the object class estimation result and the correct object class information becomes small. The cross-entropy error E between the object class estimation result and the correct object class information is as shown in the following formula (2).

Number

[0037] Thus, the update unit 125 updates the parameters of the object detection model so that, for example, the L1 distance between the object position estimation result and the correct object position information becomes small, and the cross-entropy error between the object class estimation result and the correct object class information becomes small.

[0038] The learning process is repeated until a predetermined condition is satisfied. For example, when the L1 distance between the object position estimation result and the correct object position information is less than a predetermined first threshold, and the cross-entropy error between the object class estimation result and the correct object class information is less than a predetermined second threshold, the learning process ends. Alternatively or additionally, the learning process may end when the number of repetitions reaches a predetermined number of times.

[0039] The output unit 126 outputs the learned object detection model to the image processing unit 110. The learned object detection model is stored in the model storage unit 117 of the image processing unit 110.

[0040] In FIG. 1, the image processing unit 110 and the object detection model generation unit 120 are shown as being present in the same device, but the object detection model generation unit 120 may be present in a device different from the object detection device 100. In this case, the object detection model generated by the learning device including the object detection model generation unit 120 may be provided to the object detection device 100 via data communication or a computer-readable recording medium.

[0041] FIG. 2 schematically shows a configuration example of the object detection model. In the example shown in FIG. 2, the object detection model includes a main neural network 201, an object position estimation neural network 202, and an object class classification neural network 203.

[0042] The main neural network 201 includes an input layer where the fused information, which is the input layer of the object detection model, is input. As described above, the fused information is information obtained by fusing image features and language information. The output of the main neural network 201 is given to the object position estimation neural network 202 and the object class classification neural network 203.

[0043] The object position estimation neural network 202 includes an output layer that is a part of the output layer of the object detection model and outputs an object position estimation result. The object position estimation neural network 202 receives the output of the main neural network 201 as an input and outputs an object position estimation result. The object class classification neural network 203 includes an output layer that is the remaining part of the output layer of the object detection model and outputs an object class estimation result. The object class classification neural network 203 receives the output of the main neural network 201 as an input and outputs an object class estimation result.

[0044] In the learning process, the update unit 125 updates the parameters of the main neural network 201 and the parameters of the object position estimation neural network 202 so that the object position estimation result output from the object position estimation neural network 202 approaches the correct object position information. The update unit 125 updates the parameters of the main neural network 201 and the parameters of the object class classification neural network 203 so that the object class estimation result output from the object class classification neural network 203 approaches the correct object class information.

[0045] Figure 3 schematically shows an example of the hardware configuration of the object detection device 100. As shown in Figure 3, the object detection device 100 includes, as hardware components, a processor 301, a RAM (Random Access Memory) 302, a program memory 303, a storage device 304, and an input / output interface 305. The processor 301 is communicably connected to the RAM 302, the program memory 303, the storage device 304, and the input / output interface 305.

[0046] The processor 301 includes a general-purpose circuit such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). The RAM 302 includes a volatile memory such as an SDRAM (Synchronous Dynamic Random Access Memory). The RAM 302 is used by the processor 301 as a working memory. The program memory 303 stores programs to be executed by the processor 301, including an object detection program and a learning program. Each program includes a plurality of computer-executable instructions. As the program memory 303, for example, a ROM (Read Only Memory) or a partial area of the storage device 304 may be used.

[0047] The processor 301 expands the program stored in the program memory 303 into the RAM 302 and executes the program. When the object detection program is executed by the processor 301, the processor 301 is caused to perform a series of processes described with respect to the image processing unit 110. In other words, the processor 301 functions as the acquisition unit 111, the image feature generation unit 112, the language information generation unit 113, the fusion information generation unit 114, the object detection unit 115, and the output unit 116 according to the object detection program. When the learning program is executed by the processor 301, the processor 301 is caused to perform a series of processes described with respect to the object detection model generation unit 120. In other words, the processor 301 functions as the selection unit 121, the language information generation unit 122, the fusion information generation unit 123, the object detection unit 124, the update unit 125, and the output unit 126 according to the learning program.

[0048] The program may be provided to the object detection device 100 in a state stored in a computer-readable recording medium. In this case, the object detection device 100 includes a drive for reading data from the recording medium and acquires the program from the recording medium. Examples of the recording medium include magnetic disks, optical disks (such as CD-ROM, CD-R, DVD-ROM, DVD-R), magneto-optical disks (such as MO), and semiconductor memories. Further, the program may be distributed through a network. Specifically, the program may be stored in a server on the network and the object detection device 100 may download the program from the server.

[0049] The storage device 304 includes a non-volatile memory such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive). The storage device 304 stores data such as learning data and object detection models. The storage device 304 functions as the model storage unit 117, the learning data storage unit 127, and the model storage unit 128.

[0050] The input / output interface 305 is an interface for communicating with an external device. The input / output interface 305 may include a wireless module. Examples of external devices include a display device, a keyboard, a mouse, a router, etc. The processor 301 may transmit image data including an object detection result to the display device via the input / output interface 305. The processor 301 may also transmit the object detection result to another computer such as a server via the input / output interface 305 and a router.

[0051] [Operation] FIG. 4 schematically shows a learning method executed by the object detection model generation unit 120 of the object detection device 100.

[0052] In step S401 of FIG. 4, the selection unit 121 selects, as target data, data related to an arbitrary one reference image from the learning data set. Here, for simplicity of explanation, it is assumed that data related to one reference image is selected, but data related to a plurality of reference images may be selected.

[0053] In step S402, the language information generation unit 122 generates language information indicating the characteristics of the caption included in the target data. For example, the language information generation unit 122 extracts words from the caption, converts the words into word vectors using word2vec, and averages the word vectors to obtain language information that is a vector of length v.

[0054] In step S403, the fusion information generation unit 123 fuses the image features included in the target data and the language information generated by the language information generation unit 122 to generate fusion information. For example, the image features are a tensor based on c images of size w×h, and the fusion information generation unit 123 generates w×h copies of the vector that is the language information, generates a tensor by combining these copies, and superimposes the generated tensor and the tensor that is the image features in the channel direction to obtain the fusion information.

[0055] In step S404, the object detection unit 124 uses the object detection model stored in the model storage unit 128 to generate an object position estimation result and an object class estimation result regarding a reference image corresponding to the target data based on the fusion information generated by the fusion information generation unit 123. For example, the object detection unit 124 inputs the fusion information into the object detection model, and obtains the object position estimation result and the object class estimation result output from the object detection model as the object position estimation result and the object class estimation result regarding the reference image corresponding to the target data.

[0056] In step S405, the update unit 125 updates the object detection model based on the object position estimation result and the object class estimation result obtained by the object detection unit 124, as well as the correct object position information and the correct object class information included in the target data. For example, the update unit 125 updates the parameters of the object detection model based on the comparison between the object position estimation result and the correct object position information and the comparison between the object class estimation result and the correct object class information. For example, the update unit 125 updates the parameters of the object detection model so that the L1 distance between the object position estimation result and the correct object position information becomes smaller, and the cross-entropy error between the object class estimation result and the correct object class information becomes smaller.

[0057] In step S406, it is determined whether the learning end condition is satisfied. For example, when the L1 distance calculated in step S405 is less than a predetermined first threshold and the cross-entropy error calculated in step S405 is less than a predetermined second threshold, it is determined that the learning end condition is satisfied; otherwise, it is determined that the learning end condition is not satisfied. When the learning end condition is not satisfied (step S406; No), the process returns to step S401, and the processes shown in steps S401 to S405 are repeated.

[0058] When the learning end condition is satisfied (step S406; Yes), the process ends. The output unit 116 reads out the learned object detection model from the model storage unit 128 and outputs the learned object detection model to the image processing unit 110.

[0059] FIG. 5 schematically shows an object detection method executed by the image processing unit 110 of the object detection device 100. In the flow shown in FIG. 5, the object detection model generated by the learning method described with reference to FIG. 4 is used.

[0060] In step S501, the acquisition unit 111 acquires an input image and a caption that describes the input image.

[0061] In step S502, the image feature generation unit 112 generates an image feature indicating the feature of the input image. For example, the image feature generation unit 112 applies c filters to the input image to generate c feature maps each having a size of w×h, and the image feature is a tensor having w×h×c pixel values included in the c feature maps as elements.

[0062] In step S503, the language information generation unit 113 generates language information indicating the feature of the caption. For example, the language information generation unit 113 extracts words from the caption, converts the words into word vectors using word2vec, averages the word vectors, and thereby obtains a vector of length v as the language information.

[0063] In step S504, the fusion information generation unit 114 fuses the image feature and the language information to generate fusion information. For example, the fusion information generation unit 114 generates w×h copies of the vector that is the language information, generates a tensor obtained by combining the w×h copies, and superimposes the generated tensor and the tensor that is the image feature in the channel direction, thereby obtaining a tensor of size w×h×(c + v) as the fusion information.

[0064] In step S505, the object detection unit 115 uses the object detection model to generate an object position estimation result and an object class estimation result regarding the input image based on the fusion information. For example, the object detection unit 115 inputs the fusion information to the object detection model, and obtains the object position estimation result and the object class estimation result output from the object detection model as the object position estimation result and the object class estimation result regarding the input image.

[0065] In step S506, the output unit 116 outputs an object detection result based on the object position estimation result and the object class estimation result regarding the input image.

[0066] [Effect] The object detection model according to the present embodiment is a neural network configured to receive, as an input, language information indicating features of a caption that describes an image, together with image features indicating features of the image. Through supervised learning for the object detection model, an ability to infer important objects for understanding the situation in the image from the language information is implicitly acquired. Thereby, the object detection model according to the present embodiment enables more accurate detection of important objects for understanding the situation in the image. Specifically, it becomes possible to more reliably detect an object highly relevant to the situation depicted in the image, and to suppress detection of an object less relevant to the situation depicted in the image. That is, it becomes possible to prevent detection omission and over-detection.

[0067] The object detection model generation unit 120 uses a learning data set including a plurality of data, each data including image features generated from a reference image, a caption that describes the reference image, ground-truth object position information indicating the positions of objects existing in the reference image, and ground-truth object class information indicating the classes of objects existing in the reference image, to train the object detection model. Specifically, the object detection model generation unit 120 generates fusion information including image features indicating features of the reference image and language information indicating features of a caption that describes the reference image, inputs the fusion information into the object detection model, obtains an object position estimation result and an object class estimation result output from the object detection model, and updates the object detection model based on the object position estimation result, the obtained object class estimation result, the ground-truth object position information, and the ground-truth object class information. According to this configuration, it becomes possible to train a neural network that accurately detects important objects inferred from language information. The object detection model trained in this way enables more accurate detection of important objects for understanding the situation in the image from the image.

[0068] The image processing unit 110 generates fusion information including image features indicating the features of the input image and language information indicating the features of the caption that describes the input image, inputs the fusion information into the object detection model generated by the object detection model generation unit 120, obtains the position estimation result and the object class estimation result output from the object detection model, and outputs the object detection result based on the object position estimation result and the object class estimation result. According to this configuration, it is possible to detect more accurately the objects important for understanding the situation in the input image from the input image.

[0069] The image processing unit 110 may obtain language information by converting the words included in the caption into vectors using a predetermined vector space model such as word2vec. According to this configuration, it is possible to obtain information (numerical values) in a format that can be input to the object detection model and that indicates the features of the caption from the caption (sentence).

[0070] The image processing unit 110 applies c filters to the input image to generate c feature maps as image features, generates c copies of the language information, and generates fusion information including the c feature maps and the c copies of the language information. According to this configuration, language information is given to each channel (each feature map), and the language information is considered for each channel. Thereby, it is possible to detect more accurately the objects important for understanding the situation in the input image from the input image.

[0071] <Second Embodiment> [Configuration] FIG. 6 schematically shows an object detection apparatus 600 according to the second embodiment of the present invention. In FIG. 6, the same parts as those shown in FIG. 1 are denoted by the same reference numerals, and detailed descriptions thereof are omitted.

[0072] As shown in FIG. 6, the object detection device 600 includes an image processing unit 610 and an object detection model generation unit 620. The object detection model generation unit 620 generates an object detection model for detecting an object from an image by a machine learning method. The object detection model is implemented by a neural network architecture. The image processing unit 610 uses the object detection model generated by the object detection model generation unit 620 to detect an object from the image input to the object detection device 600.

[0073] The object detection model generation unit 620 includes a word importance map generation unit 621, a selection unit 121, an important word extraction unit 623, a language information generation unit 624, a fusion information generation unit 123, an object detection unit 124, an update unit 125, an output unit 126, a learning data storage unit 127, a word importance map storage unit 622, and a model storage unit 128.

[0074] The word importance map generation unit 621 generates a word importance map as word importance information associating a word with a word importance from the captions included in the learning data set stored in the learning data storage unit 127. For example, the word importance map generation unit 621 receives all the captions included in the learning data set from the learning data storage unit 127, calculates the word importance for each word from the received all captions, and generates information indicating the word importance for each word as a word importance map. All the captions included in the learning data set are a set of captions for each of all the reference images. The word importance may be any score as long as it is a score (value) indicating the importance of the word. In the present embodiment, the word importance is defined such that the higher the importance, the higher the score. For example, as the word importance, an idf score indicating the rarity in all the captions may be used. The idf score is a score calculated by the following formula (3).

Equation

[0075] Note that the word importance may be set manually in advance. For example, the word importance of a verb may be set to 0.5, the word importance of a proper noun may be set to 0.9, and the word importance of a common noun may be set to 0.7.

[0076] The word importance map generation unit 621 stores the word importance map in the word importance map storage unit 622 and passes the word importance map to the image processing unit 610.

[0077] The selection unit 121 randomly selects data from the learning dataset as target data. The selection unit 121 passes the caption included in the target data to the important word extraction unit 623, passes the image features included in the target data to the fusion information generation unit 123, and passes the correct object position information and correct object class information included in the target data to the update unit 125.

[0078] The important word extraction unit 623 extracts important words from the caption included in the target data based on the word importance map stored in the word importance map storage unit 622. For example, the important word extraction unit 623 extracts words from the caption and identifies the word importance of each word by referring to the word importance map. The important word extraction unit 623 selects at least one word from the words as an important word by applying the word importance of each word to a predetermined criterion. In one example, the important word extraction unit 623 may select the top 2 words with the highest word importance (the word with the highest word importance and the word with the second highest word importance) as the important words. In another example, the important word extraction unit 623 may select words whose word importance exceeds a threshold value (for example, 0.5) as the important words. If there are no words whose word importance exceeds the threshold value, the important word extraction unit 623 may select the word with the highest word importance as the important word. In any example, at least one important word extracted from the caption includes the word with the highest word importance.

[0079] The language information generation unit 624 generates language information from the important words extracted by the important word extraction unit 623. The language information can indicate the characteristics of the caption. The language information may be any information as long as it contains values generated from the important words. For example, the language information includes a plurality of numerical values that depend on the important words. In this embodiment, the language information is a vector. For example, the language information generation unit 624 uses a vector space model such as word2vec to convert individual important words into word vectors, and obtains a vector obtained by averaging or concatenating the word vectors as the language information. When one important word is extracted from the caption, the language information generation unit 624 obtains the word vector obtained by converting the important word as the language information.

[0080] The fusion information generation unit 123 fuses the image features included in the target data and the language information generated by the language information generation unit 624 to generate fusion information. The object detection unit 124 uses the object detection model stored in the model storage unit 128 to generate an object position estimation result and an object class estimation result regarding the reference image corresponding to the target data based on the fusion information generation unit 123. The update unit 125 updates the object detection model based on the object position estimation result and the object class estimation result obtained by the object detection unit 124, as well as the correct object position information and the correct object class information included in the target data. The output unit 126 outputs the learned object detection model to the image processing unit 610.

[0081] The image processing unit 610 includes an acquisition unit 111, an image feature generation unit 112, an important word extraction unit 611, a language information generation unit 612, a fusion information generation unit 114, an object detection unit 115, an output unit 116, a word importance map storage unit 613, and a model storage unit 117.

[0082] The word importance map storage unit 613 stores the word importance map generated by the object detection model generation unit 620. The model storage unit 117 stores the object detection model generated by the object detection model generation unit 620.

[0083] The acquisition unit 111 acquires an input image and a caption that describes the input image. The image feature generation unit 112 generates image features from the input image.

[0084] The important word extraction unit 611 and the language information generation unit 612 perform the same processing as described in relation to the important word extraction unit 623 and the language information generation unit 624. Therefore, a detailed description of the important word extraction unit 611 and the language information generation unit 612 is omitted. The important word extraction unit 611 extracts important words from the caption acquired by the acquisition unit 111 using the word importance map stored in the word importance map storage unit 613. The language information generation unit 612 generates language information from the important words extracted by the important word extraction unit 611.

[0085] The fusion information generation unit 114 fuses the image features generated by the image feature generation unit 112 and the language information generated by the language information generation unit 612 to generate fusion information. The fusion information includes image features and language information. In this embodiment, the fusion information is a tensor. For example, the fusion information generation unit 114 replicates a vector that is language information w×h times, generates a tensor obtained by combining the w×h replications, and superimposes the generated tensor and the tensor of the image features in the channel direction to obtain the fusion information. The fusion information is a tensor of size w×h×(c + v). Here, v is the length of the vector that is language information.

[0086] The object detection unit 115 uses the object detection model stored in the model storage unit 117 to obtain an object position estimation result and an object class estimation result regarding the input image based on the fusion information generated by the fusion information generation unit 114. The output unit 116 outputs an object detection result based on the object position estimation result and the object class estimation result regarding the input image obtained by the object detection unit 115.

[0087] The object detection device 600 can have the same hardware configuration as that shown in FIG. 3. Specifically, the object detection device 600 includes, as hardware components, a processor 301, a RAM 302, a program memory 303, a storage device 304, and an input / output interface 305.

[0088] The program memory 303 stores programs including an object detection program and a learning program. The processor 301 expands the programs stored in the program memory 303 into the RAM 302 and executes the programs. When the object detection program is executed by the processor 301, it causes the processor 301 to perform a series of processes described for the image processing unit 610. In other words, the processor 301 functions as an acquisition unit 111, an image feature generation unit 112, an important word extraction unit 611, a language information generation unit 612, a fusion information generation unit 114, an object detection unit 115, and an output unit 116 according to the object detection program. When the learning program is executed by the processor 301, it causes the processor 301 to perform a series of processes described for the object detection model generation unit 620. In other words, the processor 301 functions as a word importance map generation unit 621, a selection unit 121, an important word extraction unit 623, a language information generation unit 624, a fusion information generation unit 123, an object detection unit 124, an update unit 125, and an output unit 126 according to the learning program. The storage device 304 functions as a word importance map storage unit 613, a model storage unit 117, a learning data storage unit 127, a word importance map storage unit 622, and a model storage unit 128.

[0089] [Operation] FIG. 7 schematically shows a learning method executed by the object detection model generation unit 620 of the object detection device 600. Since the processes of steps S702, S705 to S708 shown in FIG. 7 are the same as the processes of steps S401, S403 to S406 shown in FIG. 4, detailed descriptions of the processes of steps S702, S705 to S708 are omitted.

[0090] In step S701 of FIG. 7, the word importance map generation unit 621 generates a word importance map associating words with word importance. For example, the word importance map generation unit 621 calculates the word importance (e.g., idf score) for each individual word included in all the captions included in the learning dataset stored in the learning data storage unit 127, and generates information indicating the word importance for each word as a word importance map.

[0091] In step S702, the selection unit 121 selects, as target data, data regarding at least one reference image from the learning dataset. Here, for simplicity of explanation, it is assumed that data regarding one reference image is selected.

[0092] In step S703, the important word extraction unit 623 extracts important words from the caption included in the target data. For example, the important word extraction unit 623 extracts words from the caption and specifies the word importance for each individual word with reference to the word importance map. For example, the important word extraction unit 623 selects, as important words, words whose word importance exceeds a predetermined threshold from among the words. Here, it is assumed that a plurality of important words are extracted.

[0093] In step S704, the language information generation unit 624 generates language information from the important words. For example, the language information generation unit 624 uses word2vec to convert the important words into word vectors, and obtains language information by averaging the word vectors.

[0094] In step S705, the fusion information generation unit 123 generates fusion information including the image features included in the target data and the language information generated by the language information generation unit 624.

[0095] In step S706, the object detection unit 124 uses the object detection model stored in the model storage unit 128 to generate an object position estimation result and an object class estimation result regarding the reference image based on the fusion information generated by the fusion information generation unit 123.

[0096] In step S707, the update unit 125 updates the object detection model based on the object position estimation result and object class estimation result obtained by the object detection unit 124, and the correct object position information and correct object class information included in the target data.

[0097] In step S708, it is determined whether the learning end condition is satisfied. If the learning end condition is not satisfied (step S708; No), the process returns to step S702, and the processes shown in steps S702 to S707 are repeated.

[0098] If the learning end condition is satisfied (step S708; Yes), the process ends. The output unit 116 reads out the learned object detection model from the model storage unit 128 and outputs the learned object detection model to the image processing unit 610.

[0099] FIG. 8 schematically shows an object detection method executed by the image processing unit 610 of the object detection apparatus 600. Since the processes of steps S801, S802, S805 to S807 shown in FIG. 8 are the same as the processes of steps S501, S502, S504 to S506 shown in FIG. 5, detailed descriptions of the processes of steps S801, S802, S805 to S807 are omitted.

[0100] In step S801, the acquisition unit 111 acquires an input image and a caption explaining the input image. In step S802, the image feature generation unit 112 generates image features from the input image.

[0101] In step S803, the important word extraction unit 611 extracts important words from the caption. For example, the important word extraction unit 611 extracts words from the caption and specifies the word importance for each individual word with reference to the word importance map. For example, the important word extraction unit 611 selects, as important words, words whose word importance exceeds a predetermined threshold from among the words. Here, it is assumed that a plurality of important words are extracted.

[0102] In step S804, the language information generation unit 612 generates language information from important words. For example, the language information generation unit 612 uses word2vec to convert important words into word vectors, and obtains language information by averaging the word vectors.

[0103] In step S805, the fusion information generation unit 114 generates fusion information including image features and language information. In step S806, the object detection unit 115 uses an object detection model to generate an object position estimation result and an object class estimation result regarding the input image based on the fusion information. In step S807, the output unit 116 outputs an object detection result based on the object position estimation result and the object class estimation result regarding the input image.

[0104] [Effect] The object detection device 600 can obtain the same effects as the object detection device 100 according to the first embodiment. The object detection device 600 extracts important words from a caption that describes an image, and generates language features from the important words. According to this configuration, language features that more accurately reflect the situation in the image are generated. Thereby, it becomes possible to detect more accurately the important objects for understanding the situation in the image from the image.

[0105] [Modification Example] In each of the above-described embodiments, the learning dataset includes image features generated from a reference image. The learning dataset may include a reference image instead of image features, and may be provided with an image feature generation unit that generates image features from the reference image in the object detection model generation unit.

[0106] Note that the present invention is not limited to the above-described embodiments, and various modifications can be made without departing from the gist thereof at the implementation stage. Also, the respective embodiments may be implemented in appropriate combination, and in that case, the combined effects can be obtained. Furthermore, the above-described embodiments include various inventions, and various inventions can be extracted by combinations selected from the plurality of disclosed components. For example, even if some components are deleted from all the components shown in the embodiments, if the problem can be solved and the effects can be obtained, the configuration from which these components are deleted can be extracted as an invention.

Explanation of Reference Numerals

[0107] 100…Object detection device 110…Image processing unit 111…Acquisition unit 112…Image feature generation unit 113…Language information generation unit 114…Fusion information generation unit 115…Object detection unit 116…Output unit 117…Model storage unit 120…Object detection model generation unit 121…Selection unit 122…Language information generation unit 123…Fusion information generation unit 124…Object detection unit 125…Update unit 126…Output unit 127…Learning data storage unit 128…Model storage unit 201…Main neural network 202…Object position estimation neural network 203…Object class classification neural network 301…Processor 302…RAM 303…Program memory 304…Storage device 305…Input / output interface 600…Object detection device 610…Image processing unit 611…Important word extraction section 612…Language information generation section 613…Word importance map storage section 620…Object detection model generation section 621…Word importance map generation section 622…Word importance map storage section 623…Important word extraction section 624…Language information generation section

Claims

1. An acquisition unit that acquires an image and a caption for explaining the image; An image feature generation unit that generates an image feature indicating a feature of the image; A language information generation unit that generates language information indicating a feature of the caption; A fusion information generation unit that generates fusion information including the image feature and the language information; Input the generated fusion information into an object detection model that is learned in advance to estimate the position and class of an object based on the fusion information by machine learning, and obtain an object position estimation result indicating a region where the object exists in the image and an object class estimation result indicating the probability that the object belongs to each class output from the object detection model; an object detection unit; An output unit that outputs an object detection result based on the object position estimation result and the object class estimation result; An object detection apparatus comprising:

2. The language information generation unit obtains the language information by converting words included in the caption into vectors using a predetermined vector space model. The object detection apparatus according to claim 1.

3. Further comprising an important word extraction unit that extracts at least one word including the word with the highest word importance from the caption using word importance information associating words with word importance as important words, The language information generation unit obtains the language information by converting the important words into vectors using the predetermined vector space model. The object detection apparatus according to claim 2.

4. The image feature generation unit applies c filters to the image to generate c feature maps, and the image feature includes a plurality of pixel values included in the c feature maps. The fusion information generation unit generates c copies of the language information, and generates the fusion information including the plurality of pixel values included in the c feature maps and the c copies of the language information. The object detection device according to any one of claims 1 to 3.

5. Using a learning data set including a plurality of data, each data including an image feature indicating a feature of a first image, a caption explaining the first image, ground truth object position information indicating a position of an object existing in the first image, and ground truth object class information indicating a class of the object existing in the first image, a learning device that learns an object detection model configured to receive, as an input, fusion information including an image feature indicating a feature of a second image and language information indicating a feature of a caption explaining the second image, and output an object position estimation result indicating a region where an object exists in the second image and an object class estimation result indicating a probability that the object in the second image belongs to each class, A language information generation unit that generates language information indicating a feature of a caption included in first data among the plurality of data; A fusion information generation unit that generates fusion information including the image feature included in the first data and the generated language information; An object detection unit that inputs the generated fusion information into the object detection model and obtains an object position estimation result and an object class estimation result output from the object detection model; An update unit that updates the object detection model based on the obtained object position estimation result, the obtained object class estimation result, the ground truth object position information, and the ground truth object class information included in the first data; A learning device comprising.

6. Obtaining an image and a caption explaining the image; Generating an image feature indicating a feature of the image; Generating language information indicating a feature of the caption; Generating fusion information including the image feature and the language information; Input the generated fusion information into an object detection model that is learned to estimate the position and class of an object based on the fusion information by machine learning and is pre-generated, and obtain an object position estimation result indicating the region where the object in the image exists and an object class estimation result indicating the probability that the object belongs to each class, which are output from the object detection model. Output an object detection result based on the object position estimation result and the object class estimation result. An object detection method comprising the above.

7. A learning method for learning an object detection model configured to receive, as an input, fusion information including image features indicating features of a second image and language information indicating features of a caption explaining the second image, using a learning dataset including a plurality of data each including image features indicating features of a first image, a caption explaining the first image, ground truth object position information indicating the position of an object existing in the first image, and ground truth object class information indicating the class of the object existing in the first image, and output an object position estimation result indicating the region where the object in the second image exists and an object class estimation result indicating the probability that the object in the second image belongs to each class, the method comprising: Generate language information indicating features of a caption included in first data among the plurality of data. Generate fusion information including the image features included in the first data and the generated language information. Input the generated fusion information into the object detection model and obtain an object position estimation result and an object class estimation result output from the object detection model. Update the object detection model based on the obtained object position estimation result, the obtained object class estimation result, the ground truth object position information, and the ground truth object class information included in the first data. A learning method comprising the above.

8. A program for causing a computer to function as the object detection device according to any one of claims 1 to 4 or the learning device according to claim 5.

Citation Information

Patent Citations

  • Subject word extraction device and program

    JP2016122398A

  • Image processing using artificial neural network

    JP2019220174A

  • Method, device, electronic apparatus, and computer storage medium for cross-modality processing

    JP2021163456A

  • Object detector and object detection method

    JP2021506017A

  • Learning method, learning program, and learning device

    WO2021084590A1