Object detection method and object detection system
The object detection method and system address the challenge of identifying damaged buildings in disaster scenarios by using machine learning inference with image and text inputs to generate composite images and determine object relationships, achieving effective damage assessment regardless of imaging direction.
Patent Information
- Application Number
- JP2023194078
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-27
AI Technical Summary
Existing object detection methods struggle to identify damaged buildings in disaster scenarios, especially when the building itself does not collapse, such as in floods, or when imaging angles are not directly downward, making it difficult to determine damage from aerial images.
An object detection method and system using machine learning inference with images and text inputs, where the system generates a composite image based on detected objects, extracts image and text feature amounts, and performs similarity-based inference to identify objects regardless of imaging direction.
Enables reliable identification of objects from their positional relationships with surrounding objects, independent of imaging direction, thereby improving the accuracy of disaster damage assessment.
Smart Images

Figure 2025080803000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an object detection method and an object detection system.
Background Art
[0002] In a disaster-stricken area, from the perspective of saving lives, quick grasp of the damage situation is required. When grasping the situation of the disaster-stricken area, it is ideal to identify the disaster sites based on images obtained by imaging the city from a high altitude using a high-resolution camera equipped on an artificial satellite or the like. However, the arrival cycle of an artificial satellite is at least one day even at the fastest, and it cannot quickly arrive directly above the disaster-stricken area, so it is not suitable for quickly grasping disasters. In addition, at the time of a disaster, there are problems such as clouds covering the ground due to bad weather, making it difficult to recognize with satellite images.
[0003] On the other hand, a UAV (Unmanned Air Vehicle) is suitable for quickly grasping the damage situation in that it can be immediately flown in a disaster-stricken area. In addition, since a UAV can fly below the clouds, it is easy to acquire information on the ground surface as an image.
[0004] Conventionally, there is known a technique aimed at extracting a building area from an image taken from above and identifying the damage situation by paying attention to the image texture inside the area of each building (see, for example, Patent Document 1). Patent Document 1 describes that "a damaged house is detected using an image taken from above of a disaster-stricken area and a house polygon acquired before the disaster. Color information corresponding to an impermeable sheet covering the roof of the damaged house is specified in advance. The image is regionally segmented based on color, and a sheet covering region having the color of the sheet is extracted from the segmented regions. A house corresponding to a house polygon having an overlapping portion with the sheet covering region is determined to be a damaged house."
[0005] According to this prior art, for example, in the case of a large-scale disaster where the building itself collapses, such as a major earthquake or a tornado, it is possible to identify (grasp) the damage situation of the building from the image texture inside the building area.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] By the way, when the disaster is, for example, a flood, the building itself does not collapse. Therefore, from the appearance seen from above, it is impossible to recognize whether the building has been damaged. Also, in cases where fallen trees fall into the building, there are many cases where the building itself does not seem to be damaged. To confirm whether the building has actually been damaged, it is necessary to clarify not only the area of the building but also the relationship with the surrounding area of the building and surrounding objects.
[0008] From the perspective of the relationship between the building and surrounding objects, it is considered effective to regularize the arrangement state of both. For example, from an image captured directly downward from above, the areas in the image are classified into classes such as buildings, water, and trees, and the impact of the disaster on the building is identified from the relative arrangement relationship of each area with respect to the building. However, in order to make this determination, it is necessary to capture the image directly downward. If this assumption breaks down, the relative arrangement relationship will be broken, so it is impossible to perform the disaster damage determination from an aerial image taken by an obliquely downward camera such as an aerial view point.
[0009] In order to quickly grasp the damage at the actual disaster site, the UAV (unmanned aerial vehicle) needs to enable recognition without restrictions on the angle of view. Therefore, it is necessary to perform a disaster damage determination without restricting the imaging angle. In the above, the building at the disaster site of the disaster has been described as an example of the object to be detected, but it is not limited to this.
[0010] The present invention has been made in view of such a situation, and an object detection method and an object detection system are provided that can identify an object to be detected from the positional relationship between the object to be detected and surrounding objects regardless of the imaging direction of the image.
Means for Solving the Problems
[0011] The object detection method of the present invention for solving the above problems is an object detection method using machine learning inference with an image and text corresponding thereto as inputs, generating an image based on the region inferred for the input image and the input text, extracting an image feature amount from the generated image, extracting a text feature amount from the input text, and performing machine learning inference based on the similarity between the image feature amount and the text feature amount to identify the object to be detected.
[0012] Further, the object detection system of the present invention for solving the above problems is an object detection system using machine learning inference with an image and text corresponding thereto as inputs, including a region division inference unit that infers a region corresponding to the input text for the input image, an image generation unit that generates an image based on the region inferred by the region division inference unit and the input text, an image feature amount extraction unit that extracts an image feature amount from the image generated by the image generation unit, and a text feature amount extraction unit that extracts a text feature amount from the input text, and performing machine learning inference based on the similarity between the image feature amount and the text feature amount to identify the object to be detected.
Effects of the Invention
[0013] According to the present invention, the object to be detected can be more reliably identified from the positional relationship between the object to be detected and surrounding objects regardless of the imaging direction of the image.
[0014] Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments for carrying out the invention (hereinafter referred to as embodiments).
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Modes for Carrying Out the Invention
[0016] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the embodiments described below are examples for explaining the present invention, and are not intended to limit the scope of the present invention only to those embodiments. Those skilled in the art can implement the present invention in various other modes without departing from the scope of the present invention.
[0017] In the configuration of the invention described below, the same reference numerals are commonly used for the same parts or parts having similar functions among different drawings, and redundant descriptions may be omitted. Also, the positions, sizes, shapes, and ranges of the respective configurations shown in the drawings and the like may not represent the actual positions, sizes, shapes, and ranges in order to facilitate the understanding of the present invention. For this reason, the present invention is not limited to the positions, sizes, shapes, and ranges disclosed in the drawings and the like. Also, components represented in the singular form in this specification may be plural as well, unless clearly indicated in the context.
[0018] <Summary of the Present Invention> The present invention relates to an object detection method and an object detection system for detecting (identifying) an object to be detected from the arrangement relationship between the object to be detected and surrounding objects regardless of the imaging direction of an image.
[0019] In recent years, VL (Vision Language) models that take images and text as inputs, typified by CLIP (Contrastive Language-Image Pre-training) proposed, are discriminators obtained as a result of learning large-scale image and text pairs. For example, in the case of CLIP, it holds an image encoder and a corresponding text encoder inside, extracts a group of image feature amounts from a group of input images, and at the same time, extracts a group of text feature amounts from the input text group respectively, and calculates the similarity between the two by an inner product or the like, so that it is possible to relatively evaluate which text is close to which input image.
[0020] Therefore, for example, for an image of a disaster-stricken area affected by landslides, it is possible to make a determination that the similarity of the former is high between the text "houses damaged by landslides" and the text "a serene suburban landscape". However, since the feature extraction performed by the image encoder of CLIP, which is used as an image feature extractor in many VL models, is the extraction of feature quantities for the entire image, there is a problem that it is unclear which specific part of the image corresponds to the mention in the text. Therefore, even if it is known that the captured image is a disaster-stricken image, there is a problem that it is unknown which part of the image is affected.
[0021] Therefore, in the present invention, in addition to this problem, instead of directly inputting an image to the VL model, after detecting each object (for example, a building) from the image, a composite image based on the detection result is superimposed on the original image and input. Specifically, for example, for an image in which a plurality of buildings are reflected, by inputting a composite image that highlights only one of the buildings, the aim is for the VL model to focus on that building and extract image feature quantities. Then, by comparing the image feature quantities extracted by focusing on a specific area with the text feature quantities, it is determined whether each building (house) has been affected. Here, although a building is exemplified as the object to be detected, the object to be detected is not limited to a building, and can be an artificial object represented by a building or a person in need of rescue.
[0022] <Overall System Configuration> FIG. 1 is a block diagram schematically showing a configuration example of an object detection system to which the object detection method of the present invention is applied. The object detection system shown in FIG. 1 is, as an example, a system for grasping the situation in a disaster-stricken area, and has a system configuration including a mobile body sensor unit 110 and an information processing device 120. The system (object detection system) for grasping the situation in a disaster-stricken area can be a disaster detection system.
[0023] [Mobile Body Sensor Unit] The mobile sensor unit 110 may be a mobile sensor that can acquire images from above, for example, in a disaster-stricken area, such as a drone or a satellite. The mobile sensor unit 110 has at least an imaging unit 111, a communication unit 112, a drive unit 113, and an attitude sensor 114, and these may be in a connection relationship with each other.
[0024] The imaging unit 111 may be a sensor such as a camera. The communication unit 112 has a function of being able to transmit the data acquired by the imaging unit 111 to the information processing device 120, and conversely, a function of being able to receive a signal transmitted from the information processing device 120. The drive unit 113 may be, for example, something like a motor, and has a function of driving the imaging unit 111 based on a signal received from the information processing device 120 via, for example, the communication unit 112. The attitude sensor 114 is a sensor such as a gyro. The data acquired by the attitude sensor 114 is transmitted to the information processing device 120 via the communication unit 112.
[0025] [Information Processing Device] The information processing device 120 can be configured by a general server. For example, as hardware, the information processing device 120 can be configured by a server including an input device, an output device, a processing device, and a storage device. In this embodiment, the processing device reads the program stored in the storage device and executes the read program, thereby realizing various functions while cooperating with other hardware as necessary.
[0026] Note that the information processing device 120 may be configured by a single server, or any of the input device, output device, processing device, and storage device may be configured by another computer system connected by a network.
[0027] Functionally, the information processing device 120 has, for example, a communication unit 121, a storage device 122, a processor 123, a display unit 124, and a control unit 125, and these may be in a connection relationship with each other.
[0028] The communication unit 121 has a function of receiving data and signals coming from the mobile sensor unit 110 and transmitting signals to the mobile sensor unit 110. The memory device 122 may be composed of, for example, a RAM (Random Access Memory). The processor 123 may be composed of, for example, an FPGA (Field Programmable Gate Array) or a GPU (Graphic Processing Unit).
[0029] The method for disaster detection described later may be realized on the memory device 122 and the processor 123. The display unit 124 may display, for example, an image acquired by the imaging unit 111 and received by the communication unit 121 from the mobile sensor unit 110, information of the attitude sensor 114, or a disaster area estimation result inferred by the object detection method for disasters described later based on these in a form such as a GUI (Graphical User Interface). The control unit 125 may control the drive unit 113 via, for example, the communication unit 121 and the communication unit 112.
[0030] <Example 1> Example 1 is an example of realizing, for example, an object detection method for disasters.
[0031] FIG. 2 is a block diagram showing an example of the functional configuration for realizing the object detection method for disasters according to Example 1 of the present invention. However, the technology of the present invention is not limited to a disaster-oriented system in terms of the application destination.
[0032] The functional configuration for realizing the object detection method for disasters according to Embodiment 1 may be realized on the memory device 122 and the processor 123 described above. As an example, the functional configuration includes an image text DB 201, an image preprocessing unit 202, a text preprocessing unit 203, a first image feature extraction unit 204, a noun phrase feature extraction unit 205, a region division inference unit 206, a layout generation unit 207, a text generation unit 208, a text feature extraction unit 209, an image generation unit 210, a second image feature extraction unit 211, an image text similarity calculation unit 212, and a region selection unit 213.
[0033] The image text DB 201 is a database including data composed of, for example, images and text groups corresponding thereto, and is used for learning the object detection method according to the present Embodiment 1. The images in the image text DB 201 may be, for example, known image data acquired on SNS such as the COCO (Common Objects in COntext) dataset. Also, the text groups corresponding to the images may be those in which descriptions are given for each object included in the images of the COCO dataset, such as the known refCOCO dataset. For example, if there are three people in the image, descriptions such as "a person looking at a computer", "a person sitting on a chair while looking at a computer", and "a passerby" may be given to each person according to the image. However, multiple descriptions may be given to one object, or there may be an object without a description.
[0034] This database is publicly disclosed as public information for solving the problem of REC (Referring Expression Comprehension). For a general object detection problem by a normal object detection method represented by, for example, YOLO (You Only Look Once), three people in the image may be detected as people. However, for the text "a person sitting on a chair while looking at a computer", it is a database for a conditional detection problem of detecting only the corresponding person.
[0035] The image text DB 201 is composed of, for example, for each sample, one image, an indefinite number of explanatory texts, and a teacher mask in binary mask form such as a segmentation mask, etc., so that the target object in the image corresponding to each text is clear.
[0036] The image preprocessing unit 202 performs, for example, mean-variance processing on pixel values for the image of the sample selected by the image text DB 201 for numerical calculation.
[0037] The text preprocessing unit 203 performs, for example, syntactic structure analysis such as known Dependency parsing and Constituency Parsing, and performs processing to extract noun phrases from the text, for example. Since the extracted noun phrases are the explanatory texts of the aforementioned images, they have the feature of being objects existing in the images. For example, noun phrases such as "computer, chair, person" may be extracted from the text "a person sitting on a chair while looking at a computer" mentioned above.
[0038] The first image feature extraction unit 204 takes the image processed by the image preprocessing unit 202 as input and performs processing to extract image feature amounts. As the first image feature extraction unit 204, for example, a feature extraction encoder for image recognition such as a known ViT (Vision Transformer) can be used.
[0039] As the noun phrase feature extraction unit 205, for example, a text encoder included in known BERT (Bidirectional Encoder Representations from Transformers) or CLIP (Contrastive Language-Image Pre-training) can be used. The noun phrase feature extraction unit 205 performs processing to extract noun phrase feature amounts from each of the noun phrases extracted from the text preprocessing unit 203.
[0040] The region segmentation inference unit 206 is an image recognition region segmentation model executed by a method such as the well-known MaskFormer. The region segmentation inference unit 206 performs a process of performing region segmentation on the image extracted by the first image feature extraction unit 204 and inferring a group of segmentation masks M. According to this region segmentation inference unit 206, for example, for object detection for disasters, it is possible to infer a region corresponding to the input text for the input image.
[0041] In recent years, as the region segmentation inference unit 206, for example, Open-Vocabulary format segmentation models such as X-Decoder and OpenSeeD have been proposed. The region segmentation inference unit 206 performs a process of outputting, as a group of segmentation masks M, the regions in the image corresponding to the noun phrase text feature amounts extracted by the noun phrase feature extraction unit 205.
[0042] The type of segmentation may be, in particular, instance segmentation that can be distinguished for each object, or output by panoptic segmentation. The region segmentation inference unit 206 preferably has a segmentation function that can accurately infer regions, but may also be rectangular detection such as YOLO (You Only Look Once), which is also object detection, and is not particularly limited.
[0043] The layout generation unit 207 performs a process of generating a group of layouts L to be input from the group of segmentation masks M to the image generation unit 210. Regarding the layout, during learning, when there is a subject for which there is no corresponding input text, the layout generation unit 207 generates text without using a predicate expressing a relationship using the subject and object corresponding to the region inferred by the region segmentation inference unit 206 from the input image, and uses it for conditioning. Details of this process will be described later.
[0044] The text generation unit 208 performs a process of generating text corresponding to each layout using, for example, the result of syntactic analysis performed by the text preprocessing unit 203 and the layout group L generated by the layout generation unit 207. Specific processes for generating this text will be described later.
[0045] The text feature extraction unit 209 performs a process of extracting text feature amounts from the text generated by the text generation unit 208. As the text feature extraction unit 209, the same text encoder used in the noun phrase feature extraction unit 205 can be used. According to this text feature extraction unit 209, for example, text feature amounts can be extracted from the input text for object detection for disasters.
[0046] The image generation unit 210 performs a process of generating an image group G from the input original image, the segmentation mask group M inferred by the region division inference unit 206, and the text feature amounts extracted by the text feature extraction unit 209. According to this image generation unit 210, for example, for object detection for disasters, a layout for image generation is created based on the input text and the output of the region division inference unit 206, and at least an image conditioned by the layout and the input text can be generated. The process of this image generation unit 210 will also be described later.
[0047] The second image feature extraction unit 211 may be an image encoder included in CLIP. It inputs the image group G generated by the image generation unit 210 and performs a feature extraction process. According to this second image feature extraction unit 211, for example, for object detection for disasters, image feature amounts can be extracted from the images generated by the image generation unit 210.
[0048] The image-text similarity calculation unit 212 performs a process of calculating the similarity of feature amounts between heterogeneous data extracted by the text feature extraction unit 209 and the second image feature extraction unit 211 using, for example, cosine similarity.
[0049] The region selection unit 213 performs a process of selecting, from the image group G, the image most similar to each text constituting the text group based on the scores of the image group G and the text group calculated by the image-text similarity calculation unit 212. Then, the region selection unit 213 selects, from the segmentation mask group M of the region division inference unit 206, the region that contributed to the generation of the selected generated image as the detection target, and uses it as the final output of object detection. That is, the region selection unit 213 selects the layout that was the generation source by selecting the image closest to the input text from the images generated based on the layout, and uses the region of the main body constituting the selected layout as the final output of object detection.
[0050] According to the object detection system having each of the above-described functional units, machine learning inference can be performed based on the similarity between the image feature amount and the text feature amount to detect (identify) the object to be detected.
[0051] [Functional configuration example of the image generation unit] FIG. 3 is a block diagram schematically showing an example of the functional configuration of the image generation unit 210 in FIG. 2. As shown in FIG. 3, the image generation unit 210 has a configuration including functional units such as an image feature extraction unit 301, a conditional feature extraction unit 302, an image restoration unit 303, and an image synthesis unit 304, for example.
[0052] The image feature extraction unit 301 includes, for example, a linear layer, and performs a process of extracting feature amounts for inputting feature amounts to the subsequent conditional feature extraction unit 302. The image feature extraction unit 301 may have the same configuration as the first image feature extraction unit 204.
[0053] The conditional feature extraction unit 302 may include, for example, an Attention mechanism in deep learning, and conditions the image feature amount by the text feature amount using, for example, a Cross Attention mechanism, from the image feature amount extracted by the image feature extraction unit 301 and the text feature amount extracted from the text generated by the text feature extraction unit 209, for example.
[0054] In addition, the conditional feature extraction unit 302 may be provided with a normalization mechanism such as AdaIN (Adaptive Instance Normalization), and may perform conditional feature extraction on the image feature amount based on the mask information inferred by the region division inference unit 206. In particular, when conditioning the entire feature amount, a configuration such as AdaIN may be used, and when conditioning for each location of the feature amount, a Cross Attention structure may be used.
[0055] The image restoration unit 303 corresponds to the decoder part in a structure such as U-net, for example, and may include a CNN (Convolutional Neural network) structure and an UpSampling structure. The image restoration unit 303 performs a process of restoring the feature amount extracted from the conditional feature extraction unit 302 into an RGB image so as to be, for example, a 3ch image, and directs it to the image synthesis unit 304.
[0056] The image synthesis unit 304 may perform an addition process, for example, in a superimposing sense, on the RGB image restored by the image restoration unit 303 and the original RGB image corresponding to the image output from the image preprocessing unit 202, and may synthesize the images in a form such as alpha blending. Alpha blending is an example, and other methods may be used, but it is required to be an implementation that can be learned in a form capable of at least backpropagation of the error described later.
[0057] [Example of processing of image generation unit] FIG. 4 is a flowchart showing an example of the processing executed by the image generation unit 210 in FIG. 2.
[0058] In step S401, if it is during learning, the image extracted by the image text DB 201 is used as the input image X, and if it is during inference, the actually captured image is used as the input image X. After performing processes such as normalization by the image preprocessing unit 202, a process of extracting the image feature amount V by the image feature extraction unit 301 is performed.
[0059] In step S402, the layout generation unit 207 creates a layout group L from the segmentation mask group M obtained from the region division inference unit 206, the text generation unit 208 generates a text group T which is the text corresponding to each layout, and the text feature extraction unit 209 extracts the feature amount extracted from the text group T.
[0060] Under the process of extracting this feature amount, for the conditional feature, a process of extracting the conditioned image feature amount V' is performed based on the feature amounts extracted from the layout group L and the text group T. When extracting features from the layout group L, for example, a CNN may be used, and when conditioning, for example, a method such as SPADE (Spatially-Adaptive Normalization) may be used.
[0061] Regarding the specific process of the layout generation unit 207 that creates the layout group L from the segmentation mask group M and also generates the text corresponding to the layout in the process of step S402, it will be described later.
[0062] In step S403, the image restoration unit 303 performs a process of generating the image Y from the image feature amount V'. The image feature extraction unit 301, the conditional feature extraction unit 302, and the image restoration unit 303 that constitute the image generation unit 210 may have a structure such as a U-net, for example, and may have a configuration similar to that used in an image generation model such as Stable Diffusion. The order of conditioning and the order of processing described in this embodiment are examples and are not particularly limited.
[0063] The image Y generated from the image feature amount V' is not necessarily a meaningful image when viewed by humans. In step S404, a process is performed in which the image Y generated from the image feature amount V' is superimposed on the input image X by an addition process such as alpha blending to generate an image group G. Note that the image Y itself may be used as the image group G as it is. The image group G is expected to be an image that contributes to feature extraction for a specific region in the image during feature amount extraction by the second image feature extraction unit 211 in the subsequent stage, and may be an image referred to by a technique called Visual Prompting in recent years.
[0064] [Example of preprocessing information input to the image generation unit] FIG. 5 is a flowchart showing an example of a specific process for preprocessing information input to the image generation unit 210 of FIG. 2. Specifically, in FIG. 5, in step S402 of FIG. 4, an example of a processing flow during learning in which the layout generation unit 207 and the text generation unit 208 generate a layout group L and a text group T input to the image generation unit 210 is shown.
[0065] That is, the series of processes in FIG. 5 are processes by the image generation unit 210. The image generation unit 210 first performs a process of explicitly separating the subject and the object from the input text and then creating a layout based on the output of the region division inference unit 206. Also, when the region division inference unit 206 infers regions corresponding to a plurality of subjects from the input image, the image generation unit 210 generates a layout for each subject and performs a process of expressing the region representing the object and the regions for each subject in a separable form. This process corresponds to the processes in steps S520 to S522 described later.
[0066] Hereinafter, the process executed by the image generation unit 210 will be specifically described. FIG. 6 will be appropriately used in the description of FIG. 5. FIG. 6 is a diagram showing a specific example for explaining the preprocessing in FIG. 5.
[0067] The segmentation mask group M is inferred by the region segmentation inference unit 206 from the images sampled by the image text DB 201. Here, note that due to the characteristics of the image text DB 201, the above-mentioned image is associated with the accompanying explanatory text and the teacher information mask indicating the region corresponding to the explanatory text. For example, in the example of FIG. 6, there are two text groups 602 corresponding to the input image 601, namely (a) and (b). Here, the region targeting a person indicated by each of them is included in the image text DB 201 as the teacher information mask.
[0068] In step S501, based on the similarity calculation between the image feature amounts corresponding to the segmentation mask regions inferred inside the region segmentation inference unit 206 and the noun phrase feature amounts extracted by the noun phrase feature extraction unit 205 based on the noun phrases extracted, for example, by syntactic analysis processing in the text preprocessing unit 203, for the noun class group defined in advance or the text accompanying the input image, the association process of which noun each mask constituting the segmentation mask group M corresponds to is performed.
[0069] This association process is a known process performed after giving a noun phrase group in the above-mentioned Open-Vocabulary form segmentation method, and it can be used. For example, in FIG. 6, the text group 602 corresponding to the input image 601 is given. During learning, due to the nature of the image text DB 201, these are paired with each other, but during testing, for the input image 601, the user may input the explanatory text of the target to be detected as text.
[0070] In the example of FIG. 6, for the text group 602 of the explanatory text, from the result of syntactic analysis performed by the text preprocessing unit 203, a segmentation mask corresponding to a noun phrase corresponding to the subject of the text is selected. For example, in syntactic analysis, for texts such as (a) and (b) that make up the text group 602, noun phrases such as "person, laptop, chair" can be extracted, and the region division inference unit 206 extracts regions corresponding to each noun phrase, for example, masks 610, 611, 612, 613, 614 as a segmentation mask group M. Then, masks 610, 611, 612 corresponding to "person", mask 614 corresponding to "laptop", and mask 813 corresponding to "chair" can be detected. In this example, since the word "suitcase" is not included in the text group 602, the mask 615 corresponding to "suitcase" cannot be detected, but this does not necessarily need to be detected.
[0071] In step S502, from the syntactic analysis performed in the process of step S501, when the text is particularly in English, due to the nature that the subject comes before the text, it is possible to determine that "person" is the subject and "laptop, chair" are the objects. Therefore, among the segmentation mask group M, masks 610, 611, 612 corresponding to "person" are used as the subject mask candidate group S.
[0072] In step S503, a process of selecting a mask group O corresponding to the object of the input text is performed. In the above example, "laptop" and "chair" are the objects, and among the segmentation mask group M, the corresponding segmentation masks 613, 614 are used as the object mask group O.
[0073] In step S504, a process of creating a layout group L from the subject candidate mask group S and the object mask group O is performed. In the example of FIG. 6, for example, the subject candidate mask group S consists of three candidates, and the object mask group O consists of two elements. At this time, as an example, the layout group L may be created according to a rule of selecting one from the subject candidate mask group S and selecting all from the object mask group O. Under this rule, the object mask group 0 may be grouped together as one mask in the form of, for example, the object mask group 616.
[0074] Also, the object mask group 616 is combined with the components of each subject candidate mask group S to create layout masks 620, 621, and 622, and these three may be used as the layout group L. The layout in this embodiment needs to be created in a form where the subject and object can be distinguished. For example, if the layout L is a 2-channel image, the area corresponding to the subject candidate mask group S may be described with the value 1 in the first channel, the area corresponding to the object mask group O may be described with the value 1 in the second channel, and the value 0 may be stored in the other areas.
[0075] In the above example of creating a layout, the layout is created in a distinguishable form for each subject candidate, but since all the objects are grouped together, it is impossible to distinguish between the objects. Although the layout may be created in the form of distinguishable segmentation masks 613 and 614 without grouping the objects in the form of the object mask group 616, a combinatorial explosion will occur as the number of combination candidates increases. Therefore, in this embodiment, the above combination is used.
[0076] In step S505, during learning, a process of associating text with the layout group L created in the process of step S504 is performed. In the image text DB 201, the main area of the image corresponding to the text is given as teacher information. Therefore, since the subject candidates corresponding to the text are clear during learning by comparing the area of the subject candidates with the area given as teacher information, text may be associated with the layout corresponding to the teacher information. In the previous example, the text group 602(a) may be associated with the layout mask 620, and the text group 602(b) may be associated with the layout mask 621.
[0077] In step S506, a process of checking whether there is a layout in the layout group for which text has not been associated in the process of step S505 is performed. In the image text DB 201, corresponding text is not necessarily prepared for all subject candidates.
[0078] In step S507, for the layout for which text has not been associated in the process of step S506, the text generation unit 208 performs a process of generating and associating text. In the previous example, since no text is associated with the layout mask 622, it is necessary to generate this. When generating text, for example, noun phrases corresponding to the masks included in the object mask group 616 constituting the layout mask 622 may be used. For example, the generated text may be "person, laptop and chair".
[0079] Since the purpose of this embodiment is to recognize the relationship between objects, for a layout in which no relationship exists, vocabulary indicating the relationship is not set. In this text generation example, unlike the layout masks 620 and 621, verbs such as "watching, sitting on" that represent the relationship between the subject and the object are not included, and the layout mask 622 can be identified as different from other layouts in the subsequent functional configuration.
[0080] In step S508, a layout is created using only the object mask group O, added to the layout group L, and the corresponding text is created. For example, since there is no subject candidate in the object mask group 616, this can be used as it is, and the text generation unit 208 may set the text corresponding to this layout to, for example, "chair and laptop". This corresponds to data generation for making such a determination when there is no subject candidate corresponding to the input text during inference. That is, here, regarding the creation of the layout, even when there is no subject in the image area, a layout is generated using only the objects of the input text, and when the subject of the input text does not exist in the input image during inference, a process is performed to determine that there is no subject and no detection target exists.
[0081] In the processing during testing, the most appropriate one among the layout group L is selected for the input text. However, if the one corresponding to the subject of the text input during testing does not exist in the image, the layout group L becomes an empty set. To avoid this, the object mask group 616 is deliberately added in step S508.
[0082] In step S509, a process is performed to make the generated layout group and the text group appended so far in one-to-one correspondence. In FIG. 6, finally, (object mask group 616, text 633), (layout mask 620, text 630), (layout mask 621, text 631), (layout mask 622, text 632) are in the form of one-to-one correspondence. Then, the image generation unit 210 generates an image conditioned by each layout and text. The image generated here does not necessarily have a meaning visually understandable to humans.
[0083] Here, an example of learning the image generation unit 210 is described, but the image generation unit 210 is not limited to being obtained by learning. In an extreme example, for instance, the text generation unit 208 may interpret parts other than the layout as the background and superimpose a dark image on the original image with respect to the background. In this case, by making the background area difficult to see for the subsequent image discriminator, it is considered to have the effect of relatively emphasizing the foreground part corresponding to the layout and inducing recognition.
[0084] [Example of Machine Learning Inference Processing] FIG. 7 is a flowchart showing an example of a process for performing machine learning inference processing on the image generated by the image generation unit 210 in FIG. 2. In this process, a process of calculating the similarity between the image feature amount and the text feature amount is performed on the image group G generated by the image generation unit 210.
[0085] In step S701, first, the second image feature extraction unit 211 performs a process of extracting the image feature amount V' from each of the image groups G. On the other hand, the text feature extraction unit 209 extracts the text feature amount from each text for the text generated by the text generation unit 208. During learning, the image group G is generated by the image generation unit 210 under the condition of the text group corresponding to the layout group L generated by the layout generation unit 207. Therefore, each generated image corresponds to the text and layout used for the conditioning.
[0086] In step S702, a process of calculating the similarity is performed between the image feature amount V' and the text feature amount. For example, if the second image feature extraction unit 211 and the text feature extraction unit 209 are learned by CLIP or the like, the similarity of both feature amounts can be calculated by the inner product, and the similarity can be expressed as a score.
[0087] When this process is a learning process, in step S703, using the correspondence between the correct text and the image as teacher information, for the score calculated in the process of step S702, a loss calculation process is performed using, for example, the cross entropy function. The loss calculated here may update the parameters of the modules having learnable parameters among the functional configurations shown in FIG. 2 by the error backpropagation method. For example, the learnable parameters included in the conditional feature extraction unit 302 and the image restoration unit 303 that constitute the image generation unit 210 may be updated.
[0088] On the other hand, since the second image feature extraction unit 211, the noun phrase feature extraction unit 205, and the text feature extraction unit 209 are learned with large-scale data such as CLIP, it can be said that there is an advantage in using the parameters as they are. Therefore, the parameters may be fixed and not updated. On the other hand, although the parameters constituting the first image feature extraction unit 204 and the region division inference unit 206 are originally sufficiently learned, the parameters may be fixed. However, depending on the newly introduced learning data, the vocabulary and region division may not be supported. Therefore, as needed, they may be targeted for error backpropagation and the parameters may be updated.
[0089] When it is an inference process, in step S704, a process is performed in which the region selection unit 213 selects the correct answer from the similarity scores. For example, in the text example of "person watching a laptop, sitting on a chair" in FIG. 6, since the similarity score for the image generated by the layout mask 621 corresponding to this text is larger than that from the layout mask 620, the subject candidate mask 611 that contributed to the generation of the layout mask 621 may be used as the correct answer.
[0090] On the one hand, if there is no subject candidate corresponding to the text in the input image, the region division inference unit 206 cannot detect the subject candidate. Therefore, only those derived from the object mask group 616 exist in the layout group L, and it is determined that the similarity is the highest for the input text. However, since there is no corresponding subject mask, it may be determined that there is no object to be detected. On the other hand, when it is desired to detect the subject corresponding to "person watching a laptop", in the example of FIG. 6, it is necessary to detect two corresponding to the layout masks 620 and 621.
[0091] In this way, when it is necessary to detect a plurality of objects, for example, for the generated image derived from the layout mask 620, it is determined which of the two texts, "person watching a laptop" and "person and laptop", has a higher similarity. If it is closer to the former, it is set as the detection target. This process is also performed for the image group G derived from the layout mask 621. By this determination method, it is possible to handle cases where there are a plurality of detection targets. This is particularly contributed by designing a pair of layout and text that recognizes the relationship during learning. Since the image group G derived from the layout mask 622 is learned without giving a verb indicating a relationship such as "watching", like "person and laptop" or "person, laptop and chair", a difference appears as the similarity depending on the presence or absence of "watching" indicating the relationship in the input text during testing.
[0092] <Example 2> Example 2 is an example of a learning method for building identification in a disaster detection application.
[0093] In the description of Example 1, for example, the data sample with "person" as the subject was used for the image text DB 201, but the technology of the present invention does not limit the target to a specific object.
[0094] FIG. 8 is a diagram showing an example of a sample when, as the image text DB 201, for example, an aerial photograph 801 of a disaster-stricken area and a text group 802 as its explanatory text are collected. At the same time as collecting the text group 802, a teacher mask for the area corresponding to each text is prepared. Since the subject is "house" and the objects are "tree, flood" in these data, as the layout group L taking into account the subject and objects, layouts 810, 811, and 812 can be created by the same method as described in FIG. 6, and text groups 820, 821, and 822 corresponding to the respective layouts can be prepared in association with them.
[0095] Therefore, it becomes possible to learn building detection based on, for example, the relationship between a house and other objects for building detection in a disaster-stricken area by the same learning method as in the first embodiment. However, the subject is not limited to a house and may be a car or a person. As a business need for disaster detection, conditional detection such as "a person trapped in rubble" rather than just a person may be required, and machine learning inference to address this is possible if data can be prepared.
[0096] In the first and second embodiments, the data samples used in the description are different, but it is also possible to combine them into the image text DB 201. The object detection method of the present invention can be applied without being limited to daily scenes such as in the first embodiment or disaster scenes such as in the second embodiment according to the learned data. Also, in operation, it is necessary to design the text of the target to be detected. When it is desired to detect an object limited by a vocabulary indicating a relationship such as "house damaged by tree", it is necessary to deliberately prepare a text consisting only of noun phrases such as "house" or "house and tree". This corresponds to the point that the relationship is explicitly learned in step S507 of FIG. 5 during learning.
[0097] In particular, especially during actual operation, it is also possible to prepare multiple texts for the object to be detected. For disaster detection, by preparing a text group that extends the text group 802 in FIG. 8, it becomes possible to detect houses and, for example, people in various disaster patterns limited by the text. Specifically, when making an inference for object detection, create a list that limits the object to be detected by at least one of an adjective and a vocabulary indicating a relationship. Then, detect the object corresponding to the multiple limited nouns for object detection.
[0098] [Examples that can be determined without an object] FIG. 9 is a diagram for explaining that the object detection method of the present invention can also be applied to a text input without an object.
[0099] For example, in the case of disaster detection, when there are an input image 900 and an input text 910, extract "house" corresponding to a noun from the input text 910, and create a layout group L from only the subject mask groups, namely masks 901 and 902, without using the object mask group, in the same process as in FIG. 6. By using the similarity between the image feature amount extracted from the image group G generated by the image generation unit 210 from the layout group L and the feature amount extracted from, for example, "damaged house" and "house", object detection with the attribute of "damaged" can be achieved through the learning process in step S704 and the inference process in step S705. The detection target is not limited to houses. The same machine learning inference can be performed in the same way even when, for example, for an object like a running person as in image 920 and text 930, instead of a relationship, an object containing an intransitive verb is to be identified.
[0100] [Correspondence to special domains] In the object detection method of the present invention, it can be applied to detecting objects such as damaged houses and artificial objects in disasters. Generally, if it is a business operator, it is possible to create an image text DB 201 for disaster recognition. However, there are limitations in the existing image data in the disaster area, especially, and there is a problem that the domain is limited because the imaging locations are biased towards the disaster area. Although it is natural to aim to enhance the image text DB 201 by newly collecting data, it is possible to expand the data using generative AI represented by, for example, ChatGPT and Stable Diffusion.
[0101] The method of expansion will be described using the flowchart of FIG. 10. FIG. 10 is a flowchart showing an example of a process for expanding learning data for realizing an object detection method (disaster detection method) for disasters according to Example 2 of the present invention.
[0102] In step S1001, a process of collecting text indicating the nominative relationship regarding the disaster to be detected is performed. Regarding this text, it may be created by humans based on past disaster knowledge. Specifically, without being limited to the house area, text in a form that explicitly shows the nominative relationship regarding the relationship with surrounding objects and areas may be created, such as "house damaged by a fallen tree" or "house surrounded by the flood".
[0103] Here, a sample targeting a house is exemplified, but the subject may also be a car or a person, and is not limited to a house. Also, such text may be exemplified and the same text may be created in a language generation model represented by ChatGPT. The text list created by the same process may be visually confirmed by a human again for filtering.
[0104] In step S1002, for each list of texts created in the process of step S1001, a model that generates images from texts, such as Stable Diffusion, may be used to generate disaster-related images. Regarding the generated group of images, for example, after visual inspection, inappropriate images may be deleted.
[0105] In step S1003, in the process of the generation process in step S1002, attention is paid to, for example, the Cross Attention module that performs text-based conditioning when generating an image in the image generation model, and a process is performed to extract, for example, where a noun phrase such as "house" in the text is generated in the image. In this extraction process, a matrix representing the correlation between the feature amounts for each position of the image and the feature amounts for each token (token) constituting the text may be used, and extraction may be performed from the weight values of the matrix. For this process, for example, a technique called Prompt to Prompt can be used. By this process, an image region corresponding to the noun phrase constituting the text can be obtained. This is known to correspond to a mask indicating the position of the noun phrase.
[0106] Through the processes of steps S1001 - S1003 above, texts, images, and masks as teacher information can be obtained. Since this satisfies the specifications of the elements constituting the image text DB201, for example, after creating a database specialized for a special domain such as a disaster, by applying the learning method according to this embodiment, it becomes possible to learn a disaster detector.
[0107] As described above, for the learning data for realizing object detection, at least the conditions of the detection target are generated as a group of texts where the subject and object are clear, and based on the assumption of image generation from the group of texts, the generated image, the input image, and the input text, the region corresponding to the noun phrase constituting the group of texts is saved, so that learning data for learning machine learning inference can be generated.
Explanation of Signs
[0108] 110… Mobile body sensor unit, 111… Imaging unit, 112… Communication unit, 113… Driving unit, 120… Information processing device, 121… Communication unit, 122… Memory device, 123… Processor, 124… Display unit, 125… Control unit, 201… Image text DB, 202… Image preprocessing unit, 203… Text preprocessing unit, 204… First image feature extraction unit, 205… Noun phrase feature extraction unit, 206… Region division inference unit, 207… Layout generation unit, 208… Text generation unit, 209… Text feature extraction unit, 210… Image generation unit, 211… Second image extraction unit, 212… Image text similarity calculation unit, 213… Region selection unit
Claims
1. An object detection method using machine learning inference with an image and corresponding text as inputs, comprising: generating an image based on the region inferred for the input image and the input text; extracting image feature quantities from the generated image and extracting text feature quantities from the input text; performing the machine learning inference based on the similarity between the image feature quantities and the text feature quantities to identify the object to be detected. An object detection method characterized by the above.
2. In the machine learning inference, perform text recognition of noun phrases and the relationship between the subject and object for the input text, perform region division on the input image based on the recognized noun phrases, generate a plurality of images based on the relationship between the subject and object obtained by the text recognition and the result of the region division, and identify the object to be detected based on the similarity between each generated image and the input text. The object detection method according to claim 1, characterized by the above.
3. The object to be detected is an artificial object or a person in need of rescue represented by a building in a disaster-stricken area. The object detection method according to claim 1, characterized by the above.
4. At the time of inference of the object detection, create a list that limits the object to be detected by at least one of an adjective and a vocabulary indicating a relationship, and detect an object corresponding to a plurality of limited nouns. The object detection method according to claim 1, characterized by the above.
5. Regarding the learning data for realizing the object detection, at least generate the conditions of the object to be detected as a text group in which the subject and object are clear, and save the region corresponding to the noun phrase constituting the text group in the process of generating an image based on the image generated from the text group, the input image, and the input text, thereby generating learning data for learning the machine learning inference. The object detection method according to claim 1, characterized by the above.
6. An object detection system using machine learning inference with an image and corresponding text as inputs, comprising: a region division inference unit that infers a region corresponding to the input text for the input image; an image generation unit that generates an image based on the region inferred by the region division inference unit and the input text; an image feature extraction unit that extracts image feature quantities from the image generated by the image generation unit; a text feature extraction unit that extracts text feature quantities from the input text. Comprising, and performing the machine learning inference based on the similarity between the image feature amount and the text feature amount to identify the object to be detected An object detection system characterized by the above.
7. The image generation unit creates a layout for image generation based on the input text and the output of the region division inference unit, and generates at least an image conditioned by the layout and the input text The object detection system according to claim 6, characterized by the above.
8. The image generation unit explicitly separates the subject and the object from the input text, and then creates the layout based on the output of the region division inference unit The object detection system according to claim 7, characterized by the above.
9. When the region division inference unit infers regions corresponding to a plurality of subjects from the input image, the image generation unit generates the layout for each subject, and expresses the region representing the object and the regions for each subject in a separable form The object detection system according to claim 7, characterized by the above.
10. A region selection unit is further provided, which selects the image closest to the input text from the images generated based on the layout, thereby selecting the layout that is the generation source, and using the region of the subject constituting the selected layout as the output region for object detection The object detection system according to claim 7, characterized by the above.
11. Regarding the layout, during learning, when there is a subject for which the corresponding input text does not exist, text is generated using the subject and object corresponding to the region inferred by the region division inference unit from the input image, without using a predicate expressing a relationship, and a text generation unit for use in the conditioning is further provided The object detection system according to claim 7, characterized by the above.
12. The text generation unit generates a layout using only the object of the input text even when there is no subject in the image region regarding the creation of the layout, and when the subject of the input text does not exist in the input image during inference, determines that there is no subject and no object to be detected The object detection system according to claim 11, characterized by the above.
13. Comprising a mobile body sensor having an imaging unit Taking an image of the disaster-stricken area from above by the imaging unit as the input image, and detecting artificial objects represented by buildings and disaster victims in the disaster-stricken area as the objects to be detected The object detection system according to claim 6, characterized by the above.
Citation Information
Patent Citations
Damaged house detection method
JP2017220175A