Object detection method and object detection system
The object detection method addresses the challenge of identifying damaged buildings in disaster scenarios by using machine learning inference with image and text inputs to generate composite images, allowing for accurate detection regardless of image capture angles.
Patent Information
- Application Number
- PCT/JP2024/039330
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-15
- Filing Date
- 2024-11-05
- Publication Date
- 2025-05-22
AI Technical Summary
Existing object detection methods struggle to accurately identify damaged buildings in disaster scenarios, especially when the building itself does not collapse, and when images are captured from non-perpendicular angles.
The proposed object detection method uses machine learning inference with images and corresponding text as input, generating a composite image based on inferred areas to extract image features and text features, allowing for object detection regardless of the image capture direction.
This method enables more reliable identification of damaged buildings by analyzing the positional relationship between the building and surrounding objects, regardless of the image capture angle, thus facilitating rapid damage assessment in disaster situations.
Smart Images

Figure JP2024039330_22052025_PF_FP_ABST
Abstract
Description
Object detection method and object detection system
[0001] The present invention relates to an object detection method and an object detection system.
[0002] In disaster-stricken areas, a rapid assessment of the damage situation is essential from the perspective of saving lives. The ideal way to assess the situation in disaster-stricken areas would be to identify the affected areas based on images obtained by capturing images of the city from high altitudes using high-resolution cameras mounted on satellites. However, satellites arrive at an arrival cycle of at most one day, meaning they cannot arrive directly above the affected area quickly, making them unsuitable for rapid assessment of the damage. Furthermore, during disasters, factors such as clouds obscuring the ground due to bad weather often make it difficult to identify areas using satellite images.
[0003] On the other hand, UAVs (Unmanned Air Vehicles) are suitable for quickly assessing the damage situation because they can be flown immediately in disaster areas. UAVs can also fly below cloud cover, making it easy to obtain images of the ground surface.
[0004] A technique has been known in the past that extracts building regions from an image captured from the air and identifies the state of damage by focusing on the image texture inside each building region (see, for example, Patent Document 1). Patent Document 1 describes the following: "Damaged houses are detected using an image captured from the air of a disaster-stricken area and house polygons acquired before the disaster. Color information corresponding to an impermeable sheet covering the roof of a damaged house is specified in advance. The image is divided into regions based on color, and sheet-covered regions having the color of the sheet are extracted from the divided regions. Houses corresponding to house polygons that have an overlapping portion with the sheet-covered region are determined to be damaged houses."
[0005] According to this conventional technology, in the event of a large-scale disaster such as a major earthquake or tornado that causes buildings to collapse, it is possible to identify (understand) the damage status of the building from the image texture inside the building area.
[0006] Japanese Patent Application Laid-Open No. 2017-220175
[0007] However, when a disaster occurs, such as a flood, the building itself does not collapse, so it is not possible to determine whether or not a building has been damaged simply by looking at it from above. Even when a fallen tree falls on a building, the building itself often appears undamaged. Therefore, in order to confirm whether a building has actually been damaged, it is necessary to clarify not only the area of the building, but also its relationship with the surrounding area and surrounding objects.
[0008] From the perspective of the relationship between a building and surrounding objects, it is considered effective to establish rules for the positional state of both. For example, from an image taken from the sky looking directly downward, areas in the image are classified as buildings, water, trees, etc., and the impact of damage on the building is identified based on the relative positional relationship of each area to the building. However, in order to make this judgment, it is necessary to take an image looking directly downward, and if this assumption is broken, the relative positional relationship will be broken, so damage judgment cannot be made based on aerial images taken from a camera looking diagonally downward, such as a bird's-eye view.
[0009] In order to quickly grasp damage at actual disaster sites, UAVs (unmanned aerial vehicles) need to be able to recognize objects without restrictions on the angle of view. Therefore, it is necessary to perform damage assessment without restrictions on the imaging angle. Note that, although the above description has been given using buildings in disaster-stricken areas as an example of objects to be detected, this is not limiting.
[0010] The present invention has been made in consideration of the above circumstances, and aims to provide an object detection method and an object detection system that can identify a target object based on the relative positioning of the target object and surrounding objects, regardless of the image capture direction.
[0011] The object detection method of the present invention for solving the above problem is an object detection method that uses machine learning inference with an image and corresponding text as input, and is characterized in that it generates an image based on an inferred region for the input image and the input text, extracts image features from this generated image, extracts text features from the input text, and performs machine learning inference based on the similarity between the image features and the text features to identify the object to be detected.
[0012] In addition, the object detection system of the present invention for solving the above-mentioned problems is an object detection system that uses machine learning inference with an image and corresponding text as input, and is characterized by comprising an area segmentation inference unit that infers an area corresponding to the input text for the input image, an image generation unit that generates an image based on the area inferred by the area segmentation inference unit and the input text, an image feature extraction unit that extracts image features from the image generated by the image generation unit, and a text feature extraction unit that extracts text features from the input text, and is characterized by performing machine learning inference based on the similarity between the image features and the text features to identify the object to be detected.
[0013] According to the present invention, it is possible to more reliably identify a detection target object from the positional relationship between the detection target object and surrounding objects, regardless of the image capturing direction.
[0014] Problems, configurations, and effects other than those described above will become apparent from the following description of the mode for carrying out the invention (hereinafter referred to as the embodiment).
[0015] 5 is a block diagram schematically showing an example of the configuration of an object detection system to which the object detection method of the present invention is applied. FIG. 6 is a block diagram showing an example of a functional configuration for realizing an object detection method for disasters according to a first embodiment of the present invention. FIG. 7 is a block diagram schematically showing an example of the functional configuration of the image generation unit of FIG. 2. FIG. 8 is a flowchart showing an example of processing executed by the image generation unit of FIG. 2. FIG. 9 is a flowchart showing an example of specific processing for pre-processing information to be input to the image generation unit of FIG. 2. FIG. 10 is a diagram showing a specific example for explaining the pre-processing of FIG. 5. FIG. 11 is a flowchart showing an example of processing for machine learning inference processing on images generated by the image generation unit of FIG. 2. FIG. 12 is a diagram showing an example of samples when aerial images of a disaster area and text groups as their descriptions are collected. FIG. 13 is a diagram explaining that the object detection method of the present invention can also be applied to text input in which no objects exist. FIG. 14 is a flowchart showing an example of processing for expanding learning data to realize an object detection method for disasters according to a second embodiment of the present invention.
[0016] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the embodiments described below are merely examples for explaining the present invention, and are not intended to limit the scope of the present invention to these embodiments. Those skilled in the art can implement the present invention in various other forms without departing from the scope of the present invention.
[0017] In the configurations of the invention described below, the same reference numerals are used in different drawings for the same parts or parts having similar functions, and redundant descriptions may be omitted. Furthermore, the position, size, shape, and range of each component shown in the drawings may not represent the actual position, size, shape, and range in order to facilitate understanding of the present invention. Therefore, the present invention is not limited to the position, size, shape, and range disclosed in the drawings. Furthermore, in this specification, components expressed in the singular may be plural unless otherwise clearly indicated in the context.
[0018] <Summary of the Invention> The present invention relates to an object detection method and an object detection system for detecting (identifying) a target object based on the positional relationship between the target object and surrounding objects, regardless of the image capturing direction.
[0019] Recently, a VL (Vision Language) model, typified by CLIP (Contrastive Language-Image Pre-training), has been proposed, which uses images and text as input. This VL model is a classifier obtained by training a large number of image and text pairs. For example, CLIP has an internal image encoder and a corresponding text encoder, and is capable of extracting image features from input images and simultaneously extracting text features from input texts. By calculating the similarity between the two using an inner product or the like, the VL model can relatively evaluate which text is closest to which input image.
[0020] Therefore, for example, when an image of a landslide-affected area contains text such as "a house damaged by the landslide" and text such as "a peaceful suburban landscape," it can be determined that the former has a higher similarity. However, the feature extraction performed by the image encoder of CLIP, which is used as an image feature extractor in many VL models, extracts features from the entire image, which poses a problem in that it is unclear which part of the image specifically corresponds to the reference in the text. Therefore, even if it is determined that a captured image is an image of a disaster, it is still unclear which part of the image is damaged.
[0021] To address this issue, the present invention does not simply input an image to the VL model. Instead, each object (e.g., a building) is detected from the image, and a composite image based on the detection results is then superimposed on the original image. Specifically, for example, in an image containing multiple buildings, a composite image that highlights only one building is input, allowing the VL model to focus on that building and extract image features. Then, by comparing the image features extracted by focusing on a specific region with the text features, it is determined whether each building (house) has been damaged. While buildings are used as an example of objects to be detected here, the objects to be detected are not limited to buildings and can include man-made objects such as buildings or people in need of rescue.
[0022] <Overall System Configuration> Fig. 1 is a block diagram showing a schematic configuration example of an object detection system to which the object detection method of the present invention is applied. The object detection system shown in Fig. 1 is, as an example, a system for assessing the situation in a disaster-stricken area, and has a system configuration including a mobile sensor unit 110 and an information processing device 120. The system for assessing the situation in a disaster-stricken area (object detection system) can be a disaster detection system.
[0023] The mobile sensor unit 110 may be a mobile sensor, such as a drone or a satellite, that can acquire images from above a disaster-stricken area. The mobile sensor unit 110 includes at least an imaging unit 111, a communication unit 112, a drive unit 113, and an attitude sensor 114, which may be connected to one another.
[0024] The imaging unit 111 may be a sensor such as a camera. The communication unit 112 has a function of transmitting data acquired by the imaging unit 111 to the information processing device 120, and conversely, a function of receiving signals transmitted from the information processing device 120. The drive unit 113 may be, for example, a motor, and has a function of driving the imaging unit 111 based on signals received from the information processing device 120 via the communication unit 112. The attitude sensor 114 is a sensor such as a gyro. Data acquired by the attitude sensor 114 is transmitted to the information processing device 120 via the communication unit 112.
[0025] [Information Processing Device] The information processing device 120 can be configured as a general server. For example, the information processing device 120 can be configured as a server including, as hardware, an input device, an output device, a processing device, and a storage device. In this embodiment, the processing device reads a program stored in the storage device and executes the read program, thereby realizing various functions in cooperation with other hardware as necessary.
[0026] The information processing device 120 may be configured as a standalone server, or any of the input device, output device, processing device, and storage device may be configured as another computer system connected via a network.
[0027] The information processing device 120 has, as its functional configuration, for example, a communication unit 121, a storage device 122, a processor 123, a display unit 124, and a control unit 125, which may be connected to one another.
[0028] The communication unit 121 has a function of receiving data and signals from the mobile sensor unit 110 and transmitting signals to the mobile sensor unit 110. The storage device 122 may be configured, for example, as a random access memory (RAM). The processor 123 may be configured, for example, as a field programmable gate array (FPGA) or a graphic processing unit (GPU).
[0029] A disaster detection method described below may be realized on the storage device 122 and the processor 123. The display unit 124 may display, for example, an image acquired by the imaging unit 111 and received by the communication unit 121 from the mobile sensor unit 110, information from the attitude sensor 114, or a disaster area estimation result inferred by a disaster-oriented object detection method described below based on these, in the form of, for example, a GUI (Graphical User Interface). The control unit 125 may control the drive unit 113 via the communication units 121 and 112, for example.
[0030] First Embodiment A first embodiment is an example of realizing an object detection method for disasters, for example.
[0031] 2 is a block diagram showing an example of a functional configuration for realizing an object detection method for disasters according to a first embodiment of the present invention. However, the application of the technology of the present invention is not limited to systems for disasters.
[0032] The functional configuration for realizing the disaster-oriented object detection method according to the first embodiment may be realized on the above-described storage device 122 and processor 123. As an example, the functional configuration includes the following functional units: an image text DB 201, an image pre-processing unit 202, a text pre-processing unit 203, a first image feature extraction unit 204, a noun phrase feature extraction unit 205, a region segmentation inference unit 206, a layout generation unit 207, a text generation unit 208, a text feature extraction unit 209, an image generation unit 210, a second image feature extraction unit 211, an image-text similarity calculation unit 212, and a region selection unit 213.
[0033] The image text DB 201 is a database containing data consisting of, for example, images and corresponding text groups, and is used for training the object detection method according to the first embodiment. The images in the image text DB 201 may be publicly known image data acquired from a social networking site (SNS), such as the COCO (Common Objects in Context) dataset. Furthermore, the text groups corresponding to the images may be, for example, the publicly known refCOCO dataset, in which a description is provided for each object included in the image. For example, if an image contains three people, a description may be provided for each person, such as "person looking at a computer," "person sitting in a chair while looking at a computer," or "passerby." However, multiple descriptions may be provided for a single object, or an object may have no description provided.
[0034] This database was made publicly available as publicly known information in order to solve a problem known as REC (Referring Expression Comprehension). In a typical object detection problem using an object detection method such as YOLO (You Only Look Once), it is sufficient to detect three people in an image as people. However, this database is designed for a conditional detection problem in which the text "person sitting in a chair looking at a computer" is used to detect only the person who matches the text.
[0035] The image text DB 201 is composed of, for example, one image for each sample, an indefinite number of explanatory texts, and a teacher mask in the form of a binary mask, such as a segmentation mask, that clearly identifies the target object in the image that corresponds to each of the texts.
[0036] The image pre-processing unit 202 performs, for example, average variance processing on pixel values of sample images selected from the image text DB 201 in preparation for numerical calculations.
[0037] The text pre-processing unit 203 performs phrase structure analysis such as well-known dependency parsing or constituency parsing, and extracts noun phrases from the text. The extracted noun phrases are descriptions of the images described above, and therefore have the characteristic of representing objects present in the images. For example, a noun phrase such as "person, chair, person" may be extracted from the text "person sitting in a chair looking at a computer."
[0038] The first image feature extraction unit 204 receives the image processed by the image pre-processing unit 202 as input and performs processing to extract image features. As the first image feature extraction unit 204, for example, a feature extraction encoder for image recognition such as a known Vision Transformer (ViT) can be used.
[0039] For example, a text encoder included in the well-known BERT (Bidirectional Encoder Representations from Transformers) or CLIP (Contrastive Language-Image Pre-training) can be used as the noun phrase feature extraction unit 205. The noun phrase feature extraction unit 205 performs processing to extract noun phrase features from each of the noun phrases extracted by the text pre-processing unit 203.
[0040] The region segmentation inference unit 206 is an image recognition region segmentation model implemented by, for example, a known method such as MaskFormer. The region segmentation inference unit 206 performs processing to segment the image extracted by the first image feature extraction unit 204 and infer a segmentation mask group M. The region segmentation inference unit 206 can infer regions corresponding to input text from an input image, for example, for object detection in disaster situations.
[0041] In recent years, for example, Open-Vocabulary format segmentation models such as X-Decoder and OpenSeeD have been proposed as the region segmentation inference unit 206. The region segmentation inference unit 206 performs processing to output regions in the image that correspond to the noun phrase text features extracted by the noun phrase feature extraction unit 205 as a segmentation mask group M.
[0042] The type of segmentation may be output by instance segmentation, which can identify each object, or panoptic segmentation. The region division inference unit 206 preferably has a segmentation function that can accurately infer regions, but may also be object detection, such as rectangle detection such as YOLO (You Only Look Once), and is not particularly limited.
[0043] The layout generation unit 207 performs processing to generate a layout group L to be input to the image generation unit 210 from the segmentation mask group M. When there is a subject for which no corresponding input text exists during layout learning, the layout generation unit 207 generates text without using a predicate expressing a relationship using the subject and object corresponding to the area inferred from the input image by the area segmentation inference unit 206, and uses the generated text for conditioning. Details of this processing will be described later.
[0044] The text generation unit 208 performs processing to generate text corresponding to each layout, for example, using the results of phrase structure analysis performed by the text pre-processing unit 203 and the layout group L generated by the layout generation unit 207. The specific processing for generating this text will be described later.
[0045] The text feature extraction unit 209 performs processing to extract text features from the text generated by the text generation unit 208. The same text encoder as used in the noun phrase feature extraction unit 205 can be used as the text feature extraction unit 209. This text feature extraction unit 209 can extract text features from input text for, for example, object detection for disasters.
[0046] The image generation unit 210 performs processing to generate an image group G from the input original image, the segmentation mask group M inferred by the region segmentation inference unit 206, and the text feature amount extracted by the text feature extraction unit 209. This image generation unit 210 can create a layout for image generation based on the input text and the output of the region segmentation inference unit 206, for example, for object detection in disaster situations, and generate images conditioned by at least the layout and the input text. The process of this image generation unit 210 will also be described later.
[0047] The second image feature extraction unit 211 may be an image encoder included in CLIP, and receives the image group G generated by the image generation unit 210 and performs feature extraction processing. The second image feature extraction unit 211 can extract image features from the images generated by the image generation unit 210, for example, to detect objects in a disaster.
[0048] The image text similarity calculation unit 212 performs a process of calculating the similarity of the feature amounts between the different types of data extracted by the text feature extraction unit 209 and the second image feature extraction unit 211 using, for example, cosine similarity.
[0049] The region selection unit 213 performs a process of selecting from the image group G an image that is most similar to each piece of text that constitutes the text group, based on the scores of the image group G and the text group calculated by the image-text similarity calculation unit 212. Then, the region selection unit 213 selects, as a detection target, a region that contributed to the generation of the selected generated image from the segmentation mask group M of the region segmentation inference unit 206, and sets this as the final output of the object detection. In other words, the region selection unit 213 selects an image that is closest to the input text from images generated based on the layout, thereby selecting the layout that served as the source of the generation, and sets the region of the main component that constitutes the selected layout as the final output of the object detection.
[0050] According to an object detection system having the above-described functional units, it is possible to detect (identify) a target object by performing machine learning inference based on the similarity between image features and text features.
[0051] [Example of Functional Configuration of Image Generation Unit] Fig. 3 is a block diagram schematically showing an example of the functional configuration of the image generation unit 210 in Fig. 2. As shown in Fig. 3, the image generation unit 210 has functional units, for example, an image feature extraction unit 301, a conditioned feature extraction unit 302, an image restoration unit 303, and an image synthesis unit 304.
[0052] The image feature extraction unit 301 includes, for example, a linear layer, and performs processing to extract features in order to input the features to the subsequent conditioning feature extraction unit 302. The image feature extraction unit 301 may have the same configuration as the first image feature extraction unit 204.
[0053] The conditioning feature extraction unit 302 may be equipped with, for example, an attention mechanism in deep learning, and performs processing to condition the image features by the text features, using, for example, a cross attention mechanism, from the image features extracted by the image feature extraction unit 301 and, for example, the text features extracted from the text generated by the text feature extraction unit 209.
[0054] The conditioned feature extraction unit 302 may also include a normalization mechanism such as AdaIN (Adaptive Instance Normalization), and may condition and extract features from image features based on mask information inferred by the region segmentation inference unit 206. In particular, when conditioning is applied to the entire feature, a configuration such as AdaIN may be used, and when conditioning is applied to each location of the feature, a Cross Attention structure may be used.
[0055] The image restoration unit 303 corresponds to the decoder portion of a structure such as U-net, and may include, for example, a convolutional neural network (CNN) structure and an upsampling structure. The image restoration unit 303 sends the feature amounts extracted from the conditioned feature extraction unit 302 to the image synthesis unit 304, and performs processing to restore the feature amounts to an RGB image so as to create, for example, a 3ch image.
[0056] The image synthesis unit 304 may perform an addition process, for example in the sense of superimposing, between the RGB image restored by the image restoration unit 303 and the original RGB image corresponding to the image output from the image pre-processing unit 202, and may synthesize the images in a manner such as alpha blending. Alpha blending is just one example, and other methods may also be used, but an implementation that is at least capable of learning in a form that allows error backpropagation, which will be described later, is required.
[0057] [Example of Processing by Image Generation Unit] FIG. 4 is a flowchart showing an example of processing executed by the image generation unit 210 in FIG.
[0058] In step S401, during learning, an image extracted by the image text DB 201 is used as the input image X, and during inference, an actually captured image is used as the input image X. After normalization and other processing is performed by the image pre-processing unit 202, the image feature extraction unit 301 extracts the image feature amount V.
[0059] In step S402, the layout generation unit 207 creates a layout group L from the segmentation mask group M obtained from the region division inference unit 206, the text generation unit 208 generates a text group T which is text corresponding to each layout, and the text feature extraction unit 209 extracts features from the text group T.
[0060] In this feature extraction process, a conditioned feature is extracted as a conditioned image feature V′ based on the features extracted from the layout group L and the text group T. A CNN may be used to extract the feature from the layout group L, and a technique such as Spatially-Adaptive Normalization (SPADE) may be used to condition the feature.
[0061] In the process of step S402, the layout group L is created from the segmentation mask group M, and the specific process of the layout generation unit 207 that generates text corresponding to the layout will be described later.
[0062] In step S403, the image restoration unit 303 generates an image Y from the image feature V'. The image feature extraction unit 301, the conditioning feature extraction unit 302, and the image restoration unit 303 that make up the image generation unit 210 may have a structure similar to that of U-net, for example, or may have a configuration similar to that used in an image generation model such as stable diffusion. The conditioning order and processing order described in this embodiment are merely examples and are not particularly limited.
[0063] The image Y generated from the image feature V' is not necessarily an image that is meaningful to humans. In step S404, the image Y generated from the image feature V' is superimposed on the input image X using an additive process such as alpha blending to generate an image group G. Note that the image Y itself may be used as the image group G as is. The image group G is expected to be images that contribute to feature extraction related to a specific region in an image when the second image feature extraction unit 211 extracts features in a subsequent stage, and may be images referred to in a technology recently referred to as visual prompting.
[0064] [Example of Pre-Processing of Information to be Input to Image Generation Unit] Fig. 5 is a flowchart showing an example of specific processing for pre-processing information to be input to the image generation unit 210 of Fig. 2. Specifically, Fig. 5 shows an example of the processing flow during learning in step S402 of Fig. 4, in which the layout generation unit 207 and the text generation unit 208 generate a layout group L and a text group T to be input to the image generation unit 210.
[0065] 5 is performed by the image generation unit 210. The image generation unit 210 first explicitly separates subjects and objects from the input text, and then performs processing to create a layout based on the output of the region division inference unit 206. Furthermore, when the region division inference unit 206 infers regions corresponding to multiple subjects from the input image, the image generation unit 210 generates a layout for each subject and performs processing to express the region representing the object and the region for each subject in a separable form. This processing corresponds to the processing of steps S501 to S503, which will be described later.
[0066] The processing executed by the image generating unit 210 will be specifically described below. Fig. 6 will be used as appropriate in the description of Fig. 5. Fig. 6 is a diagram showing a specific example for explaining the pre-processing of Fig. 5.
[0067] The segmentation mask group M is inferred by the region segmentation inference unit 206 from images sampled by the image text DB 201. Note that, due to the characteristics of the image text DB 201, the aforementioned images are associated with accompanying explanatory text and training information masks that indicate the regions corresponding to the explanatory text. For example, in the example of FIG. 6, there are two text groups 602, (a) and (b), corresponding to the input image 601. Here, the regions corresponding to people, indicated by (a) and (b), are included in the image text DB 201 as training information masks.
[0068] In step S501, a linking process is performed to determine which nouns each mask constituting the segmentation mask group M corresponds to, based on a similarity calculation between the image features corresponding to each segmentation mask region inferred within the region division inference unit 206 and the noun phrase features extracted by the noun phrase feature extraction unit 205 based on, for example, a predefined noun class group or a noun phrase extracted by the text pre-processing unit 203 through phrase structure analysis processing for the text accompanying the input image.
[0069] This linking process is a well-known process that is performed after providing a group of noun phrases in the aforementioned Open-Vocabulary segmentation method, and this process may be used. For example, in FIG. 6, a group of texts 602 corresponding to an input image 601 is provided. During training, these are paired due to the nature of the image text DB 201, but during testing, the user may input a description of the object they wish to detect as text for the input image 601.
[0070] In the example of FIG. 6 , the text preprocessing unit 203 performs phrase structure analysis on the explanatory text group 602, and selects a segmentation mask corresponding to a noun phrase corresponding to the subject of the text. For example, phrase structure analysis can extract noun phrases such as "person, laptop, chair" from text such as (a) and (b) that make up the text group 602. The region segmentation inference unit 206 extracts regions corresponding to each noun phrase, such as masks 610, 611, 612, 613, and 614, as a segmentation mask group M. Masks 610, 611, and 612 corresponding to "person," mask 614 corresponding to "laptop," and mask 613 corresponding to "chair" can be detected. In this example, because the word "suitcase" is not included in the text group 602, mask 615 corresponding to "suitcase" cannot be detected, but this does not necessarily mean that it cannot be detected.
[0071] In step S502, the phrase structure analysis performed in step S501 reveals that, particularly when the text is in English, "person" can be determined as the subject and "laptop, chair" can be determined as the object because the subject comes before the text. Therefore, among the segmentation mask group M, masks 610, 611, and 612 corresponding to "person" are set as a subject mask candidate group S.
[0072] In step S503, a process of selecting a mask group O corresponding to objects of the input text is performed. In the above example, “laptop” and “chair” are objects, and the segmentation masks 613 and 614 corresponding to these objects in the segmentation mask group M are set as the object mask group O.
[0073] In step S504, a process is performed to create a layout group L from the subject candidate mask group S and the object mask group O. In the example of FIG. 6, for example, the subject candidate mask group S consists of three candidates, and the object mask group O consists of two elements. In this case, for example, the layout group L may be created according to a rule that one mask is selected from the subject candidate mask group S and all masks are selected from the object mask group O. Under this rule, the object mask group O may be grouped together into a single mask, such as an object mask group 616.
[0074] Furthermore, the object mask group 616 may be combined with the components of each of the subject candidate mask groups S to create layout masks 620, 621, and 622, and these three may be used as the layout group L. The layout in this embodiment needs to be created in a form that allows the subject object to be identified, and for example, if the layout L is a 2ch image, the area corresponding to the subject candidate mask group S may be assigned a value of 1 in the first channel, the area corresponding to the object mask group O may be assigned a value of 1 in the second channel, and the other areas may be assigned a value of 0.
[0075] In the above layout creation example, the layout is created in a way that allows each subject candidate to be identified, but the objects are all grouped together, making it impossible to identify one from another. It is also possible to create a layout in the form of identifiable segmentation masks 613 and 614 for the objects, rather than grouping them into a form like object mask group 616. However, since an explosion of combinations can occur if the number of combination candidates increases, the above-mentioned combinations are used in this embodiment.
[0076] In step S505, during learning, a process is performed in which text is linked to the layout group L created in the process of step S504. In the image text DB 201, the subject area of the image corresponding to the text is provided as training information. Therefore, subject candidates corresponding to the text are clear during learning by comparing the subject candidate area with the area provided as training information, so text may be linked to the layout corresponding to the training information. In the previous example, text group 602(a) may be linked to layout mask 620, and text group 602(b) may be linked to layout mask 621.
[0077] In step S506, a process is performed to check whether there is a layout in the layout group to which no text has been linked in the process of step S505. The image text DB 201 does not necessarily have text corresponding to all subject candidates.
[0078] In step S507, the text generation unit 208 generates and links text to layouts to which no text was linked in step S506. In the previous example, no text is linked to the layout mask 622, so it is necessary to generate text. When generating text, for example, a noun phrase corresponding to a mask included in the object mask group 616 that constitutes the layout mask 622 may be used. For example, the generated text may be "person, laptop and chair."
[0079] Since the purpose of this embodiment is to recognize the relationship between objects, no vocabulary indicating the relationship is set for layouts in which no relationship exists. In this text generation example, unlike layout masks 620 and 621, verbs such as "watching" and "sitting on" that indicate the relationship between subject objects are not included, and layout mask 622 is distinguished from other layouts, making it possible to identify the subsequent functional configuration.
[0080] In step S508, a layout is created using only the object mask group O and added to the layout group L, and corresponding text is also created. For example, since the object mask group 616 does not contain any subject candidates, it may be used as is, and the text generation unit 208 may set the text corresponding to this layout as, for example, "chair and laptop." This corresponds to data generation for determining whether or not a subject candidate corresponding to the input text exists during inference. That is, in this case, with regard to layout creation, even if no subject exists in the image area, a layout is generated using only the objects of the input text, and if the subject of the input text does not exist in the input image during inference, a process is performed to determine whether or not a subject exists and therefore no detection target exists.
[0081] In the test process, the most appropriate layout from the layout group L is selected for the input text, but if there is no layout in the image that corresponds to the subject of the text input during the test, the layout group L will be an empty set. To avoid this, an object mask group 616 is added in step S508.
[0082] In step S509, a process is performed to establish a one-to-one correspondence between the generated layout group and the text group that has been added up to that point. In FIG. 6, the final result is a one-to-one correspondence between (object mask group 616, text 633), (layout mask 620, text 630), (layout mask 621, text 631), and (layout mask 622, text 632). Then, the image generation unit 210 generates images conditioned by each layout and text. The images generated here do not necessarily have visual meaning to humans.
[0083] Here, an example in which the image generation unit 210 is trained is described, but the image generation unit 210 is not limited to one obtained by training. In an extreme example, the image generation unit 210 may interpret the area other than the layout as the background and superimpose an image that is darker than the background on the original image. In this case, making the background area less visible to the image classifier in the subsequent stage is thought to have the effect of relatively emphasizing the foreground area corresponding to the layout, thereby guiding recognition.
[0084] [Example of Machine Learning Inference Processing] Fig. 7 is a flowchart showing an example of processing for performing machine learning inference processing on images generated by the image generation unit 210 in Fig. 2. In this processing, a process for calculating the similarity between image features and text features is performed on the image group G generated by the image generation unit 210.
[0085] In step S701, first, the second image feature extraction unit 211 performs a process of extracting an image feature V' from each of the image group G. Meanwhile, the text feature extraction unit 209 extracts text features from each of the texts generated by the text generation unit 208. During learning, the image generation unit 210 generates the image group G by conditioning the layout group L generated by the layout generation unit 207 and the corresponding text group. Therefore, each generated image corresponds to the text and layout used for conditioning.
[0086] In step S702, a process of calculating the similarity between the image feature V' and the text feature is performed. For example, if the second image feature extraction unit 211 and the text feature extraction unit 209 have been trained using CLIP or the like, the similarity between the two features can be calculated using an inner product, and the similarity can be expressed as a score. Furthermore, in step S703, it is determined whether this process is a learning process or an inference process.
[0087] If this process is a learning process, in step S704, the correct correspondence between text and image is used as training information to perform a loss calculation process using, for example, a cross entropy function for the score calculated in the process of step S702. The loss calculated here may be used by backpropagation to update parameters of modules having learnable parameters among the functional configuration shown in FIG. 2, for example, to update learnable parameters included in the conditioned feature extraction unit 302 and the image restoration unit 303 that constitute the image generation unit 210.
[0088] On the other hand, since the second image feature extraction unit 211, the noun phrase feature extraction unit 205, and the text feature extraction unit 209 have been trained using large amounts of data such as CLIP, it can be said that there is an advantage in using the parameters as they are, and therefore the parameters do not need to be fixed and updated. On the other hand, the parameters constituting the first image feature extraction unit 204 and the region segmentation inference unit 206 have also been trained sufficiently, so the parameters may be fixed, but since the vocabulary and region segmentation may not be compatible depending on newly introduced training data, the parameters may be subject to error backpropagation as necessary and updated.
[0089] In the case of inference processing, in step S705, the region selection unit 213 selects a correct answer from the similarity score. For example, in the example text of "person watching a laptop, sitting on a chair" in Fig. 6, the similarity score for the image generated by the layout mask 621 corresponding to this text is larger than that derived from the layout mask 620, so the subject candidate mask 611 that contributed to the generation of the layout mask 621 may be determined to be the correct answer.
[0090] On the other hand, if there is no subject candidate corresponding to the text in the input image, the region segmentation inference unit 206 cannot detect a subject candidate, and therefore only those from the object mask group 616 exist in the layout group L, which are determined to have the highest similarity to the input text. However, since there is no subject mask corresponding to this, it may be determined that there is no object to be detected. On the other hand, if it is desired to detect a subject corresponding to "person watching a laptop," in the example of Figure 6, it is necessary to detect two corresponding to layout masks 620 and 621.
[0091] In this way, when it is necessary to detect multiple objects, for example, a generated image derived from the layout mask 620 is judged to have a higher similarity to two texts, "person watching a laptop" or "person and laptop," and if it is closer to the former, it is designated as the detection target. This process is also performed on the image group G derived from the layout mask 621. This judgment method makes it possible to handle multiple detection targets. This is due to the fact that, particularly during learning, a layout and text pair is designed that recognizes relationships. Since the image group G derived from the layout mask 622 is learned using "person and laptop" or "person, laptop and chair" without being given a verb that indicates a relationship, such as "watching," differences in similarity will appear during testing depending on whether or not the input text contains "watching," which indicates a relationship.
[0092] Example 2 Example 2 is an example of a learning method for identifying buildings in a disaster detection application.
[0093] In the description of the first embodiment, for example, data samples in the image text DB 201 are mainly composed of "persons." However, the technology of the present invention does not limit the target to a specific object.
[0094] 8 is a diagram showing an example of a sample of the image text DB 201, for example, when aerial photographs 801 of disaster-stricken areas and text groups 802 as explanatory text are collected. At the same time as collecting the text groups 802, teacher masks are prepared for the areas corresponding to each piece of text. Since the subject of this data is "house" and the objects are "tree, flood," layouts 810, 811, and 812 can be created as a layout group L that takes the subject object into account using the same method as described in FIG. 6, and text groups 820, 821, and 822 corresponding to each layout can be prepared in association with each other.
[0095] Therefore, it is possible to learn to detect buildings in disaster areas based on the relationship between houses and other objects using the same learning method as in Example 1. However, the subject is not limited to houses, but may be a car or a person, and business demands for damage detection may call for conditional detection, such as "a person trapped in rubble," rather than just a person, and machine learning inference to address this is possible if data is available.
[0096] Although the data samples used for the explanations in Example 1 and Example 2 are different, these can also be collectively used as the image text DB 201. The object detection method of the present invention can be used in a variety of situations, without being limited to everyday scenes such as those in Example 1 or disaster scenes such as those in Example 2, depending on the learned data. Furthermore, when using the method, it is necessary to design the text of the object to be detected. For example, if it is desired to detect an object limited by vocabulary indicating a relationship such as "house damaged by tree," it is necessary to prepare text containing only noun phrases such as "house" or "house and tree." This corresponds to the point in step S507 of FIG. 5 where the relationship is explicitly learned during learning.
[0097] In particular, in actual operation, it is possible to prepare multiple texts about objects to be detected, and for disaster detection, by preparing a text group that is an extension of the text group 802 in Fig. 8, it becomes possible to detect houses or, for example, people in various disaster patterns limited by the text. Specifically, during object detection inference, by creating a list that limits the objects to be detected by at least one of adjectives and vocabulary indicating relationships, objects corresponding to multiple limited nouns can be detected.
[0098] [Example in which determination is possible without an object] FIG. 9 is a diagram for explaining that the object detection method of the present invention can be applied to a text input in which no object exists.
[0099] For example, in the case of disaster detection, given an input image 900 and input text 910, the noun "house" may be extracted from the input text 910, and a layout group L may be created from masks 901 and 902 as a subject candidate mask group S using the same process as in FIG. 6 , without using the object mask group O. Objects with the attribute "damaged" can be detected by the learning process in step S704 and the inference process in step S705, using the similarity between image features extracted from an image group G generated by the image generation unit 210 from the layout group L and features extracted from, for example, "damaged house" and "house." This detection target is not limited to houses. Similar machine learning inference is also possible in cases where it is desired to identify objects containing intransitive verbs rather than relationships, such as image 920 and text 930, such as detecting a running person.
[0100] [Support for Special Domains] The object detection method of the present invention can be applied to detection targets such as damaged houses and man-made objects in disasters. Generally, any business can create an image-text DB 201 for disaster recognition. However, there is a limit to the amount of image data currently available, particularly of disaster-stricken areas, and the image capture locations are concentrated in disaster-stricken areas, which limits the domain. While it is natural to aim to enhance the image-text DB 201 by collecting new data, it is also possible to expand the data using, for example, generative AI such as ChatGPT or Stable Diffusion.
[0101] The expansion method will be described with reference to the flowchart of Fig. 10. Fig. 10 is a flowchart showing an example of processing for expanding learning data in order to realize an object detection method for disasters (disaster detection method) according to the second embodiment of the present invention.
[0102] In step S1001, a process is performed to collect text indicating a subject-object relationship related to the disaster to be detected. This text may be created by a human being based on knowledge of past disasters. Specifically, text may be created that clearly indicates a subject-object relationship between the subject and the surrounding objects or areas, rather than being limited to the house area, such as "house damaged by a fallen tree" or "house surrounded by the flood."
[0103] Here, a sample targeting a house is shown as an example, but the subject matter is not limited to houses and can be a car or a person. Furthermore, such text may be used as an example, and a language generation model such as ChatGPT may be used to generate similar text. The text list generated by this process may be visually reviewed by a human for further filtering.
[0104] In step S1002, for each list of text created in step S1001, a model for generating images from text, such as Stable Diffusion, may be used to generate images related to the disaster. The generated images may be visually inspected, and inappropriate images may be deleted.
[0105] In step S1003, during the generation process of step S1002, a process is performed to extract the position in the image where a noun phrase such as "house" in the text will be generated, focusing on, for example, a Cross Attention module that performs text-based conditioning when generating an image in the image generation model. This extraction process may use a matrix that represents the correlation between the feature values for each position in the image and the feature values for each token that constitutes the text, and extraction may be performed from the weight values of the matrix. For example, a technique called "Prompt to Prompt" may be used for this process. This process results in an image region corresponding to the noun phrase that constitutes the text. This is known to be equivalent to a mask that indicates the position of the noun phrase.
[0106] By the processing of steps S1001 to S1003, it is possible to obtain text, images, and masks as training information. This satisfies the specifications of the elements that make up the image text DB 201. For example, after creating a database specialized for a specific domain such as disasters, it is possible to train a disaster detector by performing the processing of the learning method according to this embodiment.
[0107] As described above, for the training data that realizes object detection, at least the conditions of the detection target are generated as a text group with clear subject objects, and by saving the images generated from the text group and the areas corresponding to the noun phrases that make up the text group under the assumption of image generation based on the input image and input text, training data for learning machine learning inference can be generated.
[0108] 110...mobile sensor unit, 111...imaging unit, 112...communication unit, 113...drive unit, 120...information processing device, 121...communication unit, 122...storage device, 123...processor, 124...display unit, 125...control unit, 201...image text DB, 202...image pre-processing unit, 203...text pre-processing unit, 204...first image feature extraction unit, 205...noun phrase feature extraction unit, 206...region segmentation inference unit, 207...layout generation unit, 208...text generation unit, 209...text feature extraction unit, 210...image generation unit, 211...second image feature extraction unit, 212...image text similarity calculation unit, 213...region selection unit
Claims
1. An object detection method using machine learning inference that takes an image and corresponding text as input, the method comprising: generating an image based on an inferred area for the input image and the input text; extracting image features from the generated image and text features from the input text; and performing the machine learning inference based on the similarity between the image features and the text features to identify the object to be detected.
2. The object detection method according to claim 1, characterized in that, in the machine learning inference, text recognition of noun phrases and subject-object relationships is performed for the input text, the input image is segmented based on the recognized noun phrases, a plurality of images are generated based on the subject-object relationships obtained by the text recognition and the results of the segmentation, and the object to be detected is identified based on the similarity between each of the generated images and the input text.
3. The object detection method according to claim 1, characterized in that the object to be detected is an artificial object such as a building in a disaster-stricken area or a person in need of rescue.
4. The object detection method according to claim 1, characterized in that, during inference for the object detection, a list is created in which the objects to be detected are limited by at least one of adjectives and vocabulary indicating relationships, thereby detecting objects corresponding to a plurality of limited nouns.
5. The object detection method according to claim 1, characterized in that, for the learning data that realizes the object detection, at least the conditions of the detection target are generated as a text group with clear subject objects, and learning data for training the machine learning inference is generated by preserving images generated from the text group and areas corresponding to noun phrases that constitute the text group during the process of image generation based on the input image and the input text.
6. An object detection system using machine learning inference that receives an image and corresponding text as input, comprising: an area segmentation inference unit that infers an area in the input image corresponding to the input text; an image generation unit that generates an image based on the area inferred by the area segmentation inference unit and the input text; an image feature extraction unit that extracts image features from the image generated by the image generation unit; and a text feature extraction unit that extracts text features from the input text, wherein the object detection system performs the machine learning inference based on the similarity between the image features and the text features to identify an object to be detected.
7. The object detection system according to claim 6, characterized in that the image generation unit creates a layout for image generation based on the input text and the output of the region segmentation inference unit, and generates an image conditioned by at least the layout and the input text.
8. The object detection system according to claim 7, wherein the image generation unit explicitly separates subjects and objects from the input text and then creates the layout based on the output of the region segmentation inference unit.
9. The object detection system described in claim 7, characterized in that when the area division inference unit infers areas corresponding to multiple subjects from the input image, the image generation unit generates the layout for each subject and represents the area representing the object and the area for each subject in a separable form.
10. The object detection system according to claim 7, further comprising an area selection unit which selects the layout from which the image was generated by selecting an image that is closest to the input text from among the images generated based on the layout, and which selects the main area constituting the selected layout as the output of the object detection.
11. The object detection system of claim 7, further comprising a text generation unit that, when learning about the layout, if there is a subject for which no corresponding input text exists, generates text without using a predicate expressing a relationship using a subject and object corresponding to an area inferred by the area segmentation inference unit from the input image, and uses the text for the conditioning.
12. The object detection system of claim 11, wherein the text generation unit, in creating the layout, generates a layout using only the objects of the input text even if no subject is present in the image area, and when the subject of the input text is not present in the input image during inference, determines that no subject exists and that no detection target exists.
13. The object detection system according to claim 6, further comprising a mobile sensor having an imaging unit, wherein an image of a disaster-stricken area captured by said imaging unit from above the disaster-stricken area is used as the input image, and man-made objects such as buildings and people in need of rescue in the disaster-stricken area are detected as the detection target objects.
Citation Information
Patent Citations
Method, apparatus, electronic device, computer-readable storage medium, and computer program for image-based data processing
JP2020135852A
Image processing device, image processing method, and program
WO2023084833A1