Learning device, learning method, and recording medium

WO2025094302A1PCT designated stage expired Publication Date: 2025-05-08NEC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/039389
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

In the modal segmentation model, the estimation accuracy of hidden areas is low, making it difficult to effectively solve this problem.

Method used

By applying simulated occlusion processing in the learning device, the input image is occluded to generate occlusion images, and through components such as object representation extractor, correspondence factor determiner, loss calculator, and parameter updater, the object representation extractor is trained so that it can establish a correspondence between the input image and the occluded image, thereby improving the estimation accuracy of the hidden area.

Benefits of technology

Through simulation occlusion processing and training of object representation extractor, the accuracy of the modal segmentation model in hidden area estimation is significantly improved, and the ability to recognize hidden object shapes in the image is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023039389_08052025_PF_FP_ABST
    Figure JP2023039389_08052025_PF_FP_ABST
Patent Text Reader

Abstract

In this learning device, an occlusion processing means applies occlusion processing to a first image to generate a second image. An object representation extraction means extracts first object representations, which are object representations of objects included in the first image, and second object representations, which are object representations of objects included in the second image. A correspondence relationship determination means determines a correspondence relationship between the first object representations and the second object representations. A loss calculation means calculates a first loss on the basis of the first object representations, the second object representations, and the correspondence relationship. A parameter update means updates parameters of the object representation extraction means on the basis of the first loss. The object representation extraction means thus obtained through learning can be used to optimize the estimation of hidden areas in an input image.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, learning method, and recording medium

[0001] The present disclosure relates to amodal segmentation.

[0002] Segmentation is known as one of the image recognition techniques using AI (Artificial Intelligence). In particular, the task of segmenting an image by estimating even hidden areas of objects is called amodal segmentation. In the fields of robot control and video surveillance, amodal segmentation is used to estimate the entire shape of an object hidden behind an object.

[0003] Patent Document 1 describes a technique for estimating the three-dimensional shape of a photographed object using machine learning (deep learning).

[0004] International Publication No. WO2019 / 035155

[0005] Amodal segmentation models have the problem of low accuracy in estimating occluded regions.

[0006] One objective of the present disclosure is to improve the accuracy of occlusion region estimation using an amodal segmentation model.

[0007] In one aspect of the present disclosure, a learning device includes: an occlusion processing means that applies occlusion processing to a first image to generate a second image; an object representation extraction means that extracts a first object representation that is an object representation of each object included in the first image and a second object representation that is an object representation of each object included in the second image; a correspondence determination means that determines a correspondence between the first object representation and the second object representation; a loss calculation means that calculates a first loss based on the first object representation, the second object representation, and the correspondence; and a parameter update means that updates a parameter of the object representation extraction means based on the first loss.

[0008] In another aspect of the present disclosure, a learning method executed by a computer includes: applying occlusion processing to a first image to generate a second image; extracting, using an object representation extractor, first object representations that are object representations of each object included in the first image and second object representations that are object representations of each object included in the second image; determining a correspondence between the first object representations and the second object representations; performing a loss calculation that calculates a first loss based on the first object representations, the second object representations, and the correspondence; and updating parameters of the object representation extractor based on the first loss.

[0009] In yet another aspect of the present disclosure, a recording medium has recorded thereon a program that causes a computer to execute the following processes: generating a second image by applying occlusion processing to a first image; extracting, using an object representation extractor, first object representations that are object representations of each object included in the first image and second object representations that are object representations of each object included in the second image; determining a correspondence between the first object representations and the second object representations; performing loss calculation that calculates a first loss based on the first object representations, the second object representations, and the correspondence; and updating parameters of the object representation extractor based on the first loss.

[0010] 1 illustrates a concept of a learning device according to the present disclosure. FIG. 2 is a block diagram illustrating a hardware configuration of a learning device. FIG. 3 is a block diagram illustrating a functional configuration of a learning device. FIG. 4 illustrates an example of a reconstructed image and an alpha mask. FIG. 5 illustrates an example of a configuration for placing an occluder using an alpha mask. FIG. 6 is a flowchart of a learning process by a learning device. FIG. 7 is a block diagram illustrating a schematic configuration of a learning device according to another example. FIG. 8 is a block diagram illustrating a schematic configuration of a learning device according to another example. FIG. 9 is a block diagram illustrating a functional configuration of a mask generation device. FIG. 10 is a flowchart of a mask generation process by a mask generation device. FIG. 11 is a block diagram illustrating a schematic configuration of a learning device according to another example. FIG. 12 is a block diagram illustrating a configuration of a learning device according to another example. FIG. 13 is a flowchart of processing by a learning device according to another example.

[0011] Preferred embodiments of the present disclosure will now be described with reference to the drawings. First Embodiment [Learning Device] (Basic Concept) FIG. 1 shows the concept of a learning device according to a first embodiment. The learning device 100 learns an amodal segmentation model based on an input image. In an embodiment of the present disclosure, the amodal segmentation model is configured using a neural network. The trained amodal segmentation model generated by the learning device 100 can be used during inference to generate an amodal segmentation mask for an unknown target image.

[0012] The learning device 100 according to this embodiment applies artificial occlusion processing to an input image to generate an occlusion image in which a portion of an object in the image is artificially concealed. The learning device 100 then trains an object representation extraction unit so that the object representations extracted from the input image and the occlusion image match. Using the object representation extraction unit thus obtained makes it possible to improve the accuracy of estimating hidden areas in an image.

[0013] 2 is a block diagram showing the hardware configuration of learning device 100. As shown in the figure, learning device 100 includes an interface (IF) 12, a processor 13, a memory 14, a recording medium 15, a database (DB) 16, a display unit 17, and an input unit 18.

[0014] The IF 12 acquires an input image from an external device. The input image is data used for training an amodal segmentation model by the training device 100. The IF 12 also outputs the amodal segmentation model obtained by training to an external device.

[0015] The processor 13 is a computer such as a CPU (Central Processing Unit) and executes a pre-prepared program to control the entire learning device 100. The processor 13 may be a CPU, a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an MPU (Micro Processing Unit), an FPU (Floating Point number Processing Unit), a PPU (Physics Processing Unit), a TPU (Tensor Processing Unit), a quantum processor, a microcontroller, or a combination thereof. The processor 13 executes the learning process described below.

[0016] The memory 14 is composed of a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The memory 14 stores various programs executed by the processor 13. The memory 14 is also used as a working memory while the processor 13 is executing various processes.

[0017] Recording medium 15 is a non-volatile, non-transitory recording medium such as a disk-shaped recording medium or semiconductor memory, and is configured to be detachable from learning device 100. Recording medium 15 records various programs executed by processor 13. When learning device 100 executes various processes, the programs recorded on recording medium 15 are loaded into memory 14 and executed by processor 13.

[0018] DB16 stores, as necessary, input images input through IF12 during learning, generated trained amodal segmentation models, and the like.

[0019] Display unit 17 is configured with, for example, a liquid crystal display, etc. Input unit 18 includes, for example, a keyboard, a mouse, etc. Display unit 17 and input unit 18 are used, for example, by an operator of learning device 100 to input necessary operations and to display and view input images and generated masks.

[0020] 3 is a block diagram showing the functional configuration of the learning device 100. Functionally, the learning device 100 includes object representation extraction units 21 and 31, an object reconstruction unit 22, an image combination unit 23, a reconstruction loss calculation unit 24, an occlusion processing unit 25, a correspondence determination unit 26, a correspondence loss calculation unit 27, and a parameter update unit 28. The object representation extraction units 21 and 31 and the object reconstruction unit 22 are configured using neural networks.

[0021] An input image IP is input to the object representation extraction unit 21. In the example of Fig. 3, the input image IP is an RGB image including objects A to C. The input image IP may be a moving image or a still image. As shown in the figure, in the input image IP, the objects A to C are positioned so as not to overlap one another.

[0022] The object representation extraction unit 21 extracts object representation vectors of each of the objects A to C from the input image IP. The object representation vector is a feature vector indicating the characteristics of each object, and is an example of an object representation in the present disclosure. The object representation vector of each object is data in which information such as the position, posture, shape, and color of the object is encoded. The object representation extraction unit 21 extracts object representation vectors V11 to V13 for the objects A to C, and outputs them to the object reconstruction unit 22 and the correspondence determination unit 26. Note that, hereinafter, it is assumed that the object representation vector V11 corresponds to object A, the object representation vector V12 corresponds to object B, and the object representation vector V13 corresponds to object C.

[0023] The object reconstruction unit 22 acquires object representation vectors V11 to V13 of objects A to C and generates reconstructed images RPa to RPc and alpha masks Ma to Mc for each object. FIG. 4 shows examples of the reconstructed images and alpha masks. The object reconstruction unit 22 reconstructs an image of object A based on the object representation vector V11 of object A to generate a reconstructed image RPa. The object reconstruction unit 22 also generates an alpha mask Ma based on the reconstructed image RPa. The alpha mask Ma masks the area of ​​the reconstructed image RPa other than the area of ​​object A. Similarly, the object reconstruction unit 22 generates a reconstructed image RPb and an alpha mask Mb based on the object representation vector V12 of object B. The object reconstruction unit 22c also generates a reconstructed image RPc and an alpha mask Mc based on the object representation vector V13 of object C. The object reconstruction unit 22 outputs the generated reconstructed images RPa to RPc and alpha masks Ma to Mc to the image combination unit 23.

[0024] The object reconstruction unit 22 also estimates the depth of each object based on the object representation vectors V11 to V13 of objects A to C. As described above, information such as the position of the object is encoded in the object representation vector, so the object reconstruction unit 22 can extract depth information for each object from the object representation vector. Note that "depth" indicates the foreground-background relationship of each object in the image.

[0025] Here, a depth estimation method will be described. Let us assume that the number of objects is K. The object reconstruction unit 22 receives K object representation vectors from the object representation extraction unit 21. Each object representation vector has d dimensions. The object reconstruction unit 22 combines the K input object representation vectors into a single (K×d)-dimensional vector. The object reconstruction unit 22 then inputs the generated single (K×d)-dimensional vector into a multi-layer perceptron (MLP) to obtain K outputs. The obtained K outputs are depth values ​​for each of the K objects. In this case, the output K values ​​are unnormalized numerical values, and the closer an object is to the foreground of the image, the larger the value. Note that the output K values ​​may be normalized using Softmax or the like. Furthermore, in the above example, the K vectors are combined into a (K×d)-dimensional vector. Alternatively, the K vectors may be individually input into the MLP to estimate the depth of each object separately. In addition to the above method, the object reconstructing unit 22 may use a neural network for estimating depth, such as DepthNet. The object reconstructing unit 22 outputs the calculated depth to the image combining unit 23.

[0026] The image combining unit 23 generates a reconstructed image RIP by combining the reconstructed images RPa to RPc of each object, the alpha masks Ma to Mc, and the depth. Specifically, the image combining unit 23 weights the alpha masks Ma to Mc using the depth of each object. The image combining unit 235 then combines the reconstructed images RPa to RPc using the weighted alpha masks Ma to Mc to generate a reconstructed image RIP. In the following description, the reconstructed images corresponding to individual objects are referred to as "individual reconstructed images RP," and the reconstructed image combined by the image combining unit 23 is referred to as the "overall reconstructed image RIP" to distinguish between the two.

[0027] Specifically, the image combining unit 23 calculates a weight value used for weighting the alpha mask for each object based on the depth of each object. The image combining unit 23 calculates the weight value based on the depth of each object so that the alpha mask of an object located further back in the image has a smaller weight value and the alpha mask of an object located closer to the image has a larger weight value. The image combining unit 23 then combines the reconstructed images using the alpha masks weighted using the calculated weight values ​​to generate a reconstructed image RIP.

[0028] A specific example of processing by the image combination unit 23 will be described below. First, the image combination unit 23 corrects the alpha mask M using the depth of each object input from the object reconstruction unit 22 as a weighting value. Now, assuming that the number of objects is K, the image combination unit 23 multiplies each of the K objects as a whole by the K depths output from the object reconstruction unit 22. Next, the image combination unit 23 normalizes the K alpha masks after multiplication by the depths for each pixel so that the sum of the values ​​of the K alpha masks is 1, thereby generating K normalized alpha masks. In this way, an alpha mask weighted by the depth of each object and further normalized on a pixel-by-pixel basis is obtained.

[0029] Next, the image combination unit 23 generates an overall reconstructed image RIP by performing so-called alpha blending of the individual reconstructed image RP of each object with the normalized alpha mask M. Specifically, for the K objects, the image combination unit 23 multiplies the individual reconstructed image RP of each object by the value of the normalized alpha mask M, and adds up the reconstructed images of each object to generate an overall reconstructed image RIP. The image combination unit 23 outputs the generated reconstructed image RIP to the reconstruction loss calculation unit 24.

[0030] In the above example, the object reconstruction unit 22 estimates the depth of the object. However, instead of explicitly estimating the depth, the depth may be learned as information incorporated into the object mask. In this case, when the object reconstruction unit 22 constructs the mask, the mask of the hidden area is configured to have a small weight. Therefore, the image combination unit 23 can generate the reconstructed image RIP by calculating the weighted sum of the mask and the reconstructed image without calculating the weight for each pixel.

[0031] The reconstruction loss calculation unit 24 calculates the difference between the reconstructed image RIP and the input image IP as a reconstruction loss LS2, and outputs it to the parameter update unit 28. The reconstruction loss LS2 is an example of a second loss.

[0032] The occlusion processing unit 25 applies occlusion processing to the input image IP to generate an occlusion image OP. Specifically, the occlusion processing unit 25 conceals at least one of a plurality of objects included in the input image IP using an occluder OC, thereby generating an occlusion image OP in which at least a portion of the object is hidden. In the example of FIG. 3 , the occlusion processing unit 25 places a rectangular occluder OC on the object C in the input image IP to generate an occlusion image OP in which a portion of the object C is hidden. The occlusion processing unit 25 outputs the generated occlusion image OP to the object representation extraction unit 31. Note that the input image IP is an example of a first image, and the occlusion image OP is an example of a second image.

[0033] The occluder OC may be a specific shape such as a rectangle, a mosaic, an image of another object cut out from another image, or the like. The type of occluder OC is not limited to a specific one. Furthermore, the occlusion processing unit 25 may arrange multiple occluders OC for one input image IP to generate an occlusion image in which multiple objects are hidden.

[0034] The occlusion processing unit 25 places the occluder OC at a specific position in the input image IP, a random position, a position where an object is likely to be, a position where an occlusion state is likely to occur, or the like. If a position where an occlusion state is likely to occur is known in advance based on the type of input image IP, the shooting environment, etc., the occlusion processing unit 25 may place the occluder OC at that position. For example, if it is known that an occlusion state is likely to occur at the top and bottom edges of an image in an image shot in a certain environment, the occlusion processing unit 25 may place the occluder OC at the top and bottom edges of the input image IP.

[0035] Alternatively, the position of each object may be detected from the input image IP using an object detection model or the like, and the occlusion processing unit 25 may place an occluder OC at a position overlapping one of the detected objects. Furthermore, if the input image IP is a video, the movement of the object may be detected using background differences between multiple frame images, and the occlusion processing unit 25 may place an occluder OC at a location where the object is moving. Alternatively, the position of the object may be identified using the aforementioned alpha mask, and the occlusion processing unit 25 may place the occluder OC at the identified position. FIG. 5 shows a configuration for placing an occluder using an alpha mask. As shown in the figure, the object reconstruction unit 22 generates an alpha mask such as that shown in FIG. 4 and transmits it to the occlusion processing unit 25 as indicated by the dashed line 81. The occlusion processing unit 25 receives the transmitted alpha mask, identifies the position of each of objects A to C, and places an occluder OC at a position overlapping one of the objects.

[0036] When the input image IP is a moving image, the occlusion processing unit 25 may change the shape of the occluder OC over time. When the input image IP is a moving image, the occlusion processing unit 25 may fix the position of the occluder OC or may move the occluder OC. Furthermore, when the input image IP is a moving image, the occlusion processing unit 25 may place the occluder OC for the entire period of the input image IP, or may place the occluder OC for a portion of the period of the input image IP. When the occlusion processing unit 25 places the occluder OC for a portion of the period of the input image IP, for example, the occlusion processing unit 25 may place the occluder OC halfway through the input image IP, or may place the occluder OC at random timing.

[0037] The object representation extraction unit 31 has the same configuration as the object representation extraction unit 21. Specifically, the object representation extraction unit 31 is configured with a neural network having the same configuration as the object representation extraction unit 21, and the same parameters as those of the object representation extraction unit 21 are set. The object representation extraction unit 31 extracts an object representation vector for each object from the occlusion image OP, and outputs the vectors to the correspondence determination unit 26. In the example of FIG. 3 , the object representation extraction unit 31 extracts four object representation vectors V21 to V24 corresponding to four objects including objects A to C included in the original input image IP and an occluder OC added by the occlusion process, and outputs the vectors to the correspondence determination unit 26.

[0038] The correspondence determination unit 26 determines the correspondence between the object representation vectors extracted from the input image IP and the object representation vectors extracted from the occlusion image OP. In the example of FIG. 3 , the correspondence determination unit 26 determines the correspondence between the object representation vectors V11 to V13 output from the object representation extraction unit 21 and the object representation vectors V21 to V24 output from the object representation extraction unit 31. As described above, the object representation vectors V11 to V13 correspond to the objects A to C included in the input image IP, respectively. On the other hand, the object representation vectors V21 to V24 correspond to the objects A to C and the occluder OC included in the occlusion image OP. The correspondence determination unit 26 creates pairs of corresponding object representation vectors (hereinafter also referred to as "object pairs"). In the example of FIG. 3 , the correspondence determination unit 26 creates object pairs for each of the objects A to C. Note that the object representation vector corresponding to the occluder OC does not belong to any object pair.

[0039] Here, if it is known whether each of the object representation vectors V21 to V24 corresponds to one of the objects A to C or the occluder OC, the correspondence determination unit 26 can create corresponding object pairs based on the known information. For example, if the input image IP is a video and the object position of each object is given by a rectangle in an early frame image of the video, the correspondence determination unit 26 can determine the correspondence between the object representation vectors V11 to V13 and the object representation vectors V21 to V24 based on the position of the given rectangle.

[0040] On the other hand, when it is unknown whether each object representation vector V21 to V24 corresponds to one of the objects A to C or the occluder OC, the correspondence determination unit 26 estimates the correspondence between each object representation vector V21 to V24 and the objects A to C and the occluder OC. For example, the correspondence determination unit 26 calculates the distance (similarity) between the object representation vectors in the feature space, and estimates that two object representation vectors with a short distance (closest positions in the feature space) are object representation vectors of the same object. In this way, the correspondence determination unit 26 estimates the correspondence between the object representation vectors V11 to V13 and the object representation vectors V21 to V24. Note that the distance in this case may be, for example, the L2 distance (Euclidean distance) or the cosine distance.

[0041] The correspondence determination unit 26 may estimate the correspondence between each object representation vector by solving a Hungarian algorithm that associates object representation vectors with each other so as to reduce the L2 distance. The correspondence determination unit 26 can estimate, as the object representation vector corresponding to the occluder OC, an object representation vector that does not belong to any of a plurality of object pairs, i.e., an object representation vector for which no corresponding object is found among the object representation vectors V11 to V13, among the object representation vectors V21 to V24 extracted from the occlusion image OP. The correspondence determination unit 26 outputs information on the determined correspondence (information on the object pair) together with each object representation vector to the correspondence loss calculation unit 27.

[0042] The correspondence loss calculation unit 27 uses the input correspondence relationship information and each object representation vector to calculate, for each object pair, the loss between the object representation vector extracted from the input image IP and the object representation vector extracted from the occlusion image OP. In the example of FIG. 3 , the correspondence loss calculation unit 27 calculates the loss for each object pair corresponding to each of objects A to C. The correspondence loss calculation unit 27 then calculates the sum of the losses for each object pair as the correspondence loss LS1 and outputs it to the parameter update unit 28. The correspondence loss LS1 is an example of a first loss. The correspondence loss calculation unit 27 can use the L2 distance, the cosine distance, or the like as the correspondence loss LS1. As described above, the object representation vector corresponding to the occluder OC does not belong to any object pair and is therefore not subject to loss calculation by the correspondence loss calculation unit 27.

[0043] Parameter updater 28 uses the sum of correspondence loss LS1 input from correspondence loss calculator 27 and reconstruction loss LS2 input from reconstruction loss calculator 24 as the total loss, and updates the parameters of object representation extraction unit 21 and object reconstruction unit 22. As described above, object representation extraction unit 21 and object representation extraction unit 31 are neural networks with the same configuration and have the same parameters set thereto, and therefore parameter updater 28 sets the same parameters for object representation extraction unit 21 and object representation extraction unit 31. Note that object representation extraction units 21 and 31 may be configured using machine learning models other than neural networks, but even in this case, parameter updater 28 updates the parameters so that the parameters of each machine learning model are the same.

[0044] (Learning Process) Fig. 6 is a flowchart of the learning process by the learning device 100. This process is realized by the processor 13 shown in Fig. 2 executing a program prepared in advance and operating as each element shown in Fig. 3.

[0045] First, the object representation extraction unit 21 extracts an object representation vector for each object from the input image IP (step S11). Next, the object reconstruction unit 22 generates a reconstructed image RP and an alpha mask M for each object from the object representation vector for each object (step S12). The object reconstruction unit 22 also estimates the depth of each object from the object representation vector for each object.

[0046] Next, the image combining unit 23 weights the alpha mask M based on the depth of each object, and combines the reconstructed images RP of each object using the weighted alpha mask to generate a reconstructed image RIP (step S13). Next, the reconstruction loss calculation unit 24 calculates the reconstruction loss LS2 from the reconstructed image RIP and the input image IP, and outputs the calculated loss to the parameter update unit 28 (step S14).

[0047] Meanwhile, the occlusion processing unit 25 applies occlusion processing to the input image IP to generate an occlusion image OP (step S15). Next, the object representation extraction unit 31 extracts an object representation vector for each object from the occlusion image OP (step S16). Next, the correspondence determination unit 26 determines the correspondence between the object representation vector extracted from the input image IP and the object representation vector extracted from the occlusion image OP (step S17). Next, the correspondence loss calculation unit 27 calculates the loss for each corresponding object pair and outputs the sum as the correspondence loss LS1.

[0048] Next, the parameter update unit 28 updates the parameters of the object representation extraction units 21 and 31 and the object reconstruction unit 22 using the sum of the input correspondence loss LS1 and reconstruction loss LS2 as the total loss (step S19).

[0049] The learning device 100 repeats the above steps S11 to S19 until a predetermined learning end condition is met, and when the learning end condition is met (step S20: Yes), ends the learning process.

[0050] As described above, in this embodiment, the object representation extraction units 21 and 31 are trained so that the object representation vectors of each object extracted from the input image IP and the object representation vectors of each object extracted from the artificially generated occlusion image OP are similar, and therefore the object representation extraction units 21 and 31 are trained to be able to estimate the shapes of objects that are hidden in the input image with high accuracy. Therefore, by using a trained object representation extraction unit, it is possible to obtain an amodal segmentation model that can estimate hidden areas with high accuracy.

[0051] (Modifications) Below, modifications of the learning device according to the first embodiment will be described. (1) First Modification FIG. 7 is a block diagram showing the schematic configuration of a learning device 100b according to the first modification. The learning device 100b according to the first modification also uses the loss of a reconstructed image generated from an occlusion image OP to train the object representation extraction units 21 and 31. As can be seen from a comparison with FIG. 3 , the learning device 100b according to the first modification includes, in addition to the configuration of the learning device 100a, an object reconstruction unit 32, an image combination unit 33, and a reconstruction loss calculation unit 34. The occlusion image OP is input to the reconstruction loss calculation unit 34. The object reconstruction unit 32, the image combination unit 33, and the reconstruction loss calculation unit 34 have the same configurations as the object reconstruction unit 22, the image combination unit 23, and the reconstruction loss calculation unit 24, respectively, and operate in the same manner. The correspondence relationship determination unit 26 and the correspondence loss calculation unit 27 have the same configurations as those in the first embodiment and operate in the same manner. Descriptions of parts similar to those in the first embodiment will be omitted where appropriate.

[0052] The object representation extraction unit 31 outputs the object representation vectors V21 to V24 extracted from the occlusion image OP to the correspondence determination unit 26 and the object reconstruction unit 32. The object reconstruction unit 32 generates reconstructed images ORPa to ORPC and alpha masks OMa to OMc for the object representation vectors corresponding to the objects A to C, i.e., the object representation vectors excluding the object representation vector of the occluder OC, from the input object representation vectors V21 to V24, and outputs them to the image combination unit 33. The image combination unit 33 generates a reconstructed image ORIP from the reconstructed images ORPa to ORPc and the alpha masks OMa to OMc, and outputs it to the reconstruction loss calculation unit 34. The reconstruction loss calculation unit 34 calculates the reconstruction loss LS3 between the reconstructed image ORIP and the occlusion image OP, and outputs it to the parameter update unit 28. The parameter update unit 28 uses the sum of the correspondence loss LS1, the reconstruction loss LS2, and the reconstruction loss LS3 as the total loss to update the parameters of the object representation extraction units 21 and 31. In this way, according to the first modification, the object representation extraction units 21 and 31 can be trained using occlusion images as well.

[0053] (2) Second Modification FIG. 8 is a block diagram showing a schematic configuration of a learning device 100c according to the second modification. The learning device 100c according to the second modification performs inpainting training. Inpainting refers to complementing hidden portions of an image. As with the second modification, the learning device 100c according to the second modification also uses the loss of a reconstructed image generated from an occlusion image OP to train the object representation extraction units 21 and 31. As can be seen from a comparison with FIG. 7 , the learning device 100c according to the second modification basically has the same configuration as the learning device 100b according to the first modification. However, in the learning device 100c according to the second modification, an input image IP is input to the reconstruction loss calculation unit 34 instead of the occlusion image OP.

[0054] The reconstruction loss calculation unit 34 calculates the reconstruction loss LS4 between the reconstructed image ORIP and the input image IP, and outputs the calculated loss to the parameter update unit 28. The parameter update unit 28 uses the sum of the correspondence loss LS1, the reconstruction loss LS2, and the reconstruction loss LS4 as the total loss to update the parameters of the object representation extraction units 21 and 31. In this way, according to the second modification, the object representation extraction units 21 and 31 can be trained to restore the original input image from the occlusion image.

[0055] In the occlusion process, the occluder may be added as a background representation vector. In this case, the background representation vector is excluded from the objects of correspondence determination by the correspondence relationship determination unit 26 and also excluded from the loss calculation by the correspondence loss calculation unit 27.

[0056] (Fourth Modification) In the above embodiment, an occlusion image in which an object is partially hidden by occlusion processing is generated, but an object representation cannot be extracted from an object that is completely hidden by an occluder, and therefore the object is excluded from the loss calculation by the correspondence loss calculation unit 27. However, if the input image is a video and the object is not completely hidden by occlusion at the beginning of the video but becomes completely hidden by occlusion halfway through the video, the position of the object is known, and therefore the object may be included in the correspondence by the correspondence relationship determination unit 26 and in the loss calculation by the correspondence loss calculation unit 27.

[0057] [Mask Generation Device] Next, a mask generation device (inference device) according to the first embodiment will be described. Fig. 9 is a block diagram showing the functional configuration of a mask generation device 200 according to the first embodiment. Note that the hardware configuration of the mask generation device is basically the same as the configuration of the learning device 100 shown in Fig. 2, so a description thereof will be omitted.

[0058] 9, the mask generation device 200 includes an object representation extraction unit 21x and an object reconstructing unit 22x. The object representation extraction unit 21x and the object reconstructing unit 22x have been trained using any of the above-described learning devices 100, 100a to 100c.

[0059] When generating a mask, i.e., during inference, the mask generation device 200 first generates, as a first step, a reconstructed image RP of each object from a target image TP containing multiple objects, and then, as a second step, generates an alpha mask M from the reconstructed image RP of each object.

[0060] Specifically, in the first step, an unknown target image TP is input to the object representation extraction unit 21x. The object representation extraction unit 21x extracts object representation vectors V31 to V33 for each object included in the target image TP and outputs them to the object reconstruction unit 22x. The object reconstruction unit 22x generates reconstructed images RPa to RPc from the object representations of each object.

[0061] Next, in the second step, the obtained reconstructed image RP of each object is input individually to the object representation extraction unit 21x. The object representation extraction unit 21x extracts an object representation from each reconstructed image RP and outputs it to the object reconstruction unit 22x. As a result, the object reconstruction unit 22 creates and outputs a reconstructed image RP and an alpha mask M for each object. The alpha mask M created in this way also includes the shape of hidden areas of the object in the target image and can be used as a mask for amodal segmentation.

[0062] (Mask Generation Processing) Fig. 10 is a flowchart of the mask generation processing by the mask generation device 200. This processing is realized by the processor 13 shown in Fig. 2 executing a program prepared in advance and operating as each element shown in Fig. 9.

[0063] First, the object representation extraction unit 21x extracts an object representation vector for each object from the input target image TP (step S21). Next, the object reconstruction unit 22x generates a reconstructed image RP and an alpha mask M for each object from the object representation vector for each object (step S22). Next, the reconstructed images RP for each object are input separately to the object representation extraction unit 21x, and the object representation extraction unit 21x extracts an object representation vector from the reconstructed image RP for each object (step S23). Next, the object reconstruction unit 22x generates and outputs a reconstructed image RP and an alpha mask M from the object representation vector for each object (step S24). Then, the processing ends.

[0064] As described above, the mask generation device 200 of the first embodiment can generate an alpha mask from an unknown target image TP containing a mixture of multiple objects by using the object representation extraction unit 21x and the object reconstruction unit 22x that have been trained by the learning devices 100, 100a to 100c. The alpha mask thus generated also represents the shapes of hidden areas of objects included in the target image TP, and can be used as a mask for amodal segmentation.

[0065] (Application Example) The amodal segmentation model of the embodiment can be applied to the fields of medicine and healthcare. For example, the obtained amodal segmentation mask can be applied to the control of a surgical robot in a medical setting and used to estimate the regions of scalpels, medical instruments, etc. from images of the surgical environment. This can support doctors' decision-making during surgery and treatment.

[0066] Second Embodiment Next, a second embodiment will be described. In the first embodiment, a reconstructed image is generated from an object representation vector extracted from an input image IP, and this image is compared with the input image IP, thereby performing unsupervised learning of the object representation extraction unit. In contrast, in the second embodiment, supervised learning of the object representation extraction unit is performed.

[0067] (First Example) Fig. 11 is a block diagram showing the schematic configuration of a learning device according to a second example of the second embodiment. In the first example, ground truth data corresponding to each object included in an input image IP is prepared, and supervised learning of the object representation extraction unit is performed. As shown in Fig. 11, a learning device 110 according to the first example of the second embodiment includes object representation extraction units 21 and 31, an occlusion processing unit 25, a correspondence relationship determination unit 26, a correspondence loss calculation unit 27, a loss calculation unit 41, and a parameter update unit 42.

[0068] Here, the object representation extraction units 21 and 31, the occlusion processing unit 25, the correspondence determination unit 26, and the correspondence loss calculation unit 27 are the same as those in the first embodiment. Therefore, the correspondence determination unit 26 determines the correspondence between the object representation vectors V11 to V13 extracted from the input image IP and, of the object representation vectors V21 to V24 extracted from the occlusion image OP, those other than the object representation vector corresponding to the occluder OC. The correspondence loss calculation unit 27 calculates the correspondence loss LS1 based on the obtained correspondence, and outputs it to the parameter update unit 42.

[0069] Meanwhile, correct answer data of object representation vectors prepared in advance based on the input image IP is input to the loss calculation unit 41. For example, for the input image IP, object representation vectors corresponding to each of objects A to C are prepared as correct answer data and input to the loss calculation unit 41. The loss calculation unit 41 calculates the error between the object representation vectors V11 to V13 and the correct answer data for each corresponding object as a loss, and outputs the sum of these as a loss LS5 to the parameter update unit 42. The parameter update unit 42 updates the parameters of the object representation extraction units 21 and 31 by using the sum of the corresponding loss LS1 input from the corresponding loss calculation unit 27 and the loss LS5 input from the loss calculation unit 41 as the total loss. In this way, in the first example of the second embodiment, object representation vectors are used as object representations, and supervised learning of the object representation extraction unit is performed.

[0070] Second Example Fig. 12 is a block diagram showing a schematic configuration of a second example of the second embodiment. In the second example, too, ground truth data corresponding to each object included in the input image IP is prepared, and supervised learning is performed on the object representation extraction unit. However, whereas the first example uses an object representation vector as an object representation, the second example uses an object mask as an object representation.

[0071] 12, a learning device 120 of the second example of the second embodiment includes an occlusion processing unit 25, object representation extraction units 51 and 52, a correspondence relationship determination unit 53, a correspondence loss calculation unit 54, a loss calculation unit 55, and a parameter update unit 56. Here, the occlusion processing unit 25 is the same as in the first embodiment.

[0072] The object representation extraction unit 51 extracts object masks M11 to M13 corresponding to each object from the input image IP. The object representation extraction unit 52 extracts object masks M21 to M24 corresponding to each object from the occlusion image OP. The object representation extraction units 51 and 52 may be, for example, trained amodal segmentation models. The correspondence determination unit 53 determines the correspondence between the object masks M11 to M13 and the object masks M21 to M24 other than the object mask corresponding to the occluder OC, and outputs the determined correspondence to the correspondence loss calculation unit 54. The correspondence loss calculation unit 54 calculates the loss between the object masks for each object pair based on the input correspondence, and outputs the sum of the determined correspondence loss to the parameter update unit 56 as the correspondence loss LS6.

[0073] Meanwhile, ground truth data of object masks prepared in advance based on the input image IP is input to the loss calculation unit 55. For example, for the input image IP, alpha masks corresponding to each of objects A to C are prepared as ground truth data and input to the loss calculation unit 55. The loss calculation unit 55 calculates the error between the object masks M11 to M13 and the ground truth data for each corresponding object as a loss, and outputs the sum of these as a loss LS7 to the parameter update unit 56. The parameter update unit 56 updates the parameters of the object representation extraction units 51 and 52 by using the sum of the correspondence loss LS6 input from the correspondence loss calculation unit 54 and the loss LS7 input from the loss calculation unit 55 as the total loss. In this way, in the second example of the second embodiment, object masks are used as object representations, and supervised learning of the object representation extraction unit is performed.

[0074] 13 is a block diagram showing the configuration of a learning device according to the third embodiment. A learning device 70 according to the third embodiment includes an occlusion processing unit 71, an object representation extraction unit 72, a correspondence determination unit 73, a loss calculation unit 74, and a parameter update unit 75.

[0075] FIG. 14 is a flowchart of processing by the learning device 70 according to the third embodiment. The occlusion processing means 71 applies occlusion processing to a first image to generate a second image (step S71). The object representation extraction means 72 extracts first object representations, which are object representations of each object included in the first image, and second object representations, which are object representations of each object included in the second image (step S72). The correspondence determination means 73 determines the correspondence between the first object representations and the second object representations (step S73). The loss calculation means 74 calculates a first loss based on the first object representations, the second object representations, and the correspondence (step S74). The parameter update means 75 updates the parameters of the object representation extraction means based on the first loss (step S75).

[0076] According to the learning device 70 of the third embodiment, it is possible to improve the accuracy of estimating an occluded region using an amodal segmentation model.

[0077] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.

[0078] (Supplementary Note 1) A learning device comprising: an occlusion processing means for applying occlusion processing to a first image to generate a second image; an object representation extraction means for extracting a first object representation that is an object representation of each object included in the first image and a second object representation that is an object representation of each object included in the second image; a correspondence determination means for determining a correspondence between the first object representation and the second object representation; a loss calculation means for calculating a first loss based on the first object representation, the second object representation, and the correspondence; and a parameter update means for updating a parameter of the object representation extraction means based on the first loss.

[0079] (Supplementary Note 2) The learning device according to Supplementary Note 1, wherein the loss calculation means calculates the loss between the first object representation and the second object representation for each corresponding object pair, and sets the sum of the losses for each object pair as the first loss.

[0080] (Supplementary Note 3) The learning device according to Supplementary Note 1, wherein the occlusion processing is processing of placing an occluder on at least one object included in the first image to conceal at least a part of the object.

[0081] (Supplementary Note 4) The learning device according to Supplementary Note 3, wherein the correspondence determination means determines a second object representation for which no corresponding first object representation is found to be an object representation corresponding to the occluder, and the loss calculation means excludes the second object representation corresponding to the occluder from being a target for calculating the first loss.

[0082] (Supplementary Note 5) The learning device according to Supplementary Note 1, wherein the object representation extraction means is a machine learning model, and the parameter update means updates parameters of the machine learning model.

[0083] (Supplementary Note 6) A learning device according to Supplementary Note 1, comprising: a first reconstruction means for reconstructing an image of each object included in the first image based on the first object representation; and a first reconstructed image generation means for generating a first reconstructed image by combining the reconstructed images of each object, wherein the loss calculation means calculates a second loss between the first image and the first reconstructed image; and the parameter update means updates parameters of the object representation extraction means based on the first loss and the second loss.

[0084] (Supplementary Note 7) The learning device according to Supplementary Note 6, wherein the occlusion processing is a process of placing an occluder on at least one object included in the first image to conceal at least a part of the object, and the occlusion processing means determines a position at which to place the occluder in the first image based on a position of the object included in the first reconstructed image.

[0085] (Supplementary Note 8) A learning device as described in Supplementary Note 6, comprising: a second reconstruction means for reconstructing an image of each object included in the second image based on the second object representation; and a second reconstructed image generation means for generating a second reconstructed image by combining the reconstructed images of each object, wherein the loss calculation means calculates a third loss between the second image and the second reconstructed image; and the parameter update means updates parameters of the object representation extraction means based on the first loss, the second loss, and the third loss.

[0086] (Supplementary Note 9) A learning device as described in Supplementary Note 6, comprising: a second reconstruction means for reconstructing an image of each object included in the second image based on the second object representation; and a second reconstructed image generation means for generating a second reconstructed image by combining the reconstructed images of each object, wherein the loss calculation means calculates a fourth loss between the first image and the second reconstructed image; and the parameter update means updates parameters of the object representation extraction means based on the first loss, the second loss, and the fourth loss.

[0087] (Supplementary Note 10) The object representation is a vector indicating the characteristics of each object, and the correspondence determination means determines the correspondence between the first object representation and the second object representation based on the distance between a first object representation of each object included in the first image and a second object representation of each object included in the second image.

[0088] (Supplementary Note 11) The learning device according to Supplementary Note 1, wherein the object representation is a mask image indicating the position of the object in the image, the loss calculation means calculates a fifth loss between the first object representation and a correct mask image, and the parameter update means updates parameters of the object representation extraction means based on the first loss and the fifth loss.

[0089] (Supplementary Note 12) A learning method executed by a computer, comprising: applying occlusion processing to a first image to generate a second image; extracting, using an object representation extractor, first object representations that are object representations of each object included in the first image and second object representations that are object representations of each object included in the second image; determining a correspondence between the first object representations and the second object representations; performing a loss calculation that calculates a first loss based on the first object representations, the second object representations, and the correspondence; and updating parameters of the object representation extractor based on the first loss.

[0090] (Supplementary Note 13) A recording medium having recorded thereon a program that causes a computer to execute the following processes: generating a second image by applying occlusion processing to a first image; extracting, using an object representation extractor, first object representations that are object representations of each object included in the first image and second object representations that are object representations of each object included in the second image; determining a correspondence between the first object representations and the second object representations; performing loss calculation that calculates a first loss based on the first object representations, the second object representations, and the correspondence; and updating parameters of the object representation extractor based on the first loss.

[0091] Although the present disclosure has been described above with reference to the embodiments and examples, the present disclosure is not limited to the above-described embodiments and examples. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure.

[0092] 13 Processor 21, 31, 51, 52 Object representation extraction unit 22, 32 Object reconstruction unit 23, 33 Image combination unit 24, 34 Reconstruction loss calculation unit 26, 53 Correspondence determination unit 27, 54 Correspondence loss calculation unit 28, 42, 56 Parameter update unit 41 Loss calculation unit 100, 100a to 100c, 110, 120 Learning device 200 Mask generation device

Claims

1. A learning device comprising: an occlusion processing means for applying occlusion processing to a first image to generate a second image; an object representation extraction means for extracting a first object representation which is an object representation of each object included in the first image and a second object representation which is an object representation of each object included in the second image; a correspondence determination means for determining a correspondence between the first object representation and the second object representation; a loss calculation means for calculating a first loss based on the first object representation, the second object representation, and the correspondence; and a parameter update means for updating parameters of the object representation extraction means based on the first loss.

2. The learning device according to claim 1, wherein the loss calculation means calculates the loss between the first object representation and the second object representation for each corresponding object pair, and defines the sum of the losses for each object pair as the first loss.

3. The learning device according to claim 1, wherein the occlusion processing is a process of placing an occluder on at least one object included in the first image to conceal at least a portion of the object.

4. The learning device according to claim 3, wherein the correspondence determination means determines that a second object representation for which no corresponding first object representation was found is an object representation corresponding to the occluder, and the loss calculation means excludes the second object representation corresponding to the occluder from the calculation of the first loss.

5. The learning device according to claim 1, wherein the object representation extraction means is a machine learning model, and the parameter update means updates parameters of the machine learning model.

6. A learning device as described in claim 1, comprising: a first reconstruction means for reconstructing an image of each object contained in the first image based on the first object representation; and a first reconstructed image generation means for generating a first reconstructed image by combining the reconstructed images of each object, wherein the loss calculation means calculates a second loss between the first image and the first reconstructed image, and the parameter update means updates parameters of the object representation extraction means based on the first loss and the second loss.

7. The learning device described in claim 6, wherein the occlusion processing is a process of placing an occluder on at least one object included in the first image to conceal at least a portion of the object, and the occlusion processing means determines a position at which to place the occluder in the first image based on the position of the object included in the first reconstructed image.

8. A learning device as described in claim 6, comprising: a second reconstruction means for reconstructing an image of each object included in the second image based on the second object representation; and a second reconstructed image generation means for generating a second reconstructed image by combining the reconstructed images of each object, wherein the loss calculation means calculates a third loss between the second image and the second reconstructed image, and the parameter update means updates parameters of the object representation extraction means based on the first loss, the second loss and the third loss.

9. A learning device as described in claim 6, comprising: a second reconstruction means for reconstructing an image of each object included in the second image based on the second object representation; and a second reconstructed image generation means for generating a second reconstructed image by combining the reconstructed images of each object, wherein the loss calculation means calculates a fourth loss between the first image and the second reconstructed image, and the parameter update means updates parameters of the object representation extraction means based on the first loss, the second loss and the fourth loss.

10. The learning device described in claim 1, wherein the object representation is a vector indicating the characteristics of each object, and the correspondence determination means determines the correspondence between the first object representation and the second object representation based on the distance between a first object representation of each object included in the first image and a second object representation of each object included in the second image.

11. The learning device of claim 1, wherein the object representation is a mask image indicating the position of an object in an image, the loss calculation means calculates a fifth loss between the first object representation and a correct mask image, and the parameter update means updates parameters of the object representation extraction means based on the first loss and the fifth loss.

12. A learning method executed by a computer, comprising: applying an occlusion process to a first image to generate a second image; extracting, using an object representation extractor, a first object representation that is an object representation of each object included in the first image and a second object representation that is an object representation of each object included in the second image; determining a correspondence between the first object representation and the second object representation; performing a loss calculation to calculate a first loss based on the first object representation, the second object representation, and the correspondence; and updating parameters of the object representation extractor based on the first loss.

13. A recording medium having recorded thereon a program that causes a computer to execute the following processes: generating a second image by applying occlusion processing to a first image; extracting, using an object representation extractor, a first object representation that is an object representation of each object included in the first image and a second object representation that is an object representation of each object included in the second image; determining a correspondence between the first object representation and the second object representation; performing a loss calculation that calculates a first loss based on the correspondence between the first object representation, the second object representation, and the first loss; and updating parameters of the object representation extractor based on the first loss.

Citation Information

Patent Citations

  • Object recognition device, object recognition method, and object recognition program

    JP2021033374A

  • Learning method, learning model, detection system, detection method, and program

    JP2022072273A

  • Hierarchical occlusion inference module and unseen object instance segmentation system and method using the same

    JP2023131087A