Method and apparatus for mask prediction in instance segmentation

By focusing on the mask prediction loss function of the occluded area during training, the problem of inaccurate mask prediction in occluded scenes in the existing technology is solved, and the performance of the model in occluded scenes is improved.

CN120726313APending Publication Date: 2025-09-30ROBERT BOSCH GMBH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410384839.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing deep learning-based instance segmentation algorithms cannot generate accurate pixel-by-pixel mask prediction information when dealing with complex scenes with occlusions between objects, resulting in poor performance.

Method used

By utilizing the occlusion information between objects during training, a mask prediction loss function is determined, with special attention paid to the occluded areas, and the weight of the loss function is adjusted to improve the performance of the model in occluded scenes.

Benefits of technology

Improves the performance of the mask prediction model in complex occlusion scenes and generates more accurate mask prediction information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726313A_ABST
    Figure CN120726313A_ABST
Patent Text Reader

Abstract

The invention provides a method for training a mask prediction model. The method comprises the following steps: obtaining a training image comprising two or more shielded objects; obtaining relative position information between the two or more shielded objects; based on the relative position information, determining a mask prediction loss function for the training image; and training the mask prediction model based on the training image by using the mask prediction loss function for the training image. The present disclosure also provides a method for performing mask prediction, comprising obtaining an input image in which an estimated region of at least one object is determined; and generating mask prediction information for each of the at least one object based on the input image by using the trained mask prediction model; wherein the trained mask prediction model is obtained through the above training method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to the field of computer vision, and more particularly to methods and apparatus for mask prediction in instance segmentation. Background Art

[0002] Instance segmentation is a classic task in computer vision. It aims to separate each object instance (referred to as an object in this disclosure for simplicity) from the background in an image. Currently, instance segmentation has been applied in many fields, including geographic information systems, medical imaging, autonomous driving, and robotics.

[0003] It has been proposed to use deep learning-based instance segmentation algorithms to perform instance segmentation tasks. Instance segmentation algorithms based on deep learning algorithms generally have good performance and are easy to operate, and therefore have been widely used. However, when processing images of some complex scenes, for example, when there is occlusion between objects, the performance of existing instance segmentation algorithms may be unsatisfactory. This is because occlusion between objects may cause the loss of information in the occluded area, resulting in the inability to output accurate pixel-by-pixel mask prediction information for the occluded area. How to further improve the performance of instance segmentation algorithms when processing images of such complex scenes has become an issue of increasing concern. Summary of the Invention

[0004] This disclosure provides an improved mechanism for mask prediction in instance segmentation. This mechanism improves the training process of the mask prediction model by using occlusion between objects as a guide to encourage the training process to focus more on areas with poor performance, i.e., areas with occlusion. This improves the performance of the trained mask prediction model in complex scenarios such as occlusion.

[0005] According to one aspect of the present disclosure, a method for training a mask prediction model is provided, comprising: obtaining a training image comprising two or more occluded objects; obtaining relative position information between the two or more occluded objects; determining a mask prediction loss function for the training image based on the relative position information; and training the mask prediction model based on the training image using the mask prediction loss function for the training image.

[0006] According to another aspect of the present disclosure, a method for performing mask prediction is provided, comprising obtaining an input image in which an estimated area of ​​at least one object is determined; and generating mask prediction information for each of the at least one object based on the input image using a trained mask prediction model; wherein the trained mask prediction model is obtained by the above-mentioned training method.

[0007] According to another aspect of the present disclosure, a device for mask prediction is provided, comprising: a memory and a processor. The processor is coupled to the memory and configured to execute the method according to any one of the various embodiments of the present disclosure.

[0008] According to yet another aspect of the present disclosure, a computer-readable medium is provided, which stores a computer program including instructions. When the instructions are executed by a processor, the processor is configured to perform the method according to any one of the various embodiments of the present disclosure.

[0009] According to yet another aspect of the present disclosure, a computer program product is provided, comprising computer-executable instructions, which, when executed, cause one or more processors to perform the method according to any one of the various embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Various embodiments of the claimed subject matter will now be described by way of example with reference to the accompanying drawings, in which the same reference numerals are used throughout the different drawings to designate the same or similar components.

[0011] Figure 1 A schematic diagram illustrating the overall architecture of an instance segmentation algorithm according to an example embodiment of the present disclosure is shown.

[0012] Figures 2a to 2d A schematic diagram of determining relative position information in an occlusion scene between two objects according to an example embodiment of the present disclosure is shown.

[0013] Figure 3 A schematic diagram illustrating a principle of determining a loss function for a mask prediction model in an occlusion scenario between two objects according to an example embodiment of the present disclosure is shown.

[0014] Figures 4a to 4d A schematic diagram illustrating a principle of determining a loss function for a mask prediction model in a multi-object occlusion scenario according to an example embodiment of the present disclosure is shown.

[0015] Figure 5 An example flow chart of operations for training a mask prediction model according to an example embodiment of the present disclosure is shown.

[0016] Figure 6An example flow chart of a method for performing mask prediction according to an example embodiment of the present disclosure is shown.

[0017] Figure 7 A block diagram of an apparatus for mask prediction according to an example embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0018] In the following description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure. However, one skilled in the relevant art will recognize that the present disclosure can be practiced without one or more of the specific details, or can be practiced using alternative methods, components, etc. In some instances, well-known structures and operations are not shown or described in detail to avoid unnecessarily obscuring the present disclosure.

[0019] Some of the currently popular instance segmentation algorithms based on deep learning typically include two processing stages when performing instance segmentation tasks. The first processing stage can be called the detection stage. In the detection stage, the estimated area where a specific object is located can be located by performing the target detection task, wherein the estimated area can be represented by a bounding box surrounding the specific object. The second processing stage can be called the mask prediction stage or the segmentation stage. In the mask prediction stage, semantic segmentation can be performed within the above-mentioned estimated area to generate a pixel-by-pixel binary mask covering the visible part of the specific object. Typical examples of the above-mentioned instance segmentation algorithms based on deep learning may include Mask R-CNN, Cascade Mask R-CNN, Mask Scoring R-CNN, and the like.

[0020] The following will refer to Figure 1 The principle of the above instance segmentation algorithm is further described in detail. Figure 1 A schematic diagram illustrating an overall architecture 100 of an instance segmentation algorithm according to an example embodiment of the present disclosure is shown.

[0021] The input of instance segmentation algorithm is image data, for example, Figure 1 An input image 102 is shown. The input image 102 may be an image captured from a real scene and may include a background and one or more objects. Figure 1 Schematically, an object 104 (shown as a triangle) is shown in the input image 102. Instance segmentation algorithms aim to accurately separate the object 104 from the background.

[0022] The object detection model 106 and the mask prediction model 110 can be used to respectively perform the operations of the two processing stages discussed above. The object detection model 106 can be used to perform the operations of the detection stage, and the mask prediction model 110 can be used to perform the operations of the mask prediction stage.

[0023] The object detection model 106 can receive the input image 102 to detect the object 104 in the input image 102 and determine the estimated area where the object 104 is located. The object detection model 106 can be implemented using any algorithm known in the prior art for performing target detection tasks, such as Faster R-CNN, DETR, YOLO, etc. In one example, the output result of the object detection model 106 can be position information for the detected object 104. The position information can be used to determine the estimated area where the object 104 is located. In one example, the position information for the object 104 can be embodied as a bounding box 108 surrounding the object 104, and the area defined by the bounding box 108 is the estimated area of ​​the object 104. The position information for the object 104 generated by the object detection model 106 may include coordinates associated with the bounding box 108. As Figure 1 As shown in the example of , the bounding box 108 can be rectangular. In one example, the above coordinates can be the coordinates of the two diagonal vertices of the bounding box 108, for example, the coordinates of the upper left vertex and the coordinates of the lower right vertex of the bounding box 108. In another example, for example, the above coordinates can be the coordinates of the center point of the bounding box 108 and the length and width of the bounding box 108. It should be understood that the embodiment of the shape and coordinates of the bounding box 108 discussed above is only an example. In other examples, the bounding box 108 can have any other shape as long as the bounding box 108 can define the estimated area occupied by the detected object in the input image, and the position information for the object 104 generated by the object detection model 106 can also be embodied as any other form of coordinate information.

[0024] The mask prediction model 110 may receive an image in which an estimated region of the object 104 is determined (e.g., by a bounding box 108) to perform the operations of the mask prediction stage described above. The mask prediction model 110 may perform mask prediction within the estimated region defined by the bounding box 108 and generate mask prediction information. This mask prediction information, which may also be referred to as pixel-by-pixel mask information, indicates whether each pixel is associated with the detected object 104. Typically, the mask prediction model 110 may utilize a normalization function (e.g., a softmax function, a sigmoid function, etc.) to output a prediction value for each pixel within the estimated region of the object 104, the prediction value indicating the probability that each pixel is associated with the object 104. A thresholding operation may then be performed on the probabilities to generate the mask prediction information. For example, if the probability of a pixel is greater than a threshold, the mask prediction information may indicate that the pixel should be included as a point on the mask of the object. Conversely, if the probability of a pixel is less than a threshold, the mask prediction information may indicate that the pixel should not be included as a point on the mask of the object.

[0025] Then, based on the mask prediction information generated by the mask prediction model 110, an output of the instance segmentation task may be generated, for example, an output image 112. The output image 112 may be rendered in which the visible portion of the object 104 is covered with a mask. Figure 1 In the example shown, the input image 102 includes only one object. However, it should be understood that when the input image 102 includes multiple objects, Figure 1 The architecture shown in FIG11 is used to detect one object at a time and predict its mask. In this case, each object in the generated output image 112 can be covered with a corresponding mask. To facilitate the distinction between multiple objects, the masks of each object can be rendered with different colors. In one example, in addition to the mask, the output image 112 can also show a bounding box around each object, a classification label for each object, and other possible information associated with the instance segmentation task.

[0026] Since in the instance segmentation algorithm discussed above, the object detection model 106 and the mask prediction model 110 jointly perform the instance segmentation task, in one example, the object detection model 106 and the mask prediction model 110 can be collectively referred to as an instance segmentation model 114.

[0027] The instance segmentation model 114 is trained. During the training process, a training image can be input into the instance segmentation model so that the instance segmentation model outputs a prediction result (e.g., mask prediction information). A loss function can be defined that characterizes the difference between the prediction result output by the instance segmentation model for the training image and the true value of the training image. The training process aims to gradually converge the loss function by adjusting the model parameters (e.g., weights, biases) of the instance segmentation model. The example of training the instance segmentation model 114 discussed above (i.e., taking the object detection model 106 and the mask prediction model 110 as a whole) can also be referred to as an end-to-end training method. In one example, instead of adopting an end-to-end training method, the object detection model 106 and the mask prediction model 110 in the instance segmentation model 114 can be trained separately. For example, a loss function for the object detection model 106 can be separately defined and the loss function can be converged by adjusting the model parameters of the object detection model 106 during the training process, and a loss function for the mask prediction model 110 can be defined and the loss function can be converged by adjusting the model parameters of the mask prediction model 110 during the training process.

[0028] Through the above training process, the optimized model parameters can be determined, that is, a set of model parameters that make the loss function converge. Based on the optimized model parameters, a trained instance segmentation model can be obtained, for example, Figure 1An instance segmentation model 114 is shown for performing an instance segmentation task on an input image.

[0029] As mentioned in the background technology section, images captured from the real world often have some complex situations. For example, there may be occlusions between objects, making parts of specific objects invisible in the image. When processing such input images, the performance of existing instance segmentation models is usually unsatisfactory. This complex situation has a particularly significant impact on the above-mentioned mask prediction stage. This is because the occluded areas contain visual information from different objects and may therefore mislead the mask prediction model, resulting in the inability to generate accurate mask prediction information. How to further improve the performance of instance segmentation models in the face of these complex situations in order to obtain accurate instance segmentation results has become an important challenge.

[0030] The mechanism proposed in the present disclosure improves the training process of the mask prediction model, which uses the occlusion between objects as a guide to promote the training process to pay more attention to areas with poor performance, that is, areas with occlusion. In an example embodiment, the mechanism of the present disclosure proposes to determine the mask prediction loss function for the training image based on the relative position information between objects with occlusion. More specifically, the corresponding mask prediction loss is assigned corresponding weights based on the different spatial regions of the objects with occlusion, so that the training process can pay more attention to the areas with occlusion. In this way, the performance of the trained mask prediction model is improved when facing complex scenes such as occlusion.

[0031] Figures 2a to 2d A schematic diagram of determining relative position information in an occlusion scene between two objects according to an example embodiment of the present disclosure is shown.

[0032] Figure 2a A training image 202 is shown, which includes a pair of objects i (represented by a circle) and j (represented by a triangle) that are occluded. In this disclosure, such a scene is referred to as a two-object occlusion scene. For the sake of clarity, the pair of objects that are occluded are referred to as the first object i and the second object j, respectively. Figure 2a As shown, a portion of the first object i covers a portion of the second object j, resulting in the covered portion of the second object j being invisible in the training image 202 .

[0033] The model for performing the object detection task can be used to determine the estimated area of ​​the first object i and the estimated area of ​​the second object j based on the training image 202. In one example, the above-mentioned Figure 1 The object detection model 106 discussed is used to determine the estimated area of ​​the first object i and the estimated area of ​​the second object j, respectively. Figure 2bAs shown in the example of , the area defined by the bounding box 204 surrounding the first object i is the estimated area of ​​the first object i. Figure 2c As shown in the example of , the area defined by the bounding box 206 surrounding the second object j is the estimated area of ​​the second object j.

[0034] In another example, other methods may be used to respectively determine the estimated area of ​​the first object i and the estimated area of ​​the second object j. For example, the estimated area of ​​the first object i and the estimated area of ​​the second object j may be respectively determined by manual annotation (e.g., annotating the bounding boxes 204 and 206, respectively).

[0035] After determining the estimated areas of the first object i and the second object j, the relative position information between the first object i and the second object j can be determined. Figure 2d As shown, the bounding box 204 surrounding the first object i and the bounding box 206 surrounding the second object j may be combined, and the combined bounding box 208 may indicate relative position information between the first object i and the second object j.

[0036] refer to Figure 2d It can be seen that based on the relative position information, the estimated area surrounded by the bounding box 204 of the first object i can be divided into two sub-areas: a sub-area 210-1 represented by white and a sub-area 210-2 represented by black. More specifically, the division of the above sub-areas 210-1 and 210-2 can be achieved based on the left edge of the bounding box 206 of the second object j. Among them, there is no inter-object occlusion in the sub-area 210-1. On the contrary, the sub-area 210-1 is an area owned by the first object i, and therefore, it can be called the owned sub-area of ​​the first object i. In contrast, there is inter-object occlusion in the sub-area 210-2. In other words, the sub-area 210-2 is a shared area between the first object i and the second object j. Therefore, it can be called the shared sub-area of ​​the first object i.

[0037] Similarly, based on this relative position information, the estimated area of ​​the second object j surrounded by the bounding box 206 can be divided into two sub-areas: a sub-area 212-1 represented by black and a sub-area 212-2 represented by white. More specifically, the division of the sub-areas 212-1 and 212-2 can be achieved based on the right edge of the bounding box 204 of the first object i. There is inter-object occlusion in sub-area 212-1. In other words, this sub-area 212-1 is shared by the first object i and the second object j, and therefore can be referred to as a shared sub-area of ​​the second object j. In contrast, there is no inter-object occlusion in sub-area 212-2. Instead, this sub-area 212-2 is owned by the second object j, and therefore can be referred to as a owned sub-area of ​​the second object j.

[0038] The relative position information between the first object i and the second object j mentioned above, more specifically, the information of the sub-region division of each object can be used to determine the loss function of the mask prediction model, so that the training process of the mask prediction model pays more attention to the occluded areas, thereby improving the model performance.

[0039] Figure 3 A schematic diagram 300 shows a principle of determining a loss function for a mask prediction model in an occlusion scenario between two objects according to an example embodiment of the present disclosure.

[0040] As above combined Figure 1 As discussed above, the object detection model and the mask prediction model can be trained separately. When training the mask prediction model, a loss function for the mask prediction model can be defined. In this disclosure, the loss function for the mask prediction model can also be referred to as the mask prediction loss function for the training image. Figure 3 In the two-object occlusion scenario discussed, there are two objects in the training image: the first object i and the second object j. The mask prediction loss L for the first object i can be calculated separately. i and the mask prediction loss L for the second object j j , and then determine the mask prediction loss function L for the training image by summing the two according to the following equation (1) total :

[0041] L total = L i +L j (1)

[0042] We will first discuss the mask prediction loss L for the first object i. i How is it calculated?

[0043] The mask prediction model can be used to obtain the predicted mask of the first object i For example, if combined Figure 1 Discussed, the mask prediction model may receive a training image in which an estimated region of a first object i is determined (e.g., by a bounding box) and generate predicted mask information for the first object i. Then, a predicted mask of the first object i may be generated based on the predicted mask information for the first object i.

[0044] The true value mask m of the first object i can be obtained i The ground truth mask m can be obtained in any manner known in the art. iFor example, the ground truth mask can be automatically generated using any mask prediction model that has been trained to produce accurate ground truth masks, or can be manually generated using manual annotation, etc.

[0045] Then, the predicted mask of the first object i can be converted into the shared sub-region based on the relative position information between the first object i and the second object j, more specifically, based on the self-owned sub-region and the shared sub-region into which the estimated region of the first object i is divided. Correspondingly, it is divided into two parts: the own prediction mask corresponding to the own sub-region of the first object i and the shared prediction mask corresponding to the shared subregion of the first object i

[0046] Similarly, the ground truth mask m of the first object i can be transformed into the ground truth mask m based on the relative position information between the first object i and the second object j, more specifically, based on the self-owned sub-region and the shared sub-region into which the estimated region of the first object i is divided. i Correspondingly divided into two parts: the own ground truth mask corresponding to the own sub-region of the first object i and the shared ground-truth mask corresponding to the shared subregion of the first object i

[0047] Next, for the first object i, the own sub-region mask prediction loss associated with the first object i can be calculated as Among them, the own sub-region mask prediction loss Can represent own prediction mask With its own truth mask The difference between the own sub-region mask prediction loss The calculation of can be implemented using any algorithm for calculating loss functions known in the prior art, for example, using a binary cross entropy algorithm. Similarly, for the first object i, the shared sub-region mask prediction loss associated with the shared sub-region of the first object i can be calculated. The shared sub-region mask prediction loss Can characterize shared prediction masks Shared ground truth mask The difference between them can be used to calculate the prediction loss of its own sub-region mask The same loss function calculation algorithm is used to calculate the shared sub-region mask prediction loss

[0048] Then, the loss can be predicted based on the own sub-region mask according to the following equation (2): and shared subregion mask prediction loss To calculate the mask prediction loss L of the first object i i, where the shared sub-region mask prediction loss Weights can be applied

[0049]

[0050] The mask prediction loss applied to the shared sub-region may be determined based on the relationship between the size of the area of ​​the first object i associated with the shared sub-region and the size of the total area of ​​the first object i. Weight In one example, the prediction mask The number of pixels in the covered area corresponding to the shared sub-region, i.e., the shared prediction mask The number of pixels contained in the covered area represents the size of the area associated with the shared sub-region of the first object i. The total area of ​​the first object i is represented by the number of pixels contained in the covered area. In one example, the weight can be calculated more specifically according to the following equation (3):

[0051]

[0052] where a i Prediction mask The number of pixels contained in the covered area; Shared prediction mask The number of pixels contained in the covered area; s is a hyperparameter, which can be initially set to a specific constant and can be scaled up and down according to the model training situation. It can be seen that according to the above equation (3), as the size of the area associated with the shared sub-region of the first object i increases, the loss imposed on the shared sub-region mask prediction is Weight It will also increase, so that more attention can be paid to the shared sub-region during training, that is, the occluded part in the training image.

[0053] The mask prediction loss L for the second object j can then be calculated in a similar manner as discussed above for the first object i. j In short, the predicted mask of the second object j can be obtained and the ground truth mask m j Then, the predicted mask of the second object j can be transformed into Correspondingly, it is divided into two parts: the own prediction mask corresponding to the own sub-region of the second object j and the shared prediction mask corresponding to the shared subregion of the second object j The true value mask m of the second object j can be similarly j Correspondingly, it is divided into two parts: the own ground truth mask corresponding to the own sub-region of the second object j and the shared ground-truth mask corresponding to the shared subregion of the second object j Then, for the second object j, the own sub-region mask prediction loss can be calculated separately and shared subregion mask prediction loss Among them, the own sub-region mask prediction loss is associated with the own subregion of the second object j and represents its own prediction mask With its own truth mask The difference between the shared sub-region mask prediction loss is associated with the shared subregion of the second object j and represents the shared prediction mask Shared ground truth mask The difference between.

[0054] Then, the mask prediction loss L for the second object j can be calculated according to the following equation (4): j , where the shared sub-region mask prediction loss Weights can be applied

[0055]

[0056] In one example, the weights can be calculated according to the following equation (5):

[0057]

[0058] where a j Prediction mask The number of pixels contained in the covered area; Shared prediction mask The number of pixels contained in the covered area; s is a hyperparameter, which can be initially set to a specific constant and can be enlarged and reduced according to the model training situation.

[0059] It should be understood that the equations (3) and (5) discussed above for calculating the weights are only used as an example. In other examples, equations (3) and (5) can be replaced with other forms that can be conceived by those skilled in the art. For example, and / or Replace with the cross-sum form, such as The hyperparameter s can be removed or similar parameters can be added; the exponential function exp() used can be replaced with other nonlinear functions, and so on.

[0060] In the above discussion, the weight applied to the shared sub-region mask prediction loss can be understood as a scalar value used to balance different loss components. By dividing the mask prediction loss for each object into its own sub-region mask prediction loss and the shared sub-region mask prediction loss, and applying the above weight to the shared sub-region mask prediction loss, the training process of the mask prediction model can pay more attention to the areas with occlusions, thereby improving the performance of the trained mask prediction model in the face of complex scenes such as occlusions.

[0061] In one example, in addition to the pair of occluded objects i and j discussed above, the training image may also include one or more other pairs of occluded objects. In this case, the mask prediction losses of the other one or more pairs of occluded objects can be calculated in a similar manner as discussed above for objects i and j, and compared with the mask prediction loss L for the first object i. i and the mask prediction loss L for the second object j j Add to determine the mask prediction loss function L for the training image total .

[0062] In one example, in addition to the pair of occluded objects i and j discussed above, the training image may also include one or more non-occluded objects. In this case, the mask prediction losses of the other one or more non-occluded objects can be calculated in a conventional manner and compared with the mask prediction loss L of the first object i. i and the mask prediction loss L for the second object j j Add to determine the mask prediction loss function L for the training image total .

[0063] Generally, the mask prediction loss of each object in the training image can be calculated separately, and then the calculated mask prediction losses are summed to determine the mask prediction loss function L for the training image. total , as shown in equation (6) below:

[0064]

[0065] Where N is the number of objects contained in the training image, and L k is the mask prediction loss calculated for the k-th object.

[0066] In one example, when combined with Figure 1When using the end-to-end approach for training, we can further determine the object detection loss function associated with the object detection model and compare it with the mask prediction loss function L for the training image above. total Added together as the loss function for the end-to-end training process.

[0067] The above discussion relates to determining the mask prediction loss function for training images in a two-object occlusion scenario. In one example, an object in a training image may be occluded by two or more other objects. In this disclosure, for simplicity, such a scenario is referred to as a multi-object occlusion scenario. The above mechanism also applies to this multi-object occlusion scenario.

[0068] Figures 4a to 4d A schematic diagram showing the principle of determining a loss function for a mask prediction model in a multi-object occlusion scenario according to an exemplary embodiment of the present disclosure is shown. Figures 4a to 4d In the example scenario discussed, there is occlusion between three objects. For the sake of clarity, in this disclosure, the three objects are referred to as the first object i, the second object j, and the third object p, respectively. Figure 4a The estimated areas of the three objects are shown, where the estimated areas can be combined with the above Figure 1 and Figure 2. Figure 4a In the example shown, the estimated area of ​​the first object i may be defined by a bounding box 402, the estimated area of ​​the second object j may be defined by a bounding box 404, and the estimated area of ​​the third object p may be defined by a bounding box 406. Figure 4a As shown, the combined bounding boxes 402 , 404 , and 406 may indicate relative position information between the first object i, the second object j, and the third object p.

[0069] In the following, we first discuss how to calculate the mask prediction loss of the first object i.

[0070] Based on the relative position information between the first object i, the second object j, and the third object p, the estimated area of ​​the first object i can be divided into multiple sub-areas. Figure 4b As shown, the estimated area of ​​the first object i can be divided into: an owned sub-area 408-1, which is owned by the first object i; a first shared sub-area 410-1, which is shared by the first object i and the second object j; a second shared sub-area 410-2, which is shared by the first object i and the third object p; and a third shared sub-area 410-3, which is shared by the first object i, the second object j and the third object p.

[0071] Based on the above division of the estimated area of ​​the first object i, the above combined Figure 3The method discussed above can be used to divide the predicted mask and the true value mask of the first object i respectively. Then, the above method can be used in a similar way. Figure 3 In the manner discussed, based on the divided prediction mask and the true value mask, the own sub-region mask prediction loss associated with the own sub-region 408-1, the first shared sub-region mask prediction loss associated with the first shared sub-region 410-1, the second shared sub-region mask prediction loss associated with the second shared sub-region 410-2, and the third shared sub-region mask prediction loss associated with the third shared sub-region 410-3 are calculated respectively. Next, the mask prediction loss of the first object i can be calculated based on the above-mentioned own sub-region mask prediction loss, the first shared sub-region mask prediction loss, the second shared sub-region mask prediction loss, and the third shared sub-region mask prediction loss according to equation (2). Wherein, the first shared sub-region mask prediction loss, the second shared sub-region mask prediction loss, and the third shared sub-region mask prediction loss can each be applied with a corresponding weight.

[0072] In one example, the weight to be applied to the shared sub-region mask prediction loss associated with a shared sub-region can be determined similarly according to equation (3) based on the relationship between the size of the area associated with the shared sub-region and the size of the total area of ​​the first object i.

[0073] In one example, considering that the third shared sub-region 410-3 is formed due to the occlusion between the first object i and the second object j and the occlusion between the first object i and the third object p, it can be Figure 4a The multi-object occlusion scene shown is split into Figure 4c and Figure 4d The scene of occlusion between two objects is shown. Figure 4c shows a two-object occlusion scene between a first object i and a second object j, and Figure 4d The two-object occlusion scene between the first object i and the third object p is shown. In this case, the above Figure 3 The discussion focuses on Figure 4c Calculate the weights applied to the shared subregion mask prediction loss and for Figure 4d A weight applied to the shared sub-region mask prediction loss is calculated. The two calculated weights may then be accumulated as the weight of the third shared sub-region mask prediction loss associated with the third shared sub-region 410 - 3 .

[0074] The above has been discussed in detail about calculating the mask prediction loss of the first object i. The mask prediction loss of the second object j and the mask prediction loss of the third object p can be calculated similarly. The mask prediction losses of the various objects can then be summed to determine the mask prediction loss function for the training image.

[0075] Generally speaking, in a multi-object occlusion scenario, the mask prediction loss for each object can be calculated according to the following equation (7):

[0076]

[0077] Among them, L k The loss for mask prediction of the kth object in the training image, is the object’s own sub-region mask prediction loss, M is the number of shared sub-regions into which the estimated region of the object is divided, predict the loss for the shared sub-region mask associated with the qth shared sub-region, and is the corresponding weight applied to the shared sub-region mask prediction loss associated with the qth shared sub-region. Further, the mask prediction loss function for the training image can be determined according to the above equation (6).

[0078] Figure 5 An example flow chart 500 of operations for training a mask prediction model according to an example embodiment of the present disclosure is shown.

[0079] In step S502 , a training image including two or more objects with occlusions may be obtained.

[0080] In step S504, relative position information between the two or more occluded objects may be obtained.

[0081] In step S506, a mask prediction loss function for the training image may be determined based on the relative position information.

[0082] In step S508 , the mask prediction model may be trained based on the training image using the mask prediction loss function for the training image.

[0083] After the above training process, the trained mask prediction model can be used to perform mask prediction. Figure 6 An example flowchart 600 of a method for performing mask prediction according to an example embodiment of the present disclosure is shown.

[0084] In step S602 , an input image in which an estimated region of at least one object is determined may be obtained.

[0085] In step S604, the trained mask prediction model may be used to generate mask prediction information for each of the at least one object based on the input image. The trained mask prediction model is obtained by combining Figure 5The results are obtained by training the discussed method for training mask prediction models.

[0086] In some embodiments, the method for performing mask prediction may further include: generating a mask for each object in the at least one object based on the mask prediction information. Figure 1 As discussed, for each object, the mask prediction information can indicate whether a specific pixel should be used as a point on the mask of the object. Thus, based on the mask prediction information, a mask covering the object can be generated in the output image to clearly indicate the object detected by performing the instance segmentation task in a visual manner.

[0087] Figure 7 A block diagram of an apparatus 700 for mask prediction according to an exemplary embodiment of the present disclosure is shown. The apparatus 700 can implement the above-mentioned Figure 5 The operation of the method discussed for training the mask prediction model is combined with the above Figure 6 Operations of the method discussed for performing mask prediction.

[0088] The example apparatus 700 includes a processor 704 connected to an internal communication bus 702, the processor 704 being used to execute instructions in a memory 706 to implement the method for training a mask prediction model and the method for performing mask prediction described in detail above. Examples of the processor 704 may include a central processing unit (CPU), a microcontroller, and the like. Examples of the processor 704 may also include a graphics processing unit (GPU) dedicated to performing graphics processing operations, and the like. The memory 706 suitable for tangibly embodying computer program instructions and data includes various forms of memory, such as EPROM, EEPROM, and flash memory devices, and the like. The apparatus 700 may also include an input interface 708 and an output interface 710. The input interface 708 is used to receive input signals and data, such as the training images and input images discussed above. The output interface 710 is used to send output signals and data, such as the output images discussed above.

[0089] The computer program may include computer-executable instructions for causing the processor 704 of the apparatus 700 to perform the method for training a mask prediction model and the method for performing mask prediction disclosed herein. The program may be stored on any data storage medium, including a memory. For example, the program may be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or a combination thereof. The process / method steps described in this disclosure may be performed by a programmable processor executing the program instructions to perform the method, steps, or operations by operating on input data and generating output.

[0090] The embodiments of the present disclosure may be implemented in a computer-readable medium. The computer-readable medium may store a computer program including instructions. In one example aspect, when the instruction is executed by a processor, it may cause the processor to: obtain a training image including two or more objects with occlusions; obtain relative position information between the two or more objects with occlusions; determine a mask prediction loss function for the training image based on the relative position information; and train the mask prediction model based on the training image using the mask prediction loss function for the training image. In another example aspect, when the instruction is executed by the processor, it may cause the processor to: obtain an input image in which an estimated area of ​​at least one object is determined; and generate mask prediction information for each of the at least one object based on the input image using a trained mask prediction model; wherein the trained mask prediction model is obtained by the above-mentioned training method. In another example aspect, when the instruction is executed by the processor, it may cause the processor to implement the above-mentioned method in combination with Figure 5 Other operations discussed in the method for training mask prediction models are combined with the above Figure 6 Other operations of the method discussed for performing mask prediction.

[0091] The embodiments of the present disclosure may be implemented in a computer program product. The computer program product may include instructions. In one example aspect, when the instruction is executed, it may enable one or more processors to: obtain a training image including two or more objects with occlusions; obtain relative position information between the two or more objects with occlusions; determine a mask prediction loss function for the training image based on the relative position information; and train the mask prediction model based on the training image using the mask prediction loss function for the training image. In another example aspect, when the instruction is executed, it may enable one or more processors to: obtain an input image in which an estimated area of ​​at least one object is determined; and generate mask prediction information for each of the at least one object based on the input image using a trained mask prediction model; wherein the trained mask prediction model is obtained by the above-mentioned training method. In another example aspect, when the instruction is executed by the processor, it may enable one or more processors to implement the above combined method. Figure 5 Other operations discussed in the method for training mask prediction models are combined with the above Figure 6 Other operations of the method discussed for performing mask prediction.

[0092] In addition to what is described herein, various modifications may be made to the disclosed embodiments and implementations of the present invention without departing from the scope of the disclosed embodiments and implementations of the present invention. Therefore, the description and examples herein should be interpreted as illustrative rather than limiting. The scope of the present invention should be measured solely by reference to the claims.

Claims

1. A method for training a mask prediction model, comprising: Obtaining a training image including two or more occluded objects; Obtaining relative position information between the two or more occluded objects; Determining a mask prediction loss function for the training image based on the relative position information; as well as The mask prediction model is trained based on the training images using the mask prediction loss function for the training images.

2. The method according to claim 1, wherein The obtaining of relative position information comprises: determining an estimated area for each of the two or more occluded objects; and The relative position information is obtained based on the estimated area of ​​each object.

3. The method according to claim 2, wherein: The estimated area of ​​each object is defined by a bounding box surrounding the object.

4. The method according to claim 3, wherein The bounding box is determined by an object detection model; or The bounding boxes are determined by manual annotation.

5. The method according to claim 4, wherein The bounding box is determined by an object detection model; The object detection model is used to generate position information for each object based on the training image; and The bounding box is determined using the position information.

6. The method according to claim 1, wherein Determining the mask prediction loss function includes: Calculating a mask prediction loss for each of the two or more occluded objects; and The mask prediction losses of each object are summed to determine the mask prediction loss function for the training image.

7. The method according to claim 6, wherein: The calculation of the mask prediction loss for each object includes: Based on the relative position information, dividing the estimated area of ​​the object into a plurality of sub-areas; calculating a sub-region mask prediction loss associated with each sub-region in the plurality of sub-regions; and The mask prediction loss of the object is calculated based on the sub-region mask prediction loss.

8. The method according to claim 7, wherein: The calculating the sub-region mask prediction loss associated with each sub-region includes: Get the predicted mask and true value mask of the object; Based on the division of the estimated area, the prediction mask and the true value mask are divided respectively; and A sub-region mask prediction loss associated with the sub-region is calculated based on the predicted mask of the partition and the ground-truth mask of the partition.

9. The method according to claim 8, wherein The prediction mask is determined based on mask prediction information output by the mask prediction model for the object.

10. The method according to claim 7, wherein: The multiple sub-areas include: An owned subregion owned by the object; and One or more shared subregions that the object shares with one or more other objects.

11. The method according to claim 10, wherein: Calculating the mask prediction loss of the object based on the sub-region mask prediction loss includes: A weight is applied to the sub-region mask prediction loss associated with each of the one or more shared sub-regions.

12. The method according to claim 11, wherein The weight applied to the sub-region mask prediction loss associated with each shared sub-region is determined based on a relationship between a size of an area of ​​the object associated with the shared sub-region and a size of a total area of ​​the object.

13. The method according to claim 12, wherein The total area of ​​the object is represented by the number of pixels contained in the area covered by the predicted mask of the object; and The size of the area of ​​the object associated with the shared sub-region is characterized by the number of pixels corresponding to the shared sub-region in the area covered by the prediction mask of the object.

14. The method according to claim 1, wherein The training image also includes one or more objects without occlusion; and The mask prediction loss function for the training image is further determined based on a mask prediction loss for each of the one or more non-occluded objects.

15. The method according to claim 4, wherein The using the mask prediction loss function for the training image to train the mask prediction model based on the training image includes: determining an object detection loss function associated with the object detection model; The object detection model and the mask prediction model are trained together in an end-to-end manner using the object detection loss function and the mask prediction loss function.

16. A method for performing mask prediction, comprising: obtaining an input image in which an estimated region of at least one object is determined; as well as generating mask prediction information for each of the at least one object based on the input image using the trained mask prediction model; The trained mask prediction model is obtained by training using the method according to any one of claims 1-15.

17. The method according to claim 16, further comprising: A mask for each of the at least one object is generated based on the mask prediction information.

18. A device for mask prediction, comprising: Memory; A processor coupled to the memory, the processor being configured to perform the method according to any one of claims 1-15 or any one of claims 16-17.

19. A computer readable medium storing a computer program comprising instructions which, when executed by a processor, cause the processor to be configured to perform the method according to any one of claims 1 to 15 or any one of claims 16 to 17.

20. A computer program product comprising computer executable instructions which, when executed, cause one or more processors to perform the method according to any one of claims 1-15 or any one of claims 16-17.

Citation Information

Cited By

  • Driver direct view calculation model and test method

    CN121564690A