Methods of training a machine learning model for instance segmentation, trained machine learning models, apparatuses, systems, and computer programs for instance segmentation

CN122820733APending Publication Date: 2026-09-25CARL ZEISS MICROSCOPY GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610287834.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-03-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0010]在图1A的图像10中,识别出了对象11,但是该对象11的面积并未被完全准确地捕捉,而且,实际对象边界外的片段12A和12B也被识别为属于该对象

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820733A_ABST
    Figure CN122820733A_ABST
Patent Text Reader

Abstract

Disclosed herein is a method of training a machine learning model for instance segmentation, in which unlabelled training images of objects to be segmented are used, and at least one region is known not to contain an object to be segmented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a method for training a machine learning model for instance segmentation of objects in an image, a correspondingly trained machine learning model, an apparatus for instance segmentation with such a trained machine learning model, a system having such an apparatus, and a corresponding computer program. Background Technology

[0002] Instance segmentation is a task in computer vision that identifies and separates individual instances of objects in an image. This process involves detecting the regions or boundaries of each instance of an object and individually identifying each object within the same category. Unlike semantic segmentation, instance segmentation distinguishes between different instances of objects within the same category.

[0003] Instance segmentation is typically performed using trained machine learning models. One application of this type of instance segmentation is microscopy, where objects are identified within images acquired through a microscope. An example is the microscopic acquisition of tissue samples, where it is necessary to detect specific types of cells, such as tumor cells.

[0004] An example of this type of machine learning logic is Mask R-CNN, described by K. He et al. in "Mask R-CNN". IEEE Transactions on Pattern Analysis and Machine Intelligence February 1, 2020, Vol. 42, No. 2, pp. 386-397, where Mask R-CNN is short for Mask Region-based Convolutional Neural Network.

[0005] Here, a convolutional neural network (roughly called a "folded neural network" in German) is used for instance segmentation.

[0006] Another approach is query-based instance segmentation based on a transformer architecture. Examples include the Detection Transformer (DETR), see, for example, "End-to-End object detection with Transformers" by N. Carion et al., published at Computer Vision - ECCV 2020; the Mask Transformer, see, for example, "Per pixel classification is not all you need for semantic segmentation" by B. Cheng et al., published at NeurIPS 2021; or its subsequent development, Mask2Former, by B. Cheng et al., "Masked-attention Mask Transformer for Universal Image Segmentation" published at CVPR 2022.

[0007] These methods employ a transformer-based encoder-decoder architecture and learnable object queries (hereinafter referred to as queries) to segment objects in images.

[0008] Traditionally, machine learning models for instance segmentation are trained on images in which at least one object of a certain object category is labeled. Partially labeled images can also be used, where only a region of the image is labeled, rather than the entire image. Thus, for example, DE102022209113A1 describes a method for training a machine learning model for instance segmentation using partially labeled images, in which both the object and the background are labeled in at least a portion of the image.

[0009] When such trained machine learning models are used for inference, i.e., for image segmentation, they may segment objects that don't need to be segmented, i.e., incorrectly identifying these objects as the objects to be segmented. It's also possible that the objects are basically segmented correctly, but the object mask, i.e., the object's recognition region, contains fragments. This will be addressed in the reference... Figure 1A and 1B This will be explained in the following circumstances.

[0010] exist Figure 1A In image 10, object 11 was identified, but the area of ​​object 11 was not captured accurately, and segments 12A and 12B outside the actual object boundary were also identified as belonging to the object. Figure 1B In image 13, object 14 was incorrectly identified, even though in this example, the entire image 1B does not contain any object to be segmented.

[0011] These types of errors—namely, misidentifying fragments and misidentifying complete objects—belong to the category of false positives (Falsch-positiv), meaning they identify things that do not exist. Such errors can distort subsequent image analysis, for example, when counting objects or determining the area of ​​objects (such as cells). Summary of the Invention

[0012] Therefore, there is a need to reduce such false positive errors.

[0013] Provide methods, machine learning models, apparatus, systems, and computer programs as defined in the independent claims.

[0014] The dependent claims define other embodiments.

[0015] According to one embodiment, a computer-executed method is provided for training a machine learning model for instance segmentation of objects in an image, comprising: Provide training images (50), wherein the training images or a certain region thereof are labeled as not containing the object to be segmented, and wherein the object to be segmented is not labeled, and A machine learning model is trained based on the training images.

[0016] Therefore, training is performed using purely negative examples, where no objects to be segmented are labeled; it is only known that the training image (or regions thereof) does not contain the objects to be segmented. If objects are segmented in such regions of the training image, this will in any case be a false positive result, and the machine learning model can then be adjusted accordingly during training to avoid such false positive results.

[0017] In this process, a machine learning model should be understood as a computer-executed algorithm that utilizes machine learning techniques and kompons to perform instance segmentation. Examples of such machine learning models have been described in the introduction.

[0018] The training images can be based on actual image acquisition, such as acquisition through a microscope, or they can be synthetically generated training images.

[0019] Annotations should be understood as a type of markup, which can be performed by the user or automatically, such as when synthesizing training images.

[0020] Although the above refers to training images, it is important to understand that the above method can be applied to multiple different training images, where the training image or a region of the training image does not contain the object to be segmented.

[0021] This method is computer-executed, as is common when training machine learning models. It is carried out on one or more networked computers, each containing one or more corresponding processors and programmed accordingly.

[0022] The training may include: The training image is segmented into n segmentation objects, n>1, where each of the n segmentation objects provides class probability information and mask probability information. The class probability information indicates with what probability the corresponding segmentation object belongs to which object class, while the mask probability information indicates with what probability which region of the training image belongs to the corresponding segmentation object. Then, the value of the cost function can be determined based on the class probability information and the mask probability information, and a machine learning model can be trained based on this value.

[0023] Probability information is expressed quantitatively. The mask probability information can be represented as a so-called mask logit, while the category probability information is represented as a category logit. Here, in probability theory, the logit refers to the natural logarithm of chance, i.e., probability divided by non-probability (if p is the probability that an object identified by a query belongs to a certain category, then the category probability is ln(p / (1-p)), where ln represents the natural logarithm). When the probability is greater than 50% (p>0.5), the logit is greater than zero; when the probability is exactly 50%, the logit is equal to zero; when the probability is less than 50% (p<0.5), the logit is less than zero. Therefore, for example, when using logit, zero can be used as a threshold; when a category logit is greater than zero, it can be determined that it belongs to a certain category; or when the mask logit for an image element (such as a pixel) is greater than zero, it can be determined that it belongs to the corresponding object mask. However, other thresholds or other probability representations can also be used as probability information, with correspondingly adjusted thresholds.

[0024] The cost function can therefore include a component (Komponente) that is negatively evaluated when the segmentation object has a mask probability greater than a first threshold in the training image (if the entire training image is labeled as not containing the object to be segmented) or in a region of the training image labeled as not containing the object to be segmented. The first threshold can, for example, be zero in this case (e.g., when using masked logical values). This component can also be called the background loss, and it is also negatively evaluated when, for a given segmentation object, the mask probability indicates that, according to the label, a portion of the test image does not belong to the segmentation object. The threshold used can specifically be a threshold that also serves as a boundary in subsequent inference, indicating from which probabilities a given image region (e.g., pixels) begins to be evaluated as belonging to an object.

[0025] Training based on the cost function value can be performed using traditional methods, such as gradient descent, which minimizes the cost function. This essentially involves repeatedly executing the method, with adjustments made to the machine learning model between repetitions, for example, by changing weight factors, to minimize the cost function.

[0026] The cost function may include an additional component that is negatively evaluated when the segmented object has a class probability greater than a second threshold in the training image (if the entire training image is labeled as not containing the object to be segmented) or in a region of the training image labeled as not containing the object to be segmented. If, although the object to be segmented is not present in the training image or a region of the training image, it is detected as belonging to an object with a probability higher than the second threshold, this will be negatively evaluated in the cost function. Even in the case of class logit, the second threshold can be zero, and / or can be a threshold from which the object class is determined during inference.

[0027] This machine learning model can be, in particular, a transformer-based machine learning model for query-based instance segmentation, such as the DETR model or Mask2Former model mentioned above.

[0028] Segmenting a training image into n segmentation objects can then include segmenting the training image using m queries, where m ≥ n. If each query segments one object, then m = n. Specifically, in the current case, where the training image or regions of the training image do not contain the object to be segmented, ideally (i.e., when segmented correctly), the query results will not contain any object, i.e., the class probability information and mask probability information at their respective thresholds. This can also be achieved by introducing a "no object" object category and then assigning the corresponding query to it.

[0029] In this context, the training may also include: The training image is segmented using m queries, where m ≥ 1. Each of the m queries provides class probability information and mask probability information. The class probability information indicates the probability that a segment belonging to each query during segmentation belongs to a specific object category (or, in such an implementation, indicates the object category as "no object"). The mask probability information indicates the probability that a region of the training image belongs to a particular object, specifically the object identified by the corresponding query. The cost function can then be determined based on the class and mask probability information, and a machine learning model can be trained based on this value.

[0030] In this query-based instance segmentation scenario, the background loss can be calculated as follows:

[0031] This is the value of this component in the cost function for each query q (q=1...m), and the summation during this process is... This is performed for all pixels p, where Pn is the number of pixels to be evaluated in the test image.

[0032] If the labeled pixels do not contain the object to be segmented (i.e., in the region or the entire test image), then it applies. ,in This represents the masked probability information indicating the probability of an object being of category c, summed over all object categories, where log represents the natural logarithm. Otherwise, the appropriate approach is... In this case, the threshold is 0: In this case, the logarithm equals 0, therefore it contributes nothing to this component of the cost function; as As the value of increases, the contribution will increase accordingly according to the logarithmic function.

[0033] The above training can serve as a supplement to traditional training, where additional training images (or often multiple, especially a large number) are provided, in which the objects to be segmented are labeled. Training based on these additional training images can be performed in the conventional manner. In particular, when performing query-based instance segmentation using a transformer-based machine learning model, a so-called matching operation, such as Hungarian matching, can be performed, where the query is assigned to the labeled objects after segmentation. This matching is ignored in training images that do not contain the objects to be labeled.

[0034] The cost function can again be based on mask probability information and class probability information, using traditional loss values ​​such as label loss, mask loss, and / or, in this case, dice loss, especially when the labeled object occupies a relatively small area in the total area of ​​the second training image. The aforementioned components that are negatively evaluated when a query in a region of image data labeled as not containing the object to be segmented has a mask probability information greater than a first threshold can also be used here.

[0035] In addition, a machine learning model trained using the above method is provided for object instance segmentation.

[0036] Furthermore, an apparatus is provided having an input interface for receiving images, and a processor configured to perform instance segmentation of objects in the image using a trained machine learning model. A system is also provided comprising such an apparatus and an image acquisition device for acquiring images, wherein the image acquisition device may include a microscope.

[0037] A computer program with program code is also provided, which, when executed on a processor, provides the methods described above for training a machine learning model. This computer program can be stored on a storage medium, particularly a physical storage medium such as a memory, CD-ROM, or one or more hard disks. The trained machine learning model can also be provided on such a storage medium. Attached Figure Description

[0038] Figure 1A and Figure 1B An image is shown to illustrate pseudo-positive class segmentation.

[0039] Figure 2 This is a flowchart of an implementation method.

[0040] Figure 3 A block diagram of a system according to one implementation is shown.

[0041] Figure 4 A machine learning model is shown, which can be trained using the disclosed methods to form a trained machine learning model according to one implementation.

[0042] Figure 5A and Figure 5B These are examples of training images.

[0043] Figure 6 This is a flowchart of a method according to one implementation method. Detailed Implementation

[0044] The embodiments of this invention will be described in detail below. These embodiments should not be considered as limiting, but are for illustrative purposes only.

[0045] Unless otherwise stated, features of different implementations can be combined with each other. Variations, details, etc., described for one implementation also apply to other implementations, and therefore will not be repeated.

[0046] Figure 2 A flowchart of a method according to one embodiment is shown. In step 20, training images of the unlabeled objects to be segmented are provided.

[0047] Examples of such training images are in Figure 5A and Figure 5BAs shown in the image. Figure 5A Training image 50 is shown, in which no objects to be segmented are labeled. Furthermore, in one embodiment, the entire training image 50 is "labeled" as not containing any objects to be segmented, meaning that during training, it is known that the entire training image 50 does not contain any objects to be segmented.

[0048] Figure 5B An alternative approach is presented. In this training image 50, no objects to be segmented are labeled. Furthermore, a region 51 is labeled, but no objects to be segmented are present there. Therefore, it is determined here that no objects to be segmented are present only for a portion of the training image, but in both cases (5A and 5B), no objects to be segmented are labeled in the training image. Figure 5B It cannot be ruled out that there may be objects to be segmented outside region 51 in the rest of the training image 50, which are also not labeled here.

[0049] exist Figure 2 In step 21, a machine learning model for instance segmentation is then trained based on the training images.

[0050] In this process, the training can primarily be conducted in a traditional manner. However, unlike training using labeled training images of objects to be segmented, in this case, it may be impossible to assign potentially identifiable objects in the training images to the labeled objects. Conversely, if in this training image or labeled region ( Figure 5B If an object is identified in step 51), it is a negative evaluation for the training objective. One implementation method for this training will be discussed in a later reference. Figure 6 Please provide an explanation.

[0051] Next, we will refer to Figure 3 and Figure 4 A brief discussion of the system and possible machine learning models follows. However, as mentioned above, this approach is generally applicable to machine learning models for instance segmentation.

[0052] Figure 3 A system according to one embodiment is shown. Figure 3The system illustrates an image acquisition device for providing images of objects that need to be segmented. The image acquisition device may include one or more cameras. It may also include a microscope, where images can be acquired, for example, using the cameras through the microscope. In other types of microscopes, such as laser scanning microscopes, samples can also be scanned to obtain images. Using such microscopes, images of tissue samples can be acquired, for example. Furthermore, the system is equipped with a computing unit 31 for running machine learning models 32, 33. In the illustrated example, machine learning models 32, 33 have two stages 32 and 33. This machine learning model is trained using the methods described herein and can then be used as a trained machine learning model for inference. For example, the first stage 32 can be used for preprocessing, while the second stage 33 can include a transformer that subsequently generates a classification c. As mentioned earlier, this machine learning model can be configured for query-based instance segmentation, and can be a DETR-based model or a Mask2Former model. However, like the Mask R-CNN architecture described above, a transformer-based machine learning model is not feasible.

[0053] Taking machine learning models as an example, in Figure 4 The text demonstrates the Mask2Former model proposed by Cheng et al. in "Masked-attention MaskTransformer for Universal Image Segmentation" (as cited above). In addition to the training described here (which generates the corresponding trained model), Figure 4 The model in this example can be executed in a conventional manner, and therefore will only be described briefly. Images (training images during training and images to be analyzed during inference) are input into a backbone network 41, which can be implemented as a neural network, for example, and is designed to extract features from a low resolution 42. These images can also be 3D images from which the backbone network 41 can extract 2D features. These are then upsampled in a pixel decoder. A transformer decoder 44 receives image data processed by the pixel decoder 43 along with multiple queries 40, and then assigns class probability information 45 (e.g., in the form of class logical values) and mask probability information 46 (via a connection 47 to the image data provided by the pixel decoder 43), for example, in the form of mask logical values. The class probability information 45 to which a query is assigned an object category, and the mask probability information 46 indicates the corresponding image region.

[0054] Figure 6According to one implementation, a more detailed flowchart for training a machine learning model (in this case, query-based instance segmentation) is shown. In this process, a training image is first provided. This training image can be a conventional training image in which the objects to be segmented are labeled, or... Figure 5A and Figure 5B The training images shown do not contain object annotations, and at least one region is annotated to indicate that it does not contain the object to be segmented. Then, during training, a large number of training images are used to perform... Figure 6 The method described above includes a portion of training images that are of the second type, i.e., training images that do not contain object annotations.

[0055] In step 60, the training image is segmented n times using the machine learning model to be trained, where n≥1. For each query, class probability information and mask probability information are output.

[0056] Then, during subsequent inference using the trained machine learning model, all queries with all category probability information below the first threshold are discarded, as they are assumed to not correspond to any identified objects due to their low probability. Alternatively, as described above, queries for which no objects were detected are assigned to the "no object" category with the corresponding category probability information. The remaining queries are assigned to object categories, whose category probability information indicates the highest membership probability. A second threshold is used to obtain the object mask, i.e., the region where the object is located. All image regions, such as pixels, with mask probability information above this second threshold become part of the object mask.

[0057] During training, step 61 distinguishes whether the training image contains object annotations, i.e., whether it is a traditional training image with object annotations, or a training image without object annotations as used in the embodiment, as referenced. Figure 5A and Figure 5B If the training images contain object annotations (as is the case in 61), then conventional training is performed. To do this, in step 62, a matching algorithm, such as "Hungarian matching," is first executed to assign each annotated object in the training images to a corresponding query with a 1:1 match. In this process, each annotated object is assigned to a query that simultaneously has the most accurate object mask and the highest probability of belonging to the object category of the annotated object.

[0058] In this process, the matching algorithm is typically an algorithm whose main task is to optimally assign the objects predicted by the model (based on class probability information and mask probability information) to the actual objects in the training images, i.e., the labeled objects. Since the machine learning model makes a fixed number of object predictions (n ​​queries), it is necessary to determine which prediction corresponds to which labeled object. Matching algorithms, such as the Hungarian matcher, solve this assignment problem by minimizing a cost function that takes into account various variables, particularly the accuracy of the region assigned to the object based on mask probability information, and the class assignment of the object based on class probability information.

[0059] In step 63A, a value for a cost function (loss function), or simply loss, is determined. This loss may consist of a label loss (primarily evaluating the accuracy of class assignment), a mask loss (primarily evaluating mask accuracy), a dice loss (used in traditional methods to account for cases where the area of ​​labeled objects is often much smaller than the total image area), or other cost function components traditionally used in training machine learning models. In step 64, training is then performed based on the loss, for example using a gradient method, where the machine learning model is gradually adjusted to minimize the loss.

[0060] However, if the training images in 61 do not contain object annotations, i.e., the training images, such as the reference images... Figure 5A and Figure 5B As stated above, a match cannot be made because there are no labeled objects. Here, the value of the cost function is determined directly in step 63B. The dice loss and mask loss cannot be determined here as in 63A because the query cannot be assigned to an object, and therefore it is impossible to determine how well the mask matches the labeled object contour based on the mask probability information. As compensation, a background loss as described above can be introduced as a component, which will affect the query in areas where objects are known to be absent (e.g., Figure 5A The entire image in 50 or Figure 5B In region 51 (marked in the middle), queries are negatively evaluated using positive mask probability information (e.g., mask probability information exceeding a threshold). When using logical values ​​as category or mask probability information, the threshold can also be set to zero, which is equivalent to a probability exceeding 50%. In step 63A, this background loss can also be additionally considered because the query should not segment objects in the background region that are themselves labeled. Furthermore, in 63B, the label loss can continue to be used as a component of the cost function, where when in the query (the query may belong to region 51 (or...)... Figure 5A If the entire training image 50 is labeled as having no object to be segmented, and it can still be identified as belonging to a certain object category (due to the absence of object labeling), it will be penalized. Then, even when encountering this loss, training will be performed based on this loss at time 64.

[0061] As mentioned above, Figure 6 The method described earlier can be used with a large number of training images, a portion of which do not contain object annotations. In this way, it is possible to obtain... Figure 3 and Figure 4 Each machine learning model generates its own trained model, which can then be used for inference.

Claims

1. A computer-executed method for training a machine learning model (32, 33) for instance segmentation of objects in images, comprising: Provide training images (50), wherein the training images or a certain region thereof are labeled as not containing the object to be segmented, and wherein the object to be segmented is not labeled. The machine learning model (32, 33) is trained based on the training images.

2. The method according to claim 1, wherein the training comprises: The training image (50) is segmented into n segmentation objects, n>1, wherein the segmentation for each of the n segmentation objects provides class probability information (45) and mask probability information (46), wherein the class probability information (45) indicates with what probability each segmentation object belongs to which class of object, and the mask probability information (46) indicates with what probability which region of the training image (50) belongs to the corresponding segmentation object. Based on the category probability information (45) and mask probability information (46), the value of the cost function is determined, and Based on the values, train the machine learning model (32, 33).

3. The method according to claim 2, wherein the cost function includes a component that is negatively evaluated when the training image labeled as not containing the object to be segmented or when the segmented object in the region of the training image has a mask probability information (46) greater than a first threshold.

4. The method according to claim 2 or 3, wherein the cost function includes a component that is negatively evaluated when one of the segmented objects belonging to the training image or region that is labeled as a training image or whose region does not contain the object to be segmented has a category probability information (45) greater than a second threshold.

5. The method according to any one of claims 1 to 4, wherein the machine learning model (32, 33) is a transformer-based machine learning model for query-based instance segmentation, and wherein segmenting the training image into n segmentation objects includes segmenting the training image using m queries, where m ≥ n.

6. The method according to any one of the preceding claims further comprises: Provide additional training images in which at least one object to be segmented (13) is labeled, and Based on the additional training images, a machine learning model (32, 33) is trained.

7. A machine learning model (32; 33) for instance segmentation of objects in an image, trained using the method according to any one of claims 1 to 6.

8. A device (31) for instance segmenting objects in an image, comprising: An input interface for receiving images, and A processor configured to segment objects in a received image using the machine learning model (32, 33) as described in claim 7.

9. The system, including: Image acquisition device (30) for acquiring images, and The device (31) according to claim 8, wherein the input interface of the device (31) is coupled to the image acquisition device (30) to receive the acquired image.

10. The system according to claim 9, wherein the image acquisition device (30) includes a microscope.

11. A computer program having program code configured to perform the method according to any one of claims 1 to 6 when executed on a processor.

12. A storage medium having stored thereon the computer program of claim 11 or the machine learning model of claim 7.

Citation Information

Patent Citations

  • TRAINING OF INSTANCE SEGMENTATION ALGORITHMS WITH PARTIALLY ANNOTATED IMAGES

    DE102022209113A1