Computer-implemented method and electronic data processing device

CN122820737APending Publication Date: 2026-09-25CARL ZEISS MICROSCOPY GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610287145.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-03-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

因此,训练相应的深度机器学习模型的耗费很高

Benefits of technology

[0011]上文阐述的特征和下文描述的特征不仅能够以相应明确阐述的组合使用,而且能够以其他组合使用或单独使用,而不脱离本发明的保护范围。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820737A_ABST
    Figure CN122820737A_ABST
Patent Text Reader

Abstract

Different embodiments of the invention relate to a computer-implemented method and an electronic data processing device, the method comprising processing a latent feature vector (122) for a volumetric image data set in a voxel decoding branch (131) to obtain a voxel embedding (132) for voxels of the volumetric image data set (111). Further, the latent feature vector (122) is processed in a transformation decoding branch (142) to obtain a patch embedding (143) for a query vector. Based on a combination (145) of the voxel embedding (132) and the patch embedding (143), one or more instance-specific volumetric segmentation masks (115) are generated for one or more structures imaged in the voxel-based volumetric image data set (111).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various examples of the present invention relate to machine learning models for 3D instance segmentation in voxel-based volumetric image datasets. The various examples relate to the training and inference of the respective models. Background Technology

[0002] Instance segmentation of instances of one or more structural types in a voxel-based volumetric image dataset (hereinafter referred to as the 3-D image dataset) is important for a variety of applications.

[0003] An exemplary application requiring instance segmentation from a 3-D image dataset is the quantification of porosity in ceramic materials. Mechanical properties, thermal conductivity, and electrical conductivity all depend on the porosity of ceramic materials. Other properties dependent on the porosity of ceramic materials include, for example, the absorption of sound waves or microwaves, the growth characteristics of layers grown on the surface of the corresponding material, or the superconducting properties used in high-temperature superconductors. To obtain a 3-D image dataset of a sample (e.g., ceramic material), X-ray tomography can be used, for example. An exemplary technique for performing 3-D instance segmentation in such applications is described in: Smith, JD, et al. “Quantitative estimation of closed cell porosity in low density ceramic composites using X-ray microtomography.” Scientific Reports 13.1 (2023): 127. There, a deep convolutional network is used to: segment cross-sections of pores in 2-D slice images from a 3-D image dataset. Then, 3-D instance segmentation is created based on 2-D segmentation masks in the 2-D slice images, where it is estimated whether two instances located in adjacent 2-D slice images are connected to each other. In such techniques, so-called "oversegmentation" can occur, where a specific instance (e.g., a specific pore in a material) is described by two separate segmentation masks. This can lead to statistical distortion, for example, when determining porosity. The distribution of pore size is thus distorted.

[0004] However, determining porosity in ceramic materials is only a specific application of 3D instance segmentation. For example, it is possible to obtain 3D image datasets of membrane materials for lithium-ion batteries using X-ray tomography. Polymer fibers can then be segmented. See, for example: Villarraga-Gómez, Herminso, et al. “Assessing rechargeable batteries with 3D X-ray microscopy, computed tomography, and nanotomography.” Nondestructive Testing and Evaluation 37.5 (2022): 519-535.

[0005] The method for segmenting fibers described there is a two-step process: (1) semantic segmentation using a deep neural network or other statistical model, and (2) post-processing for separating instances (e.g., analysis using connected components). This method for instance segmentation also has specific drawbacks. The semantic segmentation process following the post-processing algorithm for separating individual fibers is error-prone because it requires careful selection of hyperparameters to find a good balance between the merging of unconnected fibers due to small gaps in segmentation predictions and the merging of unconnected fibers that are randomly close to each other. Furthermore, a solution that works well on one dataset can fail on another.

[0006] Another example of 3-D instance segmentation is described in: Abdollahzadeh, Ali, et al. “DeepACSONautomated segmentation of white matter in 3D electron microscopy.” Communications biology 4.1 (2021): 179. There, the 3-D image dataset was obtained from a 3-D electron microscope. The process involved two deep machine learning models specifically designed for different structure types. Such methods also have specific drawbacks. In particular, manually optimized or adapted solutions—i.e., specific models adapted to very specific structure types—cannot or can only generalize poorly. Therefore, training the corresponding deep machine learning models is very costly. For various structure categories, experts must create and adapt specialized models. Summary of the Invention

[0007] Therefore, there is a need for improved techniques to perform instance segmentation of instances of specific structure types in voxel-based volumetric image datasets. In particular, there is a need for techniques that address or at least mitigate some of the aforementioned limitations and drawbacks.

[0008] The following describes a technique for instance segmenting structures belonging to one or more structural types within a voxel-based volumetric image dataset (3-D image dataset). The technique described herein enables the training of appropriate machine learning models for various domains, or even applications beyond those seen during training. It robustly segments structures, even those with aspect ratios significantly deviating from 1 (e.g., long, thin structures like polymer chains). The technique avoids oversegmentation or undersegmentation. By employing the technique described herein, the implementation cost for new application scenarios can be limited.

[0009] A computer-implemented method includes acquiring a voxel-based volumetric image dataset. The method further includes generating one or more latent feature vectors for the volumetric image dataset based on an encoding branch. The method also includes processing the one or more latent feature vectors in a voxel decoding branch to obtain voxel embeddings of the voxels in the volumetric image dataset. Furthermore, the method includes processing the one or more latent feature vectors in a transform decoding branch and for a sequence of query vectors to obtain fragment embeddings of the query vectors. Additionally, the method includes generating one or more instance-specific volume segmentation masks for one or more structures imaged in the voxel-based volumetric image dataset based on a combination of voxel embeddings and fragment embeddings.

[0010] The electronic data processing device includes a processor unit and a memory. The processor unit is designed to load and execute program code from the memory, wherein the processor unit is designed, based on program code, to acquire a voxel-based volumetric image dataset. The processor unit is also designed, based on program code, to generate one or more latent feature vectors for the volumetric image dataset based on an encoding branch. The processor unit is further designed, based on program code, to process the one or more latent feature vectors in a voxel decoding branch to obtain voxel embeddings of the voxels in the volumetric image dataset. The processor unit is also designed, based on program code, to process the one or more latent feature vectors in a transform decoding branch and for a sequence of query vectors to obtain fragment embeddings of query vectors. The processor unit is also designed, based on program code, to generate one or more instance-specific volume segmentation masks for one or more structures imaged in the voxel-based volumetric image dataset based on a combination of voxel embeddings and fragment embeddings.

[0011] The features described above and those described below can be used not only in the corresponding explicitly stated combinations, but also in other combinations or individually, without departing from the scope of protection of this invention. Attached Figure Description

[0012] Figure 1 The diagram illustrates the processing of voxel-based volumetric image datasets in machine learning models, based on various examples.

[0013] Figure 2 This is a flowchart of an exemplary method.

[0014] Figure 3 The electronic data processing apparatus is illustrated schematically according to various examples. Detailed Implementation

[0015] The following describes a technique for determining 3D instance segmentation. Here, the 3D instance segmentation includes a 3D segmentation mask for each of a plurality of instances. The 3D segmentation mask is a description of each voxel in a voxel-based volumetric image dataset: whether the corresponding voxel is part of the corresponding instance. The 3D instance segmentation can optionally also include category information for each 3D segmentation mask for the corresponding instance. This can be helpful when segmenting structures of multiple structural types. For example, it is conceivable to segment 3D cell objects of different cell types.

[0016] A voxel-based volumetric image dataset (hereinafter referred to as a 3-D image dataset) is a 3-D representation of data in which volume is divided into small, uniform cubes, known as voxels. Each voxel represents a specific value or property of the volume, such as density, color, or intensity. Voxel-based volumetric image datasets are particularly distinguishable from 3-D point clouds.

[0017] The 3-D image datasets processed using the techniques described herein can be examined using various imaging devices. For example, a 3-D microscope providing magnification can be used. Light sheet microscopy is also feasible. Magnetic resonance tomography (MRI) can be used. X-ray tomography (XCT) can be used. Another example is positron emission tomography (PET). Serial section electron microscopy can also be used; here, the sample (typically a semiconductor sample) can be peeled layer by layer using a focused ion beam, and corresponding slice images of each layer can be created using an electron microscope. Generally, 3-D electron microscopy can be used to provide 3-D image datasets processed using the techniques described herein.

[0018] Due to the flexibility of the techniques described in this paper for 3-D instance segmentation on different types of 3-D image datasets, the type of structure being segmented can also vary. For example, it is possible to segment semiconductor structures, material pores, polymer fibers, etc.

[0019] The technology described in this paper is characterized in particular by the fact that the same model architecture can be applied to many different types of 3-D image datasets and also to many different structural types. The model described in this paper for performing 3-D instance segmentation is not domain-specific, but can be applied across domains.

[0020] Figure 1Model 120 is shown, which generates instance segmentation data 115, 116 for a 3-D image dataset 111. Model 120 is a machine learning model with multiple layers having corresponding machine learning weights. Model 120 corresponds in principle to the Mask2Former model described in Cheng, Bowen, et al. “Masked-attention mask transform for universal image segmentation.” Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022. However, model 120 is capable of processing 3-D image datasets, while the Mask2Former model only processes 2-D images.

[0021] To process 3-D image datasets, model 120 has an encoding branch 121 (also called the "trunk") designed to determine one or more corresponding latent feature vectors 122 for the 3-D image dataset 111. Thus, the one or more latent feature vectors 122 represent the latent representation of the corresponding 3-D image dataset. Depending on the size of the 3-D image dataset 111, i.e., the number of voxels in the 3-D image dataset 111, more or fewer latent feature vectors 122 can be determined. Here, each latent feature vector can describe a specific sub-region of the field of view covered by the 3-D image dataset. Each feature vector can be a 1-D array of scalar eigenvalues. For example, a 3-D array composed of feature vectors can be obtained; thus, the positional relationships of regions of the 3-D image dataset encoded by feature vectors can be obtained.

[0022] For example, a ResNet-based encoding branch 121 can be used, which has one or more 3D convolutional layers. Another example would be an encoding branch based on a 3D visual transformation architecture, where 3D positional encoding is used.

[0023] Then, one or more feature vectors 122 are processed together in the voxel decoding branch 131 to obtain voxel embeddings 132 of the volumetric image dataset.

[0024] Furthermore, each latent feature vector 122 is processed in the transform decoding branch 142 against a sequence 141 of query vectors to obtain a fragment embedding 143 of the query vector. The query vector can be learned, for example, but it can also be chosen differently, for example, depending on the 3-D image dataset.

[0025] Transform decoding branch 142 focuses on the multi-scale image feature map generated by voxel decoding branch 133. Each decoding layer of transform decoding branch 142 contains a self-attention sub-layer, a cross-attention sub-layer, and a feedforward sub-layer. The self-attention sub-layer detects the correlation between query vectors. The cross-attention sub-layer computes new features of the query vector by calculating the pairwise correlation between the query vector 141 and the voxel embedding 132, taking into account the voxel embeddings, where this is restricted to a specific local spatial region in the feature map (e.g., restricted to the corresponding segmentation mask). This is also known as spatially restricted attention. Spatially restricted attention reduces storage requirements. The feedforward sub-layer applies a pointwise transformation to each query embedding. Multiple layers of this operation are stacked, allowing the query or inquiry to be refined in each layer by integrating information from the feature map and making the representations mutually influential. Finally, the output of the transform decoding branch is used to generate the segmentation mask 115. To this end, a combination of voxel embedding 132 and fragment embedding 143—block 145—is used to generate instance-specific volumetric segmentation masks for one or more structures or objects imaged in the 3-D image dataset 111.

[0026] Based on fragment embedding 143, instance labels 116 can optionally be generated for each instance category. This enables differentiation between different structure types. With the help of Mask2Former-based model 120, it is therefore feasible to perform classification based on the segmentation mask data from the transform decoding branch 142 (instead of performing classification for each voxel).

[0027] exist Figure 1In the example, transform decoding branch 142 is coupled to multiple layers of voxel decoding branch 131. Furthermore—as stated above—transform decoding branch 142 also uses spatially constrained attention, for example through masking, for each of the multiple query vectors of the corresponding sequence 141. This corresponds in principle to techniques known to the Mask2Former model—for two-dimensional images. However, both architectural features are optional. Therefore, in some examples, it is possible to omit spatially constrained attention and / or the coupling of the transform decoding branch to layers of voxel decoding branch 131, as described for the MaskFormer model, for example, in Cheng, Bowen, Alex Schwing, and Alexander Kirillov. “Per-pixel classification is not all you need for semantic segmentation.” Advances in neural information processing systems 34 (2021): 17864-17875.

[0028] Next, aspects related to at least a portion of the training of the machine learning model 120 are described. The machine learning weights of model 120 are determined using one or more machine learning methods. Such machine learning methods are optimization methods, such as gradient descent. For example, the so-called backpropagation algorithm can be used to determine the machine learning weights. Here, one or more loss functions are optimized. The loss function, for example, can describe the deviation of the instance segmentation data 115, 116 predicted by model 120 for a 3-D training image dataset from the corresponding ground truth. The ground truth can, for example, have a label for a single instance of the complete 3-D training image dataset. However, it is also conceivable that only a sub-region of the 3-D training image dataset is labeled—this helps limit the labeling cost. The loss function limited to said sub-region can then be used. Outside the sub-region, no loss value is calculated, and correspondingly, no weight adaptation is performed. In principle, weights in different parts of model 120 can be adapted during training. Here, the ground truth label can describe the association between the corresponding voxel and the corresponding instance (and, if necessary, the instance category) and / or the association between the corresponding voxel and the background for each voxel in the corresponding sub-region. By using background annotation, it is particularly feasible to achieve a steep learning curve for training, even for relatively limited sub-region expansion. For example, the query vector for the corresponding sequence 141 can be determined during training. Alternatively or additionally, the weights in the transformation decoding branch 142 and / or voxel decoding branch 131 and / or encoding branch 121 can also be adapted. In some examples, it is helpful that the weights of the encoding branch 121 are fixed and trained separately from the query vector and / or decoding branches 131, 142 (“bootstrapping”). At least part of the training of the machine learning model 120 can be performed, for example, for different sample types (so-called “fine-tuning”). For example, new training can be triggered whenever a new structure type is identified in a 3-D image dataset. For example, for this purpose, anomaly detection can be performed, which knows the structure types used so far in training and identifies structure types not used so far as anomalies. This enables domain-specific adaptation. For example, it would be feasible to: in the case of relatively similar sample types, not to retrain the encoding branch 121, but to retrain the query vector sequence 141 and / or retrain the decoding branches 131, 142.

[0029] Figure 2 This is a flowchart of an exemplary method. Figure 2 The methods described herein can be executed by a processor. To this end, the processor can load program code from memory and execute the program code. Figure 2 The method described involves generating instance segmentation data for 3-D image datasets. For example, Figure 2 The method in the middle can be used Figure 1 Machine learning model 120.

[0030] In box 905, a 3-D image dataset is acquired. For example, the 3-D image dataset can be acquired from a 3-D imaging device, such as a 3-D microscope. Box 905 may include, for example, controlling the corresponding 3-D imaging device with appropriate control data to trigger image detection. However, it is also conceivable that a 3-D image dataset is loaded from an image database in box 905.

[0031] In box 910, one or more feature vectors are then generated for the 3-D image dataset from box 905. For example, multiple feature vectors can be generated, which encode (e.g., overlapping) image regions of the 3-D image dataset. This also enables the processing of particularly large 3-D image datasets with limited storage resources. The feature vectors are 1-D array structures, such that feature vector generation is performed using appropriate encoding branches that convert the 3-D data into one or more 1-D data structures. For example, a 3D convolutional network, such as a 3D convolutional network with a ResNet architecture having so-called residual connections, can be used for this purpose. The corresponding techniques have been combined above. Figure 1 The coding branch 121 was discussed.

[0032] In various examples, it is particularly possible to use domain-independent encoding branches. For example, it is possible to use encoding branches trained for many different domains.

[0033] In box 915, one or more feature vectors from box 910 are then decoded. This is achieved through a voxel decoding branch ( Figure 1 : Voxel decoding branch 131) and by means of transform decoding branch (see Figure 1 Transformation decoding branch 142) is performed in parallel. It can use either MaskFormer or Mask2Former architectures.

[0034] In various examples, domain-specific decoding branches can be used. For instance, decoding branches can be implemented within the scope of so-called "fine-tuning" training (see...). Figure 1 The weights of decoding branches 131 and 142 are domain-specific; while the weights of encoding branch 121 are fixed. This is because the encoding branches can be trained domain-independently. This technique is based on the understanding that extracting specific latent image features using encoding branch 121 can be robustly transferred to different domains, while the generation of corresponding segmentation data should be specifically adapted to different domains to achieve sufficient robustness.

[0035] In box 920, a segmentation mask is then generated, specifically by combining the results of the voxel decoding branch with the results of the transform decoding branch. The corresponding techniques have been described above. Figure 1Block 145 elaborates on this. A segmentation mask is obtained for each instance of the structure. In principle, multiple structure types can be segmented simultaneously; therefore, it is helpful to also obtain the corresponding structure category label for each segmentation mask.

[0036] In box 925, the segmentation data previously generated in box 920 can be used. For example, it is possible to manipulate a graphical user interface to output an overlay of the segmentation mask and a 3-D image dataset. For example, it is possible to compute and display 2-D slice images of the 3-D image dataset and the 3-D segmentation mask separately. It is also conceivable to perform improved statistical evaluations based on the segmentation data. For example, it is possible to calculate the size distribution of the segmented instances. For example, it is possible to calculate such a size distribution for each of multiple structure types. Alternatively or additionally, it is also possible to calculate the average distance between adjacent instances. It is possible to calculate the 3-D density of instances of a specific structure type. All such evaluations can be improved through the particularly robust and reliable generation of the segmentation mask.

[0037] With the help of Figure 2 The techniques described herein enable particularly robust 3D instance segmentation that is transferable to many different domains. Compared to conventional techniques that use 2D segmentation of sliced ​​images and subsequent concatenation of 2D segmentation masks, the techniques described herein are more robust to oversegmentation and can also be applied to specific structure types, such as those with particularly thin and elongated shapes. In particular, non-convex shapes, such as fibers, neurons, or mitochondria, can be segmented using the techniques described herein. Domain-independent segmentation models can be used, which are adapted to robustly segment structures of different shapes, geometries, and sizes.

[0038] Figure 3 An electronic data processing device 90 is schematically shown, comprising a processor unit 91 (e.g., implemented via a central processing unit (CPU) and a graphics processing unit (GPU)) and a memory 92, as well as a communication interface 94 and a user interface 93. The processor unit 91 is capable of loading program code from the memory 92 and executing the program code. When the processor unit 91 executes the program code, this causes the processor unit 91 to perform the techniques described herein, such as: executing... Figure 2 The method includes: acquiring a 3-D image dataset via communication interface 94; processing the 3-D image dataset to obtain an instance-specific volume segmentation mask; outputting the corresponding 3-D image dataset, such as a superimposed 3-D image dataset with a volume segmentation mask, via a graphical user interface implemented by means of user interface 93; and so on.

[0039] In summary, a machine learning model has been described that is particularly advantageous for 3-D instance segmentation. Traditional solutions typically have many hyperparameters—parameters that cannot be learned from the data and must be determined separately. Examples of hyperparameters include network depth, number of channels per layer, learning rate, activation function, or the type of model architecture commonly used. Furthermore, most established methods require post-processing steps to obtain instance segments from the output of the machine learning model. Determining suitable hyperparameter configurations and matching post-processing steps requires manual work and extensive experimentation by experienced experts. Techniques employed, such as those based on MaskFormer or Mask2Former architectures, reduce or eliminate this overhead by applying such models to a variety of datasets without further manual adaptation and directly providing 3-D instance segmentation without data-specific post-processing.

[0040] Of course, the features and aspects of the invention described above can be combined with each other. In particular, the features can be used not only in the described combinations, but also in other combinations or individually, without departing from the scope of the invention.

Claims

1. A computer-implemented method, the method comprising: Obtain a voxel-based volumetric image dataset (111). One or more latent feature vectors (122) are generated for the volumetric image dataset based on the encoding branch (121). The one or more latent feature vectors are processed in the voxel decoding branch (131) to obtain voxel embeddings (132) of the voxels of the volumetric image dataset (111). In the transform decoding branch (142) and for the sequence of query vectors (141), the one or more latent feature vectors are processed (122) to obtain the fragment embedding (143) of the query vector, and Based on the combination (145) of the voxel embedding (132) and the fragment embedding (143), one or more instance-specific volume segmentation masks (115) are generated for one or more structures imaged in the voxel-based volume image dataset (111).

2. The computer-implemented method according to claim 1, The transform decoding branch is coupled to multiple layers of the neuron voxel decoding branch and / or uses spatially constrained cross attention between fragment embeddings and voxel embeddings.

3. The computer-implemented method according to claim 1, wherein the method further comprises: Obtain a voxel-based volume training image dataset. Obtain ground truth labels for sub-regions limited to the voxel-based volume training dataset, and Based on the ground truth annotations, training is performed to learn at least one of the sequence of the query vector and the weights of the transform decoding branch (142). The training described therein uses a loss function limited to the sub-regions of the voxel-based volume training dataset.

4. The computer-implemented method according to claim 3, The truth value annotation in the sub-region describes the association between the corresponding voxel and one or more structural instances or backgrounds.

5. The computer-implemented method according to claim 3, wherein the method further comprises: The training is triggered whenever a new structure type is identified in the voxel-based volumetric image dataset.

6. The computer-implemented method according to any one of claims 1-5, The voxel-based volumetric image datasets described herein are detected by imaging devices selected from the group consisting of: three-dimensional microscopes, light sheet microscopes, light field microscopes, confocal microscopes, magnetic resonance tomography scanners, X-ray tomography scanners, positron emission tomography scanners, serial section electron microscopes; and 3-D electron microscope devices.

7. The computer-implemented method according to any one of claims 1-5, The encoding branch (121) includes multiple three-dimensional convolutional layers.

8. The computer-implemented method according to any one of claims 1-5, The encoding branch (121) mentioned therein is trained in a domain-independent manner. At least one of the transform decoding branch (142) and the voxel decoding branch (131) is domain-specific trained.

9. An electronic data processing device (90) comprising a processor unit (91) and a memory (92), wherein the processor unit (91) is designed to load and execute program code from the memory, wherein the processor unit (91) is designed based on the program code to perform the following steps: Obtain a voxel-based volumetric image dataset (111). One or more latent feature vectors (122) are generated for the volumetric image dataset based on the encoding branch (121). The one or more latent feature vectors are processed in the voxel decoding branch (131) to obtain voxel embeddings (132) of the voxels of the volumetric image dataset (111). In the transform decoding branch (142) and for the sequence of query vectors (141), the one or more latent feature vectors are processed (122) to obtain the fragment embedding (143) of the query vector, and Based on the combination (145) of the voxel embedding (132) and the fragment embedding (143), one or more instance-specific volume segmentation masks (115) are generated for one or more structures imaged in the voxel-based volume image dataset (111).

10. The electronic data processing apparatus according to claim 9, wherein the processor unit (91) is further designed, based on the program code, to perform a computer-implemented method according to any one of claims 1 to 8.