Generative machine learning model for generating synthetic training image pairs for inverse image reconstruction tasks
By combining generative machine learning models and forward models, domain-specific and sample-specific training data pairs are generated, solving the problem of insufficient training data in existing technologies and improving the model accuracy and applicability for inverse image reconstruction tasks.
Patent Information
- Application Number
- CN202511113000.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-08-10
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies struggle to effectively address the diversity of sample types and imaging methods in reverse image reconstruction tasks when generating training data pairs, leading to insufficient or inconsistent training data and impacting the training effectiveness and applicability of machine learning models.
The feature description of the reference image is generated by using a first machine learning model, a synthetic reference image is generated by using a second machine learning model, and a synthetic image measurement dataset is generated by applying the forward model of the reverse image reconstruction task to form training data pairs for training the reconstruction model.
It enables the rapid generation of domain-specific and sample-specific training data pairs, improving the accuracy and applicability of the reconstruction model, reducing dependence on actual measurement data, and enhancing the accuracy of inverse image reconstruction tasks.
Smart Images

Figure CN121504753A_ABST
Abstract
Description
Technical Field
[0001] Various examples of this disclosure relate to generative artificial intelligence techniques in reverse image reconstruction tasks. In particular, various examples of this disclosure relate to techniques for generating synthetic training data pairs for training machine learning models. Background Technology
[0002] In many fields, image reconstruction algorithms are used to solve so-called "inverse image reconstruction tasks." Solving the "inverse problem" involves determining the cause of a phenomenon based on observed results. In the context of image reconstruction, this means reconstructing the original appearance of a scene from image measurement data, i.e., the conditions that caused the image; in other words, reconstructing an image of the scene from image measurement data.
[0003] For example, it is well known that image reconstruction algorithms can denoise or deconvolve images. Improving image resolution, as part of image reconstruction, is also widely used in various fields. Sometimes, image measurement data is not acquired in image space, but in spatial frequency space. In this case, subsampling can be used to accelerate image acquisition, which is obvious in image space, for example, through convolution.
[0004] There are various algorithmic approaches available for image reconstruction algorithms that perform inverse image reconstruction tasks. Examples include, for instance, classical iterative numerical optimization for minimizing an error term; the error term may include, for example, the deviation between the calculated image measurement data and the actual image measurement data, wherein the calculated image measurement data is determined using a forward model of the inverse image reconstruction task; this ensures data consistency between the image measurement data and the reconstructed image.
[0005] Recently, artificial intelligence technologies, especially machine learning models, are increasingly being used to solve inverse image reconstruction tasks. For example, such machine learning models can directly convert image measurement data into images. Typical machine learning models (reconstruction models) used to solve inverse image reconstruction tasks include: deep learning models; deep neural networks, such as convolutional networks or transformer networks. However, machine learning models can also perform sub-tasks related to image reconstruction, such as implementing regularization operations for a difficult inverse image reconstruction task.
[0006] Regardless of the specific nature of the inverse image reconstruction task or the intended use of the machine learning model, training the model is essential. Only through proper training can good results be achieved in solving inverse image reconstruction tasks. Training the machine learning model requires a large number of training data pairs. These pairs consist of image measurement datasets and associated reference images used to solve the inverse image reconstruction task.
[0007] One known existing technique for generating training data pairs is to perform measurement activities. Image measurement datasets are acquired through measurement, and corresponding reference images are generated. Depending on the inverse image reconstruction task, different techniques exist for generating reference images. For example, in inverse image reconstruction tasks involving increasing image resolution (“super-resolution”), a suitable image-capturing device can be used to directly capture a reference image at the increased resolution. However, another reconstruction algorithm can also be applied to the image measurement dataset to generate reference images. One problem with such measurement activities is that sufficient image measurement data and / or reference images are often unavailable. This is due to, for example, a limited number of available samples or a limited amount of available time. Thus, training on too few training data pairs is required, resulting in limited accuracy of the reconstruction model when solving inverse image reconstruction tasks. The applicability to slightly varied inverse image reconstruction tasks (e.g., slightly different sample types) is severely limited (lack of generality).
[0008] Another known existing technique for generating training data pairs is numerical simulation. For example, optical microscopic images of a specific sample can be simulated by numerically simulating optical images of a microscope, such as using ray tracing techniques. However, this technique for generating simulated training data pairs is limited in terms of sample types. It can typically only simulate samples with simple structures, such as punctate structures or tubulin. For example, see the article “Content-aware image restoration: pushing the limits of fluorescence microscopy” by Weigert, M., Schmidt, U., Booth, T., Müller, A., Dibrov, A., Jain, A., ... & Myers, EW (2018). Nature Methods, 15(12), 1090-1097. When simulating more complex structures, the simulation results often do not match the actual observed image measurements (the so-called “reality gap”), and there is also the risk of insufficient sample variance.
[0009] Another known approach is to train the reconstruction model using image measurement data and reference images of general samples. This means that the image measurement data and the associated reference image are recorded as training data pairs for the reference samples, whose features differ from those of the samples subsequently measured. For example, certain structural types may not appear in the reference samples but can be observed in the actual measured samples. In this case, the reconstruction model works to solve the inverse image reconstruction task in a region of the input space where there is little or no training data during training. This is equivalent to "interpolation" in prediction, and the corresponding results are often inaccurate or even incorrect when solving inverse image reconstruction tasks, but in either case, there is a high degree of uncertainty. Summary of the Invention
[0010] Therefore, there is a need to improve techniques related to training machine learning models to solve the inverse image reconstruction task. In particular, there is a need for techniques to obtain training data pairs for training the corresponding machine learning models.
[0011] A computer-implemented method includes acquiring at least one reference image. The method further includes generating at least one feature description of the at least one reference image based on the at least one reference image using a first machine learning model. The method further includes generating a plurality of synthetic reference images based on the at least one feature description using a second machine learning model. The method further includes, for each of the plurality of synthetic reference images: applying a forward model of an inverse image reconstruction task to obtain a respective synthetic image measurement dataset. The method further includes initiating the training of a third machine learning model. This training is based on training data pairs, each data pair including one of the plurality of synthetic reference images and its corresponding synthetic image measurement dataset. The training ensures that the third machine learning model can be used to solve the inverse image reconstruction task.
[0012] In addition, a corresponding electronic data processing device is disclosed, including a processor and a memory. The processor is used to load and execute program code from the memory. When the processor executes the program code, it causes the processor to perform the computer-implemented method described above.
[0013] The features and characteristics described above and below can be used not only in the explicitly given combinations, but also in other combinations or individually, without departing from the scope of protection of this invention. Attached Figure Description
[0014] Figure 1 This is a flowchart of an exemplary method.
[0015] Figure 2 An example data processing pipeline is shown.
[0016] Figure 3 An electronic data processing device is illustrated schematically. Detailed Implementation
[0017] The features, characteristics, and advantages of the present invention, as well as the ways of realizing these features, characteristics, and advantages, will become clearer and more understandable from the following description of embodiments in conjunction with the accompanying drawings.
[0018] The invention will now be explained in more detail with reference to the accompanying drawings of preferred embodiments. In the drawings, the same reference numerals denote the same or similar elements. The drawings are schematic diagrams of various embodiments of the invention. Elements shown in the drawings are not necessarily shown to scale. Rather, the representation of the various elements in the drawings enables those skilled in the art to understand their function and general purpose. Connections and couplings between functional units and elements shown in the drawings can also be achieved through indirect connections or couplings. Connections or couplings can be achieved via wired or wireless means. Functional units can be implemented through hardware, software, or a combination of hardware and software.
[0019] The following sections introduce techniques related to solving inverse image reconstruction tasks. Examples of inverse image reconstruction tasks from which the techniques described here can benefit include: deconvolution; denoising; high resolution; stray light reduction; spectral unmixing; and artifact reduction. Inverse image reconstruction tasks can be problematic because there is no single solution, the solution is unstable (small changes in image measurement data can lead to large differences in the reconstructed image), or the reconstructed image does not consistently depend on the image measurement data. Inverse image reconstruction tasks can be mapped from image space to image space. However, inverse image reconstruction tasks can also be performed based on image measurement data available in spatial frequency space, spatial space, or angular space.
[0020] For example, image measurement data can be optical microscope images of the sample. However, image measurement data can also be obtained in the spatial frequency space, such as in magnetic resonance imaging. Computed tomography measurement data can be converted into slice images. It can also be ultrasound measurement data.
[0021] Image measurement data is generally obtained through measurement and represents samples such as biological samples, workpieces, or organisms. Image measurement data may have certain characteristics that need to be removed during inverse image reconstruction.
[0022] The following section specifically introduces techniques for solving inverse image reconstruction tasks using machine learning models. These machine learning models are used in conjunction with image reconstruction algorithms. Therefore, such machine learning models are referred to as reconstruction models below. Reconstruction models can fully solve the image reconstruction task, i.e., converting image measurement data into image data. However, it is also conceivable that reconstruction models only perform the parts of the task related to the image reconstruction algorithm, such as regularization operations. For example, the reconstruction model can be a deep neural network, such as a deconvolutional network, like a so-called UNet or converter network with an encoder / decoder architecture and jumpers. For example, the machine learning network can be a so-called "unrolled" network, where each layer is assigned to different iterations of numerical iterative optimization. The reconstruction model can be a so-called image-to-image reconstruction model that maps from image space to image space.
[0023] The techniques described here are all based on the understanding that the accuracy of a reconstruction model can be improved by tailoring it to a specific, domain-specific inverse image reconstruction task. In other words, this means that the training of the reconstruction model is specific to the particular domain of the inverse image reconstruction task, depending on the example. The domain of an inverse image reconstruction task is determined by the type of sample to be measured. Different sample types, i.e., samples with inherently different structures, define different domains. The domain of an inverse image reconstruction task is also determined by the imaging method. For example, when capturing optical microscope images as image measurement data, the illumination used (e.g., bright field, dark field, or oblique illumination), the objective lens used, and the magnification used can all define the imaging method. In the case of magnetic resonance imaging, sequences, subsampling used in spatial frequency space, etc., can define the imaging method.
[0024] Using the techniques revealed here, domain-specific training data pairs can be generated particularly quickly without requiring significant effort, enabling domain-specific training of the reconstruction model. This allows for the generation of corresponding training instances of the reconstruction model for each specific inverse image reconstruction task, instances well-suited to the domain. In particular, sample-specific training is also conceivable. For example, the reconstruction model can be specifically trained for particular samples—thus obtaining sample-specific instances of the reconstruction model. These sample-specific instances of the reconstruction model can then be used to solve inverse image reconstruction tasks based on image measurement datasets obtained from the same samples. Thus, different sample-specific training instances of the reconstruction model can be used and / or stored for different samples. The optimal reconstruction model for a given sample can always be loaded and used to solve the image reconstruction task.
[0025] To train the reconstruction model, synthetic training data pairs can be generated and used based on the various examples described herein. Each synthetic training data pair comprises a synthetic image measurement dataset and an associated synthetic reference image, which are linked through the inverse image reconstruction task. The synthetic reference image serves as the basis for training the reconstruction model. Based on the training data pairs, the reconstruction model can be trained using machine learning techniques known in the literature, such as backpropagation.
[0026] To generate synthetic training data pairs, at least one reference image must first be acquired. The reference image corresponds to a solved inverse image reconstruction task. Preferably, the reference image is acquired within the domain of the corresponding inverse image reconstruction task, i.e., for samples of the corresponding sample type, or even specific samples; and using a specific imaging modality.
[0027] Then, using a first machine learning model (different from the reconstruction model), at least one feature description of the at least one reference image is generated based on the at least one reference image. Specifically, such a feature description can be a textual description of the reference image. For example, an image-to-text model can be used. For example, the feature description can be in free text format, i.e., human-readable free text, indicating a large number of features of the at least one reference image. However, the feature description can also indicate machine learning features, such as features in the latent feature space of the encoding branch of the first machine learning model. Combinations between free text and machine learning features are also conceivable. The first machine learning model used to generate the feature description is referred to below as the encoding model because it encodes the reference image in the form of at least one feature description.
[0028] Then, based on the at least one feature description, a second machine learning model (different from the reconstruction model, trained separately from the first machine learning model, and with a different architecture) can be used to generate one or more synthetic reference images based on the at least one feature description. Therefore, this means that at least one feature description can be used to generate synthetic reference images via the second machine learning model, wherein these synthetic reference images are similar to the original reference images due to the consideration of the at least one feature description. Thus, the second machine learning model is a generative model that can transform text (here referring to at least one feature description) into images (here referring to synthetic reference images). By utilizing a combination of encoding and generative models, a large number of synthetic reference images can be generated based on a relatively limited number of reference images (e.g., even based on a single reference image). The various synthetic reference images differ from each other because random components (e.g., through a “seed” value) can be taken into account during the generation process and / or when generating feature descriptions. Image generation can be controlled, for example, by corresponding instructions obtained via a user interface, particularly the range of variation between synthetic reference images. In particular, the variation of features in the synthetic reference images or the range of such variation in the set of synthetic reference images can correspond to physically observed variations or ranges of variation.
[0029] Then, degradation is performed on each synthetic reference image. This degradation corresponds to the inverse image reconstruction task. Degradation is achieved by applying the forward model of the inverse image reconstruction task. The forward model of the inverse image reconstruction task translates the cause of a phenomenon into the observed effect. In the inverse image reconstruction task, this means converting the image into image measurement data. Typically, the forward model is known, and even if the inverse model has problems, there will be no issues.
[0030] In this scenario, a corresponding synthetic image measurement dataset is available for each synthetic reference image. The synthetic reference images and their corresponding image measurement datasets are combined into training data pairs. These training data pairs are used to train a reconstruction model. The reconstruction model is trained to solve the inverse image reconstruction task, or at least helps to solve it. Therefore, the reconstruction model helps to eliminate the corresponding degradations in the image measurement data.
[0031] Figure 1 This is a flowchart of an exemplary method. Figure 1 The method can be executed by one or more electronic data processing devices. Therefore, Figure 1 The method is implemented using a computer. Figure 1 The method is used for image reconstruction related to inverse image reconstruction tasks.
[0032] exist Figure 1In this method, one or more reference images are obtained (box 905). These reference images are specific to the inverse image reconstruction task. Then, one or more feature descriptors are generated for these reference images (box 915). For example, an associated feature descriptor can be created for each reference image, or a common feature descriptor can be created for multiple reference images. For this purpose, an encoding model, such as a large language model, can be used. Then, a generative model (box 925) is used to generate one or more synthetic reference images based on the one or more feature descriptors. For each of these synthetic reference images, a forward model is applied, which performs image degradation corresponding to the inverse image reconstruction task. This yields a synthetic image measurement dataset (box 950). The reconstruction model can then be trained based on training data pairs (box 955), each pair containing a synthetic reference image as a basis and a synthetic image measurement dataset. This method will be described in detail below.
[0033] In box 905, at least one reference image is acquired. For example, at least one reference image can be obtained by measurement, such as by using a measurement method not present in subsequent inference (box 975). At least one reference image can also be generated using a known reconstruction algorithm based on reference image measurement data. At least one reference image can be loaded from memory. For example, at least one reference image can be loaded from an image database.
[0034] The appearance of one or more reference images obtained in box 905 has a significant impact on the training of the reconstruction algorithm. Therefore, one or more reference images obtained in box 905 are preferably domain-specific. This means that at least one reference image describes a sample of a specific sample type in the reconstruction task and is also acquired or generated based on the corresponding imaging method. In some examples, it is even conceivable that one or more reference images in box 905 are acquired for a specific sample, and then sample-specific training of the reconstruction model is performed on that sample to obtain a sample-specific instance of the reconstruction model.
[0035] However, there may be some differences in appearance between at least one reference image from box 905 and the image to be reconstructed later as part of the image reconstruction (this will be explained in more detail below in conjunction with the generation of the synthetic reference image in box 925).
[0036] In block 910, optionally, contextual information from at least one reference image from block 905 can be acquired. For example, this contextual information may include information about the field of view. For example, the contextual information may include information about the sample type and / or imaging method. For example, the contextual information may be in free text format. For example, such contextual information can be acquired from a user through a user interface. For example, such contextual information can also be determined based on metadata from the image capturing device used to acquire image measurement data that forms the basis of the reference image. The contextual information may display the magnification used and the lens used, etc. The contextual information may be contained in the header of one or more reference images from block 905.
[0037] Generate at least one feature description from at least one reference image from box 905 in box 915.
[0038] There are many methods for generating at least one feature description from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from at least one reference image from another ...
[0039] However, it is also possible to automatically generate at least one feature description based on at least one reference image. This situation will be described below: automatically determining at least one feature description using an encoding model. The encoding model can be machine learning-based.
[0040] In one example, a corresponding image-specific feature description is obtained for each coded image. For example, if a total of five reference images are obtained in box 905, five corresponding feature descriptions can be generated. For example, it can be envisioned that for each reference image from box 905, an associated feature description is generated for that reference image. This means that each reference image has an image-specific feature description. For example, in such an image-specific feature description, one or more image-specific global features can be indicated, and / or one or more image-specific local features can be indicated, possibly along with location information. For example, information that is generally valid for the entire reference image can be specified. However, it can also be envisioned that feature information about features appearing in certain regions of the reference image can be provided specifically. Various synthetic reference images are often better generated when more information is available in the context of at least one feature description. For example, the generation of synthetic reference images is particularly good if both global feature information and image-specific local feature information are available. The following figure is a corresponding example of a feature description for a specific reference image provided in free text format (i.e., natural language). This feature description contains image-specific global features and image-specific local features.
[0041] Table 1 below shows an example of feature descriptions, which are in free text format and include both global and local image features.
[0042]
[0043] Table 1: Examples of image-specific feature descriptions for reference images in free text format. This example shows that the feature description includes global image features ("Histological image, stained with H&E") and local features ("A layer of elongated columnar epithelial cells can be seen at the top of the image...").
[0044] The above describes an example of image-specific feature description. In this variation, it is conceivable to determine an associated image-specific feature description for each image from at least one reference image from box 905. However, alternatively or supplementarily, it is also conceivable to generate a common "global" feature description for multiple reference images from box 905. In this case, such a common feature description is no longer image-specific. Instead, such a common feature description may also contain information about the variation or range of variation of one or more features among the reference images.
[0045] This information about variations or ranges of variation in one or more features helps map the corresponding variations or ranges of variation in the generated synthetic reference image. In other words, it can be envisioned that, for example, the synthetic reference image includes the corresponding variations or ranges of variation of these features, which are also visible in multiple reference images from box 905. This allows for better training of the reconstruction model because variations in individual features are also taken into account.
[0046] Specifically, the above section introduced some examples of feature description implementations, where the feature descriptions were provided in free text format. However, different feature implementations can be envisioned. Some examples will be presented below.
[0047] For example, one or more feature descriptions may use technical language. For example, one or more feature descriptions may use terms from a predefined terminology catalog. One or more feature descriptions may contain a structure conforming to a predefined data exchange format. An example (in English) is given below:
[0048]
[0049] Table 2: Examples of image-specific feature descriptions of reference images given in technical language.
[0050] Other examples include using machine learning features in feature descriptions, such as those obtained through the encoding branch of an autoencoder network. Therefore, feature vectors can be obtained in the latent feature space of the encoding model. Tokens can also be obtained, for example, through visual transformer networks.
[0051] When generating feature descriptions based on machine learning features, image-global features and / or image-local features can also be considered. For example, in a converter architecture, each reference image can be processed based on patches to extract features from each patch (i.e., so-called "patch embedding").
[0052] Depending on the type of feature description, the encoding model can employ different architectures. However, it is also possible to use encoding models capable of providing feature descriptions of different types or formats. For example, such an example could be a large language model (e.g., used in conjunction with a visual model) that, with appropriate prompts, can be guided to output feature descriptions in a specific format.
[0053] Generally, it is helpful if the encoding model is a domain-nonspecific base model for the image reconstruction task. That is, such a non-domain-specific encoding model can handle different types of reference images. Its advantage is that domain-specific training data pairs can be generated for many different domains using the same data processing pipeline, specifically the same non-domain-specific encoding model.
[0054] Examples include CLIP (Contrastive Language-Image Pre-training), Segment Anything, and Masked Autoencoder.
[0055] For CLIP, see Radford, A., Kim, JW, Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., ... and Sutskever, 205I., “Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PMLR.” (Radford, A., Kim, JW, Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., ... & Sutskever, 205I. (2021, July).
[0056] For Segment Anything, see Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., ... and Girshick, R. in their 2023 article “Segment Anything” published on the ArXiv preprint arXiv:2304.02643 (Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., ... & Girshick, R. (2023). Segment anything.arXiv preprint arXiv:2304.02643).
[0057] For Masked Autoencoders, see He, K., Chen, X., Xie, S., Li, Y., Dollár, P., & Girshick, R. (2022). Masked autoencoders are scalable vision learners. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition (pp. 16000-16009)
[0058] Image generation instructions are available in box 920. For example, such instructions can be created as prompts to the corresponding generative model. For example, the instructions may include information about the number of synthetic reference images to be generated. For example, the instructions may contain some information about the variations or corresponding ranges of variation between different synthetic reference images. These instructions may be obtained from the user through a user interface.
[0059] Generally, box 920 is optional. For example, it can be envisioned that these instructions are predefined, or that the corresponding generative model is set to not require instructions. For example, the instructions in box 920 could indicate the desired range of variation between subsequently generated synthetic reference images. It can be envisioned that such instructions could specify which feature types will change between the various subsequently generated synthetic reference images.
[0060] Then, one or more synthetic reference images are generated in box 925. The one or more synthetic reference images are generated using a generative model. In particular, the generative model takes one or more feature descriptions from box 915 as input, and optional—if any—instructions from box 920.
[0061] For example, in principle, one or more feature descriptions from box 915 may not be processed by the generative model. In this case, the generative model would typically generate a synthetic reference image that is visually very similar to one or more original reference images from box 905. However, it is also conceivable to modify the one or more feature descriptions before they are input into the generative model. For example, specific feature types, such as the type and attributes of the sample, the type and attributes of a specific image region, or the imaging mode, can be specifically changed.
[0062] For example, generative models can be used: (latent) diffusion models, such as stable diffusion, Dall-E, or MidJourney. For an introduction to stable diffusion models, see the paper "High-resolution image synthesis with latent diffusion models" by Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022). In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition (pp. 10684-10695). Other possible architectures include generative adversarial networks, such as StyleGAN or variational autoencoders.
[0063] As can be seen from the above, a large number of synthetic reference images can be generated using two machine learning models (the encoding model in box 915 and the generative model in box 925). By using corresponding base models that are domain-specific for each inverse image reconstruction task as these machine learning models, the task of generating synthetic reference images for many different inverse image reconstruction tasks can be accomplished. For many different samples, the generation of synthetic reference images can be customized to be sample-specific. Dedicated models are also used, one for generating the feature description in box 915 and another for generating the synthetic reference image in box 925. These models are typically dedicated to this task (text generation in box 915 and image generation in box 925), that is, they have different architectures and are trained separately.
[0064] Box 926 can be used to examine one or more synthetic reference images previously generated in box 925. This examination can be performed manually or automatically, for example, through another machine learning model. Examples include classification models (such as checking if the contrast type is correct), segmentation models (such as checking if the number and location of objects are correct), and detection models (such as segmentation, etc.).
[0065] In box 930, it is optional to check whether one or more iterations of at least some of the boxes in the preceding boxes should be performed (935). For example... Figure 1 As shown, for each iteration 935, block 925 can be executed, and optionally blocks 915 or 920 can be executed. This means that in each iteration 935, one or more feature descriptions can be generated, one or more image generation instructions can be generated, and one or more synthetic reference images can be generated. For example, in each iteration 935, when determining one or more feature descriptions in a specific instance of block 915, different random components can be considered each time. This allows for the generation of different feature descriptions for the same one or more reference images from block 905 in each of the different iterations 935 of block 915. Correspondingly, different random values can also be considered in the different iterations 935 of block 925. This technique allows for greater variation in the generation of synthetic reference images in block 925.
[0066] In some examples, it is also conceivable that a specific iteration 935 takes into account one or more results from a previous iteration 935. For example, based on the corresponding feature descriptions from a previous iteration 935 in box 915, new feature descriptions can be generated in subsequent iterations 935 of box 915. This can selectively consider user feedback regarding whether the synthetic reference image from the previous iteration is “good” or “bad.” The user can also mark certain image regions that need improvement or are particularly suitable for the corresponding image reconstruction task. The results of the check in box 926 can also be considered when generating subsequent feature descriptions. Therefore, it is conceivable that in iteration 935 of box 915, one or more feature descriptions are determined based on one or more synthetic reference images from a previous iteration 935 of box 925. Thus, through multiple iterations 935, iterative and potentially interactive processing can be performed to improve the quality of the synthetic reference image and / or generate additional synthetic reference images.
[0067] In box 950, an associated synthetic image measurement dataset is generated for each synthetic reference image. For this, a forward model for the inverse image reconstruction task is used. This is equivalent to degradation.
[0068] The forward model implements image degradation, and its underlying physical processes are known and generally well-described / computed. The image is artificially degraded through the physical model. Possible configurations of the forward model include: artificial scattering (introducing scattered light into the sample); convolution of the image with the (system's) point spread function; artificial noise (Gaussian, Poisson, Salt & Pepper, etc.); reduced axial / lateral resolution; color channel spectral mixing based on wavelength stacks and known excitation and emission spectra; and the introduction of known image artifacts. Examples of image artifacts include ringing or hammering artifacts. Other examples of image artifacts include edge artifacts; these artifacts occur particularly in light-sheet microscopy due to occlusion, but can also be caused by the camera in wide-field microscopy.
[0069] The reconstruction model is trained in box 955. Training can be performed locally or on a server. In this way, training data pairs are obtained. Each training data pair contains a synthetic reference image, which is associated with its respective synthetic image measurement dataset via the feedforward model of the inverse image reconstruction task.
[0070] Based on such training data pairs, training of the reconstruction model can be triggered in box 955. For example, training can be performed on a server (typically computationally expensive). However, training can also be done locally. As part of the training, the weights of the reconstruction model are optimized, making it possible to solve the inverse image reconstruction task as well as possible. The reconstruction model is then used as part of the corresponding reconstruction algorithm.
[0071] In some examples, it is conceivable that so-called fine-tuning might be performed within box 955. This means that the reconstruction model has been pre-trained, for example, on training data pairs imaged on similar samples or sample types and / or taken using the same or similar imaging methods. As part of the fine-tuning, the corresponding weights of the pre-trained machine learning reconstruction model can be adjusted, for example, to obtain sample-specific instances of the reconstruction model. However, it is also conceivable that the reconstruction model, which has not been previously trained, could be trained based on random weights.
[0072] In box 970, a pre-trained reconstruction model can be distributed, i.e., provided to multiple field devices (also known as agents). Box 970 is optional. In some examples, the reconstruction model can be trained in box 955 based on a reference image captured in the field in box 905 to perform a specific inverse image reconstruction task. In this case, training can be performed on the field devices, so distribution in box 970 is unnecessary. Therefore, box 970 is optional.
[0073] In box 975, a previously trained reconstruction model is used to solve the inverse image reconstruction task. This means that image reconstruction is to be performed.
[0074] Figure 2 The diagram illustrates data processing pipelines based on different examples. For example, Figure 2 The data processing pipeline can achieve Figure 1 The method. Figure 2 The data processing pipeline 700 enables the generation of training data pairs 732 for training the reconstruction model 735 of machine learning. Figure 2 (Illustrated schematically as an encoder-decoder architecture).
[0075] Specifically, first, obtain reference image 705 (for comparison). Figure 1 (Block 905). Based on this reference image, a feature description 715 of the reference image 705 is generated by the encoding network 710 (comparison). Figure 1 (Block 915). This feature description 715 is then used as input to a generative model 720 to generate multiple synthetic reference images 725 (for comparison). Figure 1 (Block 925). Then, based on the forward model of the inverse image reconstruction task, these synthetic reference images 725 are degraded 726 to physically and correctly degrade the images. The corresponding image measurement data 730 (which is also the image in image space here) is as follows: Figure 2 As shown (comparison) Figure 1 (Box 950). Then, training data pairs 732 are formed (each data pair includes a corresponding synthetic reference image 725 and a corresponding image measurement dataset 730) and used to train the reconstruction model 735 (comparison). Figure 1 (955 in the box)
[0076] Figure 3 An electronic data processing apparatus 90 according to various embodiments of the present invention is schematically illustrated. The electronic data processing apparatus 90 includes a processor 91 and a memory 92. The memory 92 may include one or more volatile memory modules and / or one or more non-volatile memory modules. Furthermore, the electronic data processing apparatus 90 includes a communication interface 93 through which the processor 91 receives and / or transmits data. The processor 91 can load and execute program code from the memory 92. When the processor 91 executes such program code, it causes the processor to perform the techniques disclosed herein, such as: sending control commands to an image capturing device to obtain image measurement data; applying image reconstruction algorithms; applying one or more reconstruction models in a reverse image reconstruction task; training reconstruction models based on training data; generating synthetic reference images and / or synthetic image measurement datasets, etc. For example, the processor 91 can execute according to the program code. Figure 1 The method.
[0077] Of course, the features of the embodiments and aspects of the present invention described above can be combined with each other. In particular, these features can be used not only in the described combinations, but also in other combinations or individually, without departing from the scope of the present invention.
[0078] For example, the technique of automatically generating feature descriptions through encoding models has already been introduced above. However, it is also conceivable to generate such feature descriptions at least partially based on user input. Semi-automatic techniques for generating feature descriptions are also worth considering. For example, machine learning encoding models can be used to generate feature descriptions in free text format, and then the feature descriptions can be output to the user. The user can then adjust, supplement, or refine the feature descriptions.
Claims
1. A computer-implemented method, comprising: - Obtain at least one reference image; - Generate at least one feature description of at least one reference image based on the at least one reference image using a first machine learning model; - Generate multiple synthetic reference images based on the at least one feature description using a second machine learning model; - For each of the plurality of synthetic reference images: apply the forward model of the inverse image reconstruction task to obtain the respective synthetic image measurement dataset; as well as - The training of a third machine learning model is initiated based on the training data pairs, so that the third machine learning model can be used to solve the reverse image reconstruction task. Each training data pair includes one of the plurality of synthetic reference images and its corresponding synthetic image measurement dataset.
2. The computer-implemented method according to claim 1, in, The at least one feature description indicates multiple features of the at least one reference image in a free text format.
3. The computer-implemented method according to claim 1 or 2, in, The at least one feature description indicates multiple machine learning features of the at least one reference image.
4. The computer-implemented method according to any one of the preceding claims, in, The at least one feature description indicates one or more image-specific global features, and / or The at least one feature description indicates one or more image-specific local features, which may optionally be indicated together with location information.
5. The computer-implemented method according to any one of the preceding claims, in, Generate an associated feature description for each of the at least one reference image.
6. The computer-implemented method according to any one of the preceding claims, in, Generate common feature descriptions for multiple reference images.
7. The computer-implemented method according to claim 6, The common feature description indicates the variation or range of variation of at least one feature among the plurality of reference images.
8. The computer-implemented method according to any one of the preceding claims, in, The first machine learning model further generates at least one feature description of the at least one reference image based on predetermined contextual information about the at least one reference image.
9. The computer-implemented method according to any one of the preceding claims, in, The first machine learning model and / or the second machine learning model are base models that are domain-specific for the inverse image reconstruction task.
10. The computer-implemented method according to claim 9, in, The domain of the reverse image reconstruction task is defined by at least one of the sample type of the imaging samples of the reverse image reconstruction task or the imaging mode of the image measurement dataset of the reverse image reconstruction task.
11. The computer-implemented method according to any one of the preceding claims, in, The reverse image reconstruction task is selected from the following group: deconvolution, denoising, high resolution, stray light reduction, spectral unmixing, and artifact reduction.
12. The computer-implemented method according to any one of the preceding claims, wherein the method further comprises: - Obtain instructions for the second machine learning model, wherein the instructions indicate the range of variation between the synthetic reference images.
13. The computer-implemented method according to any one of the preceding claims, in, Multiple distinct feature descriptions are generated based on varying random components, and / or The multiple different synthetic reference images are generated based on varying random components.
14. The computer-implemented method according to any one of the preceding claims, in, The third machine learning model is trained to generate training instances of the third machine learning model that are domain-specifically associated with the inverse image reconstruction task.
15. The computer-implemented method according to any one of the preceding claims, in, The at least one reference image is obtained for a specific sample, thereby obtaining sample-specific training instances for the third machine learning model through training. The method further includes -Based on the image measurement dataset obtained for the specific samples, the sample-specific training instances of the third machine learning model are applied to solve the image reconstruction task.
16. The computer-implemented method according to any one of the preceding claims, in, The reverse image reconstruction task includes a mapping from image space to image space; Optionally, the third machine learning model provides a mapping from the image space to the image space.
17. The computer-implemented method according to any one of the preceding claims, in, The third machine learning model provides optimized regularization operations associated with the reverse image reconstruction.
18. An electronic data processing apparatus comprising a processor and a memory, the processor being arranged to load and execute program code from the memory, wherein when the program code is executed, the processor performs a method according to any one of the preceding claims.