Method of performing feature detection of image features, scanning probe microscopy system and machine learning model.

A machine learning model, trained with synthetic and real images, addresses the inaccuracies in scanning probe microscopy by accurately segmenting and localizing image features, improving detection efficiency and reducing errors in scanning probe microscopy systems.

WO2025178491A1PCT designated stage Publication Date: 2025-08-28NEARFIELD INSTR BV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/NL2025/050084
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-23
Filing Date
2025-02-21
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Existing image processing methods for scanning probe microscopy systems are laborious, prone to human error, and struggle with misinterpreting image disturbances such as noise, artefacts, and substrate contaminations, leading to inaccurate feature detection.

Method used

Employ a machine learning data processing model, such as a U-net type neural network, trained with synthetic and real images to perform image segmentation, generating feature masks that accurately represent and localize image features while filtering out disturbances.

Benefits of technology

The method provides fast and reliable feature detection by significantly reducing erroneous identifications of image features, enhancing accuracy and efficiency in identifying structural features on or below the sample surface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure NL2025050084_28082025_PF_FP_ABST
    Figure NL2025050084_28082025_PF_FP_ABST
Patent Text Reader

Abstract

The invention is directed at a method of performing feature detection of image features in an image obtained from a scanning probe microscopy system, wherein the image has been obtained by scanning a surface of a sample using a probe including a probe tip, such that the image features correspond to structural features on or below the surface of the sample, wherein the method comprising the steps of receiving, by a processor from a data source associated with the scanning probe microscopy system, image data of the image, wherein the image includes the image features and image disturbances. The image data is provided, by the processor, as input data to a machine learning data processing model, wherein the machine learning data processing model has been trained to perform image segmentation. The invention is further directed at a scanning probe microscopy system and at a machine learning model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Title: Method of performing feature detection of image features, scanning probe microscopy system and machine learning model. Field of the invention The present invention is directed at a method of performing feature detection of image features in an image obtained from a scanning probe microscopy system, wherein the image has been obtained by scanning a surface of a sample using a probe including a probe tip, such that the image features correspond to structural features on or below the surface of the sample, wherein the method comprising the steps of receiving, by a processor from a data source associated with the scanning probe microscopy system, image data of the image, wherein the image includes the image features and image disturbances. The invention is further directed at a scanning probe microscopy system and at a machine learning model. Background Although the present document includes references to various other documents, no admission is made that any reference constitutes prior art. The discussion of references refers to their content as presented therein, and does not acknowledge nor confirm the accuracy or pertinency thereof. It will be understood that, although a number of prior art publications are referred to herein, this reference does not constitute an admission that any of these documents form part of the common general knowledge in the art in any country. In the era of advanced three-dimensional (3D) semiconductor nodes with shrinking dimensions, robust and fast image processing is important for metrology and inspection. Presently, new types of defects and new parameters found to be critical for yield improvement can be found by various techniques. For example, manual processing may be performed relatively well, but at the same time this method suffers from being laborious and slow, and it is still prone to human error. Traditional image processing algorithms are a lot faster, but fall short in misinterpreting certain disturbances as features or feature boundaries as well as misinterpreting noisy image features as pure noise. These disturbances may include noise signals, image artefacts and substrate contaminations, and are often present in microscopic images. The above is important in the processing of highly magnified images, such as but not limited to images from scanning probe microscopes. Summary of the invention It is an object of the present invention to provide a solution to the above problems with existing methods and systems, and in particular to provide a fast and reliable method of performing feature detection of image features in an image obtained from a scanning probe microscopy system. To this end, in accordance with a first aspect, there is provided herewith a method of performing feature detection of image features in an image obtained from a scanning probe microscopy system, wherein the image has been obtained by scanning a surface of a sample using a probe including a probe tip, such that the image features correspond to structural features on or below the surface of the sample, wherein the method comprising the steps of: receiving, by a processor from a data source associated with the scanning probe microscopy system, image data of the image, wherein the image includes the image features and image disturbances; providing, by the processor, the image data as input data to a machine learning data processing model, wherein the machine learning data processing model has been trained to perform image segmentation such as to provide, based on the input data, output data including representation data of one or more representations of the image features and localization data of locations associated with the one or more representations; obtaining, by the processor from the output data of the machine learning data processing model, the representation data which is indicative of the one or more representations of the image features and the localization data of locations associated with the one or more representations; and providing, by the processor via an output device, the representation data and the localization data, for establishing said feature detection. The present invention applies a machine learning data processing model in order to perform image segmentation. This enables to provide output data based on the images provided as input data, wherein the output data includes representation data of one or more representations of the image features and localization data of locations associated with the one or more representations. This may be done efficiently while significantly reducing the chance on erroneous identification of image features, such as missed image features, incorrect feature shapes or incorrect locations of image features. The machine learning data processing model may for example be a U-net type neural network, but it may also be of a different type that is suitable to obtain the image features as well as their locations from an image. For example, alternatively the machine learning data processing model may be a vision transformer model, for example a segmenter transformer (SETR) or a pyramid vision transformer (PVT) or a SegFormer. Also hybrid models are potentially possible, for example architectures combining convolutional and transformer blocks such as TransUNet, cost aggregation transformers (CATs) or U-net transformer (UNETR). In some embodiments, the output data of the machine learning data processing model comprises feature mask data of a feature mask image including the representation data and the localization data, and wherein the step of obtaining the representation data and the localization data comprises the steps of obtaining, by the processor, the feature mask data and obtaining the representation data and the localization data from the feature mask data. Although in the manner described, it is possible to obtain the representation data and localization data straight from the output of the machine learning data processing model, a typical convenient manner to establish this is by generating a feature mask image containing the representations of the image features in their exact and correct locations. A feature mask provides the master image of one layer including features of a semiconductor structure. From this image, the representation data and localization data may be easily obtained. An advantage of this manner of working is that feature masks may already be available for many images, which may be used as ground truth data to train a machine learning data processing model in this way. For example, applying the scanning probe microscopy to a manufacturing process for semiconductors, the designs of the semiconductor structures may be known and may have been verified in a different manner, from which it is known what image features are to be found in which locations. Such a training method thereby trains the machine learning data processing model to generate a feature mask image on the basis of an input image. In some embodiments, the feature mask image obtained in this manner is void of said image disturbances. The machine learning data processing model may well be trained to identify only the features sought for, and thus free of disturbances such as noise signals, image artefacts and substrate contaminations. The machine learning model will enable to generate at its output the desired data components without disturbances, e.g. the feature masks free of any disturbances. In accordance with various preferred embodiments, the machine learning data processing model is a U-net type machine learning data processing model. Certainly, a U-net type machine learning data processing model may well be trained to achieve the above. It is particularly suitable to obtain the feature data by means of pooling in various convolution steps, while preserving the localization data of those features by up-sampling again in combination with concatenation of data from the convolution steps. The above wording must be interpreted to encompass any variants of a U-net, such as a recurrent residual convolutional neural network based on U-Net (R2U-Net), a deep fully convolutional neural network architecture for semantic pixel- wise segmentation (SegNet), or a residual U-Net (ResU-Net), UNet++, nnU-Net, V- Net, etc. For example, in an effective implementation, in accordance with certain embodiments, the U-net type machine learning data processing model includes a contracting path of convolution steps for capturing image context from the image data for enabling recognition of the image features for providing the representation data, and further includes an expanding path of de-convolution steps for enabling localization of the image features for providing the localization data. The contraction path of the U-net type machine learning data processing model effectively is able to obtain or identify the image features relating to structural features of the substrate, but then void of the disturbances in the image. This is done via a number of pooling operations in various layers, using a pooling mask of a desired size. The image data is convolved in each layer to reduce the resolution, while at the same time increasing the number of feature channels. It allows to obtain representation data indicative of representations of the image features sought for, while filtering out the disturbances.The expansion path thereafter up-samples (also referred to as ‘up-convolves’ or‘deconvolves’) image again. At each level of the expansion path, the image at that levelis concatenated with the correspondingly cropped feature map from the contracting path. This will gain back, and hence yield the exact position information associated with the features. As a result, at the output of the U-net type machine learning data processing model, a feature mask image is obtained which includes both the representation data and the localization data and which is void of most disturbances. The application of the machine learning data processing model as per the present invention, as described herein, provides very good result in identifying the image features in the images from the scanning probe microscope. Thereby, the results are very satisfactory. However, a disadvantage is that in order to train the machine learning data processing model, it typically requires the machine learning model to be fed with a large amount of training images for which ground truth data is needed. This may be problematic in the sense that in many cases the required data is proprietary and hence not available in large quantities to a manufacturer of a scanning probe microscopy system. Apart from that, and may be more important, the training in this case may also take a lot of time, and although effectively enabling labelling and localization of the type of image features trained for, it may be less effective in correctly labelling image features that sufficiently deviate therefrom. This may be overcome by the following class of embodiments. In accordance therewith, for training of the machine learning data processing model, the method further comprises the steps of: generating, by an image generator, a plurality synthetic images including synthetic image features; generating, by the image generator, for each synthetic image of the plurality of synthetic images, an associated binary segmentation mask, such as to provide a plurality of pairs of synthetic images with associated binary segmentation masks; augmenting, by the image generator, each synthetic image of the plurality of pairs with one or more artificial image disturbances, such as to provide the plurality of pairs of augmented synthetic images with associated binary segmentation masks as synthetic training data set for training of the machine learning data processing model. In this manner, it becomes possible to generate a large data set for any arbitrary sort of image features (e.g. in terms of shape, contrast, texture, or any gradient therein). The training data set includes both synthetic images that can be used as input, as well as binary segmentation masks that may be applied as ground truth in order to train the machine learning data processing model. Optionally, in some embodiments, the method further comprises storing the synthetic training data set in the data source; although it is also possible to feed the training data set or the pairs therein directly to the machine learning data processing model, without prior storage. In some embodiments, the method further comprises a training of the machine learning data processing model, wherein the training is performed prior to the step of receiving image data of the image from the data source associated with the scanning probe microscopy system, and wherein the training comprises the steps of: providing training data from the synthetic training data set to an input of the machine learning data processing model, wherein the training data comprises augmented synthetic images from the plurality of pairs in the synthetic training data set; provide, at an output of the machine learning data processing model, candidate feature masks generated by the machine learning data processing model based on the augmented synthetic images in the training data; receiving, by the processor, ground truth data, wherein the ground truth data comprises the binary segmentation masks associated with each of the synthetic images in the training data; and training the machine learning data processing model based on the training data by comparing the candidate feature masks to the binary segmentation masks in the ground truth data, such as to enable the machine learning data processing model to generate, based on said input data, said output data including the representation data and the localization data. The training in this manner may be performed very effectively and be optimized many different conditions. Various sorts of disturbances may be added artificially, and in combination with an arbitrary number of different sorts of image features, which benefits the quality of the training and subsequently the quality of the output of the machine learning data processing model. By adding such variations in the synthetic training data, which can be done without limitation, the machine learning data processing model becomes very effective in recognizing different kinds of image features and distinguish them from disturbances. In particular, the step of training may comprise modifying one or more weighting values of one or more nodes of the machine learning data processing model based on the step of comparing the candidate feature masks to the binary segmentation masks. In other or further embodiments, the method comprises training of the machine learning data processing model, and wherein the training comprises the steps of: providing training data to the input of the machine learning data processing model, wherein the training data comprises real images obtained from the data source associated with the scanning probe microscopy system, the real images comprising real image features and real image disturbances; provide, at the output of the machine learning data processing model, candidate feature masks generated by the machine learning data processing model based on the real images in the training data; receiving, by the processor, ground truth data, wherein the ground truth data comprises the real feature masks associated with each of the real images in the training data; and training the machine learning data processing model based on the training data by comparing the candidate feature masks to the real feature masks in the ground truth data, such as to enable the machine learning data processing model to generate, based on said input data, said output data including the representation data and the localization data. In accordance with these implementations of the present concept, real image and real feature masks are used to train the machine learning data processing model. Despite the earlier mentioned disadvantages, this remains to be a good manner of training the machine learning data processing model, provided that sufficient training data is available. Again, as in the earlier training based on synthetic training data, the step of training may comprise modifying one or more weighting values of one or more nodes of the machine learning data processing model based on the step of comparing the candidate feature masks to the real feature masks. A further advantage is achieved in combining the abovementioned training methods. In this case, transfer learning may be applied in order to perform the training very effectively, and in order to keep on training the machine learning data processing model while being used on real images during normal operation. For example, in accordance with some embodiments, the training based on the synthetic training data set including the pairs of synthetic images and associated binary segmentation masks is performed as pre-training of the machine learning data processing model, while subsequently the training based on the real images and real feature masks is performed as fine-tuning of the machine learning data processing model. In this manner, a very high quality trained U-net type machine learning data processing model is obtained while only requiring a very limited amount of real images from the scanning probe microscope. In yet further embodiments, the structural features to which the image features correspond, relate to at least one of a group comprising: on-surface features, such as holes, trenches, extensions, dimples, ridges or walls; sub-surface features, such as structural features, density interfaces or cavities underneath the surface of the substrate. In fact, this present method may be applied to perform feature detection of a lot of very different types of features, and it is not limited to certain features. However, the given examples may frequently be found in various situations and applications of a scanning probe microscope, such as in a manufacturing process of semiconductor structures. In accordance with a second aspect thereof, the invention further relates to a scanning probe microscopy system comprising a probe including a probe tip, wherein the scanning probe microscopy system includes or cooperates with a substrate carrier for supporting a substrate, the scanning probe microscope being configured for moving the probe relative to a surface of the substrate in one or more directions, for establishing contact between the probe tip and the surface for mapping one or more surface structures on the surface, and for generating image data for providing an image based on said mapping, wherein the scanning probe microscope further comprises a processor communicatively connected to a data source, wherein the processor is configured to performing feature detection of image features in the image of the scanning probe microscopy system, and wherein for providing the feature detection the processor is configured for: receiving, by the processor from the data source, the image data of the image including the image features and image disturbances; providing, by the processor, the image data as input data to a machine learning data processing model, wherein the machine learning data processing model has been trained to perform image segmentation such as to provide, based on the input data, output data including representation data of one or more representations of the image features and localization data of locations associated with the one or more representations, and wherein the machine learning data processing model is a U-net type machine learning data processing model; obtaining, by the processor from the output data of the machine learning data processing model, the representation data which is indicative of the one or more representations of the image features and the localization data of locations associated with the one or more representations; and providing, by the processor via an output device, the representation data and the localization data, for establishing said feature detection. In accordance with a third aspect thereof, the invention further relates to a machine learning data processing model for use in a method according to the first aspect, wherein the machine learning data processing model has been trained to perform image segmentation such as to provide, based on the input data, output data including representation data of one or more representations of the image features and localization data of locations associated with the one or more representations. In some embodiments thereof, the machine learning data processing model is a U-net type machine learning data processing model. However, the invention is not limited to this type of machine learning data processing model, and may also be of a different type that is suitable to obtain the image features as well as their locations from an image. For example, alternatively the machine learning data processing model may be a vision transformer model, for example a segmenter transformer (SETR) or a pyramid vision transformer (PVT) or a SegFormer. Also hybrid models are potentially possible, for example architectures combining convolutional and transformer blocks such as TransUNet, cost aggregation transformers (CATs) or U-net transformer (UNETR). Brief description of the drawings The invention will further be elucidated by description of some specific embodiments thereof, making reference to the attached drawings. The detailed description provides examples of possible implementations of the invention, but is not to be regarded as describing the only embodiments falling under the scope. The scope of the invention is defined in the claims, and the description is to be regarded as illustrative without being restrictive on the invention. In the drawings: Figure 1 schematically illustrates a scanning probe microscopy (SPM) system in accordance with the invention, or in or for which the method of the present invention may be applied; Figures 2A-C show results of a conventional image processing method for identifying image features; Figures 3A and 3B show results of a conventional image processing method for identifying image features; Figures 4A and 4B show results of a conventional image processing method for identifying image features; Figure 5 schematically illustrates a U-net type machine learning data processing model in accordance with the present invention, which may be applied in a method of the present invention; Figure 6 schematically illustrates a training data generation process in accordance with an embodiment of the method of the invention; Figure 7 schematically illustrates a training process in accordance with an embodiment of the method of the invention; Figure 8 schematically illustrates a training process based on transfer learning, in accordance with an embodiment of the method of the invention; Figures 9A to 9D show a comparison between results of a conventional image processing method for identifying image features and results obtained using a method of the present invention. Detailed description Terminology used for describing particular embodiments is not intended to be limiting of the invention. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. The term "and / or" includes any and all combinations of one or more of the associated listed items. It will be understood that the terms "comprises" and / or "comprising" specify the presence of stated features but do not preclude the presence or addition of one or more other features. It will be further understood that when a particular step of a method is referred to as subsequent to another step, it can directly follow said other step or one or more intermediate steps may be carried out before carrying out the particular step, unless specified otherwise. Likewise it will be understood that when a connection between structures or components is described, this connection may be established directly or through intermediate structures or components unless specified otherwise. The invention is described more fully hereinafter with reference to the accompanying drawings, in which embodiments of the invention are shown. In the drawings, the absolute and relative sizes of systems, components, layers, and regions may be exaggerated for clarity. Embodiments may be described with reference to schematic and / or cross-section illustrations of possibly idealized embodiments and intermediate structures of the invention. In the description and drawings, like numbers refer to like elements throughout. Relative terms as well as derivatives thereof should be construed to refer to the orientation as then described or as shown in the drawing under discussion. These relative terms are for convenience of description and do not require that the system be constructed or operated in a particular orientation unless stated otherwise. Figure 1 schematically illustrates a scanning probe microscopy (SPM) system 1. The system 1 is suitable to be used in a method in accordance with the present invention. For example, the system 1 includes a processing device 26 and amemory 27 and / or 27’ which are suitable for storing instructions which, when executedby the processing device 26, cause the processing device 26 to carry out a method as described herein, in accordance with one or more embodiments of the invention. The processing device 26 in figure 1 is illustrated as a single entity in the SPM system 1. However, the skilled person may appreciate that although all the described method steps may be implemented by using a single processing device 26, the processing device 26 may be implemented by using multiple elements that together perform the described method steps. The processing device 26 may thus comprise a cluster of multiple processing devices, or may be implemented by various entities which individually perform certain (partial) steps and which cooperatively implement the invention. Furthermore, as illustrated, the memory 27 may be an internal memory 27or may be an external memory 27’ or another external entity reachable via a datacommunication network 29. To communicate with data communication network 29, the system 1 may comprise a communication unit 28. In figure 1, the data storage elements 27, data processing elements 26 and data communication elements 28 are all illustrated as being part of analyzer unit 25 of the system 1. Although this may in many cases be implemented in this manner, this may not always be the case (as already suggested above). The illustration of a single entity 25 in figure 1 is for the sole purpose of not unnecessarily complicating the figures. In the figures, elements that are technical and functional equivalents, i.e. performing a same or similar function in a same or similar manner with respect to the invention as described herein, may be designated by a same reference numeral or by asame reference numeral followed by a prime (‘) or a sub-numbering (“-1”, “-2”, …).These entities, such as data repositories 27 and 27’, may be of a same nature, of adifferent technical nature or may be implemented (e.g. connected or controlled) in a different manner, while in terms of the invention both providing the function of enabling the storing of data or operation instructions for the processing device 26. Thememory 27’ in figure 1, has been designated including prime (‘) in order to indicatethat although this element performs (or is able to perform) a function similar or evenin certain embodiments identical to the internal memory 27 of system 1’, this memory27’ (which may even be implemented as a server or as an externally stored data file ordata base) is different in the sense that it is not an internal memory but an external memory, without departing from its function in the embodiments of the present invention. The above is just an example, and may apply likewise to other entities described below. In principle, unless the contrary is specifically indicated in any part of the present document, it is to be assumed that any entity or element described may be implemented in a different manner in an alternative embodiment. The embodiments described or illustrated are not to be considered as limiting on the invention, which is only restricted by the appended claims defining the scope and spi- rit of the invention. In the system 1 of figure 1, a substrate carrier 3 is configured for supporting the substrate or sample 7 to be examined by the SPM system 1. The substrate carrier 3 is configured, by comprising or being connected to actuators (not shown), for moving the substrate 7 in a plane parallel to the carrier 3. This is typically referred to as the XY–plane of the system. Apart from the ability to move the substrate 7 in the X- and Y-directions, the substrate carrier 3 is connected to a metrology frame 5 which provides a fixed base to the system. The SPM system 1 further includes one or more scan heads 15 which are movable in the Z direction. The scan heads 15 each include a chip holder 16 enabling to hold a probe chip 13 comprising a probe 10 forming the sensing element of the SPM system 1. The probe 10 includes a cantilever 12 and probe tip 11. The probe tip 11, typically includes a very sharp tip that allows to very accurately (with nanometer accuracy) scan and take measurements at the surface 8 of a substrate 7. In addition to the above, typically the scan head 15 further include a sensor system for determining the exact position of the cantilever relative to the scan head. For example, the sensing system in figure 1 consists of an optical beam deflection (OBD) unit provided by a laser unit 20 for generating an optical beam 23, and a four quadrant optical sensor 21 that receives the reflected beam 23. The beam 23 is directed by the optical system to the backside of the probe cantilever 12, which includes a specular reflective surface. The specular reflective surface of the probe cantilever 12 reflects the beam 23 onto the optical detector 21. Any displacement of the probe tip 11 results in the location of impact of the beam 23 on the optical detector 21 to displace as well. In a four quadrant optical detector 21, the light spot formed by beam 23 preferably by default is set to be located exactly in the middle of the four different quadrants of the detector. Therefore, an equal part of the light spot falls on to each quadrant of the four quadrant optical detector 21. It may be appreciated that, even if in practice the laser beam 23 would not exactly be aligned in this manner, the principle of detecting a displacement of the light spot will be the same. All that is needed in order to detect a displacement is that each of the four quadrants of the optical detector 21 receives a fraction of the light from beam 23. For example, a relative vertical deflection of the cantilever (up or down with respect to the cantilevers equilibrium position) may be determined by obtaining a signal from the top half (T) of the detector 21 minus the signal from the bottom half (B) of the detector, i.e. the top- bottom signal (T-B signal). If the probe tip 11 displaces or bends relative to the scan head 15 comprising the laser unit 20, the light spot formed on the optical detector 21 slightly displaces such that the ratio between the different areas illuminated by the light spot on each quadrant of the optical detector 21 changes. From this, the exact position and / or orientation of the probe tip 11 with respect to the scan head 15 can be determined. Because also the Z-position of the probe 10 as applied by the Z-actuator 18 is known (from the control data of the Z-actuator 18 or from a dedicated Z-sensor), the orientation and location of the probe tip 11 in the Z-direction can be determined. Furthermore, the XY position of the probe tip 11, indicating its position relative to the sample 7 within the plane of the substrate carrier 3, is known from the actuators of the substrate carrier 3 or corresponding sensors. In this manner, in each location of the XY plane, the exact orientation of the probe tip 11 in the Z-direction can be determined in combination with the Z-level applied via the Z-actuator 18 or dedicated Z-sensor. As an alternative to the OBD sensor system described above, a different sensor system may be applied in order to determine the position of the probe tip 11. For example, alternatively a piezoresistive sensor may be applied. The manner of applying a piezoresistive sensor in an SPM system 1 is not further described here. In the system 1 illustrated in figure 1, the above information allows to very accurately determine the exact height of the surface 8 at each point in the XY plane. Therefore, the surface topography 9, consisting of a variety of different structures on the surface 8 of the sample 7, can be accurately determined and mapped such as to provide a topography map. In each point in the XY plane parallel to the sample surface 8, the exact Z-level of the surface 8 can be determined from the data provided for controlling the Z level actuator 18 and the data coming from the optical beam detector optical sensor 21 or from a dedicated Z-sensor. The data may for example be registered as three-dimensional measurement data in the memory 27 of the system 1. To provide a topography map of the 3D topography of the surface 8, in addition to the above analysis of the data of the Z-level actuator 18 or Z-sensor and the optical detector 21, a number of other processing steps have to be performed. These for example include noise reduction and the identification and removal of measurements artifacts. Thus, in the SPM system 1, the location and orientation of probe tip 11 is obtained and registered by obtaining measurements with the OBD detector formed by laser unit 20 and optical detector 21 in combination with Z-level data obtained from the actuator control data of Z-level actuator 18 or Z-sensor. Furthermore, in the system 1, optionally, subsurface measurements may be performed of any structures below the surface 8 of sample 7. To this end, the substrate carrier 3 optionally further includes a vibrational actuator that enables to apply a vibration 4 to the sample 7 from below. In alternative SPM systems, the vibrational signal 4 may be applied in different ways, for example by vibrational actuators on the surface 8 of the sample 7 or on the sides thereof. It is also possible to apply a vibrational signal via the probe 10 using vibrational transducers on the scan head 15, or via the probe tip 11 by periodic power intensity variations in the laser beam 23 provided by a laser unit 20. Such alternative manners of applying a vibrational signal 4 to the sample 7 have been described in literature and are not further discussed here. The present invention may be applied to measurements taken from SPM systems, such as the SPM system 1 performing surface topography measurements of topography 9 or performing subsurface measurements of structures below the top surface 8. The subsurface structure may for example be a preceding layer of a semiconductor element during manufacturing thereof, e.g. in order to detect whether the overlay of subsequent layers is sufficiently accurate to yield a fully functional semiconductor device, or whether the critical dimensions are as specified having the correct tolerances. In a method in accordance with the present invention, in some specific embodiments thereof, subsurface measurement is performed in order to further characterize sidewall structures that extend inward into a surface featured on surface 8. For example, subsurface measurements may enable to obtain information on a shape or depth of a sidewall recess. In the system 1 of figure 1, the vibrational acoustic input signal applied via the substrate carrier 3 to the substrate 7 is picked up at the surface 8 of the sample 7 via the probe tip 11. Due to the vibrations, the Z-level of the probe tip 11 is periodically displaced at the frequency applied via the acoustic signal 4. Various of these acoustic measurements techniques are known in the art, amongst which for example the heterodyne methods that use a very high frequency gigahertz acoustic signal including two frequencies in the gigahertz range. The different frequency between two applied frequencies in the gigahertz range is relatively small, typically in the megahertz range. By using the principal of heterodyne mixing of signals, a low frequency signal at the difference frequency (in the megahertz range) can be picked up by the probe tip 11, and can be found by analysis of the output signal obtained via optical beam deflector unit 21. Again, this measurement data needs to be further analyzed in order to obtain therefrom the three dimensional topography data. For example amplitude and phase data may be obtained, and similar to the above, the identification of measurement artifacts and noise reduction is to be performed. Figures 2A to 2C show the results of a conventional image processing method for identifying image features. In figure 2A, a section 32 of a substrate surface is illustrated, including a plurality of circular shaped image features 31. The image features 31 in figure 2A in the substrate or sample 7 relate to shallow holes in the surface 8 thereof. An image 30 thereof, obtained using a scanning probe microscope 1,is provided in figure 2B which vaguely shows these image features 31’. A conventionalimage processing algorithm based on global threshold segmentation is applied in orderto find the contours 33 of the image features 31’. The results thereof are illustrated infigure 2C. As follows evidently from figure 2C, the conventional image processing algorithm is not well able to recognize the contours 33 accurately. Comparing the contours 33 in figure 2C to the original shapes of the image features 31 that are schematically shown in figure 2A, makes clear that the algorithm falls short in this task. Figure 3A and 3B shows another example, wherein an image 35 of a scanning probe microscopy system 1 has been processed using a conventional image processing algorithm. In figure 3A, the image 35 vaguely shows the image features 36. However, in the upper right corner of image 35, a contamination 37 (in the present case, a copper particle) can be seen. The image 35 is processed with a conventional image processing algorithm based on template-matching segmentation. This algorithm uses information from the design template of the sample 7 to identify the image features 36 correctly. However, due to the high contrast of the contamination 37 in the image 35, the location of the contours 39 is not correctly identified, as follows from figure 3B. Likewise, as follows from figures 4A and 4B, which show another image 45 (figure 4A) which is processed using a conventional image processing algorithm based on template-matching segmentation (results in figure 4B), also scratches 48 in the image 45 cause the processing algorithm to fall short in correctly locating the contours 49 of the image features 46 in image 45. To overcome the above disadvantages, the invention in accordance with some embodiments thereof applies a U-net type machine learning data processing model 40 in order to perform image segmentation. This enables to provide output data based on the images provided as input data, wherein the output data includes representation data of one or more representations of the image features and localization data of locations associated with the one or more representations. This may be done efficiently while significantly reducing the chance on erroneous identification of image features, such as missed image features, incorrect feature shapes or incorrect locations of image features. Although the machine learning data processing model 40 may be a U-net type neural network, alternatively it may also be of a different type that is suitable to obtain the image features as well as their locations from an image. For example, alternatively the machine learning data processing model may be a vision transformer model, for example a segmenter transformer (SETR) or a pyramid vision transformer (PVT) or a SegFormer. Also hybrid models are potentially possible, for example architectures combining convolutional and transformer blocks such as TransUNet, cost aggregation transformers (CATs) or U-net transformer (UNETR). FIG. 5 schematically shows an example of a U-net type machine learning data processing model 40 (a specific type of neural network processor) that is suitable for localized data labeling / classification. Hence, it enables not only to label / classify image features (such as features 31, 36 and 46) from an image (e.g. such as images 30, 35 or 45), but also allows to determine the positions of these image features such as to provide localization data. The exemplary U-net type machine learning data processing model 40 comprises a contracting path 41 having stages 411, 412, 413, 414 an expansive path 42 having stages 421, 422, 423, 424 and a bridging stage 430. The contracting path 41 is typically configured as a conventional convolutional network with a sequence of layers wherein the original input image provided at a first spatial resolution is stepwise converted to a feature vector. A stage 411, 412, 413 or 414 in the contracting path may for example comprise convolutional layers and a rectified linear unit (ReLU). An output of each stage 411, 412, 413 and 414 may be provided to a corresponding stage 421, 422, 423 and 424 in the expanding path 42 as shown by the horizontal arrows in FIG. 5. Also a down sampling module, e.g. a 2x2 max pooling operation with stride 2, is provided therein to provide for a downsampled mapping of that output to the next stage in the contracting path 41. This is indicated by the arrows going from each stage 411, 412, 413, and 414 to each lower stage 412, 413, 414 and 430 respectively, in the contracting path 41. At each downsampling step the number of feature channels may be doubled. The expansive path 42 performs a backwards conversion towards a representation at a higher resolution in the n-dimensional space, so as to providelocalized labeling / classification information for the image x. Each stage 421, 422, 423and 424 in the expansive path 42 may comprise an upsampling of the feature mapfollowed by a convolution (“up-convolution”) that halves the number of feature channels, a concatenation with the correspondingly cropped feature map from thecontracting path 41, and further, convolutional layers followed by a ReLU. Theupsampling together with the concatenation, will provide for the localization data tobe preserved in the feature mask at the output The bridging stage 430 may beprovided to map each feature vector obtained from the final stage 414 in thecontracting path 41 to the desired number of classes for labeling The above type ofmachine learning data processing model, also referred to as a U net type machinelearning data processing model, is a known type of machine learning model and is for example described in Ronneberger, O., Fischer, P., Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. In: Navab, N., Hornegger, J., Wells, W., Frangi, A. (eds) Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015. MICCAI 2015. Lecture Notes in ComputerScience(), vol 9351. Springer, Cham. https: / / doi.org / 10.1007 / 978-3-319-24574-4_28. Although it is certainly possible, in accordance with some implementations of the concept described here, to train the machine learning data processing model 40using only real images 70 and real feature masks 71; to do this requires a vast data setwhich ideally includes sets of images including different varieties of image features, which differ in shape, size, orientation, etcetera. This eventually provides a well trained neural network providing good results and which is robust in the sense that it is able to recognize a variety of different image features. However, although this is certainly a possibility, it may not be practical in many circumstances in reality,because such a large data set is often not available or the desired images areproprietary data. To overcome this, in accordance with some embodiments, the present concept further is directed at generating a synthetic data set for training purposes, and in accordance with yet further special embodiments, applies transfer learning toquickly adapt the machine learning data processing model 40 to different situations.Figure 6 schematically illustrates a method 50 of data generation forgenerating a synthetic training data set 62. The steps 52 and 58 illustrated in figure 6 may be performed using an image generator (not illustrated). The image generator may be part of processing device 26 in figure 1, or may be a separate entity of the system 1 or in communicative connection therewith via communication network 29. The method, for training of the machine learning data processing model 40, comprises a first step of generating 52, by the image generator, a plurality synthetic images 54 including synthetic image features. Step 52 simultaneously or in cooperation with generating the synthetic images 54 also includes the generating, for each individual synthetic image of the plurality of synthetic images 54, of an associated binary segmentation mask 55. This may be performed by the image generator as well, for example the generator may either start with obtaining or generating a binary segmentation mask 55 and from this create a corresponding synthetic image 54, or apply the other order of sequence. Step 52 thereby yields to provide a plurality of pairs of synthetic images 54 with associated binary segmentation masks 55. Next, in step 58, each synthetic image 54 of the plurality of pairs of synthetic images 54 with associated binary segmentation masks 55 is augmented by including therein typical artificial image disturbances (e.g. such as noise, image artefacts and substrate contaminations) in order to obtain augmented synthetic images 60. This provides a plurality of pairs of augmented synthetic images 60 with associated binary segmentation masks 55 that can be used as synthetic training data set 62 for training of the machine learning data processing model 40. As may be appreciated, the number of augmented synthetic images 60 that can be generated in this manner is unlimited, and the amount of variation in the image features can be freely determined up front. Also the challenge of potentially proprietary data is circumvented in this manner. Figure 7 schematically illustrates a manner of training the machine learning data processing model 40. Whether the training is performed using the synthetic training data set 62 or another training data set is not relevant for how the training is performed. As explained the use of real images 70 with real feature masks 71 is possible for training, but advantageously a synthetic training data set 62 is applied. In another, further advantageous implementation, the machine learning data processing model is first pre-trained using a synthetic training data set 62 and thereafter, using the principles of transfer learning, is fine-tuned by training it with a limited set of real images 70 and real feature masks 71. In any event, in accordance with an embodiment, the training may be performed prior to the real implementation of the present invention in order to put it into use. The training for example comprises the steps of providing training data, such as the augmented synthetic images 60 from the synthetic training data set 62 to an input of the machine learning data processing model 40. The augmented synthetic images 60 are taken from the plurality of pairs in the synthetic training data set 62. The associated binary segmentation masks 55 are to be used as ground truth data during training. At an output of the machine learning data processing model 40, candidate feature masks 64 are generated based on the augmented synthetic images 60 in the training data. These candidate feature masks 64 are received by the processing device 26. The processing device 26 further receives the ground truth data comprising the binary segmentation masks 55 associated with each of the augmented synthetic images 60. The candidate feature masks 64 and the binary segmentation masks 55 are both provided as input to a training algorithm or program, which modifies the weighting values of nodes in the machine learning data processing model 40 based on a comparison of the candidate feature masks 64 with the binary segmentation masks 55. The machine learning data processing model 40 is thereby trained based on the training data in order to enable the machine learning data processing model 40 to generate, based on said input data (e.g. images 60), the correct output data in the form of candidate feature masks 64 that best matches the binary segmentation masks 55 used as ground truth data. The feature masks 64 obtained in this manner include the image features sought for, correctly localized at where they are really to be found in the figures. Therefore, from the feature masks 64 the representation data of image features and the localization data thereof may be readily obtained. Transfer learning uses the knowledge gained in solving one problem, say problem S, in subsequently solving a separate problem, e.g. problem T. It distinguishes from training strategies such as multi-task learning by assuming that Problem T is addressed separately after Problem S. In a formal definition, transfer learning uses the concepts of domain and task, where a domain D is defined by a feature space X and a probability distribution P(X) defined over X, and a task T isdefined by a label space Y and a prediction function f(x) = P(y|x) for x X and y Y.Considering a source domain and task (DS,TS) and a target domain and task (DT,TT), where either DS5(T and / or TS5,T. Transfer learning aims at learning fS, and subsequently learning fT by utilizing the knowledge gained in learning fS. The implementation of transfer learning in the present application is based on this principle, and is schematically illustrated in figure 8. Figure 8 shows step 65, which relates to the creation of augmented synthetic images 60 (including features of interest) using the synthetic data generation method 50 described above and illustrated in figure 6. The method 50 yields a synthetic training data set 62, including a large number of pairs of augmented synthetic images 60 with associated binary segmentation masks 55. In step 67, the the machine learning data processing model 40 is first pre-trained based on the synthetic training data set 62, for example applying the training method explained in relation to figure 7. The candidate feature masks 64 are thereby compared with the binary segmentation masks 55 associated with the augmented synthetic images 60, and based on this the weights of the nodes of machine learning data processing model 40 may be modified. For example, to optimally adapt the weights of the machine learning data processing model 40 based on the comparison between the candidate feature masks 64 and the binary segmentation masks 55, a gradient descent method may be applied. To prevent the system from reaching a local minimum instead of an optimum, the skilled person will understand that alternative methods may be applied, such as a variable fractional order gradient descent method, a stochastic gradient descent, a nonlinear conjugate gradient, a limited-memory Broyden–Fletcher–Goldfarb–Shannon algorithm, or a Levenberg-Marquardt algorithm. Many alternatives are available. In step 69, the pre-training is followed by a fine-tuning or post-training. Here, the principle of transfer learning applies, which results in the fine-tuning step to require only a small number of real images 70 with corresponding feature masks 71. For example, in practice, it would be possible to perform the pre-training of step 67 on the premises of a system supplier, and thereafter perform the fine-tuning of step 69 using real images 70 and feature masks 71 of a specific situation wherein the system will be used. In that case, the latter step 69 may be performed using proprietary data that is already available for the specific application at hand. A comparison of the performance of the machine learning data processing model 40 in accordance with the present concept, is illustrated in figures 9A to 9D. The left hand column in figure 9 relates to the application of various conventional algorithms in order to recognize and localize image features 31, 36 and 46 (see alsofigures 2-4) in images 30’, 35 and 45. The last row, figure 9D, relates to anotherexample thereof, showing image 75. All the images shown relate to identifying the locations of shallow structures or holes on a surface of a semiconductor wafer, in a manufacturing process of semiconductor structures. Although this is the primary field of application described here, more in general the method may be applied likewise to other comparable situations. The results obtained with applying the method and machine learning model 40 of the present invention are shown in the right-hand column. Image 30’’ infigure 9A shows the contours 33’ obtained which follow closely the shape of the imagefeatures 31 in figure 2A. The image feature 31, from the resulting image 30’’ in figure9A can be accurately recognized, verified and localized. Also in figure 9B, image 35’, aswell as figure 9C, image 45’, it can be seen that the contours 39’and 49’respectively,closely follow the image features that are sought, and can be used for localization andverification thereof. In image 75’of figure 9D, all of the contours in the figure areexactly mapped onto the image features sought for. In reality, the quality of course depends also on the training of machine learning data processing model 40, but it has been found that the excellent results shown in figure 9 can be obtained with a fair amount of pre-training and only a small amount of real images used in fine-tuning step 69. The present invention has been described in terms of some specific embodiments thereof. It will be appreciated that the embodiments shown in the drawings and described herein are intended for illustrated purposes only and are not by any manner or means intended to be restrictive on the invention. It is believed that the operation and construction of the present invention will be apparent from the foregoing description and drawings appended thereto. It will be clear to the skilled person that the invention is not limited to any embodiment herein described and that modifications are possible which should be considered within the scope of the appended claims. Also kinematic inversions are considered inherently disclosed and to be within the scope of the invention. Moreover, any of the components and elements of the various embodiments disclosed may be combined or may be incorporated in other embodiments where considered necessary, desired or preferred, without departing from the scope of the invention as defined in the claims. In the claims, any reference signs shall not be construed as limiting theclaim. The term 'comprising' and ‘including’ when used in this description or theappended claims should not be construed in an exclusive or exhaustive sense butrather in an inclusive sense. Thus the expression ‘comprising’ as used herein does notexclude the presence of other elements or steps in addition to those listed in any claim. Expressions such as "consisting of", when used in this description or the appended claims, should be construed not as an exhaustive enumeration but rather in aninclusive sense of "at least consisting of". Furthermore, the words ‘a’ and ‘an’ shall notbe construed as limited to ‘only one’, but instead are used to mean ‘at least one’, and donot exclude a plurality. Features that are not specifically or explicitly described or claimed may be additionally included in the structure of the invention within its scope. Any of the claimed or disclosed devices or portions thereof may be combined together or separated into further portions unless specifically stated otherwise, withoutdeparting from the claimed invention. Expressions such as: "means for ...” should beread as: "component configured for ..." or "member constructed to ..." and should be construed to include equivalents for the structures disclosed. The use of expressions like: "critical", "preferred", "especially preferred" etc. is not intended to limit the invention. Additions, deletions, and modifications within the purview of the skilled person may generally be made without departing from the spirit and scope of the invention, as is determined by the claims. The invention may be practiced otherwise then as specifically described herein, and is only limited by the appended claims.

Claims

Claims1. Method of performing feature detection of image features in an imageobtained from a scanning probe microscopy system, wherein the image has been obtained by scanning a surface of a sample using a probe including a probe tip, such that the image features correspond to structural features on or below the surface of the sample, wherein the method comprising the steps of: receiving, by a processor from a data source associated with the scanning probe microscopy system, image data of the image, wherein the image includes the image features and image disturbances; providing, by the processor, the image data as input data to a machine learning data processing model, wherein the machine learning data processing model has been trained to perform image segmentation such as to provide, based on the input data, output data including representation data of one or more representations of the image features and localization data of locations associated with the one or more representations; obtaining, by the processor from the output data of the machine learning data processing model, the representation data which is indicative of the one or more representations of the image features and the localization data of locations associated with the one or more representations; and providing, by the processor via an output device, the representation data and the localization data, for establishing said feature detection.

2. Method according to claim 1, wherein the output data of the machine learning data processing model comprises feature mask data of a feature mask image including the representation data and the localization data, and wherein the step of obtaining the representation data and the localization data comprises the steps of obtaining, by the processor, the feature mask data and obtaining the representation data and the localization data from the feature mask data.

3. Method according to claim 2, wherein the feature mask image is void of said image disturbances.

4. Method according to any one or more of the preceding claims, wherein the machine learning data processing model is a U-net type machine learning data processing model.

5. Method according to claim 4, wherein the U-net type machine learning data processing model includes a contracting path of convolution steps for capturing image context from the image data for enabling recognition of the image features for providing the representation data, and further includes an expanding path of de- convolution steps for enabling localization of the image features for providing the localization data.

6. Method according to any one or more of the preceding claims, wherein for training of the machine learning data processing model, the method further comprises the steps of: generating, by an image generator, a plurality synthetic images including synthetic image features; generating, by the image generator, for each synthetic image of the plurality of synthetic images, an associated binary segmentation mask, such as to provide a plurality of pairs of synthetic images with associated binary segmentation masks; augmenting, by the image generator, each synthetic image of the plurality of pairs with one or more artificial image disturbances, such as to provide the plurality of pairs of augmented synthetic images with associated binary segmentation masks as synthetic training data set for training of the machine learning data processing model.

7. Method according to claim 6, further comprising storing the synthetic training data set in the data source.

8. Method according to claim 6 or 7, wherein the method further comprises a training of the machine learning data processing model, wherein the training is performed prior to the step of receiving image data of the image from the data source associated with the scanning probe microscopy system, and wherein the training comprises the steps of:providing training data from the synthetic training data set to an input of the machine learning data processing model, wherein the training data comprises augmented synthetic images from the plurality of pairs in the synthetic training data set; provide, at an output of the machine learning data processing model, candidate feature masks generated by the machine learning data processing model based on the augmented synthetic images in the training data; receiving, by the processor, ground truth data, wherein the ground truth data comprises the binary segmentation masks associated with each of the synthetic images in the training data; and training the machine learning data processing model based on the training data by comparing the candidate feature masks to the binary segmentation masks in the ground truth data, such as to enable the machine learning data processing model to generate, based on said input data, said output data including the representation data and the localization data.

9. Method according to claim 8, wherein the step of training comprises modifying one or more weighting values of one or more nodes of the machine learning data processing model based on the step of comparing the candidate feature masks to the binary segmentation masks.

10. Method accord to any one or more of the preceding claims, wherein the method comprises training of the machine learning data processing model, and wherein the training comprises the steps of: providing training data to the input of the machine learning data processing model, wherein the training data comprises real images obtained from the data source associated with the scanning probe microscopy system, the real images comprising real image features and real image disturbances; provide, at the output of the machine learning data processing model, candidate feature masks generated by the machine learning data processing model based on the real images in the training data; receiving, by the processor, ground truth data, wherein the ground truth data comprises the real feature masks associated with each of the real images in the training data; andtraining the machine learning data processing model based on the training data by comparing the candidate feature masks to the real feature masks in the ground truth data, such as to enable the machine learning data processing model to generate, based on said input data, said output data including the representation data and the localization data.

11. Method according to claim 10, wherein the step of training comprises modifying one or more weighting values of one or more nodes of the machine learning data processing model based on the step of comparing the candidate feature masks to the real feature masks.

12. Method according to claims 8 or 9 in combination with any of claims 10 or 11, wherein the training in accordance with any of claims 8 or 9 is performed as pre- training of the machine learning data processing model, and wherein the training in accordance with any of claims 10 or 11 is performed as fine-tuning of the machine learning data processing model.

13. Method according to any one or more of the preceding claims, wherein the structural features to which the image features correspond, relate to at least one of a group comprising: on-surface features, such as holes, trenches, extensions, dimples, ridges or walls; sub-surface features, such as structural features, density interfaces or cavities underneath the surface of the substrate.

14. Scanning probe microscopy system comprising a probe including a probe tip, wherein the scanning probe microscopy system includes or cooperates with a substrate carrier for supporting a substrate, the scanning probe microscope being configured for moving the probe relative to a surface of the substrate in one or more directions, for establishing contact between the probe tip and the surface for mapping one or more surface structures on the surface, and for generating image data for providing an image based on said mapping, wherein the scanning probe microscope further comprises a processor communicatively connected to a data source, wherein the processor is configured to performing feature detection of image features in the image of the scanning probemicroscopy system, and wherein for providing the feature detection the processor is configured for: receiving, by the processor from the data source, the image data of the image including the image features and image disturbances; providing, by the processor, the image data as input data to a machine learning data processing model, wherein the machine learning data processing model has been trained to perform image segmentation such as to provide, based on the input data, output data including representation data of one or more representations of the image features and localization data of locations associated with the one or more representations, and wherein the machine learning data processing model is a U-net type machine learning data processing model; obtaining, by the processor from the output data of the machine learning data processing model, the representation data which is indicative of the one or more representations of the image features and the localization data of locations associated with the one or more representations; and providing, by the processor via an output device, the representation data and the localization data, for establishing said feature detection.

15. Scanning probe microscopy system according to claim 14, wherein the output data of the machine learning data processing model comprises feature mask data of a feature mask image including the representation data and the localization data, and wherein the processor is further configured for performing the step of obtaining the representation data and the localization data by: obtaining, by the processor, the feature mask data and obtaining the representation data and the localization data from the feature mask data.

16. Scanning probe microscopy system according to claim 14 or 15, wherein the processor further comprises or is operationally connected to an image generator, and wherein for training of the machine learning data processing model, the image generator is configured for: generating a plurality synthetic images including synthetic image features;generating for each synthetic image of the plurality of synthetic images, an associated binary segmentation mask, such as to provide a plurality of pairs of synthetic images with associated binary segmentation masks; and augmenting each synthetic image of the plurality of pairs with one or more artificial image disturbances, such as to provide the plurality of pairs of augmented synthetic images with associated binary segmentation masks as synthetic training data set for training of the machine learning data processing model.

17. Machine learning data processing model for use in a method according to any one or more of the claims 1-13, wherein the machine learning data processing model has been trained to perform image segmentation such as to provide, based on the input data, output data including representation data of one or more representations of the image features and localization data of locations associated with the one or more representations.

18. Machine learning data processing model according to claim 17, wherein the machine learning data processing model is a U-net type machine learning data processing model.

Citation Information

Patent Citations

  • Generating training data usable for examination of a semiconductor specimen

    US20220383488A1