Apparatus, System, and Method for Analysis and Characterization of Surface Topography

The method addresses the inadequacies of existing surface topography measurement by characterizing surface topography through scale-dependent parameters, enhancing accuracy and predictive capability by accounting for multi-scale variations.

JP2025524428APending Publication Date: 2025-07-30UNIV OF PITTSBURGH OF THE COMMONWEALTH SYST OF HIGHER EDUCATION +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024574587
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-24
Filing Date
2023-06-24
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Existing methods for measuring and characterizing surface topography are inadequate as they fail to capture the multi-scale nature of real-world surfaces, leading to insufficient description and prediction of functional characteristics.

Method used

A method and system for characterizing surface topography by determining scale-dependent parameters through statistical characterization of first and higher-order derivatives of surface height at multiple distance scales, using a scaling factor to account for varying measurement resolutions, and combining data from different measurement methods.

Benefits of technology

Provides a comprehensive and accurate description of surface topography by capturing variations across multiple scales, enabling better prediction of functional characteristics and reducing measurement artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524428000001_ABST
    Figure 2025524428000001_ABST
Patent Text Reader

Abstract

A method for characterizing the surface topology of a target surface includes determining a feature vector from one or more measurement values of the target surface, wherein a plurality of features in the feature vector represent a statistical characterization of the distribution of one or more derivatives of the surface height or h, or are determined from the statistical characterization, and the one or more derivatives are selected from the group consisting of zero-order and higher-order derivatives determined from at least one of the one or more measurement values of the target surface at each of a plurality of distance scales, and for the at least one of the one or more measurement values, the one or more derivatives of the surface height are determined using a scaling coefficient η that is multiplied by the smallest possible distance scale provided by the at least one of the one or more measurement values at the plurality of distance scales in real space, and determining at least one characteristic of the target surface via an algorithm executable via a processor system and stored in a memory system and based on the feature vector, and providing an output indicative of the at least one characteristic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Government Interest This invention was made with government support under grant numbers 1727378 and 1844739 awarded by the National Science Foundation. The government has certain rights in this invention.

[0002] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 355,281, filed Jun. 24, 2022, the disclosure of which is incorporated herein by reference.

Background Art

[0003] The following information is provided to assist the reader's understanding of the technologies disclosed below and the environments in which such technologies may be commonly used. The terms used herein are not intended to be limited to a particular narrow interpretation unless otherwise specified herein. The references described herein may facilitate the understanding of those technologies or backgrounds. The disclosure of all references cited herein is incorporated by reference.

[0004] The properties of a surface are strongly influenced by the surface topography or surface roughness. Such properties include the frictional force and the adhesive force (i.e., how strongly two surfaces stick together) between two contacting bodies. These properties are important in any industry that constructs devices with moving or contacting parts, such as, for example, the automotive, aerospace, and manufacturing industries.

[0005] Properly characterizing surface topography or associating surface topography with functional characteristics is highly desirable during device design (e.g., during research and development) or for quality assurance / quality control (QA / QC). Generally, surface topography or surface roughness can be quantified, for example, by the deviation of the surface height from a smooth reference plane or, for example, by the deviation of the surface height from the mean plane of the surface. Currently, it is common to measure surface topography at a single size scale using, for example, a stylus-type surface profiler. This type of single measurement is applied to manufactured parts, for example, in quality assurance procedures.

[0006] There are many problems with existing methods for measuring and characterizing surface topography / surface roughness. Real-world surface topography cannot be adequately described by individual measurements that capture only a limited size scale range of the surface topography. Manufactured surfaces in reality exhibit topography variations and roughness over many size scales. Furthermore, functional characteristics are topography-dependent over many or all scales. Current approaches for measuring and analyzing topography are insufficient for describing and / or predicting characteristics. SUMMARY OF THE INVENTION

[0007] In one aspect, a method for characterizing surface topography includes determining scale-dependent parameters. Each of the scale-dependent parameters represents a statistical characterization of the distribution of at least one of the first or higher-order derivatives of the surface height or h, and is determined from one or more measurements of the target surface at each of a plurality of distance scales. For at least one of the one or more derivatives, the first and higher-order derivatives of the surface height are defined via a scaling factor η, which is multiplied by the smallest possible distance scale or resolution provided by the at least one of the one or more measurements, at the plurality of distance scales in real space. At least one characteristic of the target surface may be determined from the scale-dependent parameters.

[0008] The method may further include statistically characterizing each distribution of a plurality of derivatives of the surface height of different orders at a plurality of distance scales when characterizing the surface topography. At least one of the one or more derivatives of the surface height may be, for example, a third or higher-order derivative.

[0009] At least one of the first or higher-order derivatives of the distribution may be determined via a numerical method over the plurality of distance scales, and then statistically characterized to determine a scale-dependent parameter of the distribution. The numerical method may be, for example, a finite difference method, a finite element method, a Fourier interpolation method, or other interpolation methods using a compact or spectral basis set. The scale-dependent parameter may be determined from surface topography parameters that are not determined from the distribution of the first or higher-order derivatives of the surface height determined via a numerical method when the statistical characterization is determined for the second cumulant or second moment. In the case of such surface topography parameters, the scale-dependent parameter is determined to transform the surface topography parameters into the scale-dependent parameters by applying a determined mathematical relationship to the surface topography parameters. The surface topography parameters may be selected, for example, from the group of autocorrelation function characterization, variable bandwidth method characterization, or power spectral density characterization.

[0010] At least one of the first or higher-order derivatives may be determined over a plurality of distance scales for a line of the one or more measurements of the surface, or for an area of the one or more measurements of the surface. The distribution of at least one of the first or higher-order derivatives may be determined, for example, over the plurality of distance scales for a line of the one or more measurements of the surface and averaged over the plurality of lines of the one or more measurements of the surface. In some embodiments, the derivative for a line of the one or more measurements is, for a point x on the line k given by the following formula:

Number

Number

[0011] The first or higher-order derivative function may be determined with respect to the area of the one or more measured values of the surface, and the first or higher-order derivative function is given by the following formula:

Number

Number

[0012] The evaluation of the statistical characteristics of the distribution may be determined, for example, from the second or higher-order cumulants of the distribution, or the second or higher-order moments of the distribution. In some embodiments, the evaluation of the statistical characteristics of the distribution is selected from the group consisting of variance, skewness, and kurtosis. In some embodiments, the evaluation of the statistical characteristics of the distribution may be determined from the third or higher-order cumulants of the distribution, or the third or higher-order moments of the distribution.

[0013] The distribution is given, for example, by the following formula:

Number

[0014] In some embodiments, the tip radius effect for the measurement method used in the one or more measurements is determined as a function of the minimum value of the variance in the second derivative at a particular distance scale l. For example, the critical scale l tip may be determined, and in order to minimize the tip radius effect, data at distance scales below the tip is excluded. In some embodiments, l tip is numerically estimated using the following equation tip where

Number

Number

Number

[0015] In some embodiments, the one or more measurements are used in determining the scale-dependent parameter. In some embodiments, such measurements are determined or performed by different measurement methods and / or have different possible minimum distance scales or resolutions. In some embodiments, the different measurement methods are selected from the group consisting of stylus surface profilometry, optical surface profilometry, cross-section or side-view microscopy, and reflectometry. When determining the scale-dependent parameter, data from beyond the measurements are combined across the plurality of distance scales.

[0016] In some embodiments, the method further includes determining a feature vector from one or more measurement values of the surface, wherein a plurality of features of the feature vector are determined from scale-dependent parameters and, based on the feature vector, determining at least one characteristic of the target surface.

[0017] In another aspect, a system for characterizing surface topography includes a processor system and a memory system communicatively coupled to the processor system. The memory system includes an algorithm for determining scale-dependent parameters, each of which represents a statistical characterization of at least one distribution of a first or higher-order derivative of surface height or h, and is determined from one or more measurement values of the target surface at each of a plurality of distance scales. For at least one of the one or more derivatives, the first and higher-order derivatives of surface height are determined using a scaling coefficient η that is multiplied by the smallest possible distance scale or resolution provided by the at least one of the one or more measurement values at one or more scaling factors η at the plurality of distance scales in real space.

[0018] In some embodiments, the algorithm statistically characterizes each of a plurality of derivatives of surface height of different orders at a plurality of distance scales. The statistical characterization of the distribution may be determined, for example, from the third or higher-order cumulants of the distribution, or from the third or higher-order moments of the distribution.

[0019] In some embodiments, the system further includes a measurement system communicatively coupled to the processor system for measuring surface height across the area of the surface.

[0020] In another aspect, a non-transitory computer-readable medium for characterizing surface topography is instructions stored on the non-transitory computer-readable medium that, when executed on a processor, include instructions to determine scale-dependent parameters, each of the scale-dependent parameters representing a statistical characterization of a distribution of at least one of the first or higher order derivatives of surface height or h, determined from one or more measurements of the target surface at each of a plurality of distance scales, and for at least one of the one or more derivatives, the first and higher order derivatives of surface height are defined via a scaling coefficient η at the plurality of distance scales in real space, where the scaling coefficient η is multiplied by the smallest possible distance scale or resolution provided by the at least one of the one or more measurements.

[0021] In another aspect, a method for characterizing the surface topology of a target surface includes determining a feature vector from one or more measurements of the target surface, where the plurality of features in the feature vector represent or are determined from a statistical characterization of a distribution of one or more derivatives of surface height or h, the one or more derivatives being selected from the group consisting of zero-th and higher order derivatives determined from at least one of the one or more measurements of the target surface at each of a plurality of distance scales, and for at least one of the one or more measurements, the one or more derivatives of surface height are determined using a scaling coefficient η that, at the plurality of distance scales in real space, is multiplied by the smallest possible distance scale provided by the at least one of the one or more measurements; determining at least one characteristic of the target surface via an algorithm executable via a processor system and stored within a memory system and based on the feature vector; and providing an output indicative of the at least one characteristic.

[0022] At least one of the one or more derivatives of the surface height h may be, for example, a third or higher order derivative. At least one of the one or more derivatives of the surface height h may be, for example, a fourth or higher order derivative.

[0023] The plurality of features in the feature vector are determined from an evaluation of statistical characteristics of the distribution of two or more derivatives of the surface height, and the two or more derivatives may have different orders. The one or more derivatives of the surface height may be selected from the group consisting of a zero-order derivative, a first-order derivative, a second-order derivative, a third-order derivative, and a derivative of order higher than third. In some embodiments, the one or more derivatives of the surface height are selected from the group consisting of a first or higher order derivative. In some embodiments, the one or more derivatives of the surface height are, for example, a third or higher order derivative. In some embodiments, the values of the plurality of features are normalized.

[0024] In some embodiments, the evaluation of the statistical characteristics of the distribution is determined from the second or higher order cumulants of the distribution, or the second or higher order moments of the distribution. In some embodiments, the evaluation of the statistical characteristics of the distribution is determined from the third or higher order cumulants of the distribution, or the third or higher order moments of the distribution. The evaluation of the statistical characteristics of the distribution may be selected from the group consisting of, for example, variance, skewness, and kurtosis.

[0025] The first or higher order derivative may be determined over a plurality of distance scales for a line of the one or more measurements of the surface, or an area of the one or more measurements of the surface. The distribution of the first or higher order derivative may be determined, for example, over the plurality of distance scales for a line of the one or more measurements of the surface and averaged over the plurality of lines of the one or more measurements of the surface. In some embodiments, the derivative for a line of the one or more measurements is at a point x on the line k for the following equation:

Equation

[0026] The first or higher-order derivative may be determined for the area of the one or more measurements of the surface, and the first or higher-order derivative is given by the following equation: [Number] may be provided by. Here, α and β are the orders of the derivatives in the x and y directions, respectively, [Number] represents the stencil.

[0027] The distribution, for example, is given by the following equation: [Number] where δ is the Dirac δ function and χ is the value of the derivative of order α. The δ function may be extended to individual bins, and the number of occurrences of a given derivative value is counted.

[0028] In some embodiments, the tip radius effect for the measurement method used in the one or more measurements is determined as a function of the minimum value of the variance in the second derivative at a particular distance scale l. For example, a critical scale l tip may be determined, and to minimize the tip radius effect, l tipData at distance scales below the tip is excluded. In some embodiments, l tip is numerically estimated using the following equation

Number

Number

Number

[0029] In some embodiments, two or more measurements are used in defining the statistical property evaluation, and each of the two or more measurements is created by a different measurement method and / or has a different possible minimum distance scale or resolution. The different measurement methods may be selected from the group consisting of, for example, stylus surface profilometry, scanning probe microscopy, optical surface profilometry, cross-section or side-view microscopy, and reflectometry. When determining the statistical property evaluation, data from the one or more measurements created via two or more measurement methods may be combined across the plurality of distance scales.

[0030] In some embodiments, the algorithm includes at least one machine learning model. The at least one machine learning model may be, for example, a classification model or a regression model. The classification model may include, for example, a support vector machine model, a Gaussian process classifier model, or a neural network. In some embodiments, the at least one machine learning model is trained using the features and labels of a training set of one or more measurements at each of a plurality of training surfaces. The at least one machine learning model is trained using the features and labels of a training set of one or more measurements for each of the plurality of training surfaces.

[0031] Before inputting into the at least one machine learning model, it may further include reducing the dimensionality of the feature vector. To reduce the dimensionality, for example, a principal component analysis algorithm or an autoencoder algorithm may be used. In some embodiments, the principal component analysis algorithm or the autoencoder algorithm is adapted to process data sets with missing values or different bandwidths.

[0032] In a further aspect, a system for characterizing the surface topology of a target surface includes a memory system, a processor system operably connected to the memory system, and a database system stored within the memory system. The system further includes an algorithm stored within the memory system and executable via the processor system. The algorithm determines a feature vector from one or more measurements of the target surface. A plurality of features in the feature vector represent, or are determined from, a statistical characterization of the distribution of one or more derivatives of surface height or h. The one or more derivatives are selected from the group consisting of zero-order and higher-order derivatives determined from at least one of the one or more measurements of the target surface at each of a plurality of distance scales. For at least one of the one or more measurements, the one or more derivatives of surface height are determined using a scaling coefficient η that is multiplied by the smallest possible distance scale provided by at least one of the one or more measurements at a plurality of distance scales in real space. The algorithm further determines at least one characteristic of the target surface based on the feature vector and provides an output indicative of the at least one characteristic.

[0033] In a further aspect, a non-transitory computer-readable medium for characterizing surface topography comprises instructions stored on the non-transitory computer-readable medium that, when executed on a processor, include instructions for determining a feature vector from one or more measurements of the target surface, wherein the plurality of features in the feature vector represent a statistical characterization of the distribution of one or more derivatives of surface height or h, or are determined from the statistical characterization, and wherein the one or more derivatives are selected from the group consisting of zero-order and higher-order derivatives determined from at least one of the one or more measurements of the target surface at each of a plurality of distance scales, and wherein for the at least one of the one or more measurements, the one or more derivatives of surface height are determined using a scaling coefficient η that is multiplied by the smallest possible distance scale provided by the at least one of the one or more measurements, the scaling coefficient η being a scaling coefficient η at one or more distance scales in real space. Based on the feature vector, at least one characteristic of the target surface is determined. The instructions may further provide an output indicative of the at least one characteristic when executed on a processor.

[0034] In yet another aspect, the system includes a memory system, a processor system operably connected to the memory system, and a database system stored within the memory system. The database system includes topography data associated with one or more measurements of each of a plurality of surfaces. The topography data includes a statistical characterization of the distribution of surface height or one or more derivatives of h for at least one of the one or more measurements, the one or more derivatives being selected from the group consisting of zero - order and higher - order derivatives, and being determined at a plurality of distance scales in real space using a scaling coefficient η, which is one or more scaling coefficients η at each of the plurality of distance scales and is multiplied by the smallest possible distance scale provided by at least one of the one or more measurements. The system further includes an algorithm stored within the memory system and executable via the processor system. The algorithm includes an algorithm that includes at least one machine - learning model trained using the training set of the topography data using the features and labels of the training set of the topography data.

[0035] The apparatus, system, and method of the present invention, together with its attributes and attendant advantages, will be best recognized and understood by considering the following detailed description, taken in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0036]

Figure 1

[0037]

Figure 2

[0038]

Figure 3

[0039]

Figure 4

Number

Number

[0040]

Figure 5A

[0041]

Figure 5B

[0042]

Figure 5C

[0043]

Figure 5D

[0044]

Figure 5E

[0045]

Figure 5F

[0046]

Figure 6A

[0047]

Figure 6B

[0048]

Figure 6C

[0049]

Figure 6D

[0050]

Figure 6E

[0051]

Figure 6F

[0052]

Figure 7A

[0053]

Figure 7B

[0054]

Figure 7C

[0055]

Figure 7D

[0056]

Figure 8A

[0057]

Figure 8B

[0058]

Figure 8C

[0059]

Figure 8D

[0060]

Figure 9

[0061]

Figure 10

[0062]

Figure 11

[0063]

Figure 12

[0064]

Figure 13

[0065]

Figure 14

[0066]

Figure 15

[0067]

Figure 16

[0068]

Figure 17

[0069]

Figure 18

[0070]

Figure 19

[0071]

Figure 20

[0072]

Figure 21

[0073]

Figure 22

[0074]

Figure 23

[0075]

Figure 24

[0076]

Figure 25

[0077]

Figure 26

[0078]

Figure 27

[0079]

Figure 28

[0080]

Figure 29

DETAILED DESCRIPTION OF THE INVENTION

[0081] In some embodiments, the apparatus, system, method, and composition of the present embodiment provide the analysis and characterization of surface topography or surface roughness.

[0082] It will be readily understood that the components of the embodiments, as generally described and illustrated in the drawings of this specification, may be arranged and designed in a wide variety of different configurations in addition to the exemplary embodiments described. Accordingly, the more detailed description of the exemplary embodiments shown below is not intended to limit the scope of the embodiments to what is recited in the claims, as shown in the drawings, but is merely representative of the exemplary embodiments.

[0083] References to "one embodiment" or "an embodiment" (or the like) throughout this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Thus, even if the phrases "in one embodiment" or "in an embodiment" or equivalent phrases appear in various places throughout this specification, they are not necessarily all referring to the same embodiment.

[0084] Furthermore, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to enable a thorough understanding of the embodiments. However, one of ordinary skill in the art will recognize that various embodiments may be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail so as not to obscure the understanding.

[0085] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include references to the plural unless the context clearly dictates otherwise. Thus, for example, a reference to "an algorithm" includes references to a plurality of such algorithms and equivalents thereof known to those skilled in the art, and a reference to "the algorithm" is one or more such algorithms and equivalents thereof known to those skilled in the art. The detailed description of ranges of values herein is merely intended to serve as a shorthand way of referring individually to the individual values within the range. Unless otherwise indicated herein, the individual values and intermediate ranges are incorporated herein as if individually set forth herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly indicated to the contrary in the text.

[0086] As used herein, the terms "electronic circuitry", "circuitry", or "circuit" include, but are not limited to, hardware, firmware, software, or combinations thereof for performing one or more functions or operations. For example, depending on the desired function or requirement, the circuit may include a software-controlled microprocessor, discrete logic such as an application specific integrated circuit (ASIC), or other programmed logic device. The circuit may also be fully realized as software. As used herein, "circuit" is synonymous with "logic". As used herein, the term "logic" includes, but is not limited to, hardware, firmware, software, or combinations thereof for performing one or more functions or operations, or for causing a function or operation from another component. For example, depending on the desired application or requirement, the logic may include a software-controlled microprocessor, discrete logic such as an application specific integrated circuit (ASIC), or other programmed logic device. The logic may also be fully realized as software.

[0087] As used herein, the term "processor" includes, but is not limited to, one or more of substantially any number of processor systems. A processor system may include one or more stand-alone processors, such as a microprocessor, a microcontroller, a central processing unit (CPU), a digital signal processor (DSP), etc., in any combination. A processor may be associated with various other circuits that support the operation of the processor, such as random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), a clock, a decoder, a memory controller, or an interrupt controller. These support circuits may be internal or external to the processor or its associated electronic package. The support circuits communicate operably with the processor. The support circuits need not necessarily be shown separately from the processor in a block diagram or other drawing.

[0088] As used herein, the term "software" includes, but is not limited to, one or more computer-readable instructions or executable instructions that cause a computer or other electronic device to function, operate, or perform operations in a desired manner. The instructions may be embodied in various forms, such as a program including routines, algorithms, modules, or code from a separate application or dynamically linked library. Software may also be implemented in various forms, such as a stand-alone program, a function call, a servlet, an applet, instructions stored in memory, a part of an operating system, or other types of executable instructions. One of ordinary skill in the art will understand that the form of the software depends, for example, on the requirements of the desired application, the environment in which the software is executed, or the desires of the designer / programmer.

[0089] In some embodiments, the systems, apparatuses, and methods of the present embodiment can be used to characterize surface topography by evaluating statistical characteristics of at least one distribution of first or higher order derivatives of surface height (h) determined from one or more scans of the surface at each of a plurality of distance scales, to define a scale-dependent roughness parameter or SDRP (variance), and a scale-dependent statistical parameter or SDSP, or simply a scale-dependent parameter (including variance, and higher order moments or cumulants such as skewness, kurtosis, and parameters determined from even higher order moments or cumulants, and generalizations). The SDSP or scale-dependent parameter herein is an evaluation of the statistical characteristics of gradients, curvatures, and third (or higher) derivatives, and may be referred to herein as a statistically characterized scale-dependent parameter or simply a scale-dependent parameter. In this regard, at least for one of the one or more scans, the first or higher order derivative of the surface height may be determined using a scaling factor η that is multiplied by the smallest possible distance scale provided by at least one of the one or more scans, where the scaling factor η is one or more scaling factors. In some embodiments, the method includes statistically characterizing each distribution of a plurality of derivatives of different orders at a plurality of distance scales when characterizing surface topography to determine scale-dependent parameters.

[0090] Generally, the physical properties of a surface cannot be fully understood / predicted by applying a physical model to a single measurement value. Instead, models should be applied to measurements over a variety of different size scales. Such measurements over a variety of different sizes or distance scales can be achieved, for example, using a combination of measurements (e.g., using a variety of different measurement methods). The distribution of first- or higher-order derivatives herein may be statistically characterized, for example, after being determined over a plurality of distance scales via a numerical method (e.g., the finite difference method or other methods). Statistical characterization may alternatively be determined or estimated from surface topography / roughness parameters other than scale-dependent parameters when determined from the second cumulant or second moment, and such scale-dependent parameters may be mathematically related to the surface topography / roughness parameters. Surface roughness / topography parameters other than scale-dependent parameters may be selected, for example, from the group of autocorrelation function characteristics, variable bandwidth characteristics, or power spectral density characteristics as described below. The variable bandwidth method (VBM) or the scaled windowed variance method includes a class of methods that differ in how the data is detrended. Such methods have various names such as the bridge method, roughness around the mean height (MHR) (which may also be referred to as VBM), detrended fluctuation analysis (DFA), roughness around the rms line (SLR), etc.

[0091] The scale-dependent parameter analysis herein provides a generalization of commonly used topographic metrics. The scale-dependent parameter analysis herein may be used to combine such topographic metrics into the scale-dependent parameter analysis herein and may serve to reconcile heterogeneous topographic descriptors. However, this scale-dependent parameter analysis (based on real-space measurements) offers many advantages over such other methods, particularly in that it is computationally easy, has an intuitive interpretation, can detect artifacts, can easily combine measurements from multiple measurement methods over a wide range of scales, and can determine scale-dependent parameters whose statistical characterization of the distribution is determined from third- or higher-order cumulants or third- or higher-order moments. According to the apparatus, system, and method herein, multiple measurements obtained at various different length scales and / or by various different measurement techniques (e.g., stylus surface profilometry, cross-sectional microscopy, optical surface profilometry) can be easily combined into a single statistical description of the topography of the sample. Further, as described above and below, the scale-dependent parameter analysis facilitates and / or enables the analysis of higher-order cumulants or moments that contain information about deviations from Gaussianity.

[0092] Surface roughness is mainly characterized in terms of scalar parameters, and in particular, the root mean square (rms) height and gradient are commonly cited. These are the rms deviations from the mean height and mean gradient, regardless of the presence or absence of an additional bandwidth filter. Some variations of these quantities are calculated by all surface topography instruments and are often reported in publications to describe surface topography. These quantities are useful for describing the amplitude of the spatial variation of the height and gradient of the entire measured topography. However, the core problem with these roughness parameters is that all parameters clearly depend on the scale of measurement. For example, the rms height depends on the lateral size (maximum scale) of the measurement, and the rms gradient depends on the resolution (minimum scale) of the measurement. Although some standardized formulas for obtaining these values, such as Rq in IS 4287, include high-frequency and low-frequency filtering, such values are still strongly scale-dependent, and the scale associated with this is the size of the filter, not the size of the measurement. See International Organization for Standardization, Geometrical product specifications (GPS)-Surface texture: Profile method - Terms, definitions and surface texture parameters, ISO Standard No. 4287, 1997.

[0093] The scale dependence of these values is typically a characteristic of the multi-scale nature of surface topography. As a simple example, there is the classical observation regarding the length of a coastline by Benoit Mandelbro, where the length L coast of the coastline has been shown to depend on the length l of the measuring tool / ruler used to measure it. The smaller the ruler, the more fine details can be picked up, and thus the longer the coastline becomes. In a (self-affine) fractal, L coastThe functional relationship between \(l\) and \(L\) follows a power law, and its exponent characterizes the fractal dimension of the coastline. In the case of surface topography measurements, \(L\) coast corresponds to the resolution of the scientific instrument (or filter) used for topography measurements, and the property corresponding to the coastline length is the true surface area \(S(l)\) of the topography. It has been demonstrated that \(S(l)\) (as well as the rms gradient and curvature) scales with the measurement resolution \(l\). See A. Gujrati, S.R. Khanal, L. Pastewka, T.D.B. Jacobs, Combining TEM, AFM, and profilometry for quantitative topography characterization across all scales, ACS Appl. Mater. Interf. 10 (2018) 29169; A. Gujrati, A. Sanner, S.R. Khanal, N. Moldovan, H. Zeng, L. Pastewka, T.D., B. Jacobs, Comprehensive topography characterization of polycrystalline diamond coatings, Surf. Topogr. Metrol. Prop. 9 (2021) 014003; S. Dalvi, A. Gujrati, S.R. Khanal, L. Pastewka, A. Dhinojwala, T.D.B. Jacobs, Linking energy loss in soft adhesion to surface roughness, Proc. Natl. Acad. Sci. USA 116 (2019) 25484. This scaling of the surface area is directly relevant, for example, to adhesion between soft surfaces. While many surfaces do not behave like ideal fractals, almost all surfaces exhibit some form of size dependence of the roughness parameters described above. In this regard, processes that form surfaces such as fracture, plasticity, erosion, etc. result in multi-scale fractal topography over a variety of length scales.

[0094] The present apparatus, system, and method provide a route for generalizing the above-described (and other) geometric properties of a measured topography to explicitly include the concept of a measurement scale. The individual roughness parameters are defined as a function of the scale l at which they are measured, resulting in a curve that specifies the value of the parameter as a function of l. However, l is not limited to the resolution of the measuring instrument or some fixed filter cutoff. In this analysis, the concept of the scale l is extended to denote any size at which scale-dependent parameters are calculated. In a given topographic scan, the concept of the scale l can range, for example, from the pixel size or resolution to the scan size. The resulting curves can be associated with common surface roughness characterization techniques, such as, for example, the height-difference autocorrelation function (ACF), the variable bandwidth method (VBM), and the power spectral density (PSD). The scale-dependent parameters herein are very useful, one reason being that they are easy to interpret. In this regard, while it is difficult to assign a geometric meaning to a given value of the PSD (the units can even be ambiguous), both slope and curvature have simple geometric interpretations. Since slope and curvature are also important considerations in modern contact theory between rough surfaces, the scale-dependent parameters herein are directly related to the functional properties of rough surfaces. As an example of the usefulness of scale-dependent parameters, how such parameters are used to estimate the tip radius artifact in contact-based measurements such as scanning probe microscopy and stylus profilometry is shown below.

[0095] A surface topography is generally described by the function h(x,y), where x and y are coordinates in the plane of the surface. This is sometimes referred to as the Monge representation of the surface and is an approximate representation excluding overhangs (concave surfaces). In actual measurements, a continuous function cannot be obtained, but rather height values on a set of discrete points x k and y l :

Equation

[0096] h kl is a random process, so the topography is often random and its characteristics need to be described in a statistical way. There has been much discussion about this random process model of surface roughness, but the most commonly used roughness parameters remain simple.

[0097] The concepts in this specification are illustrated using a representative one - dimensional case, namely a line scan or profile. In many practical situations, areal topography measurements can also be interpreted as a series of line scans. For example, in the case of atomic force microscopy (AFM), the topography map is a composite of a series of adjacent line scans. Due to temporal (instrumental) drift, these line scans may not align perfectly, so the "scan" direction is the preferred direction for statistical evaluation. It is implicitly assumed in the discussion here that all values are obtained by averaging such consecutive scans, but this averaging is not explicitly stated in the equations shown later. It is straightforward to extend the ideas presented here to a true two - dimensional topography map and will be briefly described.

[0098] The simplest statistical characteristic is the root mean square (rms) height.

Equation

Number

[0099] A common (but not exclusive) way to calculate the discrete derivative function on experimental data is to use a finite difference approximation. The finite difference approximation locally approximates the height h(x) as a polynomial (Taylor series expansion). Then, the first derivative is

Number

[0100] This equation is referred to as the first-order right difference scheme. The symbol D is used for the discrete derivative function, and the term "order" in this specification indicates the truncation order, that is, how fast the error decays with the lattice spacing Δx. In this scheme, this error decreases linearly as Δx gets smaller. Another interpretation is that the truncation order gives the highest order in the polynomial used to interpolate between the points x and x + Δx. The derivative of linear interpolation is constant between these points and is given by Equation (4).

[0101] As will be apparent to those skilled in the art, in this methodology, right, left, or central finite differences may be used. Further, other representations of the discrete derivative function, such as those obtained from linear or higher-order finite elements known in the mathematical field, or from Fourier interpolation with other compact or spectral basis sets, can be used to determine the derivative in the apparatus, system, and method of this specification. The representative discrete equations described herein are for the finite difference scheme.

[0102] For discrete derivatives obtained using Fourier interpolation, the Fourier series representation

number

number

number

[0103] The amplitudes of the higher order derivatives can also be quantified as follows:

number

number

number

[0104] In some embodiments, first or higher order derivatives are determined over multiple distance scales for one or more scan lines of a surface, or for one or more scan areas of a surface. For a point x on a line k the discrete derivative with respect to one or more scan lines for the point x may be expressed as a weighted sum over the collocation points x [Number] as given by. Here, α is the order of the derivative, Δx is the smallest possible scale, C k is the stencil of the derivative, and the derivative is measured at a distance scale l = αηΔx. As will be apparent to those skilled in the art, in practical applications, this sum does not go to infinity. For example, the stencils for α = 1, 2, and 3 in some embodiments herein are l (α) and all other C [Number] are 0. As will be apparent to those skilled in the art, higher order derivatives lead to wider stencils. l (α) All of the discrete derivatives in the previous section are defined at the smallest possible scale given by the sample interval Δx and have a full width of αΔx. It is straightforward to assign an explicit scale to these derivatives by evaluating equation (8) at the sample interval ηΔx (integer η) instead of x.

[0105] In this specification, the coefficient η is referred to as the scale coefficient. The corresponding derivative is measured at a distance scale l = αηΔx. [Number] As described above, first or higher order derivatives may alternatively be determined for one or more scan areas (i.e., two-dimensional) of a surface. First or higher order derivatives in two dimensions are, for example, given by the following equation:

[0106] ​

Mathematics

Mathematics

[0107] Figure 1 represents the above - mentioned concept. In the case of a simple right - difference scheme given by Equation (4), the scale - dependent first - order derivative is simply the gradient of two points at distance l. In panel (a) of Figure 1, a representative line scan showing the calculation of the slope h’(l) and curvature h”(l) from finite differences is presented. By calculating these finite differences at various different distances l, a scale can be imparted to this calculation. Here, it is shown for l = 40Δx and l = 80Δx, where Δx is the sample interval. Similarly, the curvature at a finite scale l is given by fitting a quadratic function passing through three points spaced at a distance l / 2 apart. Panel (b) represents the local gradient obtained at a distance scale l = 40Δx for the line scan shown in panel (a). Since this gradient can be calculated for overlapping intervals, it is defined for each sample point. Panel (c) represents the distribution of the local gradients obtained from the gradient profile shown in panel (b). The rms gradient of this length scale is the width of this distribution. For the second - order derivative given by Equation (6), a quadratic function passing through three points over the entire interval l is fitted, and the curvature of this function becomes the scale - dependent second - order derivative.

[0108] The scale - dependent roughness parameter (SDRP) is

Mathematics

Mathematics

Mathematics

[0109] The distance scale l is clearly defined only for the stencil with the minimum truncation order. In a typical case of finite differences, for the n-th derivative, it can be interpreted as fitting a polynomial of degree n to n + 1 data points (see Figure 1, panel (a)). The n-th derivative of this polynomial is a constant over the width of the stencil. This width needs to be equal to the distance scale l. Higher truncation orders can be interpreted as fitting a polynomial of degree m > n to m + 1 data points. The n-th derivative is not constant on the stencil, and it is not clear what the corresponding length scale is. In some typical examples, only the stencil with the minimum truncation order for which the distance scale is clear is used.

[0110] In the case of non-periodic topography, care should be taken to include only the actually computable derivatives, i.e., the derivatives for which the stencil remains within the topographic domain. This is indicated by the subscript “domain” in Equation (13). The rms value as defined in Equation (13) characterizes the amplitude of the fluctuations or the width of the underlying distribution function. Instead of focusing on such a single parameter, it is also possible to determine the complete scale-dependent distribution. Formally, this distribution is (in a single dimension),

Mathematics

[0111] To represent this concept in the example of the gradient (α = 1), panel (b) shows the scale-dependent derivative of the line scan shown in panel (a) of FIG. 1 at l = 40Δx. The distribution function P1(h’, 40Δx) of the gradient at this scale is obtained by counting the occurrences of a given gradient value. The resulting distribution is shown in panel (c) of FIG. 1.

[0112] The rms parameter defined in the previous section is the second moment of this distribution [Number] which is the square root of. The second moment fully characterizes the underlying distribution only when this distribution is a Gaussian distribution. For example, as will be described later, deviations from the Gaussian distribution occur due to artifacts of the scanning probe, but if the complete distribution function is obtained, this can be easily detected.

[0113] The probability distribution of any derivative (e.g., gradient, curvature, or higher-order function) herein functions as an additional set of descriptors for the surface. This distribution is scale-dependent itself, but can be used to calculate various scale-dependent (statistical) parameters herein, such parameters including higher-order cumulants. The statistical characterization of the distribution may be, for example, its second or higher-order cumulants, or its second or higher-order moments. In some embodiments, the statistical characteristics of the distribution are selected from the group consisting of variance, skewness, and kurtosis. rms height (h rms ), skewness (sk), and kurtosis (ku) are provided in FIG. 2. In FIG. 2, x kis the k-th of N data points. In some embodiments, the commonly used variance, and the skewness sk and kurtosis ku parameters are used herein to characterize the probability distribution. Skewness is the standardized third moment, and kurtosis is the standardized fourth moment. In FIG. 2, μ is the mean and σ is the standard deviation. For a normal distribution, the skewness is 0, and the kurtosis is either 0 (Fisher's definition or excess kurtosis) or 3 (Pearson's definition or non-excess kurtosis). In a representative example herein, Fisher's definition is used. Compared to a normal distribution with the same variance, the skewness can be either positive or negative, which is related to, for example, a shift to the left or right compared to a Gaussian distribution. Kurtosis is a measure of how flat or peaked the distribution is compared to a normal distribution with the same variance.

[0114] As described above, various methods for calculating scale-dependent height (e.g., autocorrelation function (ACF), variable bandwidth method (VBM), power spectral density (PSD), etc.) can be associated with scale-dependent parameter analysis herein. Such analysis can be extended to define yet another method for calculating the scale-dependent parameters described herein. In this regard, some forms of the scale-dependent parameters herein can be calculated using such methods instead of using the definitions specified in Equation (13), and in a given embodiment, substantially equivalent results are obtained. Intuitively, the scale-dependent parameters herein can be considered as a general framework for analysis that includes ACF, VBM, and PSD as special cases.

[0115] A common method for analyzing the statistical characteristics of surface topography is the height-difference autocorrelation function, which is referred to herein as ACF or A(l) (as described above). ACF is

Equation

[0116] Some authors refer to 2A(l) as the structure function and the bare height autocorrelation function 〈h(x)h(x+l)〉 as the ACF. The height ACF and height-difference ACF are

number

[0117] Equation (16) is similar to equation (4), which is a finite difference formula for the first derivative. In fact, the ACF uses scale-dependent derivatives to

number

number

[0118] Furthermore, it can be shown that higher order derivatives can be expressed in terms of ACFs. Using the stencil of second order derivatives given in equation (6), the scale-dependent second order derivatives are

number

number

number

number

[0119] The SDRP herein may be derived based on the concept of different scales. In the discussion leading up to Equation (13), the length L of the line scan is not considered. That length is only relevant when determining the upper limit for the stencil length l = αηΔx, which is the scale concept in the measurement based on Equation (13). Alternatively, L may be interpreted as the relevant scale, and the scale-dependent roughness may be studied by varying L. According to this interpretation, a class of methods called the scaled windowed dispersion method or variable bandwidth method (VBM) emerges. The elements of this class of methods only differ in the method of removing the trend from the data, and various names have been given, such as the bridge method (derived from Mandelbrot), roughness around the mean height (MHR; sometimes also referred to as VBM), detrended fluctuation analysis (DFA), roughness around the rms line (SLR), etc.

[0120] In all cases, multiple roughness measurements are performed using different scan sizes L for the same sample (or the same material). Plotting the rms height hrms obtained from these measurements against the scan size L, or plotting the rms gradient h' against the scan resolution (the minimum measurable scale), provides insights into the multiscale nature of the surface topography.

[0121] These methods can be generalized for the analysis of a single measured value. Here, consider a line scan h(x) of length L. The scan is divided into segments of length l(ζ) = L / ζ for ζ ≥ 1 (where l ≤ L is the relevant scale). This dimensionless number ζ is referred to as the magnification herein and defines the scale. Some use a sliding window instead of exclusive segments.

[0122] In VBM, the height variation of the rms in each segment is considered. From this perspective, the height h VBM,i in segment i during expansion (ζ) is calculated, and then the scale-dependent h VBM (ζ) is calculated by taking the average of all i. Some researchers perform slope correction for individual segments. In that case, each segment is detrended by subtracting the corresponding average height and slope (obtained by linear regression of the segment data) before calculating h VBM,i (ζ). This approach is called DFA, while when slope correction is not performed, it is called MHR. In the bridge method, the connecting line between the first and last points of each segment is used for detrending.

[0123] These VBMs are similar to SDRP. When calculating the slope in SDRP, SDRP is calculated simply by connecting two boundary points of x = il(ζ) and x = (i + 1)l(ζ) with a straight line, as done in the bridge method. This method is different from DFA, which uses all data points between two boundary points to fit a straight line using linear regression. Detrending can be generalized to higher-order polynomials, but this has not been reported in the literature. The relationship between SDRP and VBM with first- or second-order detrending is conceptually shown in Figure 3, which represents the calculation of the scale-dependent roughness parameter from the variable bandwidth method (VBM). In the finite difference method, the slope is calculated between two points at a distance l, while in VBM, a trend line is fitted to a segment of width l. Similarly, for the second derivative, in the finite difference method, a quadratic function is fitted between three points, while in VBM, a quadratic trend line passing through all data points in an interval of length l is fitted.

[0124] In DFA, the trend line is simply used as a reference for calculating the fluctuations around it. The coefficients of the trend-removing polynomial can also be used to analyze how the surface gradient and curvature depend on the scale. This results in an alternative measure of the scale-dependent rms gradient h’ VBM (ζ). h’ VBM (ζ) is simply the standard deviation of the gradients obtained within all segments i at a given magnification. Below, it is shown that this scale-dependent gradient is very similar to the gradient obtained from SDPR.

[0125] To extend DFA to higher-order derivatives, the above observations can be utilized. Instead of fitting a linear polynomial in each segment, a higher-order polynomial can be used to perform trend removal. To extract the scale-dependent rms curvature, a quadratic polynomial can be fitted to the segment, and twice the coefficient of the quadratic term can be interpreted as the curvature. Then, from the standard deviation of this curvature over the segment, the scale-dependent second derivative h” VBM (ζ) is obtained. As described above, Figure 3 represents this concept, and again, for the second derivative, it is compared with SDRP that fits a quadratic function passing through only three collocation points.

[0126] An alternative route for considering VBM is to use a stencil whose number of coefficients is equal to the segment length. This stencil can be explicitly constructed from the least-squares regression of the polynomial coefficients (at each scale). And it is considered that the one closest to SDRP is each VBM that uses a sliding segment (rather than an exclusive segment). However, even in this case, the difference remains that SDRP uses a stencil with the same number of coefficients at each scale. In the present study, VBM using non-overlapping segments is used.

[0127] The above discussion suggests that various methods for calculating scale-dependent height (e.g., VBM, DFA, etc.) can be considered as special cases of SCRP analysis. That is, in such cases, scale-dependent trend removal occurs at most only for linear trend lines. Using the relationships developed in this book as described above, these analyses can be extended to define another method for calculating or estimating SDPR.

[0128] Power spectral density (PSD) is another common tool for the statistical analysis of topography. The basis of PSD is Fourier spectral analysis, which approximates a topographic map as a series expansion

Number

Number

Number

[0129] This basis leads to the spectral analysis of surface topography, and the derivative is the derivative of the basis function

Number

Number

Number

Number

Number

[0130] The effective amplitude of the variation can be obtained from Parseval's theorem in the Fourier shape, whereby the real-space average of equation (5) is changed to the sum of the wave vectors

Number

Number

Number

[0131] Fourier filtering and finite differences can be related. First, interpret the finite difference scheme from the perspective of Fourier analysis. Next, apply the finite difference operation to the Fourier basis equation (25). This results in

Number

Number

Number

Number

[0132] the c for different values of

Number

Number

[0133] Thus, it has been shown that SDRP defined in real space can also be calculated or estimated in frequency space using PSD. However, the calculation in frequency space has the drawbacks that it is required to window the non-periodic topography and to apply a filter cutoff.

[0134] The above concept was applied to synthetic self-affine topography. This topography consists of three virtual "measurements" of a large (65,536×65,536 pixel) self-affine topography generated by a Fourier filtering algorithm. See T.D.B. Jacobs, T. Junge, L. Pastewka, Quantitative characterization of surface topography using spectral analysis, Surf. Topogr. Metrol. Prop. 5 (2017) 013001 and S.B. Ramisetti, C. Campana, G. Anciaux, J.-F. Molinari, M.H. Muser, M.O. Robbins, The autocorrelation function for island areas on self-affine surfaces, J. Phys. Condens. Matter 23 (2011) 215004. In this algorithm, sine waves with uncorrelated random phases and amplitudes scaled according to the power law are superimposed. At the pixel at position

Number

Number

Number

Number

[0135] Figure 5A shows the topography maps of these three emulated measurements. The measurements are then zoomed into the center of the topography. The one-dimensional PSDs (C 1D ; Figure 5B) of the three topographies are in good agreement and show zero power below the cut-off wavelength of λ s . The PSD is shown as a function of the wavelength λ = 2π / q (where q is the wave vector), which facilitates comparison with the real-space techniques introduced above and makes the wavelength more intuitively understandable than the wave vector. Since the topography is self-affine, the PSD scales as C 1D ∝λ 1+2H as shown by the solid line.

[0136] The square root of the ACF is shown in Figure 5C. The ACF and all other scale-dependent quantities reported below are obtained from the average on adjacent line scans, i.e., from one-dimensional profiles rather than two-dimensional area scans. This is consistent with the calculation method of C 1D . The ACFs from the three measurements are lined up,

Number

Number

[0137] The scale-dependent curvature h” SDRP (l) is shown in Fig. 5E. Similar to the rms gradient, the curvature saturates for l < l s to reach the “true” small-scale value of the curvature. Since the entire surface is self-affine, the curvatures of the three individual measurements are aligned in a row again, and h” SDRP (l) ∝ l H-2 follows. The rms curvature calculated from the ACF (Equation (22)) is strictly applicable only to a periodic topography. However, in this numerical experiment, the ACF agrees within the range of the original definition of the SDRP (Equation (13)) and the line thickness. An error occurs at a large distance scale, and in principle, the value of h” SDRP may become negative, but such a situation was not observed in the numerical data shown here.

[0138] In the above derivation, alternative routes for obtaining scale-dependent roughness parameters from VBM and PSD were shown. The plus symbols (+) in FIGS. 5D and 5E indicate the rms gradient and curvature obtained using VBM, and the crosses (x) indicate the results obtained using PSD. These are in good agreement with the respective parameters obtained from the SDPR analysis and deviate only at large scales. In summary, all three routes (ACF, VBM, PSD) for obtaining scale-dependent parameters (i.e., second cumulants and / or second moments, or variances) determined from SDPR were verified and gave results that agreed with those calculated using the original definition (Equation (13)). The advantage of SDPR, ACF, and VBM over PSD is that they can be applied directly (without using a window) to aperiodic data. Furthermore, scale-dependent parameters or statistical properties determined from third- or higher-order cumulants, or third- or higher-order moments, cannot be determined from parameters such as ACF, VBM, and PSD.

[0139] In this way, four independent methods for obtaining scale-dependent gradients, curvatures, and higher-order derivatives were demonstrated. All four methods are novel uses of the underlying analytical techniques. In many of the following studies, the main tool is SDRP. The importance of using scale-dependent gradients and curvatures for "naked" ACF, VBM, or PSD is that it is easy to interpret the meaning of these parameters. For example, it is difficult to assign a geometric meaning to the value of PSD, but everyone intuitively understands the meaning of gradients and curvatures.

[0140] In the analysis of tip artifacts, the ability of SDPR to calculate the complete basis distribution for any derivative has been utilized in many of the studies here. Figure 6A shows an aperiodic topography generated on two computers with a size of 0.1 μm × 0.1 μm. The first topography is in the initial state and is generated using the aforementioned Fourier filter algorithm. Similar to the previous example, by obtaining a part of a larger (0.5 μm) periodic scan, it was confirmed that the scan is not periodic. The second topography contains tip artifacts and is obtained from the surface in the initial state using a non-linear procedure. At this point, for all positions (x i , y i ) on the topography, a sphere with a radius R tip (here 40 nm) is lowered towards the position (x i , y i , z i ) until the sphere touches any position on the original topography. Then, the resulting z position z i of the sphere is taken as the "measured" height of the topography. This topography is discussed in T.D.B. Jacobs, T. Junge, L. Pastewka, Quantitative characterization of surface topography using spectral analysis, Surf. Topogr. Metrol. Prop. 5 (2017) 013001, and its data file is available online. The two curves under the map in Figure 6A are cross-sections passing through the center of each topography.

[0141] As is clear from the data in Figure 6A, the scanning probe smooths the peaks of the topography. In fact, the curvature near the peak is -1 / R tipmust be equal to. Conversely, the valley appears like a cusp resulting from the overlap of two spheres. These cusps are sharp, and thus the curvature is thought to be a large positive value (theoretically unlimited, but actually limited by resolution and noise). Due to the tip artifact, the PSD is C 1D (q) ∝ q -4 has been observed to be possible, which is exactly the result of the cusp of the topography. In this regard, the Fourier transform of the triangle scales as q -2 and thus the PSD ∝ q -4 results. This observation has been numerically verified.

[0142] Figure 6B shows the scale-dependent gradient distribution P1(h’, l) normalized by the rms gradient at each scale. The black solid line shows a Gaussian distribution (with unit width) for reference. It is clear that for the scale-dependent gradients over the scale from 1 nm to 256 nm shown in the figure, both the initial state topography (left column) and the topography with a tip radius artifact (right column) follow a Gaussian distribution.

[0143] The situation is different for the scale-dependent curvature, as shown in Figure 6C. The surface in the initial state (left column) follows a Gaussian distribution, but for the topography with a tip radius artifact, it becomes a Gaussian distribution only at larger scales (l = 16 nm and 256 nm). There is an obvious deviation at the smallest scale, showing an exponential distribution for positive curvature values, which supports the above empirical consideration that the curvature becomes a large positive value due to the cusp (sharpness). Due to these cusps, the PSD ∝ q -4 ∝ λ 4 results. Figure 6D shows the PSD of both topographies. The surface with the artifact intersects C 1D ∝ λ 4 at wavelengths where λ is approximately 20 - 40 nm.

[0144] λ 4The intersection with it is slight and difficult to detect from the measurement data. Other measurement methods such as the ACF shown in Fig. 6E are not suitable for detecting such artifacts. C 1D ∝λ 4 The region of is represented as a linear region of the square root of the ACF (A∝l). The exponent 1 obtained from this region is too close to the exponent of H = 0.8 to be clearly distinguishable. A tip radius reliability cut-off comparing the scale-dependent rms curvature and the tip curvature has been proposed so far. Now, an additional index can be established for the purpose of more accurately detecting the occurrence of tip radius artifacts. 1 / 2 Rather than calculating the width of the distribution as in rms measurement, here, the question can be raised of what is the minimum value of the curvature found at a specific scale l. Therefore,

[0145] can be evaluated. The crosses in Fig. 6F show the said quantity on the surface in the initial state and the surface with artifacts. At small scales, it is clear that the curvature of the surface in the initial state is larger than that of the surface with artifacts. Furthermore, since the surface with artifacts settles to at l→0

Number

Number

[0146] below which the AFM data becomes unreliable. l tip is determined using the bisection algorithm tip and

Number

[0147] In another example, experimental analysis was performed on ultra-nano-crystalline diamond (UNCD) films. These UNCD films are described in detail in A. Gujrati, S.R. Khanal, L. Pastewka, T.D.B. Jacobs, Combining TEM, AFM, and profilometry for quantitative topography characterization across all scales, ACS Appl. Mater. Interf. 10 (2018) 29169. Figure 7A shows a single representative AFM scan of its surface, which is available online. Its peaks have rounded tips, similar to the synthetic scan shown in Figure 6A. Its curvature distribution (Figure 7B) also has similar characteristics to the synthetic topography (see Figure 6C). On a large scale, the distribution is approximately Gaussian (shown by the black solid line). On a smaller scale, a deviation towards higher curvature values is observed, indicating a cusp characteristic of tip artifacts. This is due to additional instrumental noise, which contributes to the small-scale features of the data.

[0148] Negative curvature hinders the determination of the tip radius from the scale-dependent tip curvature (Figure 7C). Unlike the synthetic surface, the scale-dependent tip curvature h” min (l) (Figure 6F) does not saturate to a specific value at small distances l. Instead, the radius of the AFM tip was determined from auxiliary transmission electron microscopy (TEM) measurements (Figure 7C inset). For the measured R tip = 10 nm, the region where h” min (l) > 1 / (2R tip ) can be identified as unreliable, and the lateral length scale is such that the data becomes unreliable below it.

Number

[0149] After investigating the effect of the tip radius in a single measurement, SDRP was applied to the entire experimental dataset of A. Gujrati id, and a total of 126 individual measurements obtained from three different instruments, a stylus surface profiler, AFM, and TEM, were combined to extract the surface power spectrum over 8 digits. Figures 8A - 8D show the PSD, ACF, rms gradient, and rms curvature of each individual measurement, respectively, as well as the average curve representative of the entire surface. In each tip - based measurement (stylus and AFM), the critical scale l tip was calculated using Equation (36), and data with scales below l tip were excluded. Due to the good overlap between the AFM data and the TEM data, it was confirmed that this procedure removed the tip artifacts. The entire dataset shows a distinct region where PSD C 1D ∝q -4 .

[0150] As shown in Figures 8A - 8D, all four methods can be used to "stitch together" the data from a large - scale measurement set, and as a result, the SDRP of the physical surface can be obtained. The ACF (Figure 8B) and rms gradient h’ rms (Figure 8C) of the TEM measurement bend downward at large l, and the effect is also seen (although less pronounced) in the synthetic data of Figures 5C and 5D. This is because it is the result of a slope correction that forces a zero gradient at the overall measurement size, so h’ rms is forced to decline towards zero. A more refined slope - correction scheme may be devised to remove this long - wavelength artifact, but the rms curvature h” rms is not affected by the local slope of the measurement, so there is no such artifact. Therefore, it may be important to note that instead of relying on a single approach, it may be important to combine scale - dependent analysis methods.

[0151] The novel SDPR analysis in this specification can be considered a generalization of commonly used roughness metrics. This SDPR approach, for example, helps to reconcile competing roughness descriptors. However, this approach is also superior to other methods, especially in terms of ease of calculation, intuitive interpretability, and artifact detection.

[0152] Using synthetic and experimental surfaces, several additional experiments were conducted, for example, to study the classification of rough surface topographies, etc. The synthetic surfaces were generated as described above. The experimental surfaces were obtained by three different microscopy techniques and were capable of calculating scale-dependent roughness parameters (SDRP) from the nanoscale to the millimeter scale. The measurement techniques applied were stylus surface profilometry, atomic force microscopy (AFM), and transmission electron microscopy (TEM). The set of parameters or feature vectors includes the SDRP and other scale-dependent parameters in this specification and is intended to describe topography in a general way. Therefore, this set is not optimized for only one context but can be applied to various contexts. To verify the selection of statistical parameters, synthetic surfaces with two different Hurst exponents and experimental surfaces with four different crystal coatings were classified. In a typical study, the obtained parameter sets were applied to machine learning classification methods, the support vector machine (SVM) and the Gaussian process classifier (GPC). In the context of machine learning, the parameters for classification are referred to as features, and the set of parameters indicates a feature vector. In a way equivalent to the representation feature vector, descriptive data points are commonly used.

[0153] Feature vectors are constructed as data representations suitable for machine learning algorithms. Therefore, feature vectors may be required to retain a meaningful amount of information regarding surface topography while reducing complexity compared to the entire set of measurements. Thus, a feature vector is a set of parameters extracted from surface topography. The parameters describe statistical features of height, gradient, curvature, and third derivative as a function of the distance scale l = αηΔx. The statistical properties used in representative examples are variance (which may also be referred to as SDRP in this specification), as well as skewness and kurtosis of the scale-dependent distribution (which may be collectively referred to as SDSP or scale-dependent parameters in this specification). These scale-dependent parameters are combined in the feature vector and have dimensions of R27 to R99 in numerical experiments.

[0154] Since the units of features vary (for example, the feature of height is m 2 and the feature of the third derivative is m -2 ), if features are evaluated in absolute values, some features may be overestimated or underestimated. Features with large values may have a greater impact on the model than features with small values. Since the units are different, this does not necessarily reflect the importance of those features regarding classification. Therefore, to align the features to the unit of standard deviation, standardization (also referred to as scaling of the input or normalization of the data) can be applied by the following formula

Number

Number

Number

[0155] For visualization purposes, data points in a high-dimensional space can be projected, for example, onto a two-dimensional subspace. In some studies herein, the two-dimensional subspace is defined by the first two principal components fitted according to the maximum variance of the data distribution. Further, a scree plot is provided showing how much of the total variance in the high-dimensional space is represented by the first 25 principal components. However, the principal component analysis (PCA) representation of data points in low dimensions can also be combined with classification methods. This becomes important, for example, in studies that include missing values.

[0156] In some studies herein, classification was performed using kernel-based methods, the support vector machine (SVM) and the Gaussian process classifier (GPC), as representative models. The commonly used radial basis function (rbf) kernel (also called the Gaussian kernel) was applied to both.

Number

[0157] Since the classification scores of the simple data split of the training set and the validation set depend on the random split variable, the classification scores are obtained by cross-validation. Figure 9 shows the case of 5-fold cross-validation, where the dataset is split into 5 equal bundles. One of those bundles is used as the validation set and the others are used as the training set. These folds are varied so that all folds become the validation set. By doing so, a score of what percentage of the validation set was correctly predicted is returned for each training-validation configuration (fold). The cross-validation score is calculated by averaging the scores of the folds. Furthermore, the variance of the folds is obtained. See Murphy, K. P. (2012). Machine learning: a probabilistic perspective. MIT press.

[0158] A special case of cross-validation is the "leave one out" configuration, in which the validation set consists of only one data point and is decomposed into N folds for N data points. This approach is reasonable for very small datasets and is applied in Studies 4 and 5 described below.

[0159] In the classification study in this specification, two different methods, PCA and recursive feature elimination (RFE), were used to estimate the relevance of features. In PCA, the principal components are linear combinations of features and weights. The larger the weight, the more important the feature estimated by PCA. The weights to be evaluated are obtained from the principal components that best separate the classes in the PCA plot. This is usually the first principal component. In addition to or instead of PCA, autoencoder analysis may be used.

[0160] In addition to PCA, RFE, also known as the backward selection algorithm, was applied. RFE uses a classifier for feature evaluation. In so doing, the classifier takes in all features and repeatedly removes the features that have the least impact on classification fitness. This procedure is repeated until one feature remains, thereby achieving the ranking of features. See Friedmen et al. (supra). The SVM was applied as the classification model for feature evaluation using RFE.

[0161] Maximizing the variance of the entire data distribution in PCA involves solving the principal components sorted by the amount of projected variance. This can be efficiently solved by eigenvalue decomposition of the data covariance matrix. However, if there are missing values in the dataset, the covariance matrix will also include missing values, and eigenvalue decomposition becomes mathematically intractable. A commonly used method for handling missing values in machine learning is to apply an imputation method. In the imputation method, the missing values are replaced with information from other data points, such as the mean value of the features. In the context of multi-scale features of surface topography, the imputation method may not be very reliable. For example, if a feature is only accessed from meter-scale measurements, estimating the feature at the nanometer scale from other feature vectors may lead to incorrect results. Rather, while intersecting scales are considered, other scales may be ignored. Therefore, in some studies in this specification, an adjusted PCA algorithm that handles missing values by ignoring the missing values during fitting has been implemented.

[0162] Similar to maximizing data dispersion, data points

Number

Number

Number

Number

Number

Number

[0163] The minimization of the mean squared error is given by the difference between the original data point y j and the projected data point

Number

Number

number

[0164] The approach of alternating updates of X and W can be adapted to handle missing values in the dataset. Ilin, A. and Raiko, T. (2010), Practical approaches to principal component analysis in the presence of missing values, The Journal of Machine Learning Research, 11:1957-2000. When there are no missing values, features are assumed to have zero mean, so the bias term m can be omitted. Because there are missing values in the dataset, the features in Y cannot simply be set to zero mean, and the alternating algorithm requires maintaining the bias term m. Therefore, the partial derivative with respect to m of the minimization statement in Equation 40 is also considered. Furthermore, because matrix-vector multiplications are decomposed into summations, the actual summation is calculated by multiplying the data entries y ij is performed only for observed indices i and j. The update equation is

Number

[0165] Five studies were conducted towards the goal of classifying surface topographies characterized by multiple measurements over different scales that contain missing values within the feature vector representation. The first study was conducted using two classes of synthetic surfaces with different Hurst exponents to verify whether the concept of classification is applicable. The second study determined whether synthetic surfaces with similar power spectral density (PSD) and experimental surfaces can be distinguished. The third study tested the classification between four different diamond crystal coatings. The fourth study tested the classification of the diamond coatings in the third study using feature vectors extracted from multiple measurements obtained by various measurement techniques over various different scales. The fifth study tested whether classification can be performed even when some features within the feature vector are not observed (missing data).

[0166] As described above, the feature vector is constructed in this specification by scale-dependent parameters (SDRP and SDSP, scale-dependent roughness parameters or generalizations of SDRP). As discussed above, SDSP describes the distribution of scale-dependent derivatives in more detail.

[0167] SDRP takes into account the square root of the second moment of the distribution function that forms the basis of the distance scale. As described above, the underlying distribution function or scale-dependent distribution can be obtained by shifting the stencil of the finite difference approximation on the measurement profile. Thus, the distribution depends on the derivative approximation and the distance scale l = αηΔx. In the context of the scale-dependent roughness parameter, the variance of the scale-dependent distribution is investigated. For a Gaussian distribution surface, the variance completely describes the scale-dependent probability distribution, but not all natural surfaces follow a Gaussian distribution (see, for example, panel (b) of FIG. 1). To explore more information about the scale-dependent distribution, the third and fourth moments may also be considered from the perspective of the measurement-based skewness and kurtosis defined above.

[0168] Furthermore, the scale-dependent distribution is characterized, for example, by scalar parameters such as variance, skewness, and kurtosis. As described above, the scalar parameters of skewness and kurtosis may sometimes be referred to herein as scale-dependent statistical parameters or SDSPs. The scale-dependent statistical parameters can be obtained, for example, from gradients, curvatures, and third (or higher-order) derivatives on the scale factor η or rather on the distance scale l. The functions of skewness and kurtosis are plotted in FIG. 10, where the function of variance corresponds to the function of SDRP. Also as described above, each of SDRP and SDSP is a statistical characterization of gradients, curvatures, and third (or higher-order) derivatives, and may sometimes be referred to herein as statistically characterized scale-dependent parameters or simply scale-dependent parameters. To repeat, such scale-dependent parameters are determined by the statistical characterization of at least one of the distributions of the first or higher-order derivatives of the surface height or h.

[0169] [Correction for redundancy] In the machine learning / classification studies in this specification, various different feature sets representing the same topography were applied. In the classification studies in this specification, up to six different feature sets were investigated, and these feature sets include (i) height, gradient, curvature, and third derivative, (ii) gradient, curvature, and third derivative, and (iii) standardized and non-standardized versions of curvature and third derivative. In some experiments, the reason for excluding the height feature is due to its relationship with the gradient feature. Furthermore, the topography in the experiments was tilt-corrected, which could introduce artifacts in the height and gradient features. The tilt in the measurement appears as a result of the tilt of the measuring device. This tilt is removed or corrected by fitting a center line to the topography map and setting the tilt to zero. Since the effect of tilt correction is not clear with respect to classification, in some feature sets, the height and gradient features are omitted. Furthermore, both standardized and non-standardized feature sets were applied because it was not clear whether having unified units was better for classification in order to utilize geometric meaning or maintaining the original units was better for classification.

[0170] For each study, to better understand the data distribution, the data was projected into two dimensions by principal component analysis (PCA). Furthermore, the data was classified by cross-validation using machine learning methods, namely support vector machine (SVM) and Gaussian process classifier (GPC). Furthermore, to determine which features have relatively high relevance to classification, the features were investigated using the recursive feature elimination (RFE) method and PCA.

[0171] Study 1. Synthetic Surfaces

[0172] In this study, synthetic surfaces were generated using the same input parameters except for the Hurst exponent H. As represented by the images in Fig. 11, 100 surfaces were generated at H = 0.8 and another 100 surfaces were generated at H = 0.3. The size of each surface was 128×128 nanometers and the resolution was 1 nanometer. The distance scale l = αηΔx at which the scale-dependent parameters were obtained was 1, 4, 10, 25, 50, and 100 nm for the height and gradient features, 2, 8, 20, 50, and 100 nm for the curvature feature, and 3, 12, 30, and 75 nm for the third derivative feature.

[0173] The PCA plot in Fig. 12 shows the data distribution projected onto the two axes with the largest variance. The plot by the standardized features indicates that the surface classes are separated and there is no overlap in the two-dimensional subspace.

[0174] According to this scree plot, about 30% of the variance is shown in the plot, while about 70% is still hidden. The data set with non-standardized features is not as well separated as the classes in the plot of the standardized features. However, the class without the height feature in panel (d) can be visually distinguished better than the class with all features in panel (b). In the plot with only the curvature and third derivative features in panel (f), better separation is obtained compared to the plots in panels (b) and (d). The plots in panels (b) and (f) show about 60% of the variance of the whole data, while in panel (d), about 40% of the variance is maintained.

[0175] Table 1 shows the classification results of a support vector machine (SVM) and a Gaussian process classifier (GPC) using a radial basis function (RBF) kernel. The classification scores were obtained by 5-fold cross-validation. All classification scores were 1.0 except for the unnormalized features of height, gradient, curvature, and third derivative classified by the GPC, while the classification score for its unnormalized features was slightly lower, at 0.99. Furthermore, FIGS. 13 and 14 show the visual classification regions of the SVM and GPC trained in a two-dimensional PCA subspace, respectively. The SVM in FIG. 13 has a solid boundary line between classes, and the GPC in FIG. 14 provides a probability distribution where the data points are part of the class. [Table 1]

[0176] Figure 15 shows the individual feature weights of the first principal component related to a feature set that includes the characteristics of height, gradient, curvature, and third derivative. The estimated values of the relevance of other feature sets (those including height and gradient, and those not including them) are qualitatively the same. There are some differences between the standardized features in panel (a) and the non-standardized features in panel (b). For the standardized features, the values of variance have a higher relevance for height, gradient, curvature, and third derivative, and the features belonging to a distance scale closer to its resolution have a higher relevance than the features belonging to a larger distance scale. And the skewness parameter is estimated to have a lower relevance than the kurtosis parameter. Furthermore, for the same feature types (variance, skewness, and kurtosis) and the same distance scale, the feature estimated values of height and gradient are equal. For the non-standardized features, the variance of height at the distance scales of 25, 50, and 100 nanometers has a very high estimated validity. In addition, some single features of gradient, curvature, and third derivative have a higher estimated relevance than most of the remaining features. The evaluation by recursive feature elimination (RFE) in panel (c) of Figure 15 evaluates the variance feature higher than the skewness and kurtosis features. According to panel (a), at small distance scales (1 - 30 nm), the best features are determined for all derivatives (including height as the zero - order derivative).

[0177] According to the PCA feature evaluation and RFE estimation of the standardized dataset, the two features that are best evaluated by RFE are plotted in panel (a) of Figure 16. These classes are clearly separable, and even one of the features in panel (a) is sufficient to separate these classes with a straight line. Furthermore, since the data points are almost in a straight line along the diagonal axis of the plot, these features show a high correlation. The features of the variance in panel (b) are the most highly evaluated features from the non-standardized dataset determined by PCA, but their classes are not visually distinguishable by a straight line. Panel (c) shows a similar setting of inseparable classes, where the most highly evaluated non-variance features determined by RFE are plotted. Furthermore, the centers of the clusters are almost 0 on both axes.

[0178] Study 2. Experimental and synthetic surfaces with the same PSD

[0179] Study 2 included the analysis of ultrananocrystalline diamond (UNCD) coatings, which were measured by atomic force microscopy (AFM) to have a size of 2500×2500 nanometers and a resolution of 4.88 nanometers. From the surface topography, the power spectral density (PSD) was extracted and used to generate a synthetic surface as the variance of the amplitude of the Fourier coefficients. As a result, 100 synthetic surfaces with the same size and resolution as the UNCD surface were generated. A feature vector was obtained from each surface. Furthermore, 100 data points were generated from the experimental UNCD surface. Fig. 17 represents the process of generating data points for a two-dimensional UNCD surface. The surface was divided into 100 equal bundles consisting of five measurement profiles, and a feature vector was generated by one bundle of profiles. The five profiles were uniformly spread over a two-dimensional region where the i-th measurement profile of the i-th feature vector was between 0 and 99. The set of distance scales of the features used in this experiment were height and gradient of 4.88, 48.8, 240, 480, 878.4, 1,464, 2,196 nm, curvature of 9.76, 97.6, 480, 960, 1,756.8 nm, and third derivative of 14.64, 146.4, 720, 1,440 nm.

[0180] Fig. 18 is a PCA plot related to this study. The configurations of the features of gradient, curvature, and third derivative are omitted here because (i) in the standardized setting, they are the same as in panel (a), and (ii) in the non-standardized setting, they are equivalent to panel (d). The PCA plots of the standardized data in panels (a) and (c) of Fig. 18 are very similar except that the classes in panel (c) are separated by a straight line. In panel (a), there is a slight overlap. For the non-standardized data, the axis range in panel (b) is two orders of magnitude larger than the axis ranges of the other plots, and some of the class data points intersect. In contrast, in panel (d), the classes are very well separated. Furthermore, the class of the UNCD data points is much more spread out than the class of the synthetic surfaces.

[0181] The classification scores in Table 2 indicate that for all standardized feature sets, the scores of SVM and GPC are 1.0. For non-standardized feature sets, the score of SVM is 0.9 or slightly higher. GPC classifies the feature set without height better than SVM, and its score is 1.0. In contrast, the classification score including the height feature is relatively low at 0.510.

Table 2

[0182] Panel (a) of Figure 19 shows the estimated feature relevance of the first principal component for the standardized features. Similar to the first study, the estimates for cases with fewer standardized features are qualitatively the same as those for non-standardized data. The feature of height variance is highly evaluated, similar to the result in Panel (b) of Figure 15. The estimates of height and gradient are equivalent, and the feature of kurtosis generally has the highest estimated relevance. The feature of variance has a high estimate when the distance scale is small and a low estimate when it is large. Similar to the feature estimates of PCA, RFE generally evaluates features at low distance scales as highly relevant for classification. The top 3 features evaluated by RFE are part of the curvature and the features of the third derivative. This applies to the standardized feature set (Panel (b)) and the non-standardized feature set (Panel (c)) in Figure 19. Furthermore, some of the overall most highly evaluated features in Panel (c) are the features of kurtosis. Panel (a) of Figure 20 shows how the data is distributed with respect to the features of skewness and kurtosis (which had the best ranking in Panel (c) of Figure 19). Additionally, the top-ranked features in Panel (c) of Figure 19 are shown in Panel (b) of Figure 20, and this Panel (b) shows two partially overlapping clusters.

[0183] Study 3. Experimental Surface

[0184] In this study, four different types of experimental surfaces were compared, including microcrystalline (MCD), nanocrystalline (NCD), ultrananocrystalline (UNCD), and polished ultrananocrystalline (PUNCD) diamond coatings as described in Gujarati, A., Sanner, A., Khanal, S. R., Moldovan, N., Zeng, H., Pastewka, L., and Jacobs, T. D. (2021). Comprehensive topography characterization of poly- crystalline diamond coatings. Surface Topography: Metrology and Properties, 9(1):014003. Overall, four 2D measurements (16 surfaces in total) were performed for each of the applied coating types to generate data points for classification. The surfaces were measured by an atomic force microscope with a size of 2,500x2,500 nanometers and a resolution of 4.88 nanometers. 100 data points were extracted for each class, and 25 data points were obtained from each 2D measurement. The process of generating multiple feature vectors from one 2D surface is shown in Figure 17. In contrast to the second study, 20 one-dimensional profiles rather than 5 were used to construct the feature vectors. The distance scales of the features applied in this experiment are 4.88, 48.8, 240, 480, 878.4, 1,464, 2,196 nm for height and gradient, 9.76, 97.6, 480, 960, 1,756.8 nm for curvature, and 14.64, 146.4, 720, 1,440 nm for the third derivative.

[0185] In addition to the classification by cross-validation performed in other studies, in this study, classification with a predetermined division was performed in the training set and the validation set. In this regard, 25 feature vectors obtained from the same 2D measurement values formed sub-clusters, particularly for the MCD surface and the UNCD surface (see Figure 21). To verify that the feature vectors sampled from the measurement values not used for training can be correctly classified, the data points on the three surfaces of each class were assigned to the training set, and the data points on the remaining surfaces were assigned to the validation set.

[0186] Figure 21 shows the PCA plots of the sample points of MCD, NCD, UNCD, and PUNCD. In the plots of panels (a), (b), and (c), the classes of UNCD and PUNCD each form one cluster, and the labels of MCD and NCD each form 3 to 4 clusters. These plots show a clear separation between the clusters, except for a partial overlap of the classes of MCD and NCD in panel (c). In contrast, the clusters in panel (d) show only one cluster each, and in particular, the classes of MCD, NCD, and UNCD are qualitatively closer. Furthermore, the clusters of UNCD and PUNCD in panel (b) are more compact compared to the clusters of MCD and NCD, and the axes in panel (b) are 2 to 3 orders of magnitude higher than those in the other panels.

[0187] The classification scores of 5-fold cross-validation are shown in Table 3. The scores of the standardized data are approximately 1.0 for all the cases listed. The classification scores are lower for the non-standardized data set, but the feature set excluding the height feature still shows good classification scores of around 0.9. However, the performance of the non-standardized feature set of height, gradient, curvature, and third derivative features is 0.628 for SVM and 0.71 for GPC, which has deteriorated significantly, but is still better than the random guessing score of 0.25. Furthermore, the variance of classification for GPC is 0.126, while it is close to 0 for other cases. These results indicate that the classification scores may strongly depend on the training-validation split of the data set.

Table 3

[0188] To compare the training principles of different classifiers, visualizations of the learned models of SVM and GPC in the PCA subspace of the first two principal components were created for the standardized data set with all features. All available data points were applied for training. In this visualization, GPC was very well trained even in the two-dimensional space, but the MCD classification region of SVM overlapped with some of the NCD data points.

[0189] For the four surface measurements of each class, a standardized feature set consisting of height, gradient, curvature, and third derivative was used, and classification was performed by dividing into a training set and a validation set. As a result, scores of 0.93 for SVM and 0.84 for GPC were obtained. GPC provided the probability of belonging to all trained classes for each predicted data point, rather than returning only the predicted class labels as in the case of SVM. These probabilities are shown in Fig. 22 for some predicted data points. MCD surfaces and UNCD surfaces can be very reliably classified with probabilities of about 60% and 80%, respectively. Also, the sample points of the PUNCD class have the highest probability that the label is correct, and thus are correctly classified. Only the sample points of the NCD surfaces did not show a significantly high probability. They were misclassified in 16 out of 25 cases. However, the probabilities of the MCD class and the NCD class are significantly higher than those of the UNCD class and the PUNCD class.

[0190] Fig. 23 shows that the PCA estimated by the skewness feature is worse than the PCA estimated by the variance and kurtosis features, while RFE highly evaluates the variance feature. There was a large contradiction in the estimation of the skewness feature of the third derivative at a wavelength of 720 nm. In this regard, RFE highly evaluated this feature, while PCA evaluated it low. Furthermore, RFE estimated the curvature and third derivative features higher than the height and gradient features, while PCA evaluated the curvature feature higher than the other derivatives.

[0191] Study 4. Combination of Measurements at Multiple Scales

[0192] Similar to Study 3, in Study 4, the experimental surfaces of MCD, NCD, UNCD, and PUNCD were analyzed. However, unlike the third study, the same-sized AFM scans were not used. Instead, multiple measurements of the same surface were combined as feature vectors. Each feature vector represents 10 different measurements spanning scales from nanometers to millimeters. Since the scale ranges covered by individual measurements overlap partially, the scale-dependent parameter / SDSP at a certain distance scale is obtained by averaging the scale-dependent parameters / SDSP of the measurements covering that scale. Those measurements were obtained by three different measurement techniques (stylus profilometry, AFM, transmission electron microscopy (TEM)). In total, 30 feature vectors were generated (6 for MCD, 6 for NCD, 12 for UNCD, and 6 for PUNCD). The features of gradient, curvature, and third derivative cover all distance scales of 1, 5, 10, 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000 nm.

[0193] Figure 24 shows the PCA plot of the feature set including all features. The PCA plots of the feature sets including only curvature and third derivative features are very similar. In the standardized feature set of panel (a), the classes were clearly separable. In the non-standardized feature set of panel (b), the classes were more mixed. Among the non-standardized features, only the PUNCD class was clearly clustered. The data points of MCD and NCD are slightly spread out, while the data points of the UNCD surface are more spread out on the PCA subspace.

[0194] Since the amount of data points was small, the classification scores were obtained by leave-one-out cross-validation as described above. The relevant classification scores are shown in Table 4. The scores for the standardized data using SVM are quite good (close to 1.0), while the performance of GPC is slightly worse for larger feature sets, with a score of 0.867. For smaller feature sets (with features of curvature and third derivative), GPC gave worse results. Furthermore, the variance of the GPC scores approaches 0.2, which is very high. The scores for a single classification task strongly depend on the training-validation split. The classification scores for the non-standardized data are relatively poor (below 0.333) for both classifiers.

Table 4

[0195] Feature estimation by PCA and RFE is shown in Figure 25. Both estimate higher relevance for features at distance scales of 5000 nm and above. PCA highly evaluates the features of variance, while RFE highly evaluates the features of variance for all derivatives and also highly evaluates the features of skewness in curvature. The low evaluation of PCA for the skewness feature of curvature does not match the evaluation of RFE.

[0196] Study 5. Handling of incomplete feature vectors

[0197] In Study 5, the dataset of Study 4 with features of gradient, curvature, and third derivative was used. To simulate the situation where not all measurements were made on the same scale, some values were removed from the dataset in Study 4. By doing so, the set of features obtained at a given distance scale was removed. That set included the variance, skewness, and kurtosis regarding all derivatives. In the fourth study, it was observed that the standardized features were better than the non-standardized features. Therefore, in this study, the standardized features were used. For standardization, it was required to calculate the mean and standard deviation of each feature. When dealing with missing values in this way, the mean and standard deviation were calculated only for the observed values, and the missing values were ignored.

[0198] Since standard classification algorithms cannot easily handle missing values in the dataset, a PCA method for handling such missing values was implemented. By preprocessing the dataset containing missing values with the modified PCA method, a data point representation in the PCA subspace without missing values was obtained. To evaluate the performance of the missing values, 25%, 40%, 60%, and 75% of all values were removed.

[0199] The phrase "missing data point" or "missing value" refers to a value that does not exist or is missing within the feature vector

Number

Number

[0200] The feature vector is a reduced vector

Number

Number

Number

Number

Number

[0201] Two different feature sets with 25% missing values were constructed. As shown in panel (a) of FIG. 26, features at large distance scales (100 - 500,000 nm) were removed from one feature set, and features at small distance scales (1 - 1,000 nm) were removed from the other feature set. Removal of features at a certain length scale included features of gradient, curvature, and variance, skewness, and kurtosis of the third derivative. In this regard, if the distance scale is not covered by the measurement, all of those features are not observed. The feature vectors affected by the removal were randomly selected. The configuration of 40% missing values was created by independently removing sets of features from large distance scales (1,000 - 500,000 nm) and small distance scales (1 - 1,000 nm) (represented by the bandwidths of the left and right dashed lines in panel (b) of FIG. 26). Thus, it is possible to delete both scale regions from the feature vector, only one scale region, or no scale region. The same procedure was applied when the missing values were 60% and 75%, and further, from some feature vectors, the distance scale of 100 - 50,000 nm (the bandwidth of the solid line in panel (b) of FIG. 26) was removed. To retain some information about the underlying surface, no study removed all three scale regions from one feature vector.

[0202] Figure 27 shows the missing value PCA with 25% of the values missing in panels (a) and (b). Here, the PCA data point representation of the complete eigenvectors is included, which is marked by the black edges. The incomplete eigenvectors are still clustered within the same class. Further, the PCA data points are close to those of the complete eigenvectors. Additionally, other configurations with missing values are represented as PCA plots in Figure 28. As can be seen in Figure 28, the cases of 40% and 60% (panels (a) and (b) respectively) are still clustered within the correct class but are closer to each other than in the case of 25% missing values. The MCD class and NCD class in panel (b) are particularly very close, similar to the UNCD class and PUNCD class in panels (a) and (b) of Figure 28. For 75% missing values in panel (c) of Figure 28, the clusters overlap at many points.

[0203] Classification is performed by leave-one-out cross-validation and the scores are shown in Table 5. The two configurations with 25% missing values (removing the magnitude distance measure) have the same classification scores given in the table. Generally, SVM classifies very well for 25%, 40% and 60% missing values, showing scores close to 1.0. However, for the GPC classification, it is about 0.75 for the same cases, showing a large variation (almost 0.2) in classification. In the case of 75% missing values, both classifiers show relatively low classification scores and the variation in classification becomes large. However, in contrast to the other cases, GPC has a better score than SVM for the case of 75% missing values.

Table 5

[0204] Summarizing the results of Studies 1 - 5, in Studies 1 - 4, using either classifier SVM or GPC, successful classification scores of 1.0 or slightly lower were achieved in at least one dataset configuration. The same also applies to Study 5 which has up to 60% missing values. Compared to GPC, SVM showed slightly better classification performance than GPC in Studies 4 and 5. The high classification variance in different training - validation splits of cross - validation indicated that GPC predictions are not very reliable in that context. This classification variance may decrease with larger datasets. In contrast to SVM, GPC provides the predicted probabilities for each class label. Referring to the third experiment, the feature set in panel (a) of Figure 21 yielded very good score results by applying cross - validation, but the model based on this data may have degraded performance with new data because the data points (or sub - clusters) of MCD and NCD are closely arranged to each other in the PCA plot. Similar to Study 3, when assigning the data points of one measured surface of each class as the validation set (extracting 100 data points from 4 different measured surfaces for each class), it becomes a more difficult classification task. In this context, SVM performs slightly better as determined by the classification score, but GPC provides the predicted probabilities for each validation point of each class. The results can be evaluated for predictions as shown in Figure 22. In Figure 22, the validation points of MCD, UNCD, and PUNCD are correctly predicted, but prediction errors occur only by the NCD data points being predicted as MCD. Thus, GPC can substantially surely indicate that the feature vector belongs to either the MCD class or the NCD class. In this context, since SVM only returns class labels as correct or incorrect, GPC provides better estimations than SVM. The same analysis also applies in the context of missing values in Study 5 which has 75% missing values.Even when the missing information regarding the data distribution affects the classification, the probabilities provided by GPC can narrow down the prediction to two or three out of the four classes in order to provide sufficient information regarding the more likely class labels.

[0205] As demonstrated in Study 5, when dealing with missing values using an algorithm such as the missing value PCA method and classifying the data in a two-dimensional subspace, it functioned well. It is considered that the amount of information regarding the data distribution increases when maintaining two or more principal components. However, in this method, since the computational complexity increases with the number of principal components, it significantly increases the computational complexity as a result of the inverse matrix in Equation 43. The prediction of new data that occurs in a high-dimensional space and is not transformed into the PCA subspace can be applied to the trained model regardless of whether it contains missing values. When doing so, the data can be projected onto the principal components before being applied to the trained model. When predicting a data point with missing values, the projection can be performed by ignoring the missing values in the matrix multiplication.

[0206] Generally, this analysis is not overly sensitive to datasets / surface scans with missing data points, or with different or limited bandwidths (as a result, values are missing from the feature vectors). As a result of the bandwidth limitation of the measuring instrument, various length scales are missing from the dataset created by one or more scans of the target surface. Nevertheless, the target surface can be appropriately characterized via the apparatuses, systems, and methods herein even in the case of missing data or limited bandwidth. In some embodiments, the principal component analysis algorithm used herein is adapted to process datasets with missing values of the data or different / limited bandwidths. In some embodiments, the missing data may also be input as known in the art.

[0207] Both classifiers used in the studies herein were applied using standard hyperparameters of scikit - learn. Thus, the classification results may be further optimized. In the case of classification using non - standardized features, there may be a greater potential for optimization. Although the score may be improved by adjusting the classifier, it is desirable to perform the adjustment individually for each setting. In contrast, in the classification model with standardized features, the default hyperparameters fit well and the pre - imparting of features by standardization functions well, so further improvement or optimization may be difficult.

[0208] The classification results in this study include scores for various combinations of feature sets. Further, those features are evaluated by weights by PCA and by RFE. The classification of standardized features was always equal to or better than that without standardization. In Study 4, the scores of standardized features were particularly excellent compared to those without standardization. Further, the PCA plots of standardized features demonstrate excellent visual clustering of classes in most cases. In contrast, the PCA plots of non - standardized features in many studies seem to overestimate values larger than those shown by the scree plot (see, for example, Figure 12). Similarly, in the feature relevance estimation in the first experiment in panel (b) of Figure 15, there are features highly evaluated and features very lowly evaluated. However, highly evaluated features are not necessarily valuable features for classification (see, for example, panel (b) of Figure 16). Overall, it is recommended to use standardized features because it suppresses the over - evaluation of high values and generally results in better classification scores. Further, as described above, standardized features match better with the standard hyperparameter settings of the classifier.

[0209] Which features of the distribution (e.g., variance, skewness, kurtosis, or higher-order moments / cumulants) play an important role depends on the pool of surfaces being classified. In Study 1, as shown in panel (c) of Figure 16, the classification is mainly done with respect to the variance feature. In this regard, since the synthetic surfaces follow a Gaussian distribution, the skewness and kurtosis features vary around zero. In contrast, some highly evaluated variance features are very powerful as shown in panel (a) of Figure 16 and can classify various classes without using the information of other features. The skewness and kurtosis features were further investigated in Study 2, and due to the related power spectral density (PSD), the height variances are structurally similar. In this study, the ability to classify using the skewness and kurtosis features is shown in panel (a) of Figure 20. Furthermore, the RFE of the standardized feature set highly evaluated some variance features that were expected to be less relevant as a result of the related PSD. Panel (b) of Figure 20 demonstrates that in the case of Study 2, the surfaces can be classified by the variances of curvature and the third derivative. Without being limited to any mechanism, this result may be due to equipment noise and tip artifacts that mainly occur at small scales of curvature. However, further investigation will be needed to draw conclusions.

[0210] Overall, variance features may be relatively highly relevant in the context of classification in the studies herein, but skewness and kurtosis features also showed high evaluation in Study 4 (see, e.g., panel (b) of Figure 25). As observed in all five studies described above, there is a large variation in the highly evaluated features. Therefore, maintaining all such features is useful for maintaining a general feature set that can also be adapted for use in specific cases.

[0211] In representative studies, even when the features of height and gradient were removed from the classification, the classification scores did not deteriorate significantly. These features may not be very important for higher-order derivatives that are scale-dependent, or at least are redundant. Furthermore, since the stencils for obtaining the scale-dependent distributions of height and gradient are the same except for the division of the distance scale for the gradient, the features of height and gradient are dependent. This observation is consistent with the results of the first three studies where the scores of the standardized features were always the same regardless of the presence or absence of the height feature. Also, the estimation of the feature relevance of PCA is equivalent for both cases. In contrast, the dependence of the non-standardized features was misinterpreted as a result of overestimating large values (which are height features).

[0212] Artifacts from surface measurements caused by the unknown inclination of the measuring device are usually automatically corrected, but can affect some gradient features. On the other hand, in the features of curvature and third derivative, this effect is not prominent. Comparing the studies conducted with or without gradient features, it has been demonstrated that there is no harmful effect due to inclination from the perspective of classification. Conversely, in Study 4 where gradient features were included, it was found that the classification score increased slightly. However, this effect may be greater for measurements obtained by other measuring devices. Furthermore, since classification depends on features that provide information on class separation rather than features that do not, the features of curvature and third derivative may complement the poorly expressive gradient features from the perspective of classification. Overall, the results of this study indicate that the height feature may be excluded, but it is desirable to include the features of curvature and third derivative. Including gradient features may add some information about topography, but measurement artifacts may also be added. Therefore, such considerations are preferably evaluated on a case-by-case basis.

[0213] Most of the available bandwidth was covered by the features of Investigations 1 - 5. In Study 4, a relatively large distance scale (5,000 - 500,000 nm) was evaluated as being more relevant than a relatively small distance scale (see Fig. 25). Thus, in the context of the diamond crystal classification of Study 4, it can be concluded that these scales are more appropriate. Studies 3 and 5 showed that this classification is applicable not only to these large scales. In Study 3, it was shown that the scale of 4.88 - 2196 nm contains sufficient topographic information for classification. On the other hand, in Study 5, some data points (25% missing values) were removed from the bandwidth of 100 - 500,000 nm, while the classification at a general bandwidth of 1 - 10 nm was successful. In techniques such as measurements by AFM and stylus profilometry, the application of small scales (e.g., scales up to the measurement resolution) may lead to incorrect results regarding the similarity of the actual topography because tip artifacts of the measuring device may appear. To detect the scale at which the influence / artifacts of the tip radius become prominent, the determination of the influence of the tip radius described herein may be used to determine the reliability cut-off. Furthermore, at large scales (e.g., scales close to the measurement size), there are only a few values describing the scale-dependent distribution. Thus, the estimation of variance, skewness, and kurtosis for such large values is not very reliable. Also, using a distance scale that is too close to the limit of measurement may result in poorly represented features. Overall, the multi-scale behavior of the topography was extracted at 2 - 6 distance scales per decade by the scale-dependent parameters of this specification, but in the 4th and 5th studies, good results were obtained only at 2 distance scales per decade. Even considering only 2 distance scales per decade, there is redundancy in the context of diamond crystal classification. As demonstrated in Study 5, the classification score does not decrease significantly even if some scales at some data points are removed. Considering the very small number of distance scales, sufficient information about the topography for the analysis of other characteristics may not be obtained.Therefore, in order to describe topography extensively, it may be appropriate to have two or three distance scales per decade.

[0214] As described above, the classification in this specification is performed using a kernel-based support vector machine (SVM) and a Gaussian process classifier (GPC). Other classification models / algorithms may be used. As is known in the fields of computers and machine learning, classification represents a model for class labels, which is a quantity that can only take two values. A regression model / algorithm may also be used. Regression represents a model for a continuous quantity that can take any value. A set of "features" represented as a vector

Number

Number

Number

Number

Number

Number

[0215] The regression model generates a continuous value v, and the continuous value v may be outside the range of 0 to 1. A typical example is the coefficient of friction. And the regression algorithm maps from the feature vector to this value

Number

Number

Number

[0216] Another representative example of a regression model suitable for use in this specification, the Bayesian regression model, does not generate the parameter

Number

Number

Number

Number

[0217] Thus, similar to classification, scale-dependent parameters can be used to calculate feature vectors, and a regression model (linear or Gaussian process) can be used to calculate the predicted value of continuous characteristics when surface roughness is given. The values of interest include, for example, the coefficient of friction, wear rate, adhesion, lifespan, and many others. The concept of using non-parametric regression is not to make assumptions about the underlying physical process. And the result is some form of interpolation of the input data.

[0218] In some embodiments, the system of the present invention includes an electronic circuit, and the electronic circuit includes a memory system and a processor system. Such a system may be embodied, for example, in a cloud-based system including a remote processing / analysis center as represented in the representative embodiment of FIG. 29. The database system may be stored within the memory system. The database system includes topographic data related to one or more scans for each of a plurality of surfaces. The database system may include "raw" topographic data such as surface height data from various sources as shown in FIG. 29. For example, the topographic data may be obtained from a storage device of topographic data from various measurement systems such as a stylus surface shape measurement system, an optical surface shape measurement system, a cross-section or side microscope system, and a reflectance system.

[0219] The topographic data stored within the database system may further include an evaluation of the statistical characteristics of the distribution of one or more derivatives of the surface height for at least one of the one or more scans, where the one or more derivatives are selected from the group consisting of zero - order and higher - order derivatives and are determined using a scaling coefficient η that is multiplied by the smallest possible distance scale provided by at least one of the one or more scans at each of a plurality of distance scales in real space.

[0220] One or more algorithms are stored within a memory system executable via a processor system. In some embodiments, the (one or more) algorithms include an algorithm for the determined scale - dependent parameters. In some embodiments, the (one or more) algorithms include at least one machine - learning procedure trained using a training set of topographic data using the features / feature vectors and labels of the training set of topographic data.

[0221] Furthermore, the surface topographic data may be uploaded by a user of the system and / or determined by one or more topographic measurement systems local to the processing / analysis center. Thereby, the data of the database system can be continuously enhanced. Further, one or more of the machine - learning models described herein may be trained using additional data.

[0222] Although it is repetitive, at least one distribution of the first or higher-order derivatives is determined over a plurality of distance scales, for example, via a numerical method (e.g., the finite difference method), and then statistically characterized. This statistical characterization may alternatively be determined from surface topography / roughness parameters other than the SDSP that may be mathematically related to the SDSP. Surface roughness / topography parameters other than the SDSP may be selected from, for example, the group of autocorrelation function characteristics, variable bandwidth characteristics, or power spectral density characteristics. Surface roughness / topography parameters other than the SDSP may be selected from, for example, the group of autocorrelation function characteristics, variable bandwidth characteristics, or power spectral density characteristics. Further, the feature vector (of the training set data and data input for characterization) herein may include surface topography / roughness parameters other than the SDSP (e.g., power spectral density data, height-difference autocorrelation function, and variable bandwidth characterization data). The stored topography data may be stored as a combination of a plurality of measurements spanning a plurality of scales, including, for example, SDSP data, power spectral density data, height-difference autocorrelation function data, and variable bandwidth characteristic data.

[0223] The algorithm stored in the memory system creates a feature vector as described herein from the input data and uses one or more machine learning models of the algorithm to characterize data from a surface shape measurement system input by a user of the system (e.g., via a cloud-connected device such as a computer). For example, such characterization may include identifying similar surfaces within a database system (such as measured by other researchers / scientists / engineers). This makes it easier to identify and compare studies performed on similar samples but independently according to the apparatus, system, and method herein.

[0224] Furthermore, by analyzing, for example, combinations of multiple surface shape measurement values, the apparatus, system, and method of the present invention can better understand the specifications required for a surface and better detect surfaces that deviate from the specifications (even from measurement values with limited bandwidth obtained by a single measurement system). Further, to predict surface / part surface characterization / properties (e.g., friction, adhesion) and more fully understand how a component will operate during use, data input from a user regarding measurements of the surface topography of the manufactured part may be used. Further, by calculating surface characteristics / properties based on candidate topographies generated by a computer, the apparatus, system, and method of the present invention enable a product designer to rationally determine an optimal surface topography. Using one or more machine learning models of the apparatus, system, and method of the present invention, a user may input data / measurements of surface topographies classified in various different ways (e.g., early failure, sufficient lifespan, etc.) to identify characteristics correlated with component failure, lifespan, etc.

[0225] The project leading to this application has received funding from the European Research Council (ERC) under the European Union's Horizon 2020 research and innovation programme (grant agreement No 757343).

[0226] The foregoing description and the accompanying drawings present some representative embodiments as of the present time. Of course, various modifications, additions, and alternative designs will become apparent to those skilled in the art without departing from the scope of the present invention in light of the foregoing teachings. Here, the scope of the present invention is indicated not by the foregoing description but by the following claims. All changes and modifications that fall within the meaning and scope equivalent to the claims shall be embraced within their scope.

Claims

1. A method for characterizing the surface topology of a target surface, comprising: determining a feature vector from one or more measurement values of the target surface, wherein a plurality of features in the feature vector are determined from a statistical characterization of the distribution of one or more derivatives of the surface height or h, and the one or more derivatives are selected from the group consisting of zero-order and higher-order derivatives determined from at least one of the one or more measurement values of the target surface at each of a plurality of distance scales, and for at least one of the one or more measurement values, the one or more derivatives of the surface height are determined in real space at the plurality of distance scales using a scaling coefficient η, which is multiplied by the smallest possible distance scale provided by at least one of the one or more measurement values; determining at least one characteristic of the target surface via an algorithm storable in a memory system and executable via a processor system and based on the feature vector; providing an output indicative of the at least one characteristic. A method comprising the above steps.

2. The method according to claim 1, wherein the plurality of features in the feature vector are determined from the statistical characterization of the distribution of two or more derivatives of the surface height, and the two or more derivatives have different orders.

3. The method according to claim 1, wherein the one or more derivatives of the surface height are selected from the group consisting of a zero-order derivative, a first-order derivative, a second-order derivative, a third-order derivative, and derivatives of order higher than third order.

4. The method according to claim 3, wherein the one or more derivatives of the surface height are selected from the group consisting of first-order or higher-order derivatives.

5. The method according to claim 1, wherein the statistical characterization of the distribution is determined from second-order or higher-order cumulants of the distribution, or second-order or higher-order moments of the distribution.

6. The method according to claim 1, wherein the statistical characterization of the distribution is determined from third-order or higher-order cumulants of the distribution, or third-order or higher-order moments of the distribution.

7. The method according to claim 5, wherein the statistical characterization of the distribution is selected from the group consisting of variance, skewness, and kurtosis.

8. The method according to claim 1, wherein the values of the plurality of features are normalized.

9. The method according to claim 1, wherein the one or more derivatives are determined over a plurality of distance scales for a line of the one or more measured values of the surface or an area of the one or more measured values of the surface.

10. The derivative with respect to the line of the one or more measured values is for a point x on the line k and is given by the following formula: 【Number 1】 is provided, where α is the order, η is an integer greater than or equal to 1, Δx is the smallest possible scale, and C l (α) is the stencil of the derivative, The method according to claim 9, wherein the derivative is measured at a distance scale l = αηΔx.

11. The stencils for α = 1, 2 and 3 are 【Number 2】 and all other C l (α) is zero, the method according to claim 10.

12. The method according to claim 1, wherein the tip radius effect for the measurement method used in the one or more measurements is determined as a function of a minimum value of the variance in the second derivative at a specific distance scale l.

13. Critical scale l tip is determined, and in order to minimize the tip radius effect, l tip data at a distance scale below the tip is excluded, the method according to claim 12.

14. l tip is given by the following formula 【Number 3】 is numerically estimated using, where 【Number 4】 is the minimum value of the second derivative in the scale (l tip ), and R tip is given by the following formula 【Number 5】 is the tip radius given by, and c is an empirically determined parameter, the method according to claim 13.

15. Two or more measured values are used in defining the statistical property evaluation, and each of the two or more measured values is created by a different measurement method or has a different possible minimum distance scale or resolution, the method according to claim 1.

16. The method according to claim 15, wherein the different measurement methods are selected from the group consisting of a stylus surface profilometry method, a scanning probe microscopy method, an optical surface profilometry method, a cross-section or side-view microscopy method, and a reflectance method.

17. The method according to claim 16, wherein data from the one or more measured values created via two or more measurement methods are combined over the plurality of distance scales when determining the statistical property evaluation.

18. The method according to claim 1, wherein at least one of the one or more derivatives of the surface height h is a third-order or higher-order derivative.

19. The method according to claim 1, wherein at least one of the one or more derivatives of the surface height h is a fourth-order or higher-order derivative.

20. The method according to claim 1, wherein the algorithm includes at least one machine learning model.

21. The method according to claim 20, wherein the at least one machine learning model is a classification model or a regression model.

22. The method according to claim 21, wherein the classification model is a support vector machine model, a Gaussian process classifier model or a neural network.

23. The method according to claim 20, further comprising reducing the dimension of the feature vector before inputting into the at least one machine learning model.

24. The method according to claim 23, wherein a principal component analysis algorithm or an autoencoder algorithm is used to reduce the dimension.

25. The method according to claim 24, wherein the principal component analysis algorithm or the autoencoder is adapted to process a dataset with missing data values or different bandwidths.

26. The method according to claim 20, wherein the at least one machine learning model is trained using features and labels of a training set of one or more measured values at each of a plurality of training surfaces.

27. A memory system and A processor system operably connected to the memory system, and A database system stored in the memory system, the database system including topographic data associated with one or more measured values of each of a plurality of surfaces, the topographic data including a statistical characterization of the distribution of surface height or one or more derivatives of h for at least one of the one or more measured values, the one or more derivatives being selected from the group consisting of zero - order and higher - order derivatives, and determined using a scaling coefficient η, which is multiplied by the smallest possible distance scale provided by the at least one of the one or more measured values, at each of a plurality of distance scales in real space, a database system; An algorithm stored in the memory system and executable via the processor system, the algorithm including at least one machine learning model trained using a training set of the topographic data using features and labels of the training set of the topographic data, an algorithm comprising a system.