Atomic Force Microscope for Surface Recognition

Multidimensional images were collected through atomic force microscopy and combined with machine learning algorithms to construct a compressed database for surface classification, solving the problems of surface recognition and biological cell detection in the existing technology, and achieving efficient and accurate classification of non-invasive cancer detection.

CN113272860BActive Publication Date: 2025-07-22TRUSTEES OF TUFTS COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980080842.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-11-28
Filing Date
2019-11-07
Publication Date
2025-07-22
Estimated Expiration
2039-11-07

AI Technical Summary

Technical Problem

The prior art is difficult to effectively use multidimensional image information of atomic force microscopy to identify and classify surfaces, especially in biological cell detection, especially cancer detection without invasive procedures.

Method used

Multidimensional images are collected through atomic force microscopy and combined with machine learning algorithms to build databases for training and testing, and use decision trees or neural networks to classify surfaces, especially optimize the classification process through compression databases and reduction techniques of surface parameters.

Benefits of technology

It realizes efficient identification and classification of surfaces, especially biological cells, and can detect cancer without invasive procedures, improving the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113272860B_ABST
    Figure CN113272860B_ABST
Patent Text Reader

Abstract

A method includes: using an atomic force microscope to collect a set of images associated with a surface, and classifying the surface using a machine learning algorithm applied to the images. As a specific example, classification can be performed in a manner that depends on surface parameters derived from the images rather than directly using the images.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of U.S. Provisional Application 62 / 756,958, filed on November 7, 2018, and U.S. Provisional Application 62 / 772,327, filed on November 28, 2018, the contents of which are incorporated herein by reference. Technical field

[0003] The present invention relates to the use of atomic force microscopy and machine learning in combination with features of a surface to classify or identify the surface, and more particularly to the use of features to identify or classify biological cells. Background art

[0004] In an atomic force microscope, a probe attached to the tip of a cantilever scans the surface of a sample. In one operating mode, the probe taps the surface as it scans. When the probe scans the sample, the magnitude and direction of the force vector associated with the loading force applied by the probe to the sample can be controlled.

[0005] The deflection of the cantilever from its equilibrium position provides a signal from which a large amount of information can be extracted. For example, by keeping the loading force or the deflection of the cantilever constant, the topology of the sample can be obtained at various points on the sample. Then, the values collected at each point are organized into an array, where the rows and columns identify the position of the points in a two - dimensional coordinate system, and the values at the rows and columns represent the properties measured at that point. Thus, the resulting digital array can be regarded as a graph. This enables the creation of a graph of the sample, where each point on the graph indicates some property of the sample surface at that point. In some examples, the property is the surface height above or below a certain reference plane.

[0006] However, it is not only the image of the surface height that can be recorded during scanning. The deflection of the cantilever can be used to collect multiple images of the sample surface, where each image is a graph of a different property of the surface. Some examples of these properties include the adhesion force between the probe and the surface, the stiffness of the surface, and the visco - elastic energy loss. Summary of the invention

[0007] The present invention provides a method for using multi - dimensional images obtained by an atomic force microscope to identify a surface and using the information from those images to classify the surface into one of several categories. According to the present invention, multi - dimensional images of a surface can be obtained, where two dimensions correspond to spatial dimensions, and additional dimensions correspond to different physical and spatial properties present at the coordinates identified by the two spatial dimensions. In some embodiments, the dimensions are lateral dimensions.

[0008] The problem that arises is how to select and use these different physical and spatial properties to identify and classify surfaces. According to the present invention, the properties for the identification and classification of surfaces are not predetermined. They are calculated based on the results of machine learning applied to an image database and their corresponding classes. They are learned. Specifically, they are learned by machine learning.

[0009] Among the embodiments of the present invention are those that include using an atomic force microscope to acquire different graphs corresponding to different properties of a surface, and using a combination of these graphs or parameters derived from those graphs to identify or classify a sample surface. Such a method includes: recording atomic force microscope images of examples of surfaces belonging to well-defined classes; forming a database in which such atomic force microscope graphs are associated with the classes to which they belong; using the atomic force microscope graphs thus obtained and their combinations to learn how to classify surfaces by splitting the database into training data and test data; wherein, for example by constructing a decision tree or a neural network or a combination thereof, the training data is used to learn how to classify; and using the test data to verify that the classification thus learned is sufficiently effective to exceed a given effectiveness threshold.

[0010] Another embodiment includes reducing the graphs provided by the atomic force microscope to a set of surface parameters, the values in the set of surface parameters being defined by a mathematical function or algorithm that uses these properties as its input. In a preferred practice, each graph or image yields a surface parameter, which can then be used together with other surface parameters to classify or identify the surface. In such an embodiment, there is a classifier that classifies based on these surface parameters. However, the classifier itself is not predetermined. The classifier is learned by a machine learning program as described above.

[0011] The method is independent of the nature of the surface. For example, the method can be used to classify the surface of a painting or currency or a security document (such as a birth certificate or passport) in order to detect forgeries. But the same method can also be used to classify the surface of a living cell or other part in order to identify various diseases. For example, various cancers have cells with specific surface characteristics. Thus, the method can be used to detect various cancers.

[0012] The difficulty that arises is actually obtaining the cells to be examined. In some cases, invasive procedures are required. However, there are certain types of cells that naturally shed from the body or can be extracted from the body with only minimal invasive force. An example is the example of gently scraping the surface of the cervix in a Pap smear test. Among the cells that naturally shed, there are cells from the urinary tract (including the bladder). Thus, the method can be used to examine these cells and detect bladder cancer without invasive and expensive procedures such as cystoscopy.

[0013] The present invention features the use of an atomic force microscope, which can generate a multi-dimensional array of physical properties, for example, when using the sub-resonant tapping mode. In some practices, the set of images acquired includes: performing a nano-scale resolution scan of the surface of cells that have been collected from a body fluid using an atomic force microscope in a mode, and providing the data obtained from the atomic force microscope scanning procedure to a machine learning system; the machine learning system provides an indication of the probability that the sample is from a patient with cancer (hereinafter referred to as "cancer patient"). The method generally applies to classifying cells based on the surface properties of the cells.

[0014] Although described in the context of bladder cancer, the methods and systems disclosed herein are applicable to detecting other cancers, where cells or body fluids can be used for analysis without invasive biopsy. Examples include: upper urinary tract cancer, urethral cancer, colorectal cancer and other gastrointestinal cancers; cervical cancer; respiratory and digestive tract cancers; and other cancers with similar properties.

[0015] In addition, the methods described herein are applicable to detecting cell abnormalities other than cancer and monitoring the response of cells to various drugs. Additionally, the methods described herein can be used to classify and identify any type of surface, whether derived from a living body or a non-living substance. All that is necessary for this is that the surface is a surface that is easily scanned by an atomic force microscope.

[0016] For example, the methods described herein can be used to detect forgeries, including forgeries of currency, stock certificates, identity documents or artworks (such as paintings).

[0017] In one aspect, the present invention features: using an atomic force microscope to acquire a set of images of each of a plurality of cells obtained from a patient; processing these images to obtain a surface parameter map; and using a machine learning algorithm applied to the images to classify the cells as being derived from a cancer patient or a non-cancer patient.

[0018] In these embodiments, those are the embodiments in which the microscope is used in the sub-resonant tapping mode. In other embodiments, the microscope is used in the ringing mode.

[0019] In another aspect, the present invention features: using an atomic force microscope to acquire a set of images associated with a surface; processing the images to obtain a surface parameter map; and using a machine learning algorithm applied to the images to classify the surface.

[0020] In these practices, it includes selecting the surface of bladder cells as the surface and classifying the surface as the surface of cells derived from a cancer patient or a non-cancer patient.

[0021] In another aspect, the invention features a method that includes: using an atomic force microscope to collect a set of images associated with a surface; combining the images; and classifying the surface using a machine learning method applied to the combined images.

[0022] The method cannot be performed in the human brain with or without pencil and paper because it requires performing an atomic force microscope, and because the human brain is not a machine, the human brain cannot perform a machine learning method. The method is also performed in a non-abstract manner to achieve a technical effect, namely, classifying a surface based on the technical attributes of the surface. A description of how to perform the method in an abstract and / or non-technical manner has been purposefully omitted to avoid misinterpreting the claims as covering anything other than non-abstract and technical implementations.

[0023] In some practices, the images are images of cells. These practices also include automatically detecting that an image of a cell has an artifact and excluding the image from being used for classifying the surface, as well as practices that include dividing an image of a sample into partitions, obtaining surface parameters for each partition, and defining the surface parameters of the cell as the median of the surface parameters for each partition.

[0024] Some practices also include: processing the images to obtain surface parameters and classifying the surface using machine learning at least in part based on the surface parameters. In these practices, it also includes defining a subset of the surface parameters and generating a database based on the subset. In such practices, defining a subset of the surface parameters includes: determining the correlations between the surface parameters, comparing the correlations with a threshold to identify a set of relevant parameters, and including a subset of the set of relevant parameters in the subset of the surface parameters. Additionally, in these practices, it also includes defining a subset of the surface parameters and generating a database based on the subset. In these practices, defining a subset of the surface parameters includes: determining a correlation matrix of the surface parameters; where determining the correlation matrix includes: generating a simulated surface. In these practices, there are also those practices that include defining a subset of the surface parameters and generating a database based on the subset. In these practices, defining a subset of the surface parameters includes: combining different surface parameters of the same type from the same sample.

[0025] The practices also include those practices where collecting the set of images includes: using a multi-channel atomic force microscope in a ringing mode, where each channel of the atomic force microscope provides information indicating a corresponding surface property of the surface.

[0026] Practices of the invention also include those practices where the surface selected is the surface of cells collected from a urine sample of a subject and classifying the cells as indicating cancer or not indicating cancer.

[0027] Various ways of using a microscope are available without departing from the scope of the present invention. These include: using a multi-channel atomic force microscope, where each channel corresponds to a surface property of the surface; using the atomic force microscope in a sub-resonant tapping mode; and using the atomic force microscope to combine the acquisition of multi-channel information, each in the multi-channel corresponding to a different surface property of the surface; compressing the information provided by the channels and constructing a compressed database based on the compressed information.

[0028] In the practice of the present invention that relies on a multi-channel atomic force microscope, there are further those practices that include forming a first database based on the information provided by the channels and constructing the compressed database in any of a variety of ways. In these practices, the first database is projected onto a subspace whose dimension is lower than the dimension of the first database, and this projection defines the compressed database, where the dimension of the compressed database is less than the dimension of the first database. In these practices, there is also included a compressed database from the first database, which has fewer metrics than the first database. This can be performed, for example, by performing tensor addition to generate a tensor sum and using the tensor sum to form the compressed database, where the tensor sum combines the information from the first database along one or more segments corresponding to one or more metrics of the first database.

[0029] In some practices of the present invention, deriving the compressed database from the first database includes: defining a subset of values from the first database, where each of the values represents a corresponding element in the first database; deriving a compressed value from the values in the subset of values; and representing the corresponding elements from the first database with the compressed value; where deriving the compressed value includes: summing the values in the subset of values. This summing can be performed in various ways, including: by performing tensor addition to generate a tensor sum and using the tensor sum to form the compressed database; where the tensor sum combines the values from the first database along one or more segments corresponding to the corresponding metrics of the first database.

[0030] The practice of the present invention also includes the practice where the compressed database is derived from the first database by: defining a subset of values from the first database, where each of the values represents a corresponding element in the first database; deriving a compressed value from the values in the subset of values; and representing the corresponding elements from the first database with the compressed value; where deriving the compressed value includes: taking the average of the values in the subset of values, for example, by obtaining an arithmetic mean or a geometric mean.

[0031] In the practice of the present invention, there are also the following practices. Among them, exporting a compressed database from a first database includes: defining a subset of values from the first database, where each value represents a corresponding element in the first database; deriving a compressed value from the values in the subset of values; and representing the corresponding elements from the first database with the compressed value; wherein the compressed value is one of the maximum or minimum values of the values in the subset of values.

[0032] In other embodiments, exporting a compressed database from the first database includes: defining a subset of values from the first database, where each value represents a corresponding element in the first database; deriving a compressed value from the values in the subset of values; and representing the corresponding elements from the first database with the compressed value; wherein deriving the compressed value includes: passing the information from the first database through a surface parameter extractor to obtain a set of surface parameters. In these practices, there are practices including normalizing the surface parameters representing the set of surface parameters to be independent of the surface area of the image from which the surface parameters are derived, and practices including dividing the surface parameters by another parameter of the same dimension.

[0033] Other practices include: automatically detecting that an image of a sample has artifacts, and automatically excluding the image from being used for classifying the surface.

[0034] Still some other practices include: dividing an image of a sample into partitions, obtaining surface parameters for each partition; and defining the surface parameters of the cell as the median of the surface parameters for each partition.

[0035] Some practices of the present invention include: processing an image to obtain surface parameters, and using machine learning to classify the surface at least partially based on the surface parameters and parameters externally derived. In these practices, the surface is the surface of a body that has been derived from the collected samples, and at least one of the samples is a bodyless sample, where a bodyless sample means it has no body. In these practices, the method further includes: selecting externally derived parameters to include data indicating the absence of a body in the bodyless sample. In the practices including bodyless samples, there are those practices including assigning artificial surface parameters to the bodyless samples. In some practices, the surface is the surface of a cell derived from a sample obtained from a patient. In these practices, there are included: selecting externally derived parameters to include data indicating the probability that the patient has a specific disease. Examples of such data indicating probability include: the age of the patient, the smoking habit of the patient, and the family history of the patient.

[0036] Various machine learning methods can be used. These include random forest methods, extremely randomized forest methods, gradient boosting tree methods, using neural networks, decision tree methods, and combinations thereof.

[0037] In some embodiments, the surface is the surface of a first plurality of cells from a patient, a second plurality of cells have been classified as being from a patient with cancer, and a third plurality of cells have been classified as being from a patient without cancer. These methods include diagnosing a patient with cancer if the ratio of the second plurality to the first plurality exceeds a predetermined threshold.

[0038] In some practices, an atomic force microscope includes a cantilever and a probe disposed at the end of the cantilever. The cantilever has a resonant frequency. In these practices, using an atomic force microscope includes oscillating the distance between the probe and the surface at a frequency less than the resonant frequency.

[0039] In some practices, using an atomic force microscope includes using a microscope that has been configured to output information for a plurality of channels corresponding to different physical properties of a sample surface.

[0040] Other practices include processing an image to obtain surface parameters and using machine learning to classify the surface at least in part based on the surface parameters and parameters externally derived. In these embodiments, the surface is the surface of cells derived from a sample obtained from a patient, and at least one of the samples is a cell-free sample that does not have cells from the patient. In such practices, the method further includes selecting the externally derived parameters to include data indicating the absence of cells in the cell-free sample. In these practices, those practices also include assigning artificial surface parameters to the cell-free sample.

[0041] In another aspect, the invention features a device including an atomic force microscope and a processing system. The atomic force microscope acquires an image associated with a surface. The processing system receives a signal representing the image from the atomic force microscope and combines the images. The processing system includes a machine learning module and a classifier that classifies an unknown sample after having learned a basis for classification from the machine learning module.

[0042] In some embodiments, the processing system is configured to process the image to obtain surface parameters and use the machine learning module to classify the surface at least in part based on the surface parameters. In these embodiments, the atomic force microscope includes a multi-channel atomic force microscope, and each channel of the multi-channel atomic force microscope corresponds to a surface property of the surface. In these embodiments, a compressor is also included that compresses the information provided by the channels and constructs a compressed database based on the compressed information.

[0043] Embodiments including a compressed database also include those embodiments in which the classifier classifies an unknown sample based on the compressed database.

[0044] Various compressors can be used to construct a compressed database. Among them, the compressor constructs the compressed database by projecting the first database onto a subspace with a dimension lower than that of the first database. This projection defines the compressed database, and the compressed database has a dimension smaller than that of the first database.

[0045] As used herein, "atomic force microscope", "AFM", "scanning probe microscope", and "SPM" are considered synonymous.

[0046] The only methods described in this specification are non-abstract methods. Therefore, the claims can only relate to non-abstract embodiments. As used herein, "non-abstract" is considered to mean meeting the requirements of 35 USC 101 as of the filing date of this application.

[0047] These and other features of the present invention will become apparent from the following detailed description and the accompanying drawings, wherein: Brief Description of the Drawings

[0048] Figure 1 A simplified diagram showing an example of an atomic force microscope;

[0049] Figure 2 Showing additional details of the processing system from Figure 1 ;

[0050] Figure 3 Showing a diagnostic method performed by the atomic force microscope and the processing system shown in Figure 1 and Figure 2 ;

[0051] Figure 4 Showing a view through an optical microscope built into the atomic force microscope shown in Figure 1 ;

[0052] Figure 5 Showing a Figure 1 bladder cell captured by the atomic force microscope;

[0053] Figure 6 Showing Figure 2 details of the interaction between the database and the machine learning module in the processing system;

[0054] Figure 7 Showing details of compressing an initial large database into a compressed database with a smaller dimension, and showing Figure 2 details of the interaction between the compressed database and the machine learning module in the processing system;

[0055] Figure 8 Showing an example of a simulated surface used in combination with evaluating the correlation between different surface parameters;

[0056] Figure 9 Shows a histogram of the importance coefficients of two surface parameters;

[0057] Figure 10 Shows a binary tree;

[0058] Figure 11 Shows a machine learning method suitable for the data structure required for classification;

[0059] Figure 12 Shows an example representation of artifacts that may be caused by contamination of the cell surface;

[0060] Figure 13 Shows the dependence of the number of surface parameters on the correlation threshold;

[0061] Figure 14 Shows the hierarchy of the importance of surface parameters of height and adhesion properties calculated within the random forest method;

[0062] Figure 15 Shows the accuracy of different numbers of surface parameters and different data allocations in the training and test databases, as calculated using the random forest method for the combined channels of height and adhesion;

[0063] Figure 16 Shows the receiver operating characteristics using the random forest method for the combined channels of height and adhesion;

[0064] Figure 17 Shows a graph similar to Figure 16 shown in, but artificial data is used to confirm the reliability of the program for generating the data in Figure 16 ;

[0065] Figure 18 Shows the area under the receiver operating characteristics of Figure 17 ;

[0066] Figure 19 Shows the accuracy of different numbers of surface parameters and different ways of allocating data between training data and test data, using the random forest method for the combined channels of height and adhesion when using 5 cells per patient and 2 cells that need to be identified as coming from patients with cancer (N = 5, M = 2);

[0067] Figure 20 Shows the receiver operating characteristics calculated using the random forest method for the combined channels of height and adhesion when using 5 cells per patient and 2 cells that need to be identified as coming from patients with cancer (N = 5, M = 2); and

[0068] Figure 21It is a table showing the statistics of the confusion matrix associated with cancer diagnosis for two separate channels, where one channel is for height and the other channel is for adhesion force. Detailed Description

[0069] Figure 1 An atomic force microscope 8 with a scanner 10 is shown. The scanner 10 supports a cantilever 12 to which a probe 14 is attached. Thus, the probe 14 protrudes from the scanner 10 cantilever. The scanner 10 moves the probe 14 along a scan direction parallel to the reference plane of the sample surface 16. In doing so, the scanner 10 scans an area of the sample surface 16. When the scanner is moving the probe 14 in the scan direction, the scanner also moves the probe 14 in a vertical direction perpendicular to the reference plane of the sample surface 16. This causes the distance from the probe 14 to the surface 16 to change.

[0070] The probe 14 is typically coupled to a reflective portion of the cantilever 12. This reflective portion reflects an illumination beam 20 provided by a laser 22. The reflective portion of the cantilever 12 will be referred to herein as a mirror 18. The reflected beam 24 travels from the mirror 18 to a photodetector 26, and the output of the photodetector 26 is connected to a processor 28. In some embodiments, the processor 28 includes FPGA electronics to allow for real-time calculation of surface parameters based on physical or geometric properties of the surface.

[0071] The movement of the probe 14 is converted into the movement of the mirror 18, which subsequently causes different components of the photodetector 26 to be illuminated by the reflected beam 24. This results in a probe signal 30 indicative of the probe movement. The processor 28 calculates certain surface parameters based on the probe signal 30 using the method described below and outputs the result 33 to a storage medium 32. These results 33 include data representing any of the surface parameters described herein.

[0072] The scanner 10 is connected to the processor 28 and provides a scanner signal 34 to the processor 28 indicating the position of the scanner. This scanner signal 34 can also be used to calculate surface parameters.

[0073] Figure 2 The processing system 28 is shown in detail. The processing system 28 features a power supply 58 having an AC power source 60 connected to an inverter 62. The power supply 58 provides power for operating the various components described below. The processing system also includes a heat sink 64.

[0074] In a preferred embodiment, the processing system 28 further includes a user interface 66 to enable a person to control its operation.

[0075] The processing system 28 further includes first and second A / D converters 68, 70 which are used to receive probe signals and scanner signals and place them on the bus 72. The program storage device section 74, the working memory 76 and the CPU registers 78 are also connected to the bus 72. The CPU 80 for executing the instructions 75 from the program storage device 74 is connected to both the register 78 and the ALU 82. The non-transitory computer-readable medium stores these instructions 75. When the instructions 75 are executed, the instructions 75 cause the processing system 28 to calculate any of the foregoing parameters based on the inputs received through the first and second A / D converters 68, 70.

[0076] The processing system 28 further includes a machine learning module 84 and a database 86 which includes training data 87 and test data 89, as Figure 6 shown. The machine learning module 84 uses the training data 87 and the test data 89 to implement the methods described herein.

[0077] A specific example of the processing system 28 may include an FPGA device which includes circuitry configured to determine the property values and / or surface parameters of the above-mentioned imaging services.

[0078] Figure 3 The process of using the atomic force microscope 8 to acquire an image and provide the image to the machine learning module 84 for using the image to characterize the sample is shown. Figure 3 The process shown in includes collecting a urine sample 88 from a patient and preparing the cells 90 that have exfoliated into the urine sample 88. After they have been scanned, the atomic force microscope 8 provides an image of the bladder cells 90 for storage in the database 86.

[0079] Each image is an array, where each element of the array represents a property of the surface 16. The position in the array corresponds to the spatial position on the sample surface 16. Thus, the image defines a graph corresponding to the property. Such a graph shows the property values at different positions on the sample surface 16 in much the same way as a soil map shows different soil properties at different locations on the earth's surface. This property will be referred to as a "mapped property".

[0080] In some cases, the mapped property is a physical property. In other cases, the property is a geometric property. An example of a geometric property is the height of the surface 16. Examples of physical properties include surface adhesion, self-stiffness, and energy loss associated with contacting the surface 16.

[0081] The multi-channel atomic force microscope 8 has the ability to map different properties simultaneously. Each mapped property corresponds to a different "channel" of the microscope 8. Thus, the image can be regarded as a multi-dimensional image array M (k), where the channel index k is an integer in the interval [1, K], where K is the number of channels.

[0082] When the multi-channel atomic force microscope 8 is used in the sub-resonance tapping mode, the multi-channel atomic force microscope 8 can map the following properties: height, adhesion force, deformation, stiffness, viscoelastic loss, feedback error. This results in six channels, each corresponding to one of the six mapped properties. When the multi-channel atomic force microscope 8 is used in the ringing mode, the atomic force microscope 8 can map one or more of the following additional properties, for example, in addition to the previous six properties: recovered adhesion force, adhesion force height, detachment height, pull-off neck height, detachment distance, detachment energy loss, dynamic creep phase shift, and zero-force height. In this example, this results in a total of 14 channels, each corresponding to one of the 14 mapped properties.

[0083] The scanner 10 defines discrete pixels on the reference plane. The probe 14 of the microscope makes measurements at each pixel. For convenience, the pixels on the plane can be defined by Cartesian coordinates (x i , y j ). The value of the k-th channel measured at this pixel is z i,j (k) . Taking this into account, the image array representing the graph or image of the k-th channel can be formally represented as:

[0084] M (k) = {x i , y j , z i,j (k)} (1)

[0085] where "i" and "j" are integers in the intervals [1, Ni] and [1, Nj] respectively, where Ni and Nj are the number of pixels available for recording the image in the x and y directions respectively. The values of Ni and Nj can be different. However, the methods described herein do not significantly depend on such differences. Therefore, for the purpose of discussion, Ni = Nj = N.

[0086] The number of elements in the sample image array will be the product of the number of channels and the number of pixels. For a relatively homogeneous surface 16, only one area of the surface 16 needs to be scanned. However, for a more heterogeneous surface 16, it is preferred to scan more than one area on the surface 16. By analogy, if one wishes to examine the surface of the water in a port, most likely only one area needs to be scanned because the other areas may be similar anyway. On the other hand, if one wishes to examine the surface of the city served by the port, it would be wise to scan multiple areas.

[0087] With this in mind, the array acquires another index to identify the specific area being scanned. This increases the dimensionality of the array. Thus, the image array is represented in the form:

[0088] M (k;s) ={x i (s) ,y j (s) ,z i,j (k;s)}(2)

[0089] where the scan area index s is an integer in the range [1, S] that identifies a specific scan area within the sample. Note that this increases the number of elements in the image array for a specific sample by a multiple equal to the number of scan areas.

[0090] Preferably, the number of such scan areas is large enough to represent the entire sample. One way to converge on an appropriate number of scan areas is to compare the deviation distributions between two such scan areas. If increasing the number of scan areas does not change this in a statistically significant way, then the number of scan areas may be sufficient to represent the entire surface. Another way is to divide the test time considered reasonable by the amount of time required to scan each scan area and use the quotient as the number of areas.

[0091] In some cases, it is useful to divide each of the scan areas into sub - regions. For the case where there are P such sub - regions in each scan area, the array can be defined as:

[0092] M (k;s;p) ={x i (s;p) ,y j (s;p) ,z i,j (k;s;p)}(2a)

[0093] where the sub - region index p is an integer in the range [1, P]. In the case of a square scan area, it is convenient to divide the square into four square sub - regions, thus setting P equal to 4.

[0094] The ability to divide the scan areas into sub - regions provides a useful way to exclude image artifacts. This is particularly important for examining biological cells 90. This is because the process of preparing the cells 90 for examination can easily introduce artifacts. These artifacts should be excluded from any analysis. This allows one sub - region to be compared with the others to identify which sub - region (if any) deviates significantly and should be excluded.

[0095] On the other hand, the addition of the new index further increases the dimensionality of the array.

[0096] In order to identify the class to which a sample belongs based on an image array M acquired by an atomic force microscope 8 (k,s) the machine learning module 84 partly relies on constructing a suitable database 86 which includes images of surfaces that are known a priori to belong to a specific class C (l) This database 86 can be formally represented as follows:

[0097] D n (l;k;s;p) ={M n (k;s;p) ,C (l)} (2b)

[0098] where k is the channel index representing an attribute or a channel, s is the scan area index identifying a specific scan area, p is the partition index representing a specific partition of the s-th scan area, n is the sample index identifying a specific sample, and l is the class index identifying a specific class from a set of L classes. Thus, the total size of the array is the product of the number of classes, the number of samples, the number of scan areas, the number of partitions per scan area, and the number of channels.

[0099] Figure 3 The diagnostic method 10 is shown, which is characterized by using an atomic force microscope 8 and a machine learning module 84. The atomic force microscope 8 operates using sub-resonant tapping and uses the machine learning module 84 to examine the surface of biological cells 90 recovered from a urine sample 88 in an effort to classify a patient into one of the following two classes: having cancer and not having cancer. Since there are two classes, L = 2.

[0100] Preferred practices include using centrifugation, gravity sedimentation, or filtration to collect the cells 90, and subsequently fixing and freeze-drying or subcritically drying the cells 90.

[0101] In the example shown, the atomic force microscope 8 operates using both the sub-resonant tapping mode and the ringing mode. The sub-resonant tapping mode is, for example, PeakForce QMN implemented by Bruker, Inc., and the ringing mode is, for example, implemented by NanoScience Solutions, LLC. Both modes allow recording of height and adhesion force channels. However, the ringing mode is a substantially faster mode for image collection. As described above, these modes allow many channels to be recorded simultaneously. However, only two channels are used in the experiments described herein.

[0102] Figure 4 The cantilever 12 of the atomic force microscope and the cells 90 obtained and prepared from a patient as described above are shown. This view is taken by an optical microscope coupled to the atomic force microscope 8.

[0103] Figure 5 Shows first and second pairs of images 92, 94. The first pair of images 92 shows an image of cells 90 from a cancer-free patient. The second pair of images 94 shows an image of cells 90 from a patient with cancer. The images shown are of a square scan area with a side length of 10 micrometers and a resolution of 512 pixels in two dimensions. When scanned in a sub-resonant tapping mode (such as PeakForce QMN mode), the scan speed is 0.1 Hz, and when scanned in a ringing mode, the scan speed is 0.4 Hz. The peak force during scanning is 5 nano-newtons (nano-newton).

[0104] Now referring to Figure 6 , the machine learning module 84 trains the candidate classifier 100 based on the database 86. A specific machine learning method can be selected from a family of machine learning methods, for example, decision trees, neural networks, or a combination thereof.

[0105] Figure 6 and Figure 7 The method shown in starts with splitting the database 86 into training data 87 and test data 89. This raises the following questions: how much of the data in the database 86 should go into the training data 87, and how much of the data in the database 86 should go into the test data 89.

[0106] In some embodiments, 50% of the data in the database 86 goes into the training data 87, while the remaining 50% of the data goes into the test data 89. In other embodiments, 60% of the data in the database 86 goes into the training data 87, while the remaining 40% of the data goes into the test data 89. In other embodiments, 70% of the data in the database 86 goes into the training data 87, while the remaining 30% of the data goes into the test data 89. In other embodiments, 80% of the data in the database 86 goes into the training data 87, while the remaining 20% of the data goes into the test data 89. The candidate classifier 100 should ultimately be independent of the ratio used in the split.

[0107] In Figure 3 the example shown, 10 bladder cells 90 are collected from each patient. Standard clinical methods (including invasive biopsies and histopathology) are used to identify the presence of cancer. These methods are reliable enough so that two classes can be well defined. Thus, Figure 6 the database 86 shown can be represented as:

[0108]

[0109] where N data1 is the number of patients in the first class, N data2is the number of patients in the second category, and s (which is an integer between 1 and 10, inclusive) identifies a particular one of the 10 cells collected from a single patient. N data1 and N data2 need not be equal.

[0110] When splitting database 86 between training data 87 and test data 89, it is important to avoid partitioning image arrays of different scan regions from the same sample {M (k;1;p) ,M (k;2;p) ..M (k;S;p)} between the training and test data 87, 89. Violating this rule would result in training and testing on the same sample. This would artificially enhance the effectiveness of the classifier in a way that may not be reproducible when applying classifier 100 to independent new samples.

[0111] Machine learning module 84 uses training data 87 to construct candidate classifier 100. Depending on the type of classifier 100, training data 87 can be a learning tree, decision tree, bootstrap of trees, neural network, or a combination thereof. Classifier 100, hereinafter referred to as "AI", outputs the probability that a particular sample n belongs to a particular category l:

[0112] Prob n (k;s;p)(l) = AI(M n (k;s;p) |C (l) ) (3a)

[0113] where Prob n (k;s;p)(l) is the probability that the image or channel defined by M n (k;s;p) belongs to category C (l) .

[0114] After having been constructed, verification module 102 uses test data 89 to verify that candidate classifier 100 is actually sufficiently effective. In the embodiments described herein, verification module 102 evaluates effectiveness at least in part based on the receiver operating characteristic and the confusion matrix. The robustness of candidate classifier 100 is verified by repeatedly randomly splitting database 86 to generate different test data 89 and training data 87 and then performing the classification procedure to see if there are any differences.

[0115] If it turns out that candidate classifier 100 is not sufficiently effective, machine learning module 84 changes the parameters of the training process and generates a new candidate classifier 100. This loop continues until machine learning module 84 finally provides a candidate classifier 100 that meets the desired effectiveness threshold.

[0116] The computational load that occurs when there is more than one probability value associated with a sample n somewhat hinders the process of constructing a suitable classifier 100. In fact, due to the multi-dimensional nature of the image array, for any one sample, there will be K·S·P probabilities Prob n (k;s;p)(l) to be processed. For such a large database, the required computational load would be impractically high.

[0117] Another bottleneck in processing such large data arrays is the large number of samples required to provide reasonable training for the classifier. When constructing a decision tree, a rule of thumb requires that the number of samples be at least six times the dimension of the database. Since atomic force microscopy is a relatively slow technique, it would be impractical to obtain enough samples to construct any reasonable classifier.

[0118] As Figure 7 shown, the compressor 104 solves the aforementioned difficulties. The compressor 104 compresses the information provided by a particular channel into the space of surface parameters that embody the information about that channel. The compressor 104 receives the database 86 and generates a compressed database 106. In effect, this is equivalent to projecting a multi-dimensional matrix in a fairly high-dimensional space onto a matrix with a much smaller dimension.

[0119] The compressor 104 performs any one of a variety of database reduction procedures. Among these procedures are those that incorporate one or more of the database reduction procedures described herein. These jointly derive surface parameters from the data set that embody at least some of the information embodied in that set.

[0120] In some practices, the compressor 104 performs a first database reduction procedure. This first database reduction procedure relies on the following observation: Each image is ultimately an array that can be combined with other such arrays such that an object is produced that retains sufficient aspects of the information of the arrays that combined to produce that object to be useful in classifying samples. For example, tensor addition can be used to combine a set of images M n (k;s;p) along segments corresponding to one of its indices.

[0121] In a particular embodiment, the segment corresponds to the exponent k. In this case, the tensor sum of the images is given by:

[0122]

[0123] Thus, each element of the compressed database 106 for machine learning becomes the following:

[0124]

[0125] This specific example reduces the dimension of the database 86 by a factor of K. Thus, the classifier 100 defines probabilities as follows:

[0126]

[0127] A similar procedure can also be performed for the remaining metrics. Finally,

[0128]

[0129] where represents the tensor summation over the metrics k, s, p.

[0130] In other practices, the compressor 104 alternatively performs a second database reduction procedure. This second database reduction procedure depends on geometric or algebraic averaging of each or combinations of the exponents k, s, p respectively. Examples of specific ways to perform the second procedure include the following averaging procedures over all metrics k, s, p:

[0131]

[0132]

[0133]

[0134]

[0135] In other practices, the compressor 104 alternatively performs a third database reduction procedure. This third database reduction procedure depends on assigning the highest or lowest probabilities of an entire series to a particular exponent. For example, considering the scan area exponent s, one of the following relationships can be used:

[0136]

[0137]

[0138] Finally, if all exponents are reduced in this way

[0139]

[0140]

[0141] In some practices, the compressor 104 reduces the database D by passing each image through the surface parameter extractor Am n (l;s) to obtain a set of surface parameters P nm (k,s) . This can be represented in the following form:

[0142] Pnm (k,s) = A m {M n (k;s;p)} (4)

[0143] Among them, the surface parameter exponent m is an integer in [1, M], the channel exponent k identifies whether the figure represents height, adhesion, stiffness, or some other physical or geometric parameter, the sample exponent n identifies the sample, the scan area exponent s identifies a specific scan area in the sample, and the partition exponent p identifies a specific partition within the scan area. This program provides a way to represent the multi-dimensional tensor M n (k;s;p) as a compact form of the surface parameter vector P nm (k,s,p) .

[0144] The surface parameter vector includes sufficient residual information related to the channel, and this surface parameter vector is derived from the channel for use as the basis for classification. However, it is much smaller than the image provided by the channel. In this way, the classification program relying on the surface parameter vector maintains a much lower computational load without a corresponding loss of accuracy.

[0145] Various surface parameters can be extracted from the channel. These parameters include average roughness, root mean square, surface skewness, surface kurtosis, peak-to-peak value, ten-point height, maximum valley depth, maximum peak height, average value, average vertex curvature, texture index, root mean square gradient, area root mean square slope, surface area ratio, projected area, surface area, surface bearing index, core area liquid retention index, valley area liquid retention index, reduced vertex height, core roughness depth, reduced valley depth, 1 - h% height interval of the bearing curve, vertex density, texture direction, texture direction index, main radial wavelength, radial wave index, average half wavelength, fractal dimension, 20% correlation length, 37% correlation length, 20% texture aspect ratio, and 37% texture aspect ratio.

[0146] The list of surface parameters can be further extended by introducing algorithms or mathematical formulas. For example, the surface parameters can be normalized to the surface area of the image by dividing each parameter by a function of the surface area, and this surface area can be different for different cells.

[0147] The examples described in this article rely on three surface parameters: valley area liquid retention index ("Svi"), surface area ratio ("Sdr"), and surface area ("S3A").

[0148] The valley area liquid retention index is a surface parameter indicating the presence of large voids in the valley area. It is defined as follows:

[0149]

[0150] Where N is the number of pixels in the x direction, M is the number of pixels in the y direction, V(hx) is the void area above the bearing area ratio curve and below the horizontal line hx, and Sq is the root mean square (RMS), which is defined by the following expression:

[0151]

[0152] The surface area ratio ("Sdr") is a surface parameter that expresses the increment of the interface surface area relative to the area of the projected x, y plane. This surface parameter is defined by the following formula:

[0153]

[0154] Where N is the number of pixels in the x direction, and M is the number of pixels in the y direction.

[0155] The surface area ("S3A") is defined as follows:

[0156]

[0157] To calculate each of the above three surface parameters based on the images provided by the atomic force microscope 8, each image of the cell is first split into 4 partitions. In this case, the 4 partitions are the quadrants of a square with a side of 5 microns. Thus, each cell produces 4 sets of surface parameters, one for each quadrant.

[0158] The presence of artifacts in the cell can be addressed in any of three different ways.

[0159] The first way is to have an operator examine the artifacts of the cell and exclude any cell having one or more such artifacts based on further processing. This requires human intervention to identify the artifacts.

[0160] The second way is to provide an artifact recognition module that is capable of identifying the artifacts and automatically excluding the cells containing the artifacts. This makes the program less operator-dependent.

[0161] The third way is to use the median instead of the mean of the parameters of each cell. When using the median instead of the mean, the results described herein are nearly unchanged.

[0162] Using the same example with only two classes, the compressed database 106 would be as follows:

[0163]

[0164] In other embodiments, additional parameters can be assigned to aid in discrimination between different classes, even if the additional parameters are not directly related to the images of the atomic force microscope.

[0165] For example, when attempting to detect bladder cancer, one or more samples of urine sample 88 are very likely to not have any cells 90. A convenient way to account for such a result is to add a new "cell-free" parameter, which is either true or false. To avoid having to change the data structure to accommodate such a parameter, samples with a "cell-free" set to "true" receive an artificial value of a surface parameter that is chosen to avoid distorting the statistical results.

[0166] As another example, there are other factors that are not related to the surface parameters but are still relevant to classification. These factors include characteristics of the patient, such as age, smoking, and family history, all of which may be related to the probability that the patient has bladder cancer. These parameters can be included in a manner similar to the "cell-free" parameter, thus avoiding having to modify the data structure.

[0167] There are also other ways to use the surface parameters to reduce the size of the database 86.

[0168] One such procedure is to exclude surface parameters that are sufficiently related to each other. Some surface parameters strongly depend on various other surface parameters. So, by including surface parameters that are related to each other, little additional information is provided. These redundant surface parameters can be removed with little loss.

[0169] One way to find the correlation matrix between surface parameters is to generate simulated surfaces, examples of which are shown in Figure 8 . The various sample surfaces imaged with the atomic force microscope 8 can also be used to identify the correlations between different surface parameters.

[0170] The machine learning module 84 is independent of the nature of its input. Thus, although it is shown operating on the image array, it is fully capable of operating on the surface parameter vector. Thus, the same machine learning module 84 can be used to determine the probability that a particular surface parameter vector belongs to a particular class, i.e., for evaluating Prob n (k;s;p)(l) = AI(P n (k;s;p) |C (l) ).

[0171] Thus, after reducing the multi-dimensional image array M n (k;s;p) to the surface parameter vector P nm (k;s;p) it is possible to use the surface parameter vector P nm (k;s;p) to replace the multi-dimensional image array M n(k;s;p) , and then the machine learning module 84 learns which surface parameters are important for classification and how to use them to classify cells.

[0172] Since some surface parameters are related to each other, the dimensionality can be further reduced. This can be done without tensor summation. Instead, this reduction is performed by directly manipulating the same parameters from different images.

[0173] In addition to the method that depends on the database reduction procedures identified above as (3-1) to (3-9), it is also possible to use a classifier 100 that combines different surface parameters of the same type from the same sample. Formally, this type of classifier 100 can be represented as:

[0174] Prob n (l) = AI(P n |C (l) ) (10)

[0175] where P n = F(P nm (k;s;p) ), and where F(P nm (k;s;p) ) is a combination of different surface parameters identified by the surface parameter exponent m and belonging to the sample identified by the sample exponent n.

[0176] The relevant classifier 100 is a classifier that combines different surface parameters of the same type m of the same sample n from images of the same attribute. Such a classifier 100 can be represented as:

[0177] Prob n (k)(l) = AI(P nm (k) |C (l) ) (11)

[0178] where P nm (k) = F(P nm (k;s;p) ), and F(P nm (k;s;p) ) is a combination of different surface parameters identified by the same surface parameter exponent m of the sample identified by the sample exponent n and coming from the channel identified by the channel exponent k.

[0179] Another classifier 100 is a classifier that does not combine all parameters but only combines surface parameters through one exponent. One such classifier 100 assigns a surface parameter to a series of entire partitions p within the same image. Such a classifier 100 is formally expressed as:

[0180] Prob n (k;s)(l) =AI(P nm (k;s) |C (l) ) (12)

[0181] where P nm (k;s) =F(P nm (k,s;p) ), and F(P nm (k;s;p) ) is a combination of surface parameters, examples of surface parameters include parameters associated with the statistical distribution of P nm (k;s;p) on the partition exponent. Examples include the mean:

[0182]

[0183] and the median:

[0184] P nm (k,s) =median{P n (k;s;p)}for p=1…N (14)

[0185] When used in combination with the imaging detection of bladder cancer for multiple cells from each patient, classifier 100 relies on the mean or the median. However, preferably, classifier 100 relies on the median rather than the mean because the median is less sensitive to artifacts.

[0186] In the specific embodiments described herein, the machine learning module 84 implements any one of various machine learning methods. However, when faced with multiple parameters, the machine learning module 84 can easily become overtrained. Therefore, it is useful to use the three methods that are least likely to be overtrained, namely the Random Forest method, the Extremely Randomized Forest method, and the method of Gradient Boosting Trees.

[0187] The random forest method and the extremely randomized forest method are bootstrap unsupervised methods. The gradient boosting tree method is a supervised method for constructing trees. Appropriate classifier functions from the SCIKIT-LEARN Python machine learning package (version 0.17.1) are used to perform variable ranking, classifier training, and validation.

[0188] Both the random forest method and the extremely randomized forest method are based on growing many classification trees. Each classification tree predicts some classification. However, the votes of all the trees determine the final classification. The trees are grown on the training data 87. In a typical database 86, 70% of all the data is in the training data 87, while the remaining data is in the test data 89. In the experiments described in this article, the split between the training data 87 and the test data 89 is random and repeated many times to confirm that the classifier 100 is insensitive to the way the database 86 is split.

[0189] Each branch node depends on a randomly selected subset of the original surface parameters. In the method described in this article, the number of elements in the selected subset of the original surface parameters is the square root of the number of the originally provided surface parameters.

[0190] Then the learning process continues by identifying the best split of the tree branches given the randomly selected subset of surface parameters. The machine learning module 84 is based on an estimate of the classification error with the split threshold. Each parameter is assigned to a parameter zone for the most commonly occurring class with respect to the training data 87. In these practices, the machine learning module 84 defines the classification error as the part of the training data 87 in that zone that does not belong to the most common class:

[0191]

[0192] where p mk represents the proportion of the training data 87 that is in the m-th zone and also belongs to the k-th class. However, for practical applications, Equation (1) is not sensitive enough to avoid overgrowth of the tree. As a result, the machine learning module 84 relies on two other measurements: the Gini index and the cross entropy.

[0193] The Gini index (which is a measure of the variance across all K classes) is defined as follows:

[0194]

[0195] When p mkWhen all values of

[0196] Cross-entropy also provides a measure of node purity and is defined as:

[0197]

[0198] Similar to the Gini index, cross-entropy is small when all values of p mk are close to zero. This indicates a pure node.

[0199] The Gini index also provides a way to obtain an "importance coefficient" indicating the importance of each surface parameter. One such measure comes from adding up all values of the decrease in the Gini index for each variable at the tree nodes and averaging over all trees.

[0200] Figure 9 The histograms shown in

[0201] use error bars to represent the mean of the importance coefficients to show the extent to which they deviate from the mean by one standard deviation. These importance coefficients correspond to the various surface parameters that can be derived from a specific channel. Thus, the histogram in the first row represents the surface parameters that can be derived from the channel measuring the feature "height", while the surface parameters in the second row represent the surface parameters that can be derived from the channel measuring the feature "adhesion". Note that mnemonic means have been used to name the features, and all surface parameters derived from the "height" channel start with "h", and all surface parameters derived from the "adhesion" channel start with "a".

[0202] Similarly, in the second row, the panel in the first column shows the importance coefficients of those surface parameters derived from the "Adhesion" channel when the machine learning module 84 uses the random forest method; the panel in the second column shows the importance coefficients of those surface parameters derived from the "Adhesion" channel when the machine learning module 84 uses the extremely randomized forest method; and the panel in the third column shows the importance coefficients of those surface parameters derived from the "Adhesion" channel when the machine learning module 84 uses the gradient boosting tree method.

[0203] Figure 9 The histograms in [reference] provide an intelligent way to select those surface parameters that are most helpful in correctly classifying the samples. For example, if the machine learning module 84 is forced to select only two surface parameters from the channels measuring height, it would likely avoid selecting "h_Sy" and "h_Std", and instead might prefer to select "h_Ssc" and "h_Sfd".

[0204] Figure 9 The importance coefficients in [reference] were obtained between one hundred and three hundred trees. The maximum number of elements in the selected subset of the original surface parameters is the square root of the number of the originally provided surface parameters, and the Gini index provides a basis for evaluating the classification error. By comparing the histograms in the same row, it is clear that the choice of the machine learning procedure does not make a great difference to the importance of a particular surface parameter.

[0205] Figure 10 Shows an example of a binary tree from the population of one hundred to three hundred trees used in the bootstrap method. In the first split, the fourth variable "X[4]" with a split value of 15.0001 is selected. This results in a Gini index of 0.4992 and splits 73 samples into two bins with 30 samples and 43 samples respectively.

[0206] At the second - level split, looking at the left - hand node, the sixth variable "X[6]" with a split value of 14.8059 is selected, which results in a Gini index of 0.2778 and splits 30 samples (5 in class 1 and 25 in class 2) into two bins with 27 samples and 3 samples respectively. The splitting continues until the Gini index of the tree node is zero, indicating that only one of the two classes is present.

[0207] The extreme random tree method differs from the random forest method in its split selection. Unlike the case of using the random forest method, instead of using the Gini index to calculate the best parameters and split combinations, the machine learning module 84 of the extreme random tree method randomly selects each parameter value from the empirical range of parameters. To ensure that these random selections eventually converge to pure nodes with a Gini index of zero, the machine learning module 84 only selects the best split among the randomly uniform splits in the set of selected variables for the current tree selected for it.

[0208] In some practices, the machine learning module 84 implements the gradient boosting tree method. In this case, the machine learning module 84 constructs a series of trees, and each tree converges with respect to some cost function. The machine learning module 84 constructs each subsequent tree to minimize the deviation from the exact prediction, for example, by minimizing the mean squared error. In some cases, the machine learning module 84 relies on the Friedman process for this type of regression. The appropriate implementation of this regression process can be performed using the routine "TREEBOOST" implemented in the "SCIKIT-LEARN PYTHON" package.

[0209] Because the method of gradient boosting trees lacks a criterion for pure nodes, the machine learning module 84 predefines the size of the tree. Alternatively, the machine learning module 84 limits the number of individual regressions, thus limiting the maximum depth of the tree.

[0210] The difficulty is that a tree constructed with a predefined size can be easily overfitted. To minimize the impact of this difficulty, preferably, the machine learning module 84 imposes constraints on quantities such as the number of boosting iterations or weakens the iteration rate, for example, by using a dimensionless learning rate parameter. In an alternative practice, the machine learning module 84 limits the minimum number of terminal nodes or leaves on the tree.

[0211] In the implementation described herein that relies on the SCIKIT-LEARN PYTHON package, the machine learning module 84 sets the minimum number of leaves to be consistent and sets the maximum depth to 3. In the application described herein of classifying bladder cells collected from human subjects, the machine learning module 84 curbs its learning ability by deliberately choosing a very low learning rate of 0.01. The resulting slow learning program reduces the variance due to having a small number of human subjects and thus a small number of samples.

[0212] When creating the training data 87 and the test data 89, it is important to avoid dividing the set {M (k;1;p) , M (k;2;p) … M (k;S;p)} between the training data 87 and the test data 89. Figure 11 The procedure disclosed in [reference] avoids this situation.

[0213] In a particular embodiment of classifying bladder cells 90, a number of cells are provided for each patient, where the image of each cell 90 is divided into 4 partitions. A human observer visually inspects the partitions in an effort to detect artifacts, two of which can be seen in Figure 12 . If an artifact is detected in a partition, anyone examining the image will mark that partition as one to be ignored.

[0214] When dealing with many cells 90, this process can become lengthy. By using the classifier 100 shown in equation (10) and taking the median of the 4 partitions, this process can be automated. This significantly dilutes the contribution of the artifacts.

[0215] The machine learning module 84 randomly splits the database 86 such that S% of its data is in the training data 87 and 100 - S% of its data is in the test data 98. Experiments are performed with S set to 50%, 60%, and 70%. The machine learning module 84 splits the database 86 in such a way that all data from the same individual is kept either in the training data 87 or the test data 98 to avoid artificial overtraining that may result from the correlation between different cells 90 of the same individual.

[0216] Then, the machine learning module 84 causes the compressor 104 to further reduce the number of surface parameters on which the classification depends. In some practices, the compressor 104 does this by sorting the surface parameters based on the corresponding Gini index of the surface parameters within a particular channel and retaining some number M p of the optimal parameters of that channel. In some practices, the optimal parameters are selected based on their isolation ability and their low correlation with other surface parameters. For example, by changing the parameter correlation threshold, the number of surface parameters on which the classification depends can be changed.

[0217] Figure 13 shows how changing the threshold of the correlation coefficient affects the number of surface parameters selected using the random forest method, where the leftmost panel corresponds to surface parameters obtainable from the height channel and the middle panel corresponds to surface parameters obtainable from the adhesion channel. As is obvious from the change in the vertical scale, the rightmost panel represents the combination of the height channel and the adhesion channel. Although Figure 13 shown for the random forest method, other methods have similar curves.

[0218] Once the trees have been trained, they are suitable for testing their ability to correctly classify the test data 98, or alternatively, for testing their ability to classify unknown samples. The classification process involves obtaining the results of the tree votes and using that result as a basis for indicating the probability of the sample belonging to a class. The result is then compared with a classifier threshold set based on a tolerable error. The classifier threshold typically varies as part of constructing the receiver operating characteristic.

[0219] In one experiment, samples of urine 88 were collected from 25 cancer patients and 43 cancer - free patients. Among the cancer patients, 14 were low - grade and 11 were high - grade as defined by TURBT. The cancer - free patients were either healthy or had a history of cancer. Using an optical microscope coupled to an atomic force microscope 8, human observers randomly selected round objects that appeared to be cells.

[0220] The database is further reduced by using the data reduction process mentioned in equation (14). Thus, the possible generator 100 is P nm (k;s) = median{P mn (k;s;p)}, where p is an integer between 1 and 4 (inclusive) to correspond to the 4 partitions of each image. The resulting compressed database has two classes and can be formally represented as:

[0221]

[0222] Each patient was imaged for at least 5 cells. For simplicity, only two attributes are considered: height and adhesion.

[0223] Figure 14 A hierarchy of the importance of surface parameters of the height and adhesion attributes calculated within the random forest method is shown. The figure shows the average of the importance coefficients and the error bars indicating one standard deviation about the average. The database 86 was randomly split into training data 87 and test data 89 thousands of times.

[0224] The mapped attributes of height and adhesion are combined by tensor addition, which is essentially a data reduction method (3 - 1) for vectors suitable for surface parameters. The associated tensor addition operation is represented as:

[0225]

[0226] As in the case of Figure 9 Figure 14 ​Each surface parameter in has the standard name of the surface parameter preceded by a letter indicating the mapped attribute from which it is derived. For example, "a_Sds" refers to the "Sds" parameter derived from an image of the adhesion attribute.

[0227] A suitable statistical performance measure for the random forest method comes from examining the receiver operating characteristics and confusion matrix. The receiver operating characteristics allow the definition of ranges of sensitivity and specificity. The range of sensitivity corresponds to "accuracy" when classifying cells as coming from patients with cancer, while the specificity corresponds to "accuracy" when classifying cells as coming from people without cancer. The receiver operating characteristics make it possible to define ranges of specificity and ranges of sensitivity using the receiver operating characteristics, as follows:

[0228] sensitivity = TP / (TP+FN);

[0229] specificity = TN / (TN + FP);

[0230] accuracy=(TN+TP) / (TP+FN+TN+FP), (19)

[0231] Among them, TN, TP, FP, and FN represent true negative, true positive, false positive, and false negative, respectively.

[0232] Figure 15 Three different curves are shown, each showing the accuracy achieved by considering a different number of surface parameters, where the surface parameters are selected based on choosing different autocorrelation thresholds and importance coefficients as described above.

[0233] By performing thousands of random splits between the training data 87 and the test data 89, we can obtain Figure 15 Each of the three different curves in . The curves differ in the allocation of data to each set. The first curve corresponds to 70% of the data being allocated to the training data 87 and 30% of the data being allocated to the test data 89. The second curve corresponds to only 60% of the data being allocated to the training data 87 and 40% of the data being allocated to the test data 89. And the third curve corresponds to an even split between the training data 87 and the test data 89.

[0234] from Figure 15 It is evident from the inspection of that there is virtually no dependency on a particular threshold split. This indicates the robustness of the process performed by the machine learning module 84.

[0235] Figure 16 A family of receiver operating characteristics is shown.Figure 16 Each receiver operating characteristic in the family of characteristics shown is derived from two hundred different random splits of the database 86 into training data 87 and test data 89.

[0236] When attempting to classify between two classes, each receiver operating characteristic shows the sensitivity and specificity for different thresholds. Figure 16 The Figure Two diagonal of the equal division in Figure 16 is equivalent to a classifier that classifies by flipping a coin. Thus, the closer the receiver operating characteristic is to the

[0237] diagonal shown in Figure 21 , the worse its classifier is at classification. The fact that the curves cluster away from this diagonal and there is little variation between the individual curves indicates the effectiveness of the classifier and its insensitivity to the specific choice of training data 87 and test data 89.

[0238] In Figure 21 , each row in the table shown characterizes a specific number (N) of the cells collected and a smaller number (M) used as a diagnostic threshold. For each row, two channels are considered: height and adhesion. For each of the three machine learning methods used, this shows the average AUC and accuracy for thousands of random splits of the database into training data and test data, where 70% of the data in the database is assigned to the training data. The accuracy is the accuracy associated with the minimum error in classification. Figure 21 Each row in

[0239] also shows the sensitivity and specificity. Figure 21 Figure 21

[0240] In principle, sensitivity and specificity can also be defined around an equilibrium point where sensitivity and specificity are equal. Because the number of human subjects is limited, it is difficult to precisely define where this equilibrium point will be. Therefore, in Figure 15 , the requirement for equality is relaxed and a balance range is defined, where the magnitude of the difference between sensitivity and specificity must be less than a selected value, which for Figure 15 is 5%. Figure 15, it is obvious that using only 8 to 10 wisely selected surface parameters is sufficient to achieve a relatively high accuracy of 80%. Consider using the first 10 surface parameters to characterize the statistical behavior of the receiver operating characteristics and the confusion matrix, including the specificity, sensitivity, and accuracy of the classifier 100.

[0241] The process of classifying a cell as coming from a patient without cancer or a patient with cancer depends on averaging the probabilities obtained for that cell over all repetitions of the procedure used to obtain that probability. This is formally expressed as:

[0242] where

[0243] where a machine learning method developed on the training database 87 is used to develop the classifier AI. According to this procedure, and assuming that class 1 represents cancer cells, if Prob n (1) exceeds a specific threshold that can be obtained from the receiver operating characteristics, the cell is identified as coming from a patient with cancer.

[0244] To confirm Figure 18 and Figure 19 the authenticity of the data shown in Figure 19 and Figure 20 a control experiment was conducted using the same procedure as in Figure 19 and Figure 20 , but the samples to be classified were evenly split between cancer cells and healthy cells. Figure 17 and Figure 18 show the results of thousands of randomly selected classifications. Obviously, the accuracy has dropped to 53% ± 10%, which is consistent with expectations. This indicates Figure 19 and Figure 20 the reliability of the data shown in Figure 19 and Figure 20 as well as the resistance of the classifier to overfitting, which is a common problem that occurs when the machine learning method is made to process too many parameters.

[0245] An alternative classification method relies on more than one cell to establish a patient's diagnosis. This avoids the lack of robustness based on high sampling errors. In addition, this avoids errors caused by the fact that it cannot be ensured that the cells found in the urine sample 88 actually come from the bladder itself. Other parts of the urinary tract are fully capable of shedding cells. In addition, the urine sample 88 can contain various other cells, such as exfoliated epithelial cells from other parts of the urinary tract. One such classification method includes: if the number M of cells classified as coming from a patient with cancer in the total number N of cells classified as being from patients with cancer is greater than or equal to a predefined value, then the patient is diagnosed with cancer. This is a generalization of the case of N = M = 1 discussed previously.

[0246] The algorithms (3-2)-(3-9) or (10)-(14) can be used to assign a probability-based probability of having cancer for N cells. As a preferred method for defining the probability of classifying N test cells as being from a cancer patient (category 1), it is as follows:

[0247] where

[0248] wherein, the classifier AI is developed based on the training database 87.

[0249] Figure 19 and Figure 20 shows robustness accuracy and receiver operating characteristics similar to those in Figure 15 and Figure 16 but for the case of N = 5 and M = 2. It can be seen that the accuracy of this method can reach 94%. The above random tests show an area under the receiver operating characteristic curve of 50 ± 22% (results of thousands of random selections of the diagnostic set). These imply a lack of overtraining.

[0250] The calculation results of the confusion matrix for multiple N and M are shown in the table of Figure 20 for two single-channel (height and adhesion) examples. The robustness of the combined channels is better compared to single-channel-based diagnosis.

[0251] The above procedure can also be applied to classify cancer-free patients. In this case, the probabilities discussed above are the probabilities that the cells belong to cancer-free patients.

[0252] The present invention and its preferred embodiments have been described. For the new claims protected by the patent certificate, refer to the claims section.

Claims

1. A method for surface recognition, comprising: An atomic force microscope is used to collect a set of images associated with a surface; The images are combined; And a machine learning method applied to the combined images is used to classify the surface, wherein using the atomic force microscope includes: using a multi-channel atomic force microscope, wherein each channel corresponds to one of a plurality of attributes of the surface, wherein using a machine learning method applied to the combined images to classify the surface includes classifying using both a machine learning module and a classifier, the classifier classifying unknown samples after learning the basis for classification from the machine learning module by being trained with training data and being able to be tested with test data, the training data having been used to learn how to classify, and the test data being able to be used to verify the effectiveness of the classification performed by the machine learning module.

2. The method according to claim 1 further comprises: The images are processed to obtain surface parameters, and the surface is classified using machine learning and the classifier at least partially based on the surface parameters.

3. The method according to claim 1, wherein, Collecting the set of images includes: using the multi-channel atomic force microscope in a ringing mode, wherein each channel of the multi-channel atomic force microscope provides information indicating a corresponding one of the attributes.

4. The method according to claim 1 further comprises: A cell surface is selected as the surface, and the cell is classified as having a cell abnormality or not having a cell abnormality.

5. The method according to claim 1, wherein, Using the atomic force microscope includes: using the atomic force microscope in a sub-resonant tapping mode.

6. The method according to claim 1, wherein Using an atomic force microscope includes: collecting information from a plurality of channels, each of the plurality of channels corresponding to a different one of the attributes, and the method further includes: compressing the information provided by the channels and constructing a compressed database based on the compressed information.

7. The method according to claim 6, further comprising: A first database is formed based on the information provided by the channels, wherein constructing the compressed database includes: projecting the first database onto a subspace having a dimension lower than the dimension of the first database, the projection defining the compressed database, and the compressed database having a dimension less than the dimension of the first database.

8. The method according to claim 6, further comprising: A first database is formed based on the information provided by the channels, the first database having metrics, and the method further includes: deriving a compressed database from the first database, the compressed database having fewer metrics than the first database.

9. The method according to claim 8, wherein, Deriving the compressed database includes: performing tensor addition to generate a tensor sum, and using the tensor sum to form the compressed database; the tensor sum combines information from the first database along one or more segments corresponding to one or more metrics of the first database.

10. The method according to claim 8, wherein, Deriving the compressed database from the first database includes: defining a subset of values from the first database, each of the values representing a corresponding element in the first database; deriving compressed values from the values in the subset of values; and representing the corresponding elements from the first database with the compressed values; wherein deriving the compressed values includes: summing the values in the subset of values.

11. The method according to claim 10, wherein, Summing the values includes: performing a tensor addition to generate a tensor sum, and forming a compressed database using the tensor sum; the tensor sum combines values from the first database along one or more segments corresponding to the corresponding metrics of the first database.

12. The method according to claim 8, wherein Deriving a compressed database from the first database includes: defining a subset of values from the first database, each of the values representing a corresponding element in the first database; deriving a compressed value from the values in the subset of values; and representing the corresponding element from the first database with the compressed value; wherein, deriving the compressed value includes: averaging the values in the subset of values.

13. The method according to claim 12, wherein, Averaging the values includes: obtaining an arithmetic mean.

14. The method according to claim 12, wherein Averaging the values includes: obtaining a geometric mean.

15. The method according to claim 8, wherein, Deriving a compressed database from the first database includes: defining a subset of values from the first database, each of the values representing a corresponding element in the first database; deriving a compressed value from the values in the subset of values; and representing the corresponding element from the first database with the compressed value; wherein, the compressed value is one of the maximum or minimum values of the values in the subset of values.

16. The method according to claim 8, wherein Deriving a compressed database from the first database includes: defining a subset of values from the first database, each of the values representing a corresponding element in the first database; deriving a compressed value from the values in the subset of values; and representing the corresponding element from the first database with the compressed value; wherein, deriving the compressed value includes: passing information from the first database through a surface parameter extractor to obtain a set of surface parameters.

17. The method according to claim 16, further comprising: Normalize the surface parameters representing the set of surface parameters to be independent of the surface area of the image from which the surface parameters are derived.

18. The method according to claim 16 further comprises: Divide the surface parameters by another parameter of the same dimension.

19. The method according to claim 1, the method further comprising: Automatically detect that an image of a sample has artifacts, and automatically exclude the image from being used for classifying the surface.

20. The method according to claim 1 further comprises: Divide an image of a sample into partitions; obtain surface parameters for each partition; And define the surface parameters of the cell as the median of the surface parameters for each partition.

21. The method according to claim 1 further comprises: Process the image to obtain surface parameters, and use machine learning to classify the surface at least partially based on the surface parameters and externally derived parameters.

22. The method according to claim 21, wherein, The surface is a surface of a body derived from a collected sample, at least one of the samples being a bodyless sample without a body, and the method further includes: selecting the externally derived parameters to include data indicating the absence of a body in the bodyless sample.

23. The method according to claim 22 further comprises: Assign artificial surface parameters to the bodyless sample.

24. The method according to claim 21, wherein The surface is the surface of a cell, and the method further includes: selecting the externally derived parameters to include data indicating the probability that the cell has a cell abnormality.

25. The method according to claim 2, the method further comprising: Define a subset of the surface parameters; and generating a database based on the subset, wherein defining the subset of the surface parameters includes: determining the correlation between the surface parameters; comparing the correlation with a threshold to identify a set of relevant parameters; and including a subset of the set of relevant parameters in the subset of the surface parameters.

26. The method according to claim 2, the method further comprising: defining a subset of the surface parameters; and generating a database based on the subset; wherein defining the subset of the surface parameters includes: determining a correlation matrix between the surface parameters; wherein determining the correlation matrix includes: generating a simulated surface.

27. The method according to claim 2, the method further comprising: defining a subset of the surface parameters; and generating a database based on the subset; wherein defining the subset of the surface parameters includes: combining different surface parameters of the same type from the same sample.

28. The method according to claim 1, wherein Using machine learning methods includes: using a random forest method.

29. The method according to claim 1, wherein, Using machine learning methods includes: using an extremely randomized forest method.

30. The method according to claim 1, wherein, Using machine learning methods includes: using a gradient boosting tree method.

31. The method according to claim 1, wherein Using machine learning methods includes: using a neural network.

32. The method according to claim 1, wherein, Using machine learning methods includes: using at least two methods selected from the following: gradient boosting tree, extremely randomized forest method, and random forest method.

33. The method according to claim 1, wherein Using machine learning methods includes: using a decision tree method.

34. The method according to claim 1, wherein The surface is the surface of a first plurality of cells, wherein a second plurality of the cells have been classified as having cell abnormalities, and a third plurality of the cells have been classified as not having cell abnormalities, and the method further includes: if the ratio of the second plurality to the first plurality exceeds a predetermined threshold, classifying the cells in the first plurality of cells as having cell abnormalities.

35. The method according to claim 1, wherein The atomic force microscope includes: a cantilever and a probe disposed at an end of the cantilever; wherein the cantilever has a resonance frequency; wherein using the atomic force microscope includes: oscillating the distance between the probe and the surface at a frequency less than the resonance frequency.

36. The method according to claim 1, wherein Using the atomic force microscope includes: using a microscope that has been configured to output information of multiple channels corresponding to different physical properties of a sample surface.

37. The method according to claim 1 further comprises: Processing the image to obtain surface parameters, and using machine learning to classify the surface at least partially based on the surface parameters and externally derived parameters; wherein the surface is the surface of a cell derived from a sample obtained from a cell source, and at least one of the samples is a cell-free sample that does not have cells from the cell source, and the method further includes: selecting the externally derived parameters to include data indicating the absence of cells in the cell-free sample.

38. The method according to claim 37, further comprising: Assigning artificial surface parameters to the cell-free sample.

39. The method according to claim 1, wherein, The image is an image of a cell, and the method further includes: automatically detecting that the image of the cell has artifacts, and automatically excluding the image from being used for classifying the surface.

40. The method according to claim 1, wherein The image is an image of a cell, and the method further includes: dividing the image of the cell into partitions; obtaining surface parameters of each partition; and defining the surface parameters of the cell as the median of the surface parameters for each partition.

41. A device for surface recognition, comprising a processing system and a multi-channel atomic force microscope for acquiring an image associated with the surface, each channel of the multi-channel atomic force microscope corresponding to a property of the surface, the processing system receiving signals representing the image from the atomic force microscope and combining the images, the processing system including a machine learning module and a classifier, the classifier classifying unknown samples after having learned a basis for classification from the machine learning module. Among them, The machine learning module has been trained with training data to classify the surface such that the resulting classification is testable using test data, the training data having been used to learn how to classify, and the test data being capable of being used to verify the effectiveness of the classification.

42. The apparatus according to claim 41, wherein, The processing system is configured to process the image to obtain surface parameters and classify the surface using the machine learning module at least partially based on the surface parameters.

43. The apparatus according to claim 42, wherein, The processing system includes a compressor that compresses the information provided by the channels and constructs a compressed database based on the compressed information.

44. The device according to claim 43, further comprising a classifier that classifies unknown samples based on the compressed database.

45. The apparatus according to claim 43, wherein, The compressor is configured to construct the compressed database by projecting a first database onto a subspace having a dimension lower than the dimension of the first database, the projection defining the compressed database, the compressed database having a dimension smaller than the dimension of the first database.

Citation Information

Patent Citations

  • Method and System for Predicting Spatial and Temporal Distributions of Therapeutic Substance Carriers

    US20150036889A1

  • Probe for atomic force microscope equipped with an optomechanical resonator, and atomic force microscope comprising such a probe

    WO2018134377A1