A computer-implemented method for identifying molecules from atomic force microscope images and generating names for the molecules according to the IUPAC nomenclature
A method using multimodal neural networks addresses the limitations of AFM in molecular recognition by generating IUPAC names from AFM images, enhancing resolution and accuracy in identifying diverse molecules.
Patent Information
- Application Number
- JP2024563845
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-29
- Filing Date
- 2023-04-28
- Publication Date
- 2025-07-02
AI Technical Summary
Existing methods for molecular identification through atomic force microscopy (AFM) are limited in recognizing specific molecules due to the dependence on human visual inspection and the complexity of chemical environments, and deep learning approaches struggle with generalizing beyond trained datasets.
A computer-executed method using two pre-trained multimodal recurrent neural networks (M-RNN and AM-RNN) to generate IUPAC names from AFM images, combining convolutional and recurrent neural networks with dropout layers and an image data generator to handle diverse molecular structures.
Accurately identifies and names molecules from AFM images, including complex structures, by generating IUPAC names with high precision, even in the presence of noise and asymmetry, and improves resolution with CO-functionalized probes.
Smart Images

Figure 2025520222000001_ABST
Abstract
Description
Detailed Description of the Invention
[0001] 〔Technical Field〕 The present invention relates to a computer-executed method for identifying molecules from atomic force microscope images by using two trained multimodal recurrent neural networks to generate the name of the molecule according to the IUPAC nomenclature.
[0002] Therefore, the present invention is of interest in the field of nanotechnology, particularly in areas related to chemical reactions on surfaces, and is also of interest to users and manufacturers of atomic force microscopes.
[0003] 〔Background Art〕 Scanning probe microscopes have played an important role in the development of nanoscience as a basic tool for locally evaluating and manipulating the properties of substances with high spatial resolution. In particular, an atomic force microscope (AFM) operating in frequency modulation (FM) mode enables the evaluation and manipulation of the properties of all kinds of substances at the atomic scale. This is achieved by measuring the change in the frequency of a probe vibrating due to its interaction with the sample. When the tip of the probe is functionalized with an inert closed-shell atom or molecule, particularly a CO molecule, the resolution is dramatically improved, and access to the internal structure of the molecule becomes possible.
[0004] This distinct contrast results from the fact that the Pauli repulsion between the CO probe and the sample molecule is modified by the electrostatic interaction (ES) between the potential generated by the sample and the charge distribution associated with the lone pair of oxygen in the probe. Furthermore, due to the flexibility of the molecular probe, the saddle line of the total potential energy surface (PES) sensed by CO is emphasized. These high-resolution AFM (HR-AFM) functions enable the visualization of frontier orbitals, the determination of bond order potentials and charge distributions, and open the door to the tracking and control of surface chemical reactions. Despite such excellent results, one of the most important goals, "molecular recognition," has not yet been achieved. That is, it is to be able to recognize specific molecules solely by observing HR-AFM.
[0005] The molecules have been identified in combination with other experimental techniques such as AFM and scanning tunneling microscopy (STM) and Kelvin probe force microscopy (KPFM), or with the aid of theoretical simulations (“Non-contact atomic force microscopy: Bond imaging and beyond” Q. Zhong, X. Li, H. Zhang, L. Chi, Surf. Sci. Rep. 75, 100509 (2020)).
[0006] The chemical identification of individual atoms of semiconductor surface alloys has been achieved using reactive semiconductor tips. In this case, the maximum attractive force between the tip of the probe and the probe atoms on the sample conveys information about the chemical species involved in the covalent interaction. However, with a probe functionalized with inert CO molecules, the main contrast source in AFM is the Pauli repulsion, and the image is strongly affected by the relaxation of the probe. So far, the few attempts to identify atoms within molecules by HR-AFM have been based on differences seen in the interaction attenuation between the probe and the sample at molecular sites, or on characteristic image features related to the chemical properties of specific molecular components. For example, a substituted N atom on a hydrocarbon aromatic ring displays a sharper apex due to its lone pair. Furthermore, the attenuation of the CO-sample interaction on the substituted N atom is faster than on the adjacent C atom. Halogen atoms can also be identified in AFM images due to their elliptical shape (associated with σ holes) and a significantly stronger repulsive force compared to atoms such as nitrogen and carbon. However, even such atomic features are highly dependent on the molecular structure and can be associated not only with specific chemical species but also with sites within the molecule. Since the chemical environment is so diverse, the identification of molecules by visual inspection alone by the human eye is impossible.
[0007] Artificial intelligence (AI) techniques are precisely optimized to handle this kind of subtle correlation and vast amounts of data. Deep learning has an excellent ability to explore patterns and is now routinely used for image classification, interpretation, description, and analysis, endowing machines with capabilities specific to or exceeding those of humans. In our previous study ("A Deep Learning Approach for Molecular Classification Based on AFM Images" J.Carracedo-Cosme, C.Romero-Muniz, R.Perez, Nanomaterials 11, 1658 (2021)), we tested the potential of deep learning techniques to classify 60 types of organic molecules from their AFM images with a certain height, basically limited to flat molecules. Although reassuring, this clear success of the proof-of-concept does not provide a solution to the general problem of molecular identification. The classification approach can only identify the molecules included in the training dataset. Considering the rich complexity provided by organic chemistry, even a very large dataset that already poses a formidable computational challenge (since the output vector has as many dimensions as the number of molecules in the dataset) will not be able to classify many of the molecules already known or potentially synthesizable target molecules.
[0008] Therefore, it is necessary to develop a new method to achieve complete molecular identification (structure and composition) including non-planar structures through atomic force microscopy imaging. [Description of the Invention]
[0009] The present invention relates to a computer-executed method for identifying molecules from atomic force microscopy images by using a combination of two different pre-trained multimodal recurrent neural networks (M-RNN A , AM-RNN) to generate the names of the molecules according to the IUPAC nomenclature; each multimodal recurrent neural network (M-RNN A, the AM-RNN) includes a convolutional neural network CNN component, a recurrent neural network RNN component, and a multimodal component φ. Therefore, an object of the present invention is to provide a text (the name of a molecule according to the IUPAC nomenclature) that describes an image (a plurality of acquired FM-AFM images).
[0010] The IUPAC nomenclature is the most widely accepted and used nomenclature in the scientific community and is adopted as a molecular descriptor in the present invention. The IUPAC name clearly determines the molecular composition and structure. This is done by defining a hierarchical keyword list for naming functional groups that are described according to a systematic syntax that defines the structural position of each part or group within the molecule. To convert the IUPAC nomenclature into an appropriate computational language, the present invention defines a set of terms obtained by decomposing each IUPAC name. In the present invention, the term "set of terms into which each IUPAC name is decomposed" refers to a set of characters or symbols that indicate molecular sites, ligands, or specify the positions of atoms used in the IUPAC nomenclature in this specification. Combining these terms in a specific order generates a molecular name according to the IUPAC nomenclature.
[0011] To generate the name of a molecule according to the IUPAC nomenclature, the method of the present invention uses a combination of two different pre-trained multimodal recurrent neural networks (M-RNN A , AM-RNN). The first pre-trained M-RNN A determines the main chemical groups (main molecular sites) that make up the molecule, which define the keywords (here, IUPAC attributes) for each part. On the other hand, the second AM-RNN predicts the remaining IUPAC terms and assembles the remaining IUPAC terms in the correct order to generate the IUPAC nomenclature of the IUPAC attributes determined by the pre-trained M-RNN A and the molecular structure.
[0012] In the present invention, the "terms" are a set of letters or symbols indicating specific positions of molecular moieties, ligands, and atoms used in the IUPAC nomenclature. By combining these terms, the IUPAC name of the molecule is obtained.
[0013] In the present invention, the term "IUPAC attributes" mainly refers to a 100-element subset of IUPAC terms that describe atomic groups. Table 1 shows the terms for IUPAC decomposition. The elements above the double line are the 100-term subset considered as attributes. The gray cells do not correspond to any term but are colored to distinguish terms that separate words with a space.
Table 1
[0014] Quasar Science Resources S.L.-Universidad Autonoma de Madrid-Atomic Force Microscopy (QUAM-AFM) (https: / / doi.org / 10.21950 / UTGMZ7) is used for the training of the first pre-trained multimodal recurrent neural network M-RNN A network and the second pre-trained attribute multimodal recurrent neural network AM-RNN.
[0015] QUAM-AFM is a publicly available dataset of 165 million AFM images theoretically generated from 686,000 isolated molecules using 240 different AFM operation parameters (10 probe-sample distances, 6 oscillation amplitudes, and 4 different values of the torsional rigidity of the CO molecule, which are known to depend on the details of the attachment of the molecule to the tip of the metal probe). QUAM-AFM also provides a ball-and-stick depiction of each molecule generated from atomic coordinates. These depictions share the same scale as the scale used in the AFM images: when the two images are overlaid, each ball in the depiction is centered at the position occupied by the atom it represents in the AFM image.
[0016] The QUAM-AFM dataset contains organic molecules, excluding organic salts or inorganic compounds and compounds without a pure molecular form such as polymers. The selected molecules contain, in addition to the four basic elements of organic chemistry (carbon, hydrogen, nitrogen, oxygen), elements that are less common than these but frequent in organic compounds, such as sulfur, phosphorus, and halogen atoms (fluorine, chlorine, bromine, iodine). The largest molecule in the QUAM-AFM database has a total of 85 atoms.
[0017] Very small molecules, i.e., those containing less than 8 atoms, are not suitable for identification by AFM alone because of their very high surface mobility and very diverse adsorption forms and are not included in the QUAM-AFM database. Furthermore, very large molecules with structures that do not fit into a square-based cell with a side length of 24 Å are not included in the QUAM-AFM dataset.
[0018] The QUAM-AFM database is limited to quasi-planar molecules that display only height changes up to 1.83 Å along the z-axis because it contains aliphatic chains with carbon atoms (methyl groups) in the side chains. QUAM-AFM includes aliphatic, cyclic, and aromatic compounds, especially numerous hydrocarbons (alkanes, alkenes, alkynes, etc.) and all typical organic families (alcohols, thiols, ethers, aldehydes and ketones, carboxylic acids, amines, amides, imines, esters, nitriles, nitro and azo compounds, halocarbons, and acyl halides, etc.). 3 The IUPAC names of the QUAM-AFM images can be broken down into a total of 199 terms. The maximum length of the terms in the breakdown of the IUPAC names of QUAM-AFM is 57.
[0019]
[0020] Note that the class of a molecule is defined by the types of atomic species it contains and the respective repeating atomic numbers of these atomic species. Since the representative number of QUAM-AFM for each class is the one obtained by removing hydrogen from the chemical species list, molecules with completely different structures such as pyrazine and pyridazine, 2-ethylnitrile or butanedinitrile belong to the same class (C4N2). As a result, there are a total of 2339 classes in the molecular structure considered by QUAM-AFM. Refer to Figure 1.
[0021] The multimodal recurrent neural network generates a new text description for explaining the content of the image. The first pre-trained multimodal recurrent neural network M-RNN A is used in the method of the present invention to obtain the main chemical groups constituting the molecule, that is, the main molecular sites, and define the keywords of each site with IUPAC attributes.
[0022] The first pre-trained multimodal recurrent neural network M-RNN A A method for training, the training method including the following steps: i) A step of providing the first pre-trained multimodal recurrent neural network M-RNN A where the first pre-trained multimodal recurrent neural network M-RNN A is · A first convolutional neural network CNN / M-RNN including a block of 3D convolutional layers and one or more dropout layers A component, · A first recurrent neural network RNN / M-RNN including one or more embedding layers, one or more dropout layers, and one or more recurrent layers, where at least one recurrent layer is a gated recurrent unit (GRU), component, A component, · A first multimodal component φ / M-RNN including a connection layer with one or more dropout layers A 、 · And an image data generator, ii) Using a set of AFM images of known or predetermined molecules corresponding to a class of molecules sharing the same chemical composition (the number of different chemical species and the number of atoms of each chemical species excluding H atoms), the first recursive neural network CNN / M-RNN A The first convolutional neural network CNN / M-RNN supplied to the component A Pre-train the component, for example, with AFM images of known or predetermined molecules corresponding to a class of molecules sharing the same chemical composition in the database QUAM-AFM. The changes in the shape, contrast, and height of the images indicate the three-dimensional positions, sizes (chemical properties), and distances between the atoms within the organic molecule; iii) In step (ii), the first convolutional neural network CNN / M-RNN A After the component is pre-trained, the first pre-trained multimodal recurrent neural network M-RNN A Supply a plurality of AFM grayscale images of a certain height of known or predetermined organic molecules, such as AFM images in the database QUAM-AFM, to · RNN / M-RNN A Train the component and update its weights while keeping the weights of the CNN / M-RNN A Component fixed, · Or, while keeping the weights of the RNN / M-RNNA component fixed, train the CNN / M-RNN A Component and update its weights, Thereby predicting the IUPAC attributes of the organic molecule.
[0023] CNN / M-RNN AThe component includes blocks of 3D convolutional layers and dropout layers, and is composed of a modified Inception ResNet V2 model (C. Szegedy, S. Ioffe, V. Vanhoucke, A. A. Alemi, 31rd Proc. AAAI Conf. on Artificial Intelligence (AAAI Press, Palo Alto, CA, USA, 2017), p. 4278 - 4284). In the modified Inception ResNet V2 model, the 2D convolutional layer is replaced by a block containing two 3D convolutional layers (each with 32 filters, (3, 3, 3) kernel size, (2, 1, 1) stride) to process a stack of 10 AFM images with various probe - sample distances, followed by a dropout layer. This dropout layer is essential for generalizing to different images such as experimental images. Furthermore, the output vector v is obtained by removing the last fully - connected layer of the model that is specific to the original classification task.
[0024] RNN / M - RNN A includes an embedding layer, a subsequent dropout layer, and a terminal recurrent layer (see Figure 2). Attributes are encoded by assigning integer values from 1 to 100 to each attribute. RNN / M - RNN A has an input that is a vector of fixed size 19, which is the maximum number of different attributes included in the molecule name of QUAM - AFM(18) plus the startseq token. In the first step, only one integer value specifying the startseq token is included, and in the next step, the startseq and the integer values associated with the attributes predicted in the previous step are included. At each time step, the input is padded with zeros until it reaches a length of 19.
[0025] The embedding layer processes the input to represent each attribute in vector space, converting the input vector into a dense vector with real - valued numbers that reflect the syntactic and semantic meaning of the attribute, and placing similar attributes close to each other in vector space according to the neural connections established during training. For example, the attributes closest to bromine are chlorine, fluorine, and iodine.
[0026] The recurrent layer stores information about previous predictions in its internal state. For attribute prediction, a Gated Recurrent Unit (GRU) is used as the recurrent layer. GRUs are computationally efficient and suitable for short attribute chains of the prediction target. To avoid overfitting during training, it is necessary to introduce a dropout layer between the embedding layer and the recurrent layer.
[0027] The multimodal component φ / M-RNNA first processes the output v of the CNN through two fully connected layers, sandwiching dropout in between. The resulting vector is concatenated with the output vector of the RNN and fed into two fully connected layers that generate a vector of probabilities. This final vector has 103 components, 100 of which are related to attributes, 1 is related to padding, and 2 are related to the startseq and endseq tokens. The prediction of the new attribute is obtained from the position of the larger component within the vector.
[0028] This attribute prediction starts only with the startseq token S0 in the input (S0,0,...,0) of the RNN / M-RNNA. For a given time step t, the input (S0,Y1 A ,...,Y A ,...,Y A t-1 ,0...,0) is supplied to the RNN / M-RNN. This input concatenates S0 with all the predictions already made at the previous time steps and is padded with zeros until it reaches a length of 19 (see Figure 3 as an example). This process is repeated until the endseq token is predicted and the loop ends (see Figure 4).
[0029] Figure 5 shows the details of the layers of the RNN / M-RNN A including the operator, dimension, activation function, and the φ / M-RNN A component of the M-RNN. A
[0030] M-RNN A Training the entire complex network like this is very time-consuming, and overfitting is likely to occur in components with fewer layers. Therefore, the first pre-trained multimodal recurrent neural network M-RNNA is trained in various stages:
[0031] First, the first recurrent neural network CNN / M-RNN A The component of the first convolutional neural network CNN / M-RNN A The component is pre-trained with a set of AFM images of known or predetermined molecules corresponding to classes of molecules sharing the same chemical composition. For example, it is pre-trained with AFM images of known or predetermined molecules corresponding to classes of molecules sharing the same chemical composition in the database QUAM-AFM. The shape and contrast of the images, and their variations due to height, indicate the three-dimensional positions and sizes (chemical properties) of the atoms within the organic molecule, and the distances between said atoms.
[0032] Second, the first convolutional neural network CNN / M-RNN A The component is the first pre-trained multimodal recurrent neural network M-RNN pre-trained using a plurality of AFM grayscale images of known or predetermined organic molecules, for example, AFM images in the database QUAM-AFM A is supplied, and alternatively, ·RNN / M-RNN A The component and the CNN / M-RNN A The component is the RNN / M-RNN A During the training of the RNN / M-RNN A The weights of the CNN / M-RNN · or the weights of the RNN / M-RNN A While fixing the weights of the RNN / M-RNN, the CNN / M-RNN A The component and the RNN / M-RNN A The component is weighted, and the CNN / M-RNN A The component is trained to predict the IUPAC attributes of organic molecules.
[0033] The first trained multimodal recurrent neural network M-RNN A During the training of, for each input stack, one of the 24 combinations of AFM operation parameters available in QUAM-AFM (6 vibration amplitudes and 4 values for the torsional rigidity of the CO molecule) is randomly selected. This variation in the input data · For the first trained M-RNN A (multimodal recurrent neural network), to confirm that the parameters for which the AFM experiment was performed do not play a decisive role in successfully identifying the structure, · Prevent overfitting, · Provide the ability to generalize to the first trained multimodal recurrent neural network M-RNN A .
[0034] This variability is further improved by applying an Imaging Data Generator (IDG) that normalizes pixel values and adds various deformations (zoom, rotation, shift, inversion, shear) to the input images (Figure 6). The use of IDG is motivated because experimental images have features that cannot be captured in AFM simulations and may hinder identification. For example, experimental images do not display the perfect symmetry of organic molecules. Such differences between experimental and theoretical AFM images for certain organic molecules may be due to the presence of noise that cannot be avoided in experiments, the asymmetry of the probe that is not included in AFM image simulations, or the relaxation or deformation of organic molecules due to their interaction with the substrate, considering the ideal gas layer structure in the simulated AFM images used in learning.
[0035] The deformations provided by applying IDG during training mimic these effects and greatly contribute to endowing the first trained multimodal recurrent neural network M-RNN A with the ability to identify organic molecules from experimental images. The selection of appropriate deformation parameters for IDG is important because the right choice can significantly improve the accuracy of identification.
[0036] The first trained multimodal recurrent neural network M-RNNA includes the following (see Figure 2). · A convolutional neural network CNN / M-RNN including a block of 3D convolutional layers and one or more dropout layers configured to encode a plurality of atomic force microscope images of a certain height into an output vector v A Component. · A recurrent neural network RNN / M-RNN including one or more embedding layers, one or more dropout layers, and one or more recurrent layers A Component. At least one recurrent layer is a Gated Recurrent Unit (GRU) that handles language processes, embeds the representation of each attribute based on semantic meaning that places similar attributes close together, and second, is configured to store semantic temporal context in the recurrent layer. · CNN / M-RNN A and RNN / M-RNN A A multimodal component φ / M-RNN including a connection layer with one or more dropout layers configured to process the outputs of both CNN / M-RNN A and · An image data generator configured to apply to the theoretical AFM image various random deformations selected from pixel value normalization, zoom of [-15, 15]%, rotation of [-180, 180] degrees, vertical and horizontal shifts, random vertical and horizontal flips, shear of [-20, 20]%, and any combination thereof. This enables mimicking the characteristics of experimental images such as the presence of noise that does not exist in AFM simulations and effects related to probe asymmetry that may interfere with identification.
[0037] The second trained multimodal recurrent neural network AM-RNN is used in the method of the present invention to predict the different terms of the IUPAC name of a molecule in the correct order using QUAM-AFM and obtain the name of the molecule according to the IUPAC nomenclature.
[0038] A method for training a second trained multi-modal recurrent neural network AM-RNN, the method comprising the following steps: 1) Providing a second trained multi-modal recurrent neural network AM-RNN, the second trained multi-modal recurrent neural network AM-RNN comprising: · A second convolutional neural network CNN / AM-RNN component including a block of 3D convolutional layers and one or more dropout layers, · A second recurrent neural network RNN / AM-RNN component including one or more embedding layers, one or more dropout layers, and one or more recurrent layers, wherein at least one recurrent layer is an LSTM (Long-Short-Term Memory), · A second multi-modal component φ / M-RNN including a concatenation layer with one or more dropout layers A components, and · A second image data generator. 2) Feeding and training the second convolutional neural network CNN / M-RNN A component with a set of AFM images of known or predetermined molecules corresponding to classes of molecules sharing the same chemical composition, e.g., AFM images of known or predetermined molecules corresponding to classes of molecules sharing the same chemical composition in the database QUAM-AFM, to the first convolutional neural network CNN / M-RNN A component. The shape and contrast of the images and their variations due to height indicate the three-dimensional positions and sizes (chemical properties) of the atoms within the organic molecule and the distances between said atoms. 3) To the second trained multi-modal recurrent neural network AM-RNN in which the first convolutional neural network CNN / AM-RNN component was learned in step (2), a plurality of constant-height AFM grayscale images of known or predetermined organic molecules, e.g., AFM images of the database QUAM-AFM, and the first trained multi-modal recurrent neural network M-RNNA Supply the IUPAC attributes obtained from the training of · Weight the RNN / AM-RNN component and the CNN / AM-RNN component, and fix the weights of the CNN / AM-RNN component during the training of the RNN / AM-RNN component. · Alternatively, weight the CNN / AM-RNN component and the RNN / AM-RNN component, and train the CNN / AM-RNN component while fixing the weights of the RNN / AM-RNN component. Thereby, the IUPAC name of the organic molecule is generated.
[0039] The CNN / AM-RNN component includes a block of 3D convolutional layers and dropout layers, and is composed of a modification of the Inception ResNet V2 model (C. Szegedy, S. Ioffe, V. Vanhoucke, A. A. Alemi, 31rd Proc. AAAI Conf. on Artificial Intelligence (AAAI Press, Palo Alto, CA, USA, 2017), p. 4278-4284). In the modification of the Inception ResNet V2 model, the 2D convolutional layer is replaced by a block containing two 3D convolutional layers (each having 32 filters, a (3, 3, 3) kernel size, and a (2, 1, 1) stride) for processing a stack of 10 AFM images with various probe-sample distances, followed by a dropout layer. This dropout layer is essential for generalizing to different images such as experimental images. Further, the last fully connected layer of the model, which is specific to the original classification task, is removed to obtain the output vector v.
[0040] The RNN / AM-RNN component includes an embedding layer, followed by a dropout layer, and a final recurrent layer (see Figure 2). Terms are encrypted by assigning an integer number (from 1 to 199) to each term. The input to the RNN / AM-RNN is a vector of fixed size 76. This number comes from the sum of the maximum number of different attributes of the molecular name of QUAM-AFM and the startseq token (18 + 1 = 19), and the maximum number of terms in the IUPAC name of the molecule of QUAM-AFM is 57. Each RNN / AM-RNN input is a vector of size 76, and the first pre-trained multimodal recurrent neural network M-RNN A (padded with zeros if less than 18) is concatenated with the startseq token and the terms predicted at each previous time step (padded until a vector of length 57 is obtained). At the first step, it includes the startseq token and the integer values specifying the attributes, and at the next step, it also includes the integer values related to the terms predicted at the previous step.
[0041] The embedding layer processes the input to represent each term in vector space, converting the input vector into a dense vector with real values that reflect the syntactic and semantic meaning of the word, and placing similar inputs close to each other within the vector space according to the neural connections established during training. For example, terms representing numbers are represented close to each other. That is, the terms closest to nona are octa, deca, undeca, dodeca.
[0042] The recurrent layer stores information about previous predictions in its internal state. For term prediction, an LSTM (Long Short-Term Memory) network is used as the recurrent layer. LSTM is more accurate for longer time series than the GRU used in RNN / M-RNN A and is suitable for predicting long term strings of terms up to 57 terms. Finally, to avoid overfitting during training, it is necessary to introduce a dropout layer between the embedding layer and the recurrent layer.
[0043] The multimodal component φ / AM-RNN first processes the output v of the CNN / M-RNN A through two fully connected layers with dropout in between. The resulting vector is concatenated with the output of the RNN / AM-RNN and fed into two fully connected layers that generate a probability vector. This vector has 202 components: 199 components are related to terms, 1 is for padding, and 2 are related to the startseq and endseq tokens. The position of the larger component within the vector provides the prediction of the new term.
[0044] Term prediction starts by adding the startseq token (Y1 A ,...,Y A 18 ,startseq,0,...,0) as the input to the RNN / AM-RNN to the list of attributes predicted in step (b) (padded if necessary). At a certain time step t, the input ((Y1 A ,...,Y A 18 ,startseq,Y1 T ,···,Y T t-1 , 0…,0)) is fed into the RNN / AM-RNN. This input contains all the term predictions already made at the previous time steps and is padded with zeros to a length of 57, which is the maximum number of terms. This process is repeated until the endseq token is predicted and the loop is broken (see Figure 4). Figure 7 shows the input and output at each time step in the term prediction by the AM-RNN for the perylene-1,12-diol molecule. Figure 3 shows the representation of the state of the RNN / AM-RNN and the input vector of the AM-RNN at the fourth time step.
[0045] Figure 5 shows the details of the layers of the RNN / AM-RNN and φ / AM-RNN components of the AM-RNN (including operators, sizes, activation functions).
[0046] First, the second convolutional neural network CNN / AM-RNN component to be supplied to the second recursive neural network CNN / AM-RNN component is pre-trained with a set of AFM images of known or predetermined molecules corresponding to classes of molecules sharing the same chemical composition. For example, it is pre-trained with AFM images of known or predetermined molecules corresponding to classes of molecules sharing the same chemical composition in the database QUAM-AFM. The shape and contrast of the images, and their variations in height, indicate the three-dimensional positions and sizes (chemical properties) of the atoms within the organic molecule, and the distances between said atoms.
[0047] Second, a second pre-trained multi-modal recursive neural network AM-RNN, which has been pre-learned using a plurality of constant-height AFM grayscale images of known or predetermined organic molecules, for example, AFM images from the database QUAM-AFM, is supplied to the second convolutional neural network CNN / AM-RNN component. Alternatively, · The RNN / AM-RNN component and the CNN / AM-RNN component are weighted by fixing the weights of the CNN / AM-RNN component during the training of the RNN / AM-RNN component, · Or, the CNN / AM-RNN component and the RNN / AM-RNN component generate the IUPAC name of the organic molecule by fixing the weights of the RNN / AM-RNN component while training the CNN / AM-RNN component. The second pre-trained multi-modal recursive neural network AM-RNN includes the following (see Figure 2). · A second convolutional neural network CNN / AM-RNN including a block of 3D convolutional layers and one or more dropout layers configured to encode a plurality of constant-height atomic force microscope images obtained in step (a) into an output vector v; · A second recurrent neural network RNN / AM-RNN component including one or more embedding layers, one or more dropout layers, and one or more recurrent layers, wherein at least one recurrent layer is an LSTM (Long-Short-Term Memory) configured to process language processes, embed the representation of each IUPAC term based on the semantic meaning of placing similar terms in proximity, and second, store the semantic temporal context in the recurrent layer, the RNN / AM-RNN component; · A second multimodal φ / AM-RNN component having a combination layer with one or more dropout layers, processing by combining the outputs of CNN / AM-RNN and RNN / AM-RNN, and predicting one of the IUPAC terms; and, · For the theoretical AFM image, pixel value normalization and different random deformations selected from [-15, 15]% zoom, [-180, 180] degree rotation, vertical and horizontal shifts, random vertical and horizontal flips, [-20, 20]% shear, and any combination thereof are applied, thereby mimicking the features of experimental images such as the presence of noise and effects related to probe asymmetry that do not exist in AFM simulations and may interfere with identification.
[0048] A first aspect of the present invention relates to a computer-executed method for identifying organic molecules from atomic force microscope images and generating the names of organic molecules according to the IUPAC nomenclature, characterized in that the method includes the following steps: a) Using the tip of a functionalized metal probe, acquiring a plurality of constant-height atomic force microscope images of an organic molecule at different probe height distances above the molecule by frequency modulation atomic force microscope (FM-AFM), wherein the different probe height distances are in the range of 280 pm to 370 pm above the molecule, and the changes in the shape and contrast of the images and their probe heights indicate the three-dimensional positions and sizes (chemical properties) of the atoms in the organic molecule and the distances between the atoms, the step; b) Providing the first pre-trained multi-modal recurrent neural network M-RNN to a data processing device, where the first pre-trained multi-modal recurrent neural network M-RNN A includes the following: A - A first convolutional neural network CNN / RNN component including a block of 3D convolutional layers and one or more dropout layers, - A first recurrent neural network component RNN / M-RNN including one or more embedding layers, one or more dropout layers, and one or more recurrent layers, where at least one recurrent layer is a gated recurrent unit (GRU), and A - A first multi-modal φ / AM-RNN component including a combination layer with one or more dropout layers; A c) Supplying the data processing device with the first pre-trained multi-modal recurrent neural network M-RNN having the atomic force microscope image obtained in step (a), where the first pre-trained multi-modal recurrent neural network M-RNN A generates IUPAC attributes having syntactic and semantic meanings of organic molecules; A d) Providing a second pre-trained attribute multi-modal recurrent neural network AM-RNN to the data processing device, where the second pre-trained attribute multi-modal recurrent neural network AM-RNN includes the following: - A second convolutional neural network CNN / AM-RNN component including a block of 3D convolutional layers and one or more dropout layers, - A second recurrent neural network component RNN / AM-RNN including one or more embedding layers, one or more dropout layers, and one or more recurrent layers, where at least one recurrent layer is a long short-term memory (LSTM), - A second multimodal φ / AM-RNN component including a coupling layer having one or more dropout layers, e) Together with the IUPAC attributes obtained in step (d) and the atomic force microscope image obtained in step (a), supply a second trained multimodal recurrent neural network AM-RNN to the data processing device, and the second trained multimodal recurrent neural network AM-RNN generates the IUPAC name of the organic molecule.
[0049] An AFM operating in frequency modulation dynamic mode (FM-AFM) enables the characterization and manipulation of any substance at the atomic scale by measuring the frequency change of a vibrating probe due to its interaction with an organic molecule sample. When the probe is functionalized with an inert closed-shell molecule, especially a CO molecule, the resolution is improved and access to the internal structure of the molecule becomes possible. The excellent contrast mainly originates from the Pauli repulsion between the CO probe and the sample molecule. This contribution of the repulsive force occurs because the electron densities of the probe and the sample overlap, increasing the frequency shift. The frequency shift is the change in the vibration frequency of the cantilever holding the probe due to the interaction between the probe and the sample. The shift is observed as a bright feature in the AFM image at a certain height above the position, size, and bond (distance between atoms) of the atoms, reflecting the organic molecule structure.
[0050] As used herein, the "tip of a functionalized metal probe" refers to the tip of a probe functionalized with an inert closed-shell atom or molecule, where the metal is usually Cu, but other metals such as Ag and Pt can also be used. This inert closed-shell atom or molecule dramatically improves the resolution of the AFM image and provides access to the internal structure of the organic molecule. Examples of inert closed-shell atoms or molecules are Xe atoms and CO molecules, respectively.
[0051] The tip of a metal probe functionalized with CO is preferred in the present invention for the following reasons to improve the contrast of the AFM image: · The Pauli repulsion between the lone pair of the oxygen atom in the CO molecule and the charge density of the sample molecule is highly directional because the electronic charge associated with the lone pair is preferentially distributed along the molecular axis. · The CO molecule attached to the metal probe provides a complex electric field. This electric field has a highly localized central feature, is in front of the O atom, and provides an electrostatic interaction that is repulsive to the electrons of the sample molecule and changes rapidly both vertically and horizontally. · The tilt of the CO molecule significantly enhances the features within the molecule.
[0052] Therefore, in a preferred embodiment of the present invention, the tip of the functionalized metal probe used in step (a) is selected from Cu, Ag, or Pt. In another preferred embodiment of the method of the present invention, the tip of the functionalized metal probe used in step (a) is functionalized with an inert closed-shell atom or molecule. More preferably, the tip of the functionalized metal probe is functionalized with an Xe atom or a CO molecule.
[0053] To obtain a plurality of AFM grayscale images of a constant height, within a range between 280 pm and 370 pm, different probe height distances above the organic molecules must be obtained using FM-AFM, which can acquire images of different contrasts as a function of the probe height distance. The shape, contrast, and their variations with the probe height of the images define the 3D spatial distribution of the electronic charge resulting from the interaction of the chemical species, their chemical environment, and their relative height with respect to other atoms in the molecular arrangement. Therefore, since all the information regarding the 3D spatial distribution of the electronic charge is there, it is essential to acquire a plurality of images. Preferably, at least 10 images should be acquired to appropriately characterize the FM-AFM contrast in the probe height distance range between 280 pm and 370 pm above the organic molecule. In this probe height range, the Pauli repulsion and electrostatic interaction between the probe and the organic molecule change significantly, resulting in a strong change in the contrast of the FM-AFM images that contain features characteristic of the atoms and their molecular environment. The shape, contrast, and their variations with the probe height of the FM-AFM images contain all the information regarding the three-dimensional position, size (chemical properties), and interatomic distance of the atoms.
[0054] Therefore, in a preferred embodiment of the method of the present invention, in step (a), at least 10 AFM grayscale images of a constant height of the organic molecule are acquired.
[0055] In another preferred embodiment of the method of the present invention, step (a) is performed at at least 10 different probe height distances.
[0056] Step (b) of the method of the present invention refers to providing a first pre-trained multimodal recurrent neural network M-RNN A to a data processing device, and the first pre-trained multimodal recurrent neural network M-RNN A includes the following configuration: - A first convolutional neural network CNN / RNN including a block of 3D convolutional layers and one or more dropout layersA Component - The first recurrent neural network component RNN / M-RNN including one or more embedding layers, one or more dropout layers, and one or more recurrent layers A A component in which at least one recurrent layer is a gated recurrent unit (GRU), and - The first multimodal φ / AM-RNN component including a combination layer with one or more dropout layers Step c) supplying the data processing device with the first trained multimodal recurrent neural network M-RNN having the atomic force microscope image obtained in step (a), wherein the first trained multimodal recurrent neural network M-RNN A generates IUPAC attributes having syntactic and semantic meanings of organic molecules. As described above, the attributes are listed in Table 1. A
[0057] In step (d), a second trained attribute multimodal recurrent neural network AM-RNN is provided to the data processing device, where the second trained attribute multimodal recurrent neural network AM-RNN comprises the following. - The second convolutional neural network CNN / AM-RNN component including a block of 3D convolutional layers and one or more dropout layers - The second recurrent neural network component RNN / AM-RNN component including one or more embedding layers, one or more dropout layers, and one or more recurrent layers, wherein at least one recurrent layer is LSTM (Long-Short-Term Memory), and - The second multimodal φ / AM-RNN component including a combination layer with one or more dropout layers.
[0058] Step (e) refers to supplying a second trained multimodal recurrent neural network AM-RNN having the IUPAC attributes obtained in step (d) and the atomic force microscope image obtained in step (a) to the data processing device, and the second trained multimodal recurrent neural network AM-RNN generates the IUPAC name of the organic molecule.
[0059] Another aspect of the present invention refers to a frequency modulation atomic force microscope (FM-AFM, herein the microscope of the present invention) including the tip of a functionalized metal probe configured to perform step (a) of the method of the present invention described above, and a data processing device configured to perform steps (b) to (e) of the method of the present invention described above.
[0060] Preferably, the microscope of the present invention is connected to a data processing device and further includes a display device configured to display the name of the molecule according to IUPAC obtained in step (e) of the method of the present invention. More preferably, the display device connected to the data processing device is further configured to display the structural representation of the molecule identified in step (e) of the method of the present invention in the form of a ball-and-stick depiction.
[0061] In another preferred embodiment of the microscope of the present invention, the metal at the tip of the functionalized metal probe is selected from Cu, Ag or Pt. The tip of the metal probe has improved the resolution of the FM-AFM image.
[0062] In another preferred embodiment of the microscope of the present invention, the tip of the functionalized metal probe is functionalized with an inert closed-shell atom or molecule, and preferably, the apex of the functionalized metal probe is functionalized with an inert closed-shell Xe atom or an inert closed-shell CO molecule. The tip of the functionalized metal probe has dramatically improved the resolution of the FM-AFM image.
[0063] Another aspect of the present invention refers to a computer program (herein, the computer program of the present invention) that includes instructions for causing a data processing device to execute steps (b) to (e) of the method of the present invention described above when the program is executed by the data processing device.
[0064] The last aspect of the present invention refers to a computer-readable data carrier storing the computer program of the present invention as described above.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which the present invention belongs. Methods and materials similar or equivalent to those described herein can be used in the practice of the present invention. Throughout this specification and the claims, the word "comprise" and its variations are not intended to exclude other technical features, additives, components, or steps. Additional objects, advantages, and features of the present invention will become apparent to those skilled in the art upon examination of this specification, or can be learned by practice of the present invention. The following examples and drawings are provided by way of illustration and are not intended to limit the present invention.
Brief Description of the Drawings
[0066] Figure 1 Atomic structures belonging to different classes depending on chemical species.
[0067] Figure 2 illustrates the structure common to both the M-RNN (A) and the AM-RNN in a layer composed of three components: a CNN, an RNN, and φ. The CNN component follows the Inception ResNetV2 model, where the first 2D convolutional layer is replaced by two 3D convolutional layers (to process the image stack), followed by a dropout layer (gray). The following blocks represent the complex architecture (164 layers) of the original Inception ResNet V2 model in a pictorial form. Note that the last fully connected layer of this model is specific to the original classification task and has been removed to obtain the output vector (v) that is further processed by the φ component. The RNN component includes an embedding layer (white), a dropout layer (gray), and a recurrent layer (black). The black box represents a GRU layer in the M-RNN A and an LSTM layer in the AM-RNN. The φ component processes the output vector of the CNN component with two fully connected layers (white) and a dropout layer (gray) in between, concatenates this result with the output of the RNN component, and processes it through another two fully connected layers to generate a vector of probabilities. In the case of the M-RNN A (AM-RNN), the prediction of the new attribute (term) is obtained from the position of the larger component within the vector. Figure 5 shows further details of the RNN and φ component layers of the M-RNN A and the AM-RNN, including the operator, dimension, and activation function.
[0068] Figure 3 shows the same form of the RNN representation used in Figure 4 corresponding to the fourth time step for the perylene-1,12-diol molecule for the M-RNN A and the AM-RNN. This figure emphasizes the fact that the state of the RNN, particularly the state of the recurrent layer, depends on the previous prediction.
[0069] Figure 4 shows the architecture scheme of the method of the present invention, including a combination of two different multimodal recurrent neural networks: the data flow is as follows: (a) the M-RNN A, (b) is illustrated in the AM-RNN. The rectangular boxes represent the three components of each M-RNN, namely, the convolutional neural network CNN, the recurrent neural network RNN, and the multimodal component φ. In the RNN component, x t represents the input at time step t, h t represents the internal state of the recurrent layer at t, and o t represents the output of the RNN at t. The arrows indicate the flow of information within the model. The M-RNN A predicts attributes at each time step until the loop is terminated by the endseq token, while the AM-RNN predicts sorted terms (one at each time step) that generate the IUPAC name. Figure 7 shows an example of the input and output at each time step predicted by the M-RNN A and AM-RNN networks from a 3D image stack (obtained from the QUAM-AFM dataset) corresponding to the perylene-1,12-diol molecule. Figure 3 shows the representation of the RNN corresponding to the fourth time step of the AM-RNN for the perylene-1,12-diol molecule in the M-RNN A .
[0070] Figure 5 shows the details for each layer of the RNN and φ components integrated into the M-RNN A and AM-RNN. v represents the output vector of the CNN component. The operators are the same in both models except for the final layer of the RNN, which is GRU (attribute prediction) in the M-RNN A and LSTM (term prediction) in the AM-RNN. However, the units are different to accommodate the differences in the input and output sizes of each model.
[0071] Figure 6 shows the result of applying an image data generator (IDG) to the same AFM image of the N-(5-amino-4-methylpyridin-2-yl)-6-fluoro-1-benzothiophene-2-carboxamide molecule. The parameters selected for the IDG during training are that the rotation is [-180, 180] degrees, the zoom and vertical / horizontal shifts are [-15, 15]%, the shear is [-20, 20]%, and the vertical / horizontal flips are randomly selected. If necessary, nearest-neighbor filling is applied to the points outside the boundary of the input. By appropriately selecting the parameter range of the IDG used during training, the accuracy of the IUPAC name predicted by applying the method of the present invention to experimental images is significantly improved.
[0072] Figure 7 shows the input and output of each time step predicted by the M-RNN A and the AM-RNN network from a 3D image stack (obtained from the QUAM-AFM dataset) corresponding to the case of the perylene-1,12-diol molecule. The time step t = 4 of both networks analyzed in detail in Figure 3 is highlighted in gray.
[0073] Figure 8 is an AFM experimental image of dibenzothiophene taken at different probe-sample distances (from P. Zahl, Y. Zhang, Energy Fuels 33, 4775 (2019)). Despite the strong noise and the white line cutting across the image diagonally, the method of the present invention provides a complete prediction of the IUPAC name.
[0074] Figure 9 shows perfect prediction examples of molecules containing carbon, nitrogen, oxygen, and various halogen atoms. Each figure shows (from left to right) a ball-and-stick depiction of the molecular structure (generated with the open-source code Jmol) and five AFM images at various probe-sample distances. The IUPAC name predicted by the method of the present invention, which exactly matches the ground truth, is shown below the image. Usage example
[0075] Example 1: Prediction of the IUPAC name from the AFM experimental image of dibenzothiophene This example is based on AFM experimental images measured in the constant-height mode using a CO-functionalized probe for dibenzothiophene adsorbed on Au(111). The measurements were carried out using a Createc-based low-temperature (LT)-STM upgraded with AFM and custom GXSM control software. A Q-Plus sensor (operating at 30 kHz with a typical Q-factor of about 10,000) equipped with a PtIr probe wire sharpened by focused ion beam (FIB) and functionalized with CO molecules was used. The AFM was operated in the constant-height mode, and the feedback was off, and no correction was needed for more than about 1 hour. The XYZ drift and creep were less than 1 pm / h, with no significant change in the large XYZ offset even after at least 24 hours at 5 K. The molecules were imaged by bimodal STM and AFM under ultrahigh vacuum conditions of 5 K after evaporation using a homemade evaporator.
[0076] Regarding the operating conditions, the probe vibration amplitude was about 50 pm, and a bias voltage of 40 mV was applied. This bias minimized the electrostatic force between the Au(111) surface and the tip of the actual probe, resulting in stable imaging conditions. A new phase-amplitude convergence detector and a phase-locked loop were used for frequent frequency tracking and amplitude adjustment. An Au(111) single crystal was used as the substrate and was cleaned by usually performing Ar+ sputtering / annealing in 3 cycles before use. Probe tuning was performed by controlled collisions with the Au surface followed by functionalization with CO molecules, proximity scanning, and pickup at ultra-low bias.
[0077] Figure 8 is a set of 10 AFM images taken with varying probe-sample distances for dibenzothiophene adsorbed on Au(111). The M-RNN A trained with the QUAM-AFM dataset as described above was applied to a stack of 10 constant-height AFM images shown in Figure 8. The M-RNN A correctly generated the attributes of benz, phen, and thi.
[0078] As described above, to the AM-RNN trained with the QUAM-AFM dataset, a stack of 10 AFM images of constant height shown in FIG. 1 and the attributes (benz, phen, thi.) determined by M-RNN A were given. The AM-RNN predicted the terms of the molecular name in the correct order of di, benz, o, thi, o, phen, e. It should be noted that perfect prediction was obtained despite strong noise and white lines crossing the image diagonally.
[0079] Example 2: Prediction of IUPAC names from theoretically simulated AFM images obtained from the QUAM-AFM dataset
[0080] This second example demonstrates the ability of our approach to predict the IUPAC names of non-flat molecules with complex structures and compositions, including most of the chemical species related to organic chemistry. In this case, theoretically simulated AFM images obtained from the QUAM-AFM dataset are given to the model of the present invention. These images were calculated with a torsional stiffness of 0.40 N / m and a vibration amplitude of 40 pm. None of the three molecules examined were shown in either M-RNN A or AM-RNN during training. For each molecule, consider a stack of 10 AFM images of constant height calculated at different probe heights in the range of 280 pm to 370 pm in 10 pm increments.
[0081] FIG. 9 shows five of the stacks of 10 AFM images of constant height of each molecule under study generated using the open-source Java chemical structure viewer Jmol (http: / / www.jmol.org / ), together with a ball-and-stick depiction, graphically showing its structure and composition.
[0082] In the case of FIG. 9a, the trained M-RNN AIt correctly generated the attributes of this molecule. The attributes are (in alphabetical order): azin, hydr, ide, imidin, meth, one, oxy, phen, pyr, and yl. The trained AM-RNN predicted all terms in the correct order to form the IUPAC name in all cases. This means that the method of the present invention correctly identified all molecular sites from the image and provided the correct IUPAC name character by character.
[0083] Considering the molecule of this example, the method of the present invention can identify cyclic or aliphatic, planar hydrocarbons, but can also identify more complex structures such as structures containing nitrogen or oxygen atoms. These atoms usually appear as weak features on the image because the decay of the charge density is fast.
[0084] Halogens, which are characterized on the image by elliptical features with size and intensity proportional to the σ-hole strength, were also correctly labeled (Figs. 9b, 9d, 9e).
[0085] The method of the present invention can even recognize the presence of a fluorine element that does not induce a σ-hole and generates an AFM fingerprint very similar to that of a carbonyl group when bonded to a carbon atom (compare Fig. 9e and Fig. 9f). sp 2 Hydrogen atoms bonded to carbon atoms have a very low charge density and are hardly detected by high-resolution AFM (HR-AFM), so the positions of hydrogen atoms are often inferred.
[0086] Example 3: Comparative Example A stack of AFM images was given to a single M-RNN to predict the IUPAC name of the imaged molecule. A model completely similar to the AM-RNN was followed. The only difference is that the input of the RNN component is a vector of length 58 and contains only the startseq token in the first step. The M-RNN was trained in three stages, and the weights of the CNN and RNN components were fixed alternately. This single M-RNN predicted the IUPAC name.
Claims
Claim 1 A computer-executed method for identifying organic molecules from atomic force microscope images and generating names for the organic molecules according to the IUPAC nomenclature, the method comprising the following steps: (a) Using a frequency-mode atomic force microscope and the tip of a functionalized metal probe to obtain a plurality of constant-height atomic force microscope images of the organic molecule at different probe height distances above the organic molecule, wherein the different probe height distances are in the range between 280 pm and 370 pm, and the shape, contrast of the images and their changes due to the probe height indicate the three-dimensional positions, sizes of the atoms in the organic molecule and the distances between the atoms, step; (b) Providing the first trained multimodal recurrent neural network M-RNN A to the data processing device, the first pre-trained multimodal recurrent neural network M-RNN A is - A first convolutional neural network CNN / RNN including a block of 3D convolutional layers and one or more dropout layers A component - A first recurrent neural network component RNN / M-RNN including one or more embedding layers, one or more dropout layers, and one or more recurrent layers A A component in which at least one recurrent layer is a gated recurrent unit (GRU), and - Including a first multimodal φ / AM-RNN component including a bonding layer having one or more dropout layers, step; (c) supplying the data processing device with the first learned multimodal recurrent neural network M-RNN having the atomic force microscope image obtained in step (a), A wherein the first learned multimodal recurrent neural network M-RNN A generates IUPAC attributes having syntactic and semantic meanings of the organic molecule; (d) Providing a second pre-trained attribute multimodal recurrent neural network AM-RNN to the data processing device, wherein the second pre-trained multimodal recurrent neural network AM-RNN, - A second convolutional neural network CNN / AM-RNN component including a block of 3D convolutional layers and one or more dropout layers, - A second recurrent neural network component RNN / AM-RNN component including one or more embedding layers, one or more dropout layers, and one or more recurrent layers, wherein at least one recurrent layer is LSTM (Long-Short-Term Memory), RNN / AM-RNN component, and, - Including a second multimodal φ / AM-RNN component including a bonding layer having one or more dropout layers, step; and, e) Supplying the second pre-trained multimodal recurrent neural network AM-RNN having the IUPAC attributes obtained in step (d) and the atomic force microscope image obtained in step (a) to the data processing device, and the second pre-trained multimodal recurrent neural network AM-RNN generates the IUPAC name of the organic molecule. Claim 2 The method according to claim 1, wherein in step (a), at least 10 constant-height atomic force microscope images of the organic molecule are obtained. Claim 3 The method according to claim 1 or 2, wherein step (a) is performed at at least three different height distances, preferably at at least 10 different probe height distances.
4. The method according to any one of claims 1 to 3, wherein the tip of the functionalized metal probe used in step (a) is selected from Cu, Ag or Pt.
5. The method according to any one of claims 1 to 4, wherein the tip of the functionalized metal probe used in step (a) is functionalized with an inert closed-shell atom or molecule.
6. The method according to any one of claims 1 to 5, wherein the tip of the functionalized metal probe used in step (a) is functionalized with Xe atoms or CO molecules.
7. A frequency modulation atomic force microscope (FM-AFM) comprising a tip of a functionalized metal probe configured to perform step (a) of the method according to claims 1 to 6, and a data processing device configured to perform steps (b) and (c) of the method according to claims 1 to 6.
8. The FM-AFM according to claim 7, further comprising a display device connected to the data processing device and configured to display the name of the molecule according to IUPAC obtained in step (e) of the method according to any one of claims 1 to 6.
9. The FM-AFM according to any one of claims 7 or 8, wherein the display device connected to the data processing device is further configured to display the structural representation of the molecule identified in step (e) of the method according to any one of claims 1 to 6 in the form of a ball-and-stick depiction.
10. The FM-AFM according to any one of claims 7 to 9, wherein the metal of the tip of the functionalized metal probe is selected from Cu, Ag or Pt.
11. The FM-AFM according to any one of claims 7 to 10, wherein the tip of the functionalized metal probe is functionalized with an inert closed-shell atom or molecule.
12. The FM-AFM according to any one of claims 7 to 11, wherein the tip of the functionalized metal probe is functionalized with Xe atoms or CO molecules.
13. A computer program comprising instructions which, when executed by a data processing device, cause the data processing device to perform steps (b) to (e) according to the method according to claims 1 to 6 in the data processing device.
14. A computer-readable data carrier storing the computer program according to claim 13.