Image analysis device, image analysis method, and computer program product
By constructing an analysis model based on machine learning, using the feature quantity of the compound structure image to generate linear marker symbol information, the problem of difficulty in identifying variations of the compound structure image to write, and the adaptation to writing changes and accurate symbol information generation are achieved.
Patent Information
- Application Number
- CN202080087306.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-16
- Filing Date
- 2020-12-16
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-12-16
AI Technical Summary
The prior art is difficult to effectively identify and process writing variations and new writing methods in compound structural images, resulting in failure of recognition.
By constructing an analysis model based on machine learning, the structural image feature quantity of the object compound is used to generate symbol information represented by linear marking method, which is adapted to the changes in structural writing method.
It realizes adaptation to the changes in the image writing method of the compound structure formula, can generate accurate structural symbol information, and solves the problem of recognition failure.
Smart Images

Figure CN114846508B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image analysis device, an image analysis method, and a program, and more particularly to an image analysis device, an image analysis method, and a program for analyzing an image representing a chemical structure of a compound. Background Art
[0002] Chemical structures of compounds are often processed as image data, for example, publicly available on the Internet or incorporated into document data. However, it is difficult to retrieve chemical structures of compounds processed as image data by ordinary retrieval methods.
[0003] On the other hand, in order to be able to retrieve the chemical structure of a compound represented by an image, a technique for recognizing the chemical structure from an image of the chemical structure of a compound using computer-based automatic recognition technology has been developed. As a specific example, the techniques described in Patent Documents 1 and 2 can be cited.
[0004] The technique described in Patent Document 1 performs pattern recognition on character information (for example, atoms constituting a chemical substance) in a chemical structure diagram, and recognizes the line diagram information (for example, bonds between atoms) of the chemical structure diagram by a prescribed algorithm.
[0005] In the technique described in Patent Document 2, an image of a chemical structure of a compound is read, an attribute value representing an atomic symbol is assigned to a region (pixel) representing an atomic symbol in the image, and an attribute value representing a bond symbol is assigned to a region (pixel) representing a bond symbol.
[0006] Prior Art Documents
[0007] Patent Documents
[0008] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2013-61886
[0009] Patent Document 2: Japanese Unexamined Patent Application Publication No. 2014-182663 Summary of the Invention
[0010] Technical Problem to be Solved by the Invention
[0011] In the techniques described in Patent Documents 1 and 2, the correspondence between a part representing a partial structure (constituent element) in a chemical structure of a compound and the partial structure in an image representing the chemical structure of the compound is regularized, and the chemical structure in the image is recognized according to the rule.
[0012] However, there are many equivalent styles in the notation style of chemical structures, and in addition, the thickness and direction of bond lines in chemical structures can also vary depending on the writing method. In this case, in order to cope with different writing methods of chemical structures, it is necessary to prepare many rules for recognizing partial structures described in various writing methods.
[0013] In addition, in the technologies described in Patent Documents 1 and 2, for example, since there are no recognition rules prepared for images of structural formulas described in a new writing style, it may not be possible to perform recognition.
[0014] The present invention has been created in view of the above circumstances, and is an invention that solves the problems of the above-described prior art. Specifically, an object of the present invention is to provide an image analysis apparatus, an image analysis method, and a program for implementing the image analysis method that can correspond to changes in the writing style of a structural formula when generating character information of the structural formula from an image representing the structural formula of a compound.
[0015] Means for Solving Technical Problems
[0016] To achieve the above object, the image analysis apparatus of the present invention includes a processor that analyzes an image representing a structural formula of a compound, and is characterized in that the processor generates symbol information representing the structural formula of the target compound in linear notation based on the feature amount of the target image representing the structural formula of the target compound through an analysis model, and the analysis model is constructed by machine learning using learning images and symbol information representing the structural formulas of the compounds represented by the learning images in linear notation.
[0017] In addition, preferably, the processor detects the target image from a document including the target image, and generates symbol information of the structural formula of the target compound by inputting the detected target image into the analysis model.
[0018] Furthermore, more preferably, the processor detects the target image from the document using an object detection algorithm.
[0019] Furthermore, even more preferably, the processor detects a plurality of target images from a document including a plurality of target images, and generates symbol information of the structural formulas of the target compounds represented by the plurality of detected target images by inputting each of the detected target images into the analysis model.
[0020] In addition, it may be that the analysis model includes: a feature amount output model that outputs a feature amount by being input with the target image; and a symbol information output model that outputs symbol information corresponding to the feature amount by being input with the feature amount.
[0021] Furthermore, it may be that the feature amount output model includes a convolutional neural network, and the symbol information output model includes a recurrent neural network.
[0022] In addition, preferably, the symbol information of the structural formula of the target compound is composed of a plurality of symbols, and the symbol information output model sequentially determines the symbols constituting the symbol information corresponding to the feature amount from the beginning of the symbol information, and outputs the symbol information in which the symbols are arranged in the determined order.
[0023] In addition, it may also be that the processor generates a plurality of symbol information for the structural formula of the target compound based on the feature quantity of the object image through an analysis model. In this case, more preferably, the symbol information output model calculates the output probability of each of the plurality of symbols constituting the symbol information for each symbol information, calculates the output score of the symbol information based on the calculated output probabilities of the plurality of symbols, and outputs a predetermined number of symbol information according to the calculated output score.
[0024] In addition, it is further preferred that the processor performs a determination process for determining whether there is an abnormality in the marking for each of the symbol information output by the symbol information output model, and outputs the normal symbol information without abnormality in the symbol information output by the symbol information output model as the symbol information of the structural formula of the target compound.
[0025] In addition, it is further preferred that the processor generates first description information for describing the structural formula of the target compound using a description method different from the linear notation method based on the object image through a comparison model, generates second description information for describing the structural formula represented by the normal symbol information using the description method, compares the first description information and the second description information, and outputs the normal symbol information as the symbol information of the structural formula of the target compound according to the coincidence degree between the first description information and the second description information.
[0026] In addition, it is further preferred that the comparison model is constructed by machine learning, and the machine learning uses second learning images and description information for describing the structural formula of the compounds represented by the second learning images using the above description method.
[0027] In addition, it is further preferred that the comparison model includes: a feature quantity output model that outputs a feature quantity by being input with the object image; and a description information output model that outputs first description information corresponding to the feature quantity by being input with the feature quantity output from the feature quantity output model.
[0028] In addition, it may also be that the analysis model is constructed by machine learning, and the machine learning uses learning images, symbol information representing the structural formula of the compounds represented by the learning images using the linear notation method, and description information for describing the structural formula of the compounds represented by the learning images using a description method different from the linear notation method. In this case, it may also be that the analysis model includes: a feature quantity output model that outputs a feature quantity by being input with the object image; a description information output model that outputs description information of the structural formula of the target compound by being input with the object image; and a symbol information output model that outputs symbol information corresponding to the combined information by being input with the combined information obtained by combining the output feature quantity and description information.
[0029] In addition, preferably, the feature quantity output model outputs vectorized feature quantities, and the description information output model outputs description information composed of vectorized molecular fingerprints.
[0030] Alternatively, the linear notation may be the Simplified Molecular Input Line Entry System notation or the canonical Simplified Molecular Input Line Entry System notation.
[0031] Alternatively, the above object may be achieved by an image analysis method that analyzes an image representing the structural formula of a compound. In this method, a processor performs a step of generating symbol information representing the structural formula of the target compound in linear notation based on the feature quantities of the target image representing the structural formula of the target compound through an analysis model. The analysis model is constructed by machine learning, and the machine learning uses learning images and symbol information representing the structural formula of the compound represented by the learning images in linear notation.
[0032] Alternatively, a program may be implemented to cause a processor to perform the steps of the above image analysis method.
[0033] Advantages of the Invention
[0034] According to the present invention, it is possible to cope with changes in the writing of structural formulas and appropriately generate character information of structural formulas based on images representing the structural formulas of compounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is an explanatory diagram regarding the linear notation of structural formulas.
[0036] Figure 2 It is a conceptual diagram of the analysis model.
[0037] Figure 3 It is a diagram showing the hardware structure of an image analysis device according to an embodiment of the present invention.
[0038] Figure 4 It is a diagram showing the process flow of image analysis.
[0039] Figure 5 It is an explanatory diagram regarding molecular fingerprints.
[0040] Figure 6 It is a diagram showing a state in which multiple target images are detected from a document.
[0041] Figure 7It is a conceptual diagram of the control model.
[0042] Figure 8 It is a conceptual diagram of the analysis model related to the modified example. Detailed implementation manners
[0043] Next, with reference to the accompanying drawings, an image analysis apparatus, an image analysis method, and a program according to an embodiment of the present invention (hereinafter referred to as "the present embodiment") will be described.
[0044] In addition, the following embodiments are merely examples given for the purpose of easily understanding the present invention, and do not limit the present invention. That is, the present invention is not limited to the following embodiments, and various improvements or changes can be made without departing from the gist of the present invention. In addition, of course, equivalents thereof are also included in the present invention.
[0045] In addition, in the following description, unless otherwise specified, it is assumed that "documents" and "images" are electronic (digitized) documents and images, which are information (data) that can be processed by a computer.
[0046] <Functions of the image analysis apparatus of the present embodiment>
[0047] The image analysis apparatus of the present embodiment includes a processor that analyzes an image representing a structural formula of a compound. The main function of the image analysis apparatus of the present embodiment is to analyze an image (object image) representing the structural formula of an object compound and generate symbol information of the structural formula represented by the object image. Here, the "object compound" is a compound that is the object of generating symbol information of the structural formula. For example, an organic compound or the like that shows a structural formula in an image included in a document conforms to this.
[0048] The "image representing the structural formula" is an image representing a line drawing of the structural formula. There are various equivalent recording methods for recording the structural formula. For example, it is possible to cite the omission of the mark of the single bond of the hydrogen atom (H), the omission of the mark of the carbon atom (C) of the skeleton, and the abbreviated mark of the functional group. In addition, the line drawing can be changed according to the drawing method (for example, the thickness, length, and direction of the bond line between atoms). In addition, in the present embodiment, the resolution of the image representing the structural formula is included in the writing method of the structural formula.
[0049] "Symbol information" is information representing the structural formula of a compound using a linear notation method, which is formed by arranging multiple symbols (e.g., ASCII symbols). As the linear notation method, there are SMILES (Simplified Molecular Input Entry System) notation, canonical SMILES, SMARTS (Smiles Arbitrary Target Specification) notation, SLN (Sybyl Line Notation) notation, WLN (Wiswesser Line-Formula Notation) notation, ROSDAL (Representation of structure diagram arranged linearly) notation, InChI (International Chemical Identifier), and InChI Key (hashed InChI), etc.
[0050] Any one of the above linear notation methods can be used, but in terms of being relatively simple and widely used, the SMILES notation is preferred. In addition, in terms of uniquely determining the label considering the sequence and order of atoms within the molecule, canonical SMILES is also preferred. Furthermore, in this embodiment, a notation method for generating symbol information representing the structural formula according to the SMILES notation is adopted. Additionally, hereinafter, labeling using the SMILES notation is also referred to as SMILES labeling.
[0051] The SMILES notation is a notation method for converting the structural formula of a compound into a single-line symbol information (character information) composed of multiple symbols. The symbols used in the SMILES notation represent the type of atom (element), the bond between atoms, the branched structure, and the cutting position when the ring structure is cut to form a chain structure, etc., according to the specified rules.
[0052] In addition, as an example of the structural formula of a compound labeled using the SMILES notation, that is, symbol information, Figure 1 shows an example of (S)-bromochlorofluoromethane. Figure 1 In it, the left side represents the structural formula, and the right side represents the symbol information (the structural formula labeled with SMILES).
[0053] The image analysis device of the present embodiment performs machine learning using a learning image representing the structural formula of a compound and symbol information (correct label information) of the structural formula represented by the learning image as a learning data set. Through this machine learning, an analysis model is constructed. The above analysis model generates symbol information of the structural formula represented by an image based on the feature amount of the image representing the structural formula of the compound. The analysis model will be described in detail in a later section.
[0054] In addition, the image analysis device of the present embodiment has a function of detecting an image (object image) from a document including an image representing the structural formula of a compound. Then, by inputting the detected object image into the above analysis model, symbol information of the structural formula represented by the object image is generated.
[0055] With the functions as described above, when a document such as a paper or a patent specification includes an image representing the structural formula of a compound, the image can be detected, and the structural formula of the compound represented by the image can be converted into symbol information.
[0056] In addition, since the structural formula converted into symbol information can be used as a search keyword hereafter, it is possible to easily search for a document including an image representing the structural formula of a compound as a target.
[0057] Furthermore, the image analysis device of the present embodiment has a function of checking the correctness of the symbol information generated by the analysis model. More specifically, in the present embodiment, multiple pieces of symbol information can be obtained from the feature amount of one object image, and for each piece of symbol information, it is determined whether there is an abnormality in the labeling (for example, an incorrect label in the SMILES notation).
[0058] In addition, for each piece of symbol information (normal symbol information) for which no abnormality is found, the following comparison process is performed. Then, based on the result of the comparison process, a specified number of pieces of normal symbol information are output as the symbol information of the structural formula of the object compound.
[0059] As described above, by checking the symbol information generated by the analysis model, accurate information can be obtained as the symbol information of the structural formula of the object compound.
[0060] <Regarding the analysis model>
[0061] The analysis model (hereinafter, the analysis model M1) used in the present embodiment will be described. As Figure 2 shown, the analysis model M1 is composed of a feature amount output model Ma and a symbol information output model Mb. The analysis model M1 is constructed by machine learning that uses a learning image representing the structural formula of a compound and symbol information (correct data) of the structural formula shown in the learning image as a learning data set, and uses multiple learning data sets.
[0062] In addition, regarding the number of learning data sets for machine learning, from the viewpoint of improving learning accuracy, it is preferably large, and it is more preferably set to 50,000 or more.
[0063] In the present embodiment, the machine learning is supervised learning, and the method thereof is deep learning (i.e., a multi-layer neural network), but it is not limited thereto. Regarding the type (algorithm) of machine learning, it may also be unsupervised learning, semi-supervised learning, reinforcement learning, or transduction.
[0064] In addition, regarding the technology of machine learning, it may also be genetic programming, inductive logic programming, support vector machines, clustering, Bayesian networks, extreme learning machines (ELMs), or decision tree learning.
[0065] In addition, as a method for minimizing the objective function (loss function) in the machine learning of neural networks, the gradient descent method can be used, or the error backpropagation method can also be used.
[0066] The feature quantity output model Ma is a model that outputs the feature quantity of an object image by being input with an image (object image) representing the structural formula of an object compound. For example, it is composed of a convolutional neural network (CNN) having a convolutional layer and a pooling layer in the intermediate layer. Here, the feature quantity of the image refers to the feature quantity learned in the convolutional neural network CNN, and is the feature quantity determined in the process of general image recognition (pattern recognition). In the present embodiment, the feature quantity output model Ma outputs a vectorized feature quantity.
[0067] In addition, in the present embodiment, the feature quantity output model Ma may also use a network model for image classification. As such a model, for example, a 16-layer CNN (VGG16) of the Oxford visual geometry group, an Inception model (GoogLeNet) of Google Inc., a 152-layer CNN (Resnet) of Kaiming He, and a modified Iception model (Xception) of Chollet can be cited.
[0068] The size of the image input to the feature quantity output model Ma is not particularly limited, but for a compound image, it may be, for example, a size of 75×75 in height and width. Alternatively, for the reason of improving the output accuracy of the model, the size of the compound image may be set to a larger size (e.g., 300×300). In addition, in the case of a color image, for the reason of reducing computational processing, it may be converted into a black-and-white monochromatic image, and the monochromatic image is input into the feature quantity output model Ma.
[0069] In addition, after repeatedly arranging a convolutional layer and a pooling layer in the intermediate layer, a fully-connected layer is arranged, and a multi-dimensional vectorized feature quantity is output from the fully-connected layer. Further, the feature quantity (multi-dimensional vector) output from the fully-connected layer is input into a symbol information output model Mb after passing through a linear layer.
[0070] The symbol information output model Mb is a model that outputs symbol information (character information obtained by performing SMILES notation on a structural formula) of the structural formula of an object compound by being input with the feature quantity output from the feature quantity output model Ma. The symbol information output model Mb is constituted by, for example, an LSTM (Long Short Term Memory) network which is a type of recurrent neural network (RNN). The LSTM replaces the hidden layer of the RNN with an LSTM layer.
[0071] In addition, in the present embodiment, as Figure 2 shown, an embedding layer (Embedding layer: denoted as Wemb in Figure 2 ) can be arranged in the front stage of each LSTM layer, and an inherent vector is given to the input to each LSTM layer. Further, a softmax function (denoted as softmax in Figure 2 ) is applied to the output from each LSTM layer to convert the output from each LSTM layer into a probability. The sum of the n (n is a natural number) output probabilities to which the softmax function is applied is 1.0. In the present embodiment, the output from each LSTM layer is converted into a probability by the softmax function, and the cross entropy error is used as a loss function to obtain the loss (the difference between the learning result and the correct data).
[0072] In the present embodiment, the symbol information output model Mb is constituted by an LSTM network, but is not limited thereto, and the symbol information output model Mb can also be constituted by a GRU (Gated Recurrent Unit).
[0073] When an object image is input into the analysis model M1, the analysis model M1 configured as described above generates a plurality of pieces of symbol information for the structural formula of the object compound based on the feature quantity of the object image.
[0074] The step of generating the symbol information will be described. When the object image is input into the feature quantity output model Ma, the feature quantity output model Ma outputs the feature quantity of the object image, and this feature quantity is input into the symbol information output model Mb. The symbol information output model Mb sequentially determines symbols constituting the symbol information corresponding to the input feature quantity from the beginning of the symbol information, and outputs the symbol information in which the symbols are arranged in the determined order.
[0075] More specifically, when the symbol information output model Mb outputs symbol information composed of m (m is a natural number greater than or equal to 2) symbols, for each of the first to mth symbols, multiple candidates are output from the corresponding LSTM layer. Based on the combination of candidates determined for each of the first to mth symbols, the symbol information is determined. For example, in the case of m = 3, when there are 3 candidates for the first symbol, 4 candidates for the second symbol, and 5 candidates for the third symbol, 60 (= 3 × 4 × 5) kinds of symbol information are determined.
[0076] In addition, the number of combinations of symbols (i.e., the number of symbol information) is not limited to the number when all the multiple candidates determined for each of the first to mth symbols are combined. For example, for the purpose of reducing the load of computational processing, a search algorithm such as beam search may be applied to the multiple candidates determined for each of the first to mth symbols, and the top K (K is a natural number) symbols among the multiple candidates may be adopted.
[0077] Next, the symbol information output model Mb calculates the output probability of each of the m symbols constituting the symbol information for each symbol information. For example, when j (j is a natural number) candidates are output for the ith (i = 1 to m) symbol in the symbol information of the structural formula of the target compound, the output probability P of each of the j symbols is calculated by the softmax function i1 、P i2 、P i3 …P ij 。
[0078] After that, the symbol information output model Mb calculates the output score of each symbol information based on the calculated output probability of each symbol. Here, the output score is the sum when the output probabilities of each of the m symbols constituting each symbol information are all added together. However, it is not limited to this, and the product when the output probabilities of each of the m symbols constituting each symbol information are multiplied may also be used as the output score.
[0079] Then, the symbol information output model Mb outputs a predetermined number of symbol information according to the calculated output score. In this embodiment, Q symbol information are output in order starting from the symbol information with the highest calculated output score. Here, regarding the number Q of the output symbol information, it can also be arbitrarily determined, but it is preferably about 2 to 20. However, it is not limited to this, and regarding the structural formula of the target compound, only one symbol information with the highest output score may be output. Or, the symbol information with the number equivalent to the combination number after all the candidates of each symbol are combined may be output.
[0080] <Structure of the image analysis device of this embodiment>
[0081] Next, refer to Figure 3A structural example of the image analysis device according to this embodiment (hereinafter referred to as the image analysis device 10) will be described. In addition, in Figure 3 the external interface is described as "External I / F".
[0082] As Figure 3 shown, the image analysis device 10 is a computer in which a processor 11, a memory 12, an external interface 13, an input device 14, an output device 15, and a storage 16 are electrically connected to each other. In addition, in Figure 3 the structure shown, the image analysis device 10 is composed of one computer, but the image analysis device 10 may also be composed of multiple computers.
[0083] The processor 11 is configured to execute the program 21 described later and perform a series of processes related to image analysis. In addition, the processor 11 is composed of one or more CPUs (Central Processing Unit) and the program 21 described later.
[0084] The hardware processor constituting the processor 11 is not limited to a CPU, and may also be an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a GPU (Graphics Processing Unit), an MPU (Micro-Processing Unit), or other IC (Integrated Circuit), or may also be a hardware processor formed by combining them. In addition, the processor 11 may also be an IC (Integrated Circuit) chip represented by an SoC (System on Chip) or the like, which performs the functions of the entire image analysis device 10.
[0085] In addition, the above-mentioned hardware processor may also be a circuit (Circuitry) formed by combining circuit elements such as semiconductor elements.
[0086] The memory 12 is composed of semiconductor memories such as a ROM (Read Only Memory) and a RAM (Random Access Memory), provides a working area for the processor 11 by temporarily storing programs and data, and also temporarily stores various data generated by the processes executed by the processor 11.
[0087] The program stored in the memory 12 includes a program 21 for image analysis. The program 21 includes a program for implementing machine learning and constructing an analysis model M1, a program for detecting an object image from a document, and a program for generating symbolic information of the structural formula of an object compound based on the feature amount of the object image by the analysis model M1. In addition, in the present embodiment, the program 21 further includes a program for performing determination processing and comparison processing on the generated symbolic information.
[0088] In addition, the program 21 can be obtained by reading from a computer-readable recording medium (medium), or can also be obtained by receiving (downloading) via a network such as the Internet or an intranet.
[0089] The external interface 13 is an interface for connecting to an external device. The image analysis device 10 communicates with an external device, such as a scanner or another computer on the Internet, via the external interface 13. Through this communication, the image analysis device 10 can obtain part or all of the data for machine learning, and can also obtain a document in which an object image is posted.
[0090] The input device 14 is composed of, for example, a mouse and a keyboard, and receives user input operations. The image analysis device 10 can obtain part of the data for machine learning, for example, by the user inputting character information equivalent to symbolic information via the input device 14.
[0091] The output device 15 is composed of, for example, a display and a speaker, and is a device for displaying the symbolic information generated by the analysis model M1 or for performing audio playback.
[0092] The storage 16 is composed of, for example, a flash memory, an HDD (Hard Disc Drive), an SSD (Solid State Drive), an FD (Flexible Disc), an MO disk (Magneto-Optical disc), a CD (Compact Disc), a DVD (Digital Versatile Disc), an SD card (Secure Digital card), and a USB memory (Universal Serial Bus memory), etc. Various data including data for machine learning are stored in the storage 16. In addition, data of various models constructed by machine learning, with the analysis model M1 as the head, are also stored in the storage 16. In addition, the symbolic information of the structural formula of the object compound generated by the analysis model M1 can be stored in the storage 16 and registered in advance as a database.
[0093] In addition, in the present embodiment, the storage 16 is a device built into the image analysis device 10, but is not limited thereto. The storage 16 may also be an external device connected to the image analysis device 10, or may be an external computer (for example, a server computer for cloud services) connected in a manner that enables communication via a network.
[0094] Regarding the hardware structure of the image analysis device 10, it is not limited to the above structure, and the constituent devices can be appropriately added, omitted, and replaced according to the specific embodiment.
[0095] <Regarding the image analysis process>
[0096] Next, the image analysis process using the image analysis device 10 will be described.
[0097] In addition, in the image analysis process described below, the image analysis method of the present invention is adopted. That is, in the following description, the image analysis method of the present invention is included. In addition, each step in the image analysis process constitutes the image analysis method of the present invention.
[0098] As Figure 4 shown, the image analysis process of the present embodiment proceeds in the order of the learning stage S001, the symbol information generation stage S002, and the symbol information inspection stage S003. Hereinafter, each stage will be described.
[0099] [Learning stage]
[0100] The learning stage S001 is a stage in which machine learning is performed to build the models required in the subsequent stages. In the learning stage S001, as Figure 4 shown, the first machine learning S011, the second machine learning S012, and the third machine learning S013 are performed.
[0101] The first machine learning S011 is machine learning for building the analysis model M1. As described above, it is performed using the learning images and the symbol information of the structural formulas of the compounds represented by the learning images as the learning dataset.
[0102] The second machine learning S012 is machine learning for building the reference model used in the symbol information inspection stage S003. The reference model is a model that generates description information for describing the structural formula of the target compound using a description method different from the linear notation method.
[0103] As a description method different from the linear notation method, for example, a description method based on molecular fingerprints can be cited. Molecular fingerprints are used to identify molecules with certain characteristics, as Figure 5Convert the structural formula into a binary multi-dimensional vector indicating the presence or absence of partial structures (fragments) representing various types in the structural formula. Here, the partial structure refers to an element representing a part of the structural formula and includes multiple atoms and bonds between atoms.
[0104] The dimension of the vector constituting the molecular fingerprint can be arbitrarily determined, for example, set to several tens to several thousands of dimensions. In this embodiment, mimicking MACCS Keys, which is a representative fingerprint, a molecular fingerprint represented by a 167-dimensional vector is used.
[0105] In addition, the description method different from the linear notation method is not limited to the molecular fingerprint, and can also be other description methods, such as those based on the KEGG (Kyoto Encyclopedia of Genes and Genomes) Chemical Function format (KCF format), the input format of the chemical structure database (MACCS) operated by Molecular Design Limited, i.e., the MOL notation, and the variant of MOL, i.e., the SDF method.
[0106] The second machine learning S012 is performed using a learning image (second learning image) representing the structural formula of the compound and description information of the structural formula represented by the second learning image (specifically, description information composed of molecular fingerprints) as a learning dataset. Here, the second learning image used for the second machine learning S012 can be the same image as the learning image used in the first machine learning S011, or can also be an image separately prepared from the learning image used in the first machine learning S011.
[0107] Then, by performing the second machine learning S012 using the above learning data, a control model is constructed. The control model will be described in detail later.
[0108] The third machine learning S013 is machine learning for constructing a model (hereinafter referred to as an image detection model) for detecting an image representing the structural formula of a compound from a document. The image detection model is a model that detects an image of a structural formula from a document using an object detection algorithm. As the object detection algorithm, for example, R-CNN (Region-based CNN), Fast R-CNN, YOLO (You only Look Once), and SDD (Single Shot MultiboxDetector) can be used. In this embodiment, from the perspective of detection speed, an image detection model using YOLO is constructed.
[0109] The learning data (teacher data) for the third machine learning S013 is created by applying an annotation tool to a learning image representing the structural formula of a compound. The annotation tool is a tool that assigns correct labels (identifications) and information related to the coordinates of the object, etc. as annotations to the data to be the object. The annotation tool is started, a document including the learning image is displayed, the region representing the structural formula of the compound is surrounded by a bounding box, and this region is annotated, thereby creating the learning data.
[0110] In addition, as the annotation tool, for example, labeImg of tzutalin company and VoTT of microsoft company can be used.
[0111] Then, by using the above-mentioned learning data for the third machine learning S013, an image detection model as an object detection model in the YOLO format is constructed.
[0112] [Symbol information generation stage]
[0113] The symbol information generation stage S002 is a stage in which an image (object image) of the structural formula of the target compound included in the document is analyzed to generate symbol information of the structural formula of the target compound.
[0114] In the symbol information generation stage S002, first, the processor 11 of the image analysis device 10 applies the image detection model to the document including the object image to detect the object image in the document (S021). That is, in this step S021, the processor 11 uses an object detection algorithm (specifically, YOLO) to detect the object image from the document.
[0115] In addition, when multiple object images are included in one document, as Figure 6 shown, the processor 11 detects multiple images from the above-mentioned document (the images in the part surrounded by the dotted line in Figure 6 ).
[0116] Next, the processor 11 inputs the detected object image into the analysis model M1 (S022). In the analysis model M1, the feature amount of the object image is output in the feature amount output model Ma in the previous stage, and in the symbol information output model Mb in the latter stage, based on the feature amount of the input object image, the symbol information of the structural formula of the target compound is output. At this time, as described above, the symbol information with a high output score is output first, and the symbol information of a predetermined number is output in order. As described above, the processor 11 generates multiple symbol information for the structural formula of the target compound based on the feature amount of the object image through the analysis model M1 (S023).
[0117] In addition, when multiple object images are detected in step S021, the processor 11 inputs the detected multiple object images into the analysis model M1 for each object image. In this case, for the structural formulas of the object compounds represented by the multiple object images, multiple symbol information is generated for each object image.
[0118] [Symbol Information Checking Phase]
[0119] The symbol information checking phase S003 is a phase for performing a determination process and a comparison process for each of the multiple symbol information generated for the structural formula of the object compound in the symbol information generation phase S002.
[0120] In the symbol information checking phase S003, first, the processor 11 performs a determination process (S031). The determination process is a process for determining whether there is an abnormality in the SMILES notation for each of a specified number of symbol information output from the symbol information output model Mb of the analysis model M1.
[0121] Specifically described, in order for the processor 11 to determine whether the string forming each symbol information is the correct word order of the SMILES notation for each symbol information output from the symbol information output model Mb, it attempts to convert from this string to a structural formula. Here, if the conversion to the structural formula is successful, it is determined that there is no abnormality in the symbol information (in other words, the symbol information is normal). Hereinafter, the symbol information without abnormality is referred to as "normal symbol information".
[0122] In addition, as an algorithm for converting from a string to a structural formula, an algorithm similar to the conversion function incorporated in well-known structural formula drawing software such as ChemDraw (registered trademark) and RDKit can be used.
[0123] After performing the determination process, the processor 11 performs a comparison process (S032) on the normal symbol information. The comparison process is a process for comparing the first description information of the structural formula of the object compound generated by the comparison model and the second description information generated from the normal symbol information. The first description information is information that describes the structural formula of the object compound in the form of a molecular fingerprint. In the present embodiment, the first description information is generated by inputting the object image into Figure 7 the illustrated comparison model M2.
[0124] The comparison model M2 is constructed by the second machine learning S012, as Figure 7 shown, and includes a feature quantity output model Mc and a description information output model Md.
[0125] The feature quantity output model Mc is the same as the feature quantity output model Ma of the analysis model M1. It is a model that outputs the feature quantity of the object image by inputting an image (object image) representing the structural formula of the object compound, and is composed of a CNN in this embodiment. In addition, in this embodiment, the feature quantity output model Mc outputs vectorized feature quantities in the same way as the feature quantity output model Ma.
[0126] The description information output model Md is a model that outputs description information (specifically, description information composed of molecular fingerprints) corresponding to the feature quantity by inputting the feature quantity output from the feature quantity output model Mc. In this embodiment, the description information output model Md is composed of a neural network (NN) for example. The description information output model Md outputs description information composed of vectorized molecular fingerprints as the first description information. The description information output from the description information output model Md is the description information of the structural formula of the object compound.
[0127] In addition, as the feature quantity output model Mc of the control model M2, the feature quantity output model Ma of the analysis model M1 can also be used concurrently. That is, the weights of the intermediate layer of the CNN can be set to a common value between the feature quantity output models Ma and Mc. In this case, the second machine learning S012 fixes the weights of the intermediate layer of the CNN determined by the first machine learning S011 as they are, and determines the weights of the intermediate layer of the NN serving as the description information output model Md, which can reduce the load (computational load) of model construction. However, the control model M2 can also be composed of another CNN without using the CNN (feature quantity output model Ma) of the analysis model M1 concurrently.
[0128] The second description information is the description information that describes the structural formula represented by the normal symbol information in the form of a molecular fingerprint. In this embodiment, the second description information is generated by converting the symbol information marked with SMILES into a molecular fingerprint according to the conversion rule. The conversion rule used at this time is specified by regularizing the correspondence between the structural formula marked with SMILES and the molecular fingerprint for many compounds.
[0129] In the control process, the first description information and the second description information generated as described above are compared, and the degree of overlap between the two description information is calculated. In the case where there are multiple normal symbol information, the second description information is generated for each normal symbol information, and the degree of overlap with the first description information is calculated for each second description information. In addition, as a method for calculating the degree of overlap, a known method for calculating the similarity between molecular fingerprints can be used. For example, the calculation method of the Tanimoto coefficient can be used.
[0130] After performing the comparison process, the processor 11 performs an output process (S033). The output process is a process of finally outputting (for example, displaying) normal symbol information as symbol information of the structural formula of the target compound based on the degree of coincidence calculated in the comparison process. Here, outputting normal symbol information based on the degree of coincidence can be, for example, outputting only normal symbol information whose degree of coincidence exceeds a reference value, or it can also be outputting normal symbol information in order starting from the normal symbol information with a high degree of coincidence.
[0131] <Regarding the effectiveness of this embodiment>
[0132] The image analysis device 10 of this embodiment can use the analysis model M1 constructed by the first machine learning to generate symbol information with the structural formula marked with SMILES based on the feature amount of the target image representing the structural formula of the target compound. As a result, it is possible to appropriately correspond to changes in the writing method of the structural formula in the target image.
[0133] To describe the above effect in detail, in the prior art, the correspondence relationship between a part of the image representing the structural formula of a compound and the partial structure in the structural formula appearing in that part is regularized, and the structural formula is recognized according to the recognition rule. However, when the writing method of the structural formula changes, if there is no recognition rule adaptable to the writing method, the structural formula cannot be recognized. As a result, in the above case, it is difficult to generate symbol information of the structural formula.
[0134] In contrast, in this embodiment, the analysis model M1, which is the result of machine learning, is used to generate symbol information based on the feature amount of the target image. That is, in this embodiment, even if the writing method of the structural formula changes, the feature amount of the image representing the structural formula can be determined, and if the feature amount can be determined, symbol information can be generated based on the feature amount.
[0135] As described above, according to this embodiment, even when the writing method of the structural formula of the target compound changes, symbol information can be appropriately obtained.
[0136] <Other embodiments>
[0137] In summary, specific examples have been given to illustrate the image analysis device, image analysis method, and program of the present invention, but the above embodiments are only examples, and other embodiments can also be considered.
[0138] For example, the computer that constitutes the image analysis device may also be a server for ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service), etc. In this case, a user who uses the above ASP and other services operates a terminal (not shown) to send a document containing the object image to the server. When the server receives the document sent from the user, it detects the object image from the document and generates symbol information of the structural formula of the object compound represented by the object image based on the feature amount of the object image. Then, the server outputs (sends) the generated symbol information to the user's terminal. On the user side, the symbol information sent from the server is displayed or audio playback is performed.
[0139] In addition, in the above-described embodiment, it is assumed that a determination process for determining whether there is an abnormality in the label is performed on the symbol information generated by the analysis model M1. Further, in the above-described embodiment, it is assumed that a comparison process for comparing the molecular fingerprint (first description information) generated based on the feature amount of the object image and the molecular fingerprint (second description information) obtained by converting from the normal symbol information is performed.
[0140] However, it is not limited thereto, and either the determination process or the comparison process may be performed alone, or no process may be performed.
[0141] In addition, in the above-described embodiment, it is assumed that machine learning (first to third machine learning) for constructing various models is performed by the image analysis device 10, but it is not limited thereto. Part or all of the machine learning may also be performed by other devices (computers) different from the image analysis device 10. In this case, the image analysis device 10 acquires the models constructed by the machine learning performed by other devices.
[0142] For example, when the first machine learning is performed by other devices, the image analysis device 10 acquires the analysis model M1 constructed by the first machine learning from other devices. Then, the image analysis device 10 analyzes the object image through the acquired analysis model M1 and generates symbol information for the structural formula of the object compound represented by the image.
[0143] In addition, in the above-described embodiment, the analysis model M1 is constructed by machine learning that uses learning images and symbol information representing the structural formula of the compound represented by the learning images in linear labels. Moreover, the analysis model M1 generates symbol information of the structural formula of the object compound represented by the object image based on the feature amount of the object image.
[0144] However, it is not limited to this. As an analysis model for the symbol information of the structural formula of the target compound, other models can be considered. For example, the analysis model shown in Figure 8 can be cited (hereinafter, the analysis model M3 related to the modified example).
[0145] As shown in Figure 8 , the analysis model M3 related to the modified example has a feature quantity output model Me, a description information output model Mf, and a symbol information output model Mg. The analysis model M3 of the modified example is constructed by machine learning (hereinafter, the machine learning related to the modified example). The machine learning related to the modified example uses learning images representing the structural formula of the compound, symbol information (e.g., SMILES-labeled symbol information) of the structural formula of the compound represented by the learning images, and description information (e.g., description information composed of molecular fingerprints) of the structural formula of the compound represented by the learning images as a learning data set for training.
[0146] Similar to the feature quantity output model Ma of the analysis model M1, the feature quantity output model Me outputs the feature quantity of the object image by being input with an image (object image) representing the structural formula of the target compound. For example, it is composed of a CNN. The feature quantity output model Me outputs vectorized feature quantities (e.g., 2048-dimensional vectors).
[0147] The description information output model Mf is a model that outputs description information (specifically, description information composed of molecular fingerprints) of the structural formula of the target compound by being input with the object image. The description information output model Mf is a model based on the control model M2. For example, it is composed of a CNN and outputs description information composed of vectorized molecular fingerprints (e.g., 167-dimensional vectors).
[0148] In the analysis model M3 related to the modified example, as shown in Figure 8 , the feature quantity output from the feature quantity output model Me and the description information output from the description information output model Mf are synthesized to generate vectorized synthesized information. The vector dimension of the synthesized information is the value obtained by adding the vector dimension of the feature quantity and the vector dimension of the description information (i.e., 2215 dimensions).
[0149] The symbol information output model Mg is a model that outputs symbol information (specifically, SMILES-labeled symbol information) corresponding to the synthesized information by being input with the above-mentioned synthesized information. The symbol information output model Mg is almost the same as the symbol information output model Mb of the analysis model M1. For example, it is composed of an RNN. As an example, an LSTM network can be used.
[0150] Even when using the analysis model M3 related to the modification example configured as described above, it is also possible to generate symbol information representing the structural formula of the target compound with linear markers based on the feature amounts of the target image.
[0151] Symbol Explanation
[0152] 10 Image analysis device
[0153] 11 Processor
[0154] 12 Memory
[0155] 13 External interface
[0156] 14 Input device
[0157] 15 Output device
[0158] 16 Storage
[0159] 21 Program
[0160] M1 Analysis model
[0161] M2 Control model
[0162] M3 Analysis model related to the modification example
[0163] Ma, Mc, Me Feature amount output models
[0164] Mb, Mg Symbol information output models
[0165] Md, Mf Description information output models.
Claims
1. An image analysis device includes a processor that analyzes an image representing a structural formula of a compound, wherein the processor generates symbol information representing the structural formula of the target compound in a linear notation method based on a feature amount of a target image representing the structural formula of the target compound through an analysis model; the analysis model is constructed by machine learning using a learning image and symbol information representing the structural formula of the compound represented by the learning image in the linear notation method; the analysis model is constructed by machine learning, and the machine learning uses the learning image, the symbol information representing the structural formula of the compound represented by the learning image in the linear notation method, and description information describing the structural formula of the compound represented by the learning image in a description method different from the linear notation method; the analysis model includes: a feature amount output model that outputs the feature amount by being input with the target image; a description information output model that outputs the description information of the structural formula of the target compound by being input with the target image; and a symbol information output model that outputs the symbol information corresponding to the synthesis information by being input with the synthesis information obtained by synthesizing the output feature amount and the description information.
2. The image analysis device according to claim 1, wherein the feature amount output model outputs the vectorized feature amount; the description information output model outputs the description information composed of vectorized molecular fingerprints.
3. An image analysis device includes a processor that analyzes an image representing a structural formula of a compound, wherein the processor generates symbol information representing the structural formula of the target compound in a linear notation method based on a feature amount of a target image representing the structural formula of the target compound through an analysis model; the analysis model is constructed by machine learning using a learning image and symbol information representing the structural formula of the compound represented by the learning image in the linear notation method; the analysis model includes: a feature amount output model that outputs the feature amount by being input with the target image; and a symbol information output model that outputs the symbol information corresponding to the feature amount by being input with the feature amount, the processor performs a determination process for determining whether there is an abnormality in the marking for each of the symbol information output by the symbol information output model, outputs the normal symbol information without the abnormality in the symbol information output by the symbol information output model as the symbol information of the structural formula of the target compound, the processor generates first description information describing the structural formula of the target compound in a description method different from the linear notation method based on the target image through a comparison model, generates second description information describing the structural formula represented by the normal symbol information in the description method, compares the first description information and the second description information, Output the normal symbol information as the symbol information of the structural formula of the target compound according to the overlap degree between the first description information and the second description information.
4. The image analysis device according to claim 3, wherein the feature quantity output model includes a convolutional neural network, the symbol information output model includes a recurrent neural network.
5. The image analysis device according to claim 3, wherein the symbol information of the structural formula of the target compound is composed of multiple symbols, the symbol information output model sequentially determines the symbols constituting the symbol information corresponding to the feature quantity from the beginning of the symbol information, and outputs the symbol information in the order of the determined symbols.
6. The image analysis device according to claim 5, wherein the processor, through the analysis model, based on the feature quantity of the target image, generates multiple pieces of the symbol information for the structural formula of the target compound, the symbol information output model for each piece of the symbol information, calculates the output probability of each of the multiple symbols constituting the symbol information, and calculates the output score of the symbol information based on the calculated output probabilities of the multiple symbols, and outputs a predetermined number of pieces of the symbol information according to the calculated output score.
7. The image analysis device according to claim 3, wherein the control model is constructed by machine learning, and the machine learning uses second learning images and description information describing the structural formula of the compounds represented by the second learning images.
8. The image analysis device according to claim 3, wherein the control model includes: a feature quantity output model that outputs the feature quantity by being input with the target image; and a description information output model that outputs the first description information corresponding to the feature quantity by being input with the feature quantity output from the feature quantity output model.
9. The image analysis device according to any one of claims 1 to 8, wherein the processor detects the target image from a document containing the target image, and generates the symbol information of the structural formula of the target compound by inputting the detected target image into the analysis model.
10. The image analysis device according to claim 9, wherein the processor uses an object detection algorithm to detect the target image from the document.
11. The image analysis device according to claim 9, wherein the processor detects multiple target images from the document containing multiple target images, and generates the symbol information of the structural formula of the target compound represented by each of the multiple detected target images by inputting each of the detected multiple target images into the analysis model.
12. The image analysis device according to any one of claims 1 to 8, wherein the linear notation is the simplified molecular linear input specification notation or the canonical simplified molecular linear input specification notation.
13. An image analysis method that analyzes an image representing a structural formula of a compound, wherein a processor performs a step of generating symbol information representing the structural formula of the target compound in a linear notation method based on feature amounts of a target image representing the structural formula of the target compound through an analysis model; the analysis model is constructed by machine learning, and the machine learning uses learning images and symbol information representing the structural formula of the compound represented by the learning images in a linear notation method; the analysis model is constructed by machine learning, and the machine learning uses the learning images, the symbol information representing the structural formula of the compound represented by the learning images in the linear notation method, and description information describing the structural formula of the compound represented by the learning images in a description method different from the linear notation method; the analysis model includes: a feature amount output model that outputs the feature amounts by being input with the target image; a description information output model that outputs the description information of the structural formula of the target compound by being input with the target image; and a symbol information output model that outputs the symbol information corresponding to the synthesis information by being input with the synthesis information obtained by synthesizing the output feature amounts and the description information.
14. A computer program product including a program for causing a processor to perform the steps of the image analysis method according to claim 13.
Citation Information
Patent Citations
Chemical structure diagram recognition system and computer program for chemical structure diagram recognition system
JP2013061886A
Information processing program, information processing method and information processing device
JP2014182663A