Information Processing Apparatus, Information Processing Method, and Program
By constructing an identification model, using machine learning technology to identify the constituent elements in the compound structural image of different writing methods and use them for compound search, the problem of difficult to identify and utilize this information in the prior art is solved, and efficient identification and retrieval of compound structural image is achieved.
Patent Information
- Application Number
- CN202080089203.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-26
- Filing Date
- 2020-10-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2040-10-30
AI Technical Summary
The prior art is difficult to identify the constituent elements in the compound structural formula images of different writing methods, and it is impossible to effectively use these recognition results for compound search.
By constructing an identification model, machine learning technology is used to identify constituent elements based on the feature quantity of compound structural images, and the recognition results are stored as feature information for subsequent compound retrieval.
Regardless of the way of writing the structural formula, the constituent elements in the compound structural formula image can be accurately identified and used for compound retrieval, which improves the search efficiency and accuracy.
Smart Images

Figure CN114868192B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program, and more particularly to an information processing apparatus, an information processing method, and a program capable of retrieving information on a structural formula of a compound represented as an image. Background Art
[0002] In many cases, the structural formula of a compound is processed as image data, for example, publicly available on the Internet or incorporated into document data. However, it is difficult to retrieve the structural formula of a compound processed as image data by ordinary retrieval methods.
[0003] On the other hand, in order to be able to retrieve the structural formula of a compound represented by an image, a technique for recognizing the structural formula from an image of the structural formula of a compound using computer-based automatic recognition technology has been developed. As a specific example, the techniques described in Patent Documents 1 and 2 can be cited.
[0004] The technique described in Patent Document 1 performs pattern recognition on character information (for example, atoms constituting a chemical substance) in a chemical structure diagram, and recognizes line diagram information (for example, bonds between atoms) of the chemical structure diagram by a prescribed algorithm.
[0005] In the technique described in Patent Document 2, an image of the structural formula of a compound is read, an attribute value representing an atomic symbol is assigned to a pixel representing an atomic symbol in the image, and an attribute value representing a bond symbol is assigned to a pixel representing a bond symbol.
[0006] Prior Art Documents
[0007] Patent Documents
[0008] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2013-61886
[0009] Patent Document 2: Japanese Unexamined Patent Application Publication No. 2014-182663 Summary of the Invention
[0010] Technical Problem to be Solved by the Invention
[0011] In the techniques described in Patent Documents 1 and 2, the correspondence relationship between each region of an image representing the structural formula of a compound and the constituent elements in the structural formula represented by each region is regularized. Moreover, each constituent element in the structural formula represented by the image is recognized according to this rule.
[0012] However, there are various equivalent ways of recording structural formulas. In addition, the thickness and direction in the structural formula can vary depending on the writing method. In this case, in order to cope with different writing methods of the structural formula, it is necessary to prepare in advance many rules for identifying each constituent element in the structural formula recorded in various writing methods. For a structural formula recorded in a writing method for which the identification rules are not prepared, it is difficult to identify each constituent element contained therein.
[0013] On the other hand, when each constituent element in the structural formula is identified from an image of the structural formula representing a certain compound, information about each identified constituent element can be useful information for later retrieving the above compound.
[0014] The present invention has been completed in view of the above circumstances and is an invention that solves the above-mentioned problems of the prior art. Specifically, an object thereof is to provide an information processing apparatus, an information processing method, and a program that can identify each constituent element of a structural formula from an image representing the structural formula regardless of the writing method of the structural formula, and can use the identification result for later compound retrieval.
[0015] Means for Solving Technical Problems
[0016] In order to achieve the above object, the information processing apparatus of the present invention includes a processor, and is characterized in that the processor identifies, through an identification model and based on the feature amounts of each region in an object image representing the structural formula of an object compound, the constituent elements represented by each region among the constituent elements of the structural formula of the object compound, stores element information about the constituent elements in the structural formula of the identified object compound in association with the object compound, and the identification model is constructed by machine learning using a learning image representing one constituent element in the structural formula of a compound.
[0017] Alternatively, in machine learning, when using a plurality of learning images representing constituent elements having the same chemical structure but different recording methods, an identification model for deriving common feature amounts from the plurality of learning images may be constructed by machine learning.
[0018] In addition, preferably, the processor acquires input information related to a retrieval compound, and retrieves an object compound corresponding to the retrieval compound from the object compounds storing the element information based on the input information and the element information associated with the object compound.
[0019] In the above structure, more preferably, the processor calculates the similarity between the retrieval compound and the object compound based on the input information and the element information stored in association with the object compound, and retrieves an object compound whose similarity satisfies the retrieval condition from the object compounds storing the element information as the retrieval compound.
[0020] In addition, it is further preferable that the processor acquires input information related to the constituent elements included in the structural formula of the retrieved compound.
[0021] Alternatively, the processor may detect an object image from a document including the object image, and identify the constituent elements represented by the respective regions in the object image by inputting the detected object image into a recognition model.
[0022] In the above structure, it is more preferable that the processor detects the object image from the document using an object detection algorithm.
[0023] In addition, the element information may also include information indicating the types of the constituent elements in the structural formula of the identified object compound. At this time, the element information may further include information indicating the arrangement positions of the constituent elements in the structural formula of the identified object compound in the coordinate space set with respect to the object image.
[0024] In the above structure, the information indicating the types of the constituent elements may also be information indicating the types of atoms or bonds between atoms corresponding to the constituent elements.
[0025] Alternatively, the information indicating the types of the constituent elements may also be information indicating the chemical formulas of the functional groups corresponding to the constituent elements.
[0026] Alternatively, the information indicating the types of the constituent elements may also be information constituted by a part of a molecular fingerprint, and the molecular fingerprint indicates the presence or absence of the constituent elements in the structural formula of the object compound for each type of the constituent elements.
[0027] In addition, the object may be achieved by the following information processing method: the processor executes a step of identifying, through a recognition model, the constituent elements represented by the respective regions among the constituent elements included in the structural formula of the object compound based on the feature amounts of the respective regions in the object image representing the structural formula of the object compound; and a step of storing, in association with the object compound, element information regarding the constituent elements in the structural formula of the identified object compound, and the recognition model is constructed by machine learning using a learning image representing one constituent element in the structural formula of the compound.
[0028] Alternatively, a program for causing the processor to execute the respective steps of the above information processing method may also be implemented.
[0029] Advantages of the Invention
[0030] According to the present invention, regardless of the way of writing the structural formula, it is possible to identify the respective constituent elements of the structural formula from an image representing the structural formula, and the identification result can be used for subsequent compound retrieval. Brief Description of the Drawings
[0031] Figure 1It is an explanatory diagram of the constituent elements in the structural formula of a compound.
[0032] Figure 2 It is a diagram showing an example of a database that stores element information for each compound.
[0033] Figure 3 It is a conceptual diagram of an identification model.
[0034] Figure 4 It is an explanatory diagram regarding the difference in the description method of constituent elements.
[0035] Figure 5 It is a diagram showing the structure of an information processing apparatus according to an embodiment of the present invention.
[0036] Figure 6 It is a diagram showing the information processing flow of an information processing apparatus according to an embodiment of the present invention.
[0037] Figure 7 It is a diagram of a state in which multiple object images are detected from a single document.
[0038] Figure 8 It is a diagram showing an example of a screen for displaying a search result of an object compound.
[0039] Figure 9 It is an explanatory diagram regarding a molecular fingerprint. Detailed implementation mode
[0040] Hereinafter, with reference to the accompanying drawings, an information processing apparatus, an information processing method, and a program according to an embodiment of the present invention (hereinafter, referred to as "the present embodiment") will be described.
[0041] In addition, the following embodiments are merely examples given for easy understanding of the present invention and do not limit the present invention. That is, the present invention is not limited to the following embodiments, and various modifications or changes can be made without departing from the gist of the present invention. Of course, equivalents thereof are also included in the present invention.
[0042] In addition, in the following description, unless otherwise specified, it is assumed that "documents" and "images" are electronic (digitized) documents and images, which are information (data) that can be processed by a computer.
[0043] <Functions of the information processing apparatus of the present embodiment>
[0044] The information processing apparatus according to the present embodiment (hereinafter simply referred to as "information processing apparatus") includes a processor that can analyze an image (object image) representing a structural formula of an object compound and identify each constituent element in the structural formula. The object compound is, for example, a compound whose structural formula is represented in an image in a document and the constituent elements represented by each region in the image are identified by the information processing apparatus.
[0045] The image representing the structural formula is an image representing a line drawing of the structural formula. There are various equivalent ways of recording the structural formula. For example, omission of the single bond mark of a hydrogen atom (H), omission of the mark of a skeletal carbon atom (C), and abbreviated marks of functional groups can be cited. In addition, the line drawing can be changed according to the drawing method (for example, the thickness, length, and direction of the bond line between atoms). In addition, in the present embodiment, the resolution of the image representing the structural formula is included in the writing method of the structural formula.
[0046] The constituent elements in the structural formula refer to atoms, bond lines between atoms, or a combination thereof that constitute the structural formula. In the present embodiment, as Figure 1 shown, each atom (for example, Figure 1 "Bend C" and "O" in Figure 1 ) and each bond line (for example,
[0047] "Double" in Figure 1 ) that constitute the structural formula correspond to the constituent elements.
[0048] The information processing apparatus uses one constituent element (specifically, the label information of the constituent element) in the structural formula of the compound and a learning image representing one constituent element as a learning data set to perform machine learning. Through this machine learning, an identification model is constructed. The identification model is a model that identifies the constituent element represented by each region among the constituent elements in the structural formula based on the feature amounts of the regions of the image representing the structural formula of the compound. In addition, the identification model will be described in detail in a later section.
[0049] In addition, the information processing apparatus has a function of detecting an image (object image) from a document in which an image representing a structural formula of a compound is posted. The detected object image is input into the above-mentioned identification model. Thereby, each constituent element in the structural formula of the compound (object compound) represented by the object image is identified.
[0050] Moreover, the information processing apparatus acquires element information for each constituent element in the identified target compound. In the present embodiment, the element information includes information indicating the type of the identified constituent element and information indicating the arrangement position of the constituent element.
[0051] In the present embodiment, the information indicating the type of the constituent element is information indicating the type of atom or bond between atoms that conforms to the constituent element. In the case of the compound shown in Figure 1 "Bend C", "O", and "Double" conform.
[0052] The information indicating the arrangement position of the constituent element is information indicating the arrangement position of the constituent element in a coordinate space set with respect to the target image (for example, a two-dimensional coordinate space in which the horizontal direction of the target image is set as the X direction and the vertical direction is set as the Y direction). In the present embodiment, with a reference position (for example, the upper left vertex position) in the target image as the origin, as the arrangement position of the constituent element, the representative position and size (for example, the lengths in the X and Y directions) of the rectangular region surrounding the constituent element are expressed in pixel units.
[0053] For each of the plurality of constituent elements included in the structural formula of the target compound, element information is acquired. The acquired element information is stored in association with the target compound. For example, as shown in Figure 2 , it is stored in a state bound to a document or the like that includes an image showing the structural formula of the target compound.
[0054] In addition, in the present embodiment, the information indicating the type of the constituent element in the element information is automatically acquired by the recognition model recognizing each constituent element in the structural formula. Further, the information indicating the arrangement position of the constituent element in the element information is automatically acquired by analyzing an image (i.e., the target image) including the region representing the constituent element.
[0055] The information processing apparatus repeatedly executes the above series of processes (specifically, detecting an image from a document, recognizing each constituent element in the structural formula, and acquiring and storing the element information) for various target compounds. As a result, the element information regarding each constituent element in the structural formula of the target compound is stored as information related to the target compound. As a result, a database that separately collects the element information for each target compound is constructed (refer to Figure 2 ).
[0056] In addition, the information processing apparatus has a function of retrieving a target object compound, that is, an object compound corresponding to a retrieved compound, using the element information stored in the database as a retrieval keyword. For example, a user who performs a retrieval inputs image information representing the structural formula of a retrieved compound. The information processing apparatus acquires the image information as input information, and retrieves an object compound corresponding to the retrieved compound from the object compounds storing the element information, based on the acquired input information and the element information stored in the database.
[0057] As described above, according to the information processing apparatus, it is possible to detect an image of the structural formula of a compound included in a document such as a thesis or a patent specification, and databaseize information (element information) on each constituent element in the structural formula represented by the image. Moreover, by using the database, it is possible to easily retrieve a target compound. Thus, for example, it is possible to simply find a document that publishes an image representing the structural formula of a target compound.
[0058] <Regarding the recognition model>
[0059] The recognition model used in the present embodiment (hereinafter referred to as recognition model M1) will be described.
[0060] The recognition model M1 is a model for recognizing each constituent element included in the structural formula from an image (object image) representing the structural formula of an object compound. As Figure 3 shown, the recognition model M1 of the present embodiment is composed of a feature quantity derivation model Ma and a constituent element output model Mb.
[0061] The feature quantity derivation model Ma is a model that derives the feature quantities of each region of the object image by inputting the object image. In the present embodiment, the feature quantity derivation model Ma is composed of, for example, a convolutional neural network (CNN) having a convolutional layer and a pooling layer in an intermediate layer. As a model of CNN, for example, a 16-layer CNN (VGG16) of the Oxford visual geometry group, an Inception model (GoogLeNet) of Google Inc., a 152-layer CNN (Resnet) of Kaiming He, and a modified Iception model (Xception) of Chollet can be cited.
[0062] When deriving the feature quantities of each region in the object image by the feature quantity derivation model Ma, each region in the specific object image is specified. Specifically, each constituent element included in the structural formula represented by the object image is detected, and a region surrounding each detected constituent element is specified for each constituent element. Such a region specifying function is installed in the feature quantity derivation model Ma by machine learning described later.
[0063] The feature quantity of the image output from the feature quantity derivation model Ma is the learning feature quantity in the convolutional neural network CNN, and is a specific feature quantity in the process of general image recognition (pattern recognition). Then, the feature quantities of each region derived by the feature quantity derivation model Ma are input into the component output model Mb for each region.
[0064] The component output model Mb is a model that, by inputting the feature quantities of each region derived by the feature quantity derivation model Ma into each region, outputs the components (for example, the types of components) corresponding to the feature quantity to each region. In the present embodiment, the component output model Mb is constituted by, for example, a neural network (NN).
[0065] When the component output model Mb according to the present embodiment outputs the components corresponding to the feature quantities of the respective regions of the target image, a plurality of candidates (candidates for components) are specified for each region. The softmax function is applied to the plurality of candidates specified for each region, and the output probability is calculated for each candidate. The output probability is a numerical value indicating the certainty (accuracy) corresponding to the component represented by each region for each of the plurality of candidates. In addition, the sum of the n (n is a natural number) output probabilities to which the softmax function is applied is 1.0.
[0066] The component output model Mb outputs, as the components represented by each region, the candidate determined according to the output probability among the plurality of candidates specified for each region, for example, the candidate with the highest output probability. Thus, in the present embodiment, each component in the structural formula represented by the target image is determined based on the output probabilities of the respective candidates from among the plurality of candidates specified based on the feature quantities of the respective regions of the target image.
[0067] The recognition model M1 described above (in other words, each of the above two models Ma and Mb) uses a learning image representing a component in the structural formula of a compound and the label (correct label) of the component as a learning data set, and is constructed by machine learning using a plurality of learning data sets.
[0068] In addition, regarding the number of learning data sets for machine learning, from the viewpoint of improving learning accuracy, it is preferably large, and it is preferably set to 50,000 or more.
[0069] In the present embodiment, the machine learning is supervised learning, and the method is deep learning (that is, a multi-layer neural network), but it is not limited thereto. Regarding the type (algorithm) of machine learning, it may also be unsupervised learning, semi-supervised learning, reinforcement learning, or transduction.
[0070] In addition, regarding the technology of machine learning, it can also be genetic programming, inductive logic programming, support vector machines, clustering, Bayesian networks, extreme learning machines (ELMs), or decision tree learning.
[0071] In addition, as a method for minimizing the objective function (loss function) in the machine learning of neural networks, the gradient descent method can be used, or the error backpropagation method can also be used.
[0072] In addition, in the machine learning of the present embodiment, multiple learning images representing constituent elements with the same chemical structure but different notations are sometimes used. For example, as Figure 4 , when a certain constituent element (in Figure 4 , xylyl is illustrated) is described in an equivalent notation, it is possible to assume a case where machine learning is performed using learning images prepared according to the notations respectively. Or, it is also possible to assume a case where machine learning is performed using multiple learning images representing constituent elements with the same chemical structure but different bond line thicknesses, lengths, or directions between atoms.
[0073] In the case as described above, an identification model M1 (strictly speaking, a feature quantity derivation model Ma) for deriving a common feature quantity from multiple learning images is constructed through machine learning. For example, for each of the learning images of two xylyls with different notations shown in Figure 4 , the same label (correct label) of "xylyl" is attached and supervised learning is performed. Thus, an identification model M1 is constructed that can derive a common feature quantity from the learning images of two xylyls with different notations and output the same constituent element (xylyl) from each image.
[0074] <Structure of the information processing device of the present embodiment>
[0075] Next, a structural example of the information processing device (hereinafter referred to as information processing device 10) shown in Figure 5 will be described. In addition, in Figure 5 , the external interface is described as "External I / F".
[0076] As shown in Figure 5 , the information processing device 10 is a computer in which a processor 11, a memory 12, an external interface 13, an input device 14, an output device 15, and a storage 16 are electrically connected to each other.
[0077] In addition, in the present embodiment, the information processing device 10 is composed of one computer, but the information processing device 10 can also be composed of multiple computers.
[0078] The processor 11 is configured to execute the program 21 described below and perform processing for realizing the functions of the information processing apparatus 10 described above. In addition, the processor 11 is composed of one or more CPUs (Central Processing Unit) and the program 21 described below.
[0079] The hardware processor constituting the processor 11 is not limited to a CPU, and may also be an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a GPU (Graphics Processing Unit), an MPU (Micro-Processing Unit), or other IC (Integrated Circuit), or may also be a hardware processor formed by combining them. In addition, the processor 11 may also be an IC (Integrated Circuit) chip that represents, for example, an SoC (System on Chip) and realizes the functions of the entire information processing apparatus 10.
[0080] In addition, the above-mentioned hardware processor may also be a circuitry formed by combining circuit elements such as semiconductor elements.
[0081] The memory 12 is composed of semiconductor memories such as a ROM (Read Only Memory) and a RAM (Random Access Memory), provides a working area for the processor 11 by temporarily storing programs and data, and also temporarily stores various data generated by the processing executed by the processor 11.
[0082] Stored in the memory 12 is a program 21 for causing a computer to function as the information processing apparatus 10 of the present embodiment. The program 21 includes the following programs p1 to p5.
[0083] p1: A program for constructing an identification model M1 through machine learning
[0084] p2: A program for detecting an object image from a document in which the object image is posted
[0085] p3: A program for identifying each constituent element in the structural formula represented by the object image
[0086] p4: A program for storing element information about the identified constituent elements
[0087] p5: A program for retrieving an object compound corresponding to a retrieval compound from an object compound storing element information
[0088] In addition, the program 21 can be obtained by reading it from a computer-readable recording medium, or can also be obtained by receiving (downloading) it via a network such as the Internet or an intranet.
[0089] The external interface 13 is an interface for connecting to an external device. The information processing device 10 communicates with an external device, such as a scanner or another computer on the Internet, via the external interface 13. Through this communication, the information processing device 10 can obtain data for machine learning, and can also obtain a document carrying an object image.
[0090] The input device 14 is composed of, for example, a mouse and a keyboard, and accepts input operations from the user. The information processing device 10 can, for example, obtain data for machine learning by the user drawing constituent elements via the input device 14. In addition, when the user retrieves an object compound corresponding to a retrieval compound, the input device 14 is operated to input information related to the retrieval compound. Thereby, the information processing device 10 can obtain input information related to the retrieval compound.
[0091] The output device 15 is composed of, for example, a display and a speaker, and is a device for displaying an object compound retrieved based on input information (i.e., an object compound corresponding to the retrieval compound) or for performing audio playback. In addition, the output device 15 can output the element information stored in the database for each object compound.
[0092] The storage 16 is composed of, for example, a flash memory, an HDD (Hard Disc Drive), an SSD (Solid State Drive), an FD (Flexible Disc), an MO disc (Magneto-Optical disc), a CD (Compact Disc), a DVD (Digital Versatile Disc), an SD card (Secure Digital card), and a USB memory (Universal Serial Bus memory), etc. Various data including data for machine learning are stored in the storage 16. Moreover, various models constructed by machine learning, with the recognition model M1 as the leading one, are also stored in the storage 16.
[0093] Furthermore, in the storage 16, element information about each constituent element in the structural formula of the object compound identified by the recognition model M1 is stored in association with the object compound. As a result, a database 22 of the element information shown in Figure 2 is constructed in the storage 16.
[0094] In the database 22, for each target compound, element information regarding each constituent element included in the structural formula of the target compound is stored. Specifically, it is the type and arrangement position of the constituent elements.
[0095] As Figure 2 shown, the type of the constituent element stored in the database 22 is the type of the constituent element with the highest output probability calculated by the recognition model M1, and it is stored together with its output probability (marked as "accuracy" in the figure).
[0096] In addition, the arrangement position of the constituent element stored in the database 22 is a position represented in a coordinate space with the reference position of the target image as the origin. For example, it is represented by the representative position, the length in the X direction, and the length in the Y direction of a rectangular area surrounding the constituent element.
[0097] In addition, as Figure 2 shown, the element information regarding each constituent element in the structural formula of the target compound is stored in a bound manner with the information related to the document in which the image (target image) representing the structural formula is published. As the information related to the document, for example, the title of the paper when the document is a paper, or the publication number of the gazette when the document is a gazette, as well as the page on which the target image is published in the document and its arrangement position on that page, etc. can be cited.
[0098] Furthermore, in the present embodiment, the storage 16 is a device built into the information processing apparatus 10, but it is not limited thereto. The storage 16 may also include an external device connected to the information processing apparatus 10. In addition, the storage 16 may also include an external computer (for example, a server computer for cloud services) connected in a manner that enables communication via a network. In this case, part or all of the above database 22 may also be stored in the external computer constituting the storage 16.
[0099] Regarding the hardware structure of the information processing apparatus 10, it is not limited to the above structure, and constituent devices can be appropriately added, omitted, and replaced according to specific embodiments.
[0100] <Regarding the information processing flow>
[0101] Next, the information processing flow using the information processing apparatus 10 will be described.
[0102] In addition, in the information processing flow described below, the information processing method of the present invention is adopted. That is, each step in the following information processing flow constitutes the information processing method of the present invention.
[0103] As Figure 6As shown in the figure, the information processing flow of this embodiment proceeds in the order of the learning stage S001, the database construction stage S002, and the retrieval stage S003. Hereinafter, each stage will be described.
[0104] [Learning Stage]
[0105] The learning stage S001 is a stage where machine learning is implemented to construct a model required in later stages. In the learning stage S001, as Figure 6 shown, the first machine learning S011, the second machine learning S012, and the third machine learning S013 are implemented.
[0106] The first machine learning S011 is machine learning for constructing an identification model M1. As described above, it is implemented using a learning image representing a constituent element of the structural formula of a compound. In this embodiment, supervised learning is implemented as the first machine learning S011. In this supervised learning, a learning image and a label (correct label) of a constituent element represented by the learning image are used.
[0107] In addition, in the first machine learning S011, as described above, multiple learning images representing constituent elements with the same chemical structure but different notations are sometimes used. Thus, an identification model M1 (strictly speaking, a feature quantity derivation model Ma) that derives common feature quantities from multiple learning images is constructed.
[0108] The second machine learning S012 is machine learning for constructing a model (hereinafter referred to as an image detection model) for detecting an image of a structural formula of a compound from a document in which an image representing the structural formula of a compound is posted. The image detection model is a model for detecting an image of a structural formula from a document using an object detection algorithm. As the object detection algorithm, R-CNN (Region-based CNN), Fast R-CNN, YOLO (You only Look Once), and SDD (Single Shot MultiboxDetector) can be used. In this embodiment, from the perspective of detection speed, an image detection model using YOLO is constructed.
[0109] The learning data (teacher data) for the second machine learning S012 is produced by applying an annotation tool to a learning image representing the structural formula of a compound. The annotation tool is a tool that gives correct labels (identifications) and relevant information such as the coordinates of the object to the data to be the object as annotations. The annotation tool is started, a document including the learning image is displayed, the area representing the structural formula of the compound is surrounded by a bounding box, and this area is annotated, thereby producing the learning data.
[0110] In addition, as annotation tools, for example, labelImg of tzutalin company and VoTT of microsoft company can be used.
[0111] By using the above learning data for the second machine learning S012, an image detection model as an object detection model in the YOLO format is constructed.
[0112] The third machine learning S013 is machine learning for constructing a model (hereinafter referred to as a retrieval model) for retrieving an object compound corresponding to a retrieval compound from among a plurality of object compounds whose element information is stored in the database 22.
[0113] The retrieval model of the present embodiment is a model for retrieving, from among the object compounds whose element information is stored in the database 22, an object compound having the same or similar structural formula as the retrieval compound as the retrieval compound.
[0114] In addition, hereinafter, it is assumed that the input information is information related to each constituent element included in the structural formula of the retrieval compound. For example, it is image information representing the structural formula of the retrieval compound. However, as the input information, if it is content that can specify at least a part of the structural formula of the retrieval compound (that is, information that can be a key when retrieving the retrieval compound using the database 22), it can also be other information. For example, it can also be image information representing a part of the constituent elements in the structural formula of the retrieval compound. In addition, information equivalent to the element information (for example, information representing the type of the constituent element and the arrangement position of the constituent element in the structural formula) can be used as the input information. Moreover, a part or all of the structural formula of the retrieval compound can be drawn using a known structural formula drawing software such as ChemDraw (registered trademark) and RDKit, and the drawing data can be used as the input information.
[0115] The retrieval model is composed of a retrieval compound specific model and a similarity evaluation model. The retrieval compound specific model is a model for specifying the structural formula of the retrieval compound represented by the input information. In the present embodiment, when the image information as the input information is input to the retrieval compound specific model, information related to each constituent element in the structural formula represented by the image information (for example, information representing the type of each constituent element and the arrangement position in the structural formula) is output.
[0116] In addition, as the retrieval compound specific model, the recognition model M1 can be used instead, or transfer learning can be implemented as the machine learning in this case.
[0117] The similarity evaluation model evaluates the similarity between the structural formula of the retrieved compound specified by a specific model for the retrieved compound and the structural formula of the target compound whose element information of each constituent element is stored in the database 22. In the present embodiment, the similarity is evaluated based on the element information of the constituent elements included in the structural formula of the retrieved compound and the element information of the constituent elements included in the structural formula of the target compound.
[0118] The algorithm of the similarity evaluation model is not particularly limited. For example, a well-known algorithm for evaluating the similarity between images or the computational degree between texts can be used. For example, an algorithm that vectorizes the element information of the constituent elements included in the structural formula and calculates the similarity between vectors using an index such as the Euclidean distance can be used.
[0119] In addition, it is preferable that the similarity becomes higher between multiple structural formulas written in different notations for the same chemical substance. This is because in the structural formulas recorded in different notations for the same compound, the writing method of each functional group (for example, the direction of the bond line, etc.) and the position of each atom, etc. in each structural formula will change. Considering such differences, it is preferable to increase the similarity between the structural formulas recorded in different notations for the same compound. For example, it is only necessary to label the same label (correct label) for each of the multiple structural formulas recorded in the database 22 and written in different notations for the same compound for machine learning to construct the similarity evaluation model.
[0120] Furthermore, regarding the method for evaluating similarity, it is not limited to the method based on machine learning. For example, according to a pre-specified comparison rule, each constituent element in the structural formula can be compared between the retrieved compound and the target compound, and the similarity can be evaluated based on the comparison result. Or, the target compounds whose element information of each constituent element is stored in the database 22 can be clustered based on the element information, and the similarity can be evaluated by specifying the cluster to which the retrieved compound belongs.
[0121] The third machine learning S013 is implemented using the element information of each constituent element in the structural formula stored in the database 22 for each target compound and the learning information related to the structural formula of the compound. Here, the learning information is, for example, information indicating the types and arrangement positions, etc. of the constituent elements in the structural formula for the compounds selected for the third machine learning S013.
[0122] Then, by implementing the third machine learning, the above-mentioned retrieval model is constructed.
[0123] [Database construction phase]
[0124] In the database construction stage S002, for the structural formula of the target compound represented by the image (target image) included in the document, element information about each constituent element in the structural formula is stored, and the database 22 is constructed.
[0125] In the database construction stage S002, first, the processor 11 of the information processing device 10 applies the image detection model to the document including the target image to detect the target image in the document (S021). That is, in this step S021, the processor 11 uses an object detection algorithm (specifically, YOLO) to detect the target image from the document.
[0126] At this time, when there are multiple target images included in one document, as Figure 7 shown, the processor 11 detects multiple target images from the above-mentioned document ( Figure 7 in the figure, the images of the parts surrounded by the dotted line).
[0127] Next, the processor 11 uses the recognition model M1 to identify each constituent element in the structural formula of the target compound based on the feature amounts of the respective regions of the target image (S023).
[0128] Specifically, the processor 11 inputs the target image detected in step S021 into the recognition model M1. In the feature amount derivation model Ma at the front stage in the recognition model M1, the feature amounts of the respective regions of the target image are output. In the constituent element output model Mb at the back stage, based on the input feature amounts of the respective regions, constituent elements (strictly speaking, the types of constituent elements) are output. At this time, based on the feature amounts of the respective regions, a plurality of candidates for the constituent elements corresponding to the respective regions are specified, and an output probability is calculated for each candidate.
[0129] As described above, the constituent element output model Mb outputs the candidate with the highest output probability as the constituent element represented by each region. By outputting the constituent elements represented by the respective regions in the target image for each region, the structural formula represented by the target image (that is, the structural formula of the target compound) can be divided into constituent elements for recognition.
[0130] In addition, when multiple target images are detected in step S021, the processor 11 inputs the detected multiple target images into the recognition model M1 for each target image. Thereby, for each of the multiple target images, each constituent element in the structural formula of the target compound represented by the target image is recognized.
[0131] Next, the processor 11 acquires element information about each constituent element in the structural formula of the identified target compound, and stores the acquired element information (S023). At this time, the processor 11 stores the element information about each constituent element in association with the target compound whose structural formula includes each constituent element. In the present embodiment, the element information about each constituent element is stored in association with information such as a document of an image (target image) that publishes the structural formula constituted by each constituent element (see Figure 2 ).
[0132] For a new target compound, step S023 is repeated each time each constituent element in its structural formula is identified. As a result, element information about each constituent element in the structural formula of the target compound is stored, and a database 22 of element information is constructed. The target compound whose element information is stored in the database 22 can be retrieved using the element information as a key in the subsequent retrieval stage S003.
[0133] [Retrieval stage]
[0134] The retrieval stage S003 is a stage of retrieving a target compound corresponding to a retrieval compound from the target compounds whose element information is stored in the database 22. The "retrieval compound" is a compound for which information related to a part or all of its structural formula is obtained as input information when performing a retrieval.
[0135] In the retrieval stage S003, first, the processor 11 of the information processing device 10 acquires input information related to the retrieval compound (S031). In this step S031, the processor 11 acquires, as input information, information related to each constituent element included in the structural formula of the retrieval compound. As an example of such information, for example, image information representing the structural formula of the retrieval compound can be cited.
[0136] After acquiring the input information, the processor 11 retrieves a target compound corresponding to the retrieval compound from the target compounds whose element information is stored in the database 22 through the above-described retrieval model (S032). Specifically, the processor 11 calculates the similarity between the retrieval compound and the target compound based on the acquired input information and the element information stored in the database 22 in association with the target compound through the retrieval model. In the present embodiment, the similarity of the structural formula is calculated between the retrieval compound represented by the input information and the target compound whose element information is stored in the database 22.
[0137] After that, the processor 11 retrieves (selects) from the object compounds in which the element information is stored in the database 22 the object compounds whose calculated similarity satisfies the retrieval condition as the retrieval compounds. The retrieval condition is a condition determined in advance for selecting the object compounds corresponding to the retrieval compounds based on the calculation result of the similarity. In the present embodiment, the predetermined number of object compounds are retrieved as the retrieval compounds in descending order of similarity. However, it is not limited thereto. For example, only the object compound with the highest similarity may be retrieved as the retrieval compound. Alternatively, the object compounds with a similarity equal to or higher than the reference value may be retrieved as the retrieval compounds.
[0138] Then, the processor 11 outputs the information of the retrieved object compounds through the output device 15. For example, as Figure 8 shown, the retrieval result is displayed on the screen. As the information of the retrieved object compounds, for example, a document and a page on which an image showing the structural formula of the object compound is published can be cited. In addition, as Figure 8 shown, it is preferable to output the similarity between the retrieved object compounds and the retrieval compounds together with the retrieval result of the object compounds.
[0139] In addition, as input information related to the retrieval compounds, the case of obtaining information representing a part of the constituent elements (hereinafter, for convenience, referred to as "partial structure") included in the structural formula of the retrieval compounds can be considered. In this case, the object compounds including the partial structure are retrieved as the retrieval compounds. Specifically, for each object compound in which the element information is stored in the database 22, the similarity between the partial structure included in its structural formula and the partial structure represented by the input information is calculated. Then, the predetermined number of object compounds are retrieved as the retrieval compounds in descending order of similarity.
[0140] <Regarding the effectiveness of the present embodiment>
[0141] The information processing apparatus 10 of the present embodiment can use the recognition model M1 constructed by the first machine learning S011 to recognize each constituent element in the structural formula based on the feature amounts of the respective regions in the image (object image) representing the structural formula of the object compound. In addition, the information processing apparatus 10 of the present embodiment stores the element information regarding the recognized constituent elements in association with the object compounds, and constructs the database 22. The element information stored in the database 22 can be used as a retrieval keyword when retrieving the object compounds.
[0142] The above effects will be described in detail. In the prior art, the correspondence relationship between each region of an image representing the structural formula of a compound and the constituent elements in the structural formula appearing in each region is regularized, and each constituent element in the structural formula is identified according to this rule. However, when the writing method of the structural formula changes, if an identification rule adaptable to this writing method is not prepared, it may be impossible to identify each constituent element in the structural formula. In this case, due to reasons such as the inability to utilize the identification results of the constituent elements, the retrieval of the structural formula including this constituent element becomes difficult.
[0143] In contrast, in the present embodiment, the identification model M1, which is the result of machine learning, can be used to identify each constituent element in the structural formula based on the feature amounts of each region of the object image. That is, in the present embodiment, even if the writing method of the structural formula changes, the feature amounts of each region of the image representing the structural formula are specified. If the feature amounts can be specified, the constituent elements can be inferred (identified) based on these feature amounts. Moreover, since the element information regarding the identified constituent elements is stored in association with the object compound and databaseized, the element information can thereafter be used as a retrieval keyword to retrieve the target object compound.
[0144] As described above, according to the present embodiment, even when the writing method of the structural formula of the object compound changes, each constituent element in the structural formula can be well identified. Moreover, the element information regarding each identified constituent element can be used as a retrieval keyword to appropriately retrieve the target object compound.
[0145] <Other Embodiments>
[0146] In summary, specific examples have been given to illustrate the information processing apparatus, information processing method, and program of the present invention. However, the above-described embodiments are only examples, and other embodiments can also be considered.
[0147] For example, a computer that constitutes an information processing apparatus may also be a server for ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service), etc. In this case, a user who uses the above ASP and other services operates a terminal (not shown) and sends input information related to the retrieved compound to the server. When the server receives the input information, based on the input information, it retrieves the target compound corresponding to the retrieved compound from the target compounds storing the element information. Then, the server outputs (sends) information related to the retrieval result (i.e., the target compound corresponding to the retrieved compound) to the user's terminal. On the user side, the information sent from the server (i.e., the retrieval result) is displayed or audio playback is performed.
[0148] In addition, in the above-described embodiment, each atom and each bond between atoms included in the structural formula are used as constituent elements, but it is not limited thereto. For example, a functional group (atomic group) containing a plurality of atoms may also be used as a constituent element. In this case, it is preferable that the information indicating the type of the constituent element in the element information of the constituent element is information indicating the chemical formula of the functional group corresponding to the constituent element.
[0149] Alternatively, a plurality of functional groups adjacent to each other in the structural formula may be used as constituent elements, or each fragment when the structural formula is divided according to an arbitrary rule may be used as a constituent element.
[0150] In addition, the information indicating the type of the constituent element in the element information may also be information constituted by a part of the molecular fingerprint of the structural formula of the target compound. The molecular fingerprint is a binary-type multi-dimensional vector indicating the presence or absence of the constituent elements in the structural formula for each type of the constituent elements. For example, for Figure 9 the functional group shown on the left side, the molecular fingerprint shown on the Figure 9 right side is set.
[0151] In addition, in the above-described embodiment, it is assumed that machine learning (first to third machine learning) for constructing various models is performed by the information processing apparatus 10, but it is not limited thereto. A part or all of the machine learning may also be performed by another apparatus (computer) different from the information processing apparatus 10. In this case, the information processing apparatus 10 acquires the model constructed by the machine learning performed by the other apparatus.
[0152] For example, when the first machine learning is performed by another device, the information processing device 10 acquires the recognition model M1 from the other device, and recognizes each constituent element in the structural formula represented by the object image by using the acquired recognition model M1.
[0153] Symbol Explanation
[0154] 10 Information processing device
[0155] 11 Processor
[0156] 12 Memory
[0157] 13 External interface
[0158] 14 Input device
[0159] 15 Output device
[0160] 16 Storage
[0161] 21 Program
[0162] 22 Database
[0163] M1 Recognition model
[0164] Ma Feature quantity derivation model
[0165] Mb Constituent element output model
Claims
1. An information processing apparatus includes a processor, wherein, the processor identifies, by using an identification model and based on feature amounts of respective regions in an object image representing a structural formula of an object compound, the constituent elements represented by the respective regions among the constituent elements in the structural formula of the object compound, stores, in association with the object compound, element information regarding the constituent elements in the identified structural formula of the object compound, the identification model is constructed by machine learning using learning images each representing one constituent element in a structural formula of a compound, in the machine learning, when using a plurality of the learning images representing the same chemical structure but different notations of the constituent element, the identification model that derives common feature amounts from the plurality of the learning images is constructed by the machine learning.
2. The information processing apparatus according to claim 1, wherein, the processor acquires input information related to a retrieval compound, retrieves, from the object compounds storing the element information, the object compound corresponding to the retrieval compound based on the input information and the element information associated with the object compound.
3. The information processing apparatus according to claim 2, wherein, the processor calculates a similarity between the retrieval compound and the object compound based on the input information and the element information stored in association with the object compound, retrieves, from the object compounds storing the element information, the object compound whose similarity satisfies a retrieval condition as the retrieval compound.
4. The information processing apparatus according to claim 2, wherein, the processor acquires the input information related to the constituent elements included in the structural formula of the retrieval compound.
5. The information processing apparatus according to claim 1, wherein, the processor detects the object image from a document including the object image, identifies, by inputting the detected object image into the identification model, the constituent elements represented by the respective regions in the object image.
6. The information processing apparatus according to claim 5, wherein, the processor detects the object image from the document by using an object detection algorithm.
7. The information processing apparatus according to claim 1, wherein, the element information includes information indicating the types of the constituent elements in the identified structural formula of the object compound.
8. The information processing apparatus according to claim 1, wherein, the element information further includes information indicating the arrangement positions of the constituent elements in the identified structural formula of the object compound in a coordinate space set with respect to the object image.
9. The information processing apparatus according to claim 7, wherein, the information indicating the types of the constituent elements is information indicating the types of atoms or bonds between atoms corresponding to the constituent elements.
10. The information processing apparatus according to claim 7, wherein, the information indicating the types of the constituent elements is information indicating chemical formulas of functional groups corresponding to the constituent elements.
11. The information processing apparatus according to claim 7, wherein the information indicating the type of the constituent element is information constituted by a part of a molecular fingerprint, and the molecular fingerprint indicates the presence or absence of the constituent element in the structural formula of the target compound for each type of the constituent element.
12. An information processing method, wherein a processor performs: a step of identifying, by an identification model, the constituent element represented by each region among the constituent elements included in the structural formula of the target compound, based on the feature amount of each region in the target image representing the structural formula of the target compound; and a step of storing, in association with the target compound, element information regarding the constituent element in the structural formula of the identified target compound, wherein the identification model is constructed by machine learning using a learning image representing one constituent element in the structural formula of a compound, and in the machine learning, when using a plurality of the learning images representing the constituent elements having the same chemical structure but different notations, the identification model that derives a common feature amount from the plurality of the learning images is constructed by the machine learning.
13. A program for causing a processor to perform the steps of the information processing method according to claim 12.
Citation Information
Patent Citations
Chemical structure diagram recognition system and computer program for chemical structure diagram recognition system
JP2013061886A
Information processing program, information processing method and information processing device
JP2014182663A
Methods and systems for processing and matching information of chemical substance as well as storage system
CN102436447A