Method for Extracting Molecular Structural Formula, Computer-Readable Medium, and Computing Device
Through the methods of object detection, graph recognition, feature point marking and Markov logical network inference, the time and accuracy of extracting molecular structural formulas in the information source files in the prior art are solved, and fast, efficient and accurate molecular structural formula extraction is achieved.
Patent Information
- Application Number
- CN202010826575.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-08-17
AI Technical Summary
The prior art is difficult to quickly, efficiently and accurately extract the molecular structure formulas in information source files in batches, which are time-consuming and have low accuracy.
Using the steps of object detection, graph recognition, feature point marking and connection relationship inference, the molecular structural formula is identified and extracted using Markov logical network and predefined inference formulas.
It realizes rapid, efficient and accurate batch extraction of molecular structural formulas, simplifies subsequent processing steps, and improves calculation efficiency and accuracy.
Smart Images

Figure CN114155539B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of molecular structural formula image recognition, and more particularly to extracting molecular structural formulas from information source files. Background Art
[0002] For a long time, in research aspects such as medicinal chemistry and organic chemistry, researchers usually need to read a large number of documents and manually record the chemical structural formulas of compounds in the documents. How to automatically identify and extract the molecular structural formulas of compounds from patent documents or non-patent documents by using computer technology has always been an issue worthy of attention in this field. Batch-identifying molecular structural formula image information from a large number of documents, extracting the corresponding molecular structural formulas and storing these molecular structural formulas in a computer-readable form is of great significance for new drug discovery, chemical manufacturing, etc. Summary of the Invention
[0003] A brief overview of the present invention is given below in order to provide a basic understanding of some aspects of the present invention. However, it should be understood that this overview is not an exhaustive overview of the present invention. It is not intended to identify the key or important parts of the present invention, nor is it intended to limit the scope of the present invention. Its purpose is only to present some concepts of the present invention in a simplified form as a prelude to the more detailed description given later.
[0004] A method for extracting molecular structural formulas known to the inventors of the present invention is to train a neural network with a large number of images containing molecular structural formulas to obtain a neural network model for extracting molecular structural formulas, which is used for the recognition and extraction of molecular structural formula images. However, such a method will consume a large amount of time and computing resources and has low accuracy.
[0005] The object of the present invention is to provide a method, a computer-readable medium, and a computing device capable of quickly, efficiently, and accurately batch-extracting molecular structural formulas from information source files.
[0006] According to one aspect of the present invention, a method for extracting molecular structural formulas is provided, including: a target detection step of detecting molecular structural formulas in an information source file and segmenting a molecular structural formula image including the molecular structural formulas from the information source file; a graphic recognition step of performing graphic recognition on the molecular structural formula image to obtain a character part and a skeleton part of the molecular structural formula; a feature point marking step of marking feature points in the skeleton part as first type feature points or second type feature points, where the skeleton part includes one or more line segments, and the feature points include endpoints of the one or more line segments; and a connection relationship inference step of using a Markov logic network to infer connection relationships between the first type feature points, between the second type feature points, and between the first type feature points and the second type feature points according to a predefined inference formula, where the first type feature points represent endpoints not adjacent to characters in the character part, and the second type feature points represent endpoints adjacent to characters in the character part.
[0007] According to another aspect of the present invention, a non-transitory computer-readable medium is provided, storing program instructions that, when executed by a processor, execute the method for extracting molecular structural formulas according to the present invention.
[0008] According to another aspect of the present invention, a computing device is provided, including: a memory storing program instructions; and one or more processors configured to execute the program instructions, so that the computing device executes the method for extracting molecular structural formulas according to the present invention.
[0009] According to the present invention, molecular structural formulas in information source files can be batch-extracted quickly, efficiently, and accurately, which is helpful for aspects such as chemical research, new drug discovery, and chemical manufacturing. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings forming a part of the specification depict embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.
[0011] Referring to the drawings, the present invention can be understood more clearly according to the following detailed description, where:
[0012] Figure 1 An exemplary flowchart of the method for extracting molecular structural formulas according to an embodiment of the present invention is shown;
[0013] Figure 2 An example of an information source file including a molecular structural formula is shown;
[0014] Figure 3 A molecular structural formula image of a certain compound according to an embodiment of the present invention is shown;
[0015] Figure 4 Shows two types of feature points on the skeleton part of the molecular structural formula image of a certain compound according to an embodiment of the present invention;
[0016] Figure 5 Shows the mol file of a certain compound according to an embodiment of the present invention;
[0017] Figure 6 Shows an exemplary configuration of a computing device capable of implementing the embodiments according to the present invention. Detailed Description of the Invention
[0018] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.
[0019] Meanwhile, it should be understood that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships.
[0020] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended as a limitation on the present invention or its application or use.
[0021] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered as part of the specification.
[0022] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.
[0023] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.
[0024] (Method for Extracting Molecular Structural Formula)
[0025] Figure 1 Shows an exemplary flowchart 100 of a method for extracting a molecular structural formula according to an embodiment of the present invention. The following will refer to Figure 1 Generally describe the method for extracting a molecular structural formula according to an embodiment of the present invention. Details of the method will be further described later with reference to Figures 2 to 5 Describe further in conjunction with embodiments.
[0026] As Figure 1As shown, first, in S101, a target detection step is performed. The target detection step includes detecting the molecular structural formula in the information source file and segmenting out the molecular structural formula image including the molecular structural formula from the information source file.
[0027] Next, in S102, a graphic recognition step is performed. The graphic recognition step includes performing graphic recognition on the molecular structural formula image to obtain the character part of the molecular structural formula and the skeleton part including one or more line segments. In some embodiments, the graphic recognition step may further include recording the positions of the characters in the character part.
[0028] Next, in S103, a feature point marking step is performed. The feature point marking step includes marking the feature points in the skeleton part of the molecular structural formula as the first type of feature points or the second type of feature points. The feature points include the endpoints of one or more line segments in the skeleton part. The first type of feature points represents the endpoints not adjacent to the characters in the character part, and the second type of feature points represents the endpoints adjacent to the characters in the character part. In some embodiments, the feature point marking step may further include associating the characters and their positions in the character part with the second type of feature points to obtain the character association relationship between the characters and the second type of feature points.
[0029] Next, in S104, a connection relationship inference step is performed. The connection relationship inference step includes using a Markov logic network to infer the connection relationships between the first type of feature points, between the second type of feature points, and between the first type of feature points and the second type of feature points according to the predefined inference formula.
[0030] In addition, the method for extracting the molecular structural formula according to the embodiment of the present invention may optionally further include a mol file generation step. The mol file generation step includes generating a mol file according to the connection relationship between the feature points and the character association relationship between the feature points and the characters.
[0031] The details of the method for extracting the molecular structural formula according to the embodiment of the present invention will be further described below in conjunction with the embodiments.
[0032] (Target Detection)
[0033] Figure 2 An example of an information source file including a molecular structural formula is shown.
[0034] As Figure 2 shown, in the information source file, there are molecular structural formula images of one or more compounds. The information source file may be, for example, a picture format that can be read by a computer, or other formats that can be converted into a picture format by a computer.
[0035] In some embodiments, the information source file may include patent documents and / or non-patent documents. Patent documents may be, for example, patent documents in the chemical field, patent documents in the pharmaceutical field, etc. Non-patent documents may be, for example, journal papers, conference papers, dissertations, etc. In some embodiments, the information source file may be sourced from web pages, books, news, newspapers, drug instructions, etc.
[0036] According to the method for extracting molecular structural formulas according to an embodiment of the present invention, a target detection step is first performed. The target detection step includes detecting molecular structural formulas from the information source file and segmenting molecular structural formula images including the molecular structural formulas. For example, for Figure 2 the information source file shown, target detection is performed, and six molecular structural formula images 201, 202, 203, 204, 205, and 206 including the molecular structural formulas marked by the dashed boxes can be segmented.
[0037] In some embodiments, the Mask RCNN technique can be used to implement target detection. Mask RCNN is an algorithm framework for target detection using neural networks, which can quickly and accurately detect target images (such as the molecular structural formula images in this article) in the input file (such as the information source file in this article).
[0038] In these embodiments, a Mask RCNN model is constructed using the Mask RCNN technique. A large number (for example, 100,000) of information source files with the positions and contours of the molecular structural formula images pre-marked are used as training data to train the Mask RCNN model, so that the Mask RCNN model can identify and segment the molecular structural formula images in the information source file with a very high accuracy (for example, close to 1). The Mask RCNN model can quickly batch-detect the molecular structural formula images in the information source file and store these molecular structural formula images in the molecular structural formula image set.
[0039] Note that the Mask RCNN technique is only an example of target detection techniques, and other techniques such as YOLO, FastRCNN, SSD, etc. can also be used for the target detection step of the method for extracting molecular structural formulas according to an embodiment of the present invention.
[0040] In the method for extracting molecular structural formulas according to an embodiment of the present invention, the molecular structural formula images in the information source file are identified and segmented using target detection techniques, reducing the scale of the data to be processed in the subsequent steps, thereby improving the processing efficiency and accuracy of the subsequent steps.
[0041] In some embodiments, in addition to the molecular structural formula images, the information source file may also include other types of information, such as text, photos, graphics, symbols, etc. For example, Figure 2The information source file shown also includes a large amount of text and graphics different from molecular structural formulas. In the method for extracting molecular structural formulas according to an embodiment of the present invention, by using object detection technology to identify and segment the molecular structural formula images in such information source files, redundant information other than the molecular structural formulas can be removed, further improving the processing efficiency and accuracy.
[0042] (Graphic recognition)
[0043] The method for extracting molecular structural formulas according to an embodiment of the present invention performs a graphic recognition step after the object detection step.
[0044] Figure 3 The molecular structural formula image of compound 300 according to an embodiment of the present invention is shown. This molecular structural formula image is, for example, the molecular structural formula image obtained from the above object detection step.
[0045] As Figure 3 shown, compound 300 includes a character part 301 and a skeleton part 302. The skeleton part 302 is the chemical bond part in the molecular structural formula of compound 300 and includes one or more line segments. In the case where the skeleton part 302 includes multiple line segments, one end point of at least one line segment among the multiple line segments is connected to another line segment, and the two connected line segments form a certain angle, and the connection point between the line segments represents a carbon atom. The character part 301 usually includes one or more groups of characters. Each group of characters represents an atom or atomic group and includes at least letters, and may also include numbers, plus or minus signs, etc.
[0046] In this embodiment, the character part 301 of compound 300 includes 3 groups of characters, namely Ph, O, and H, which represent a phenyl group, an oxygen atom, and a hydrogen atom respectively. The skeleton part 302 includes four single bonds and one double bond, namely the single bond between the phenyl group (Ph) and the carbon atom, two carbon single bonds, the single bond between the carbon atom and the hydrogen atom (H), and the double bond between the carbon atom and the oxygen atom (O), where the single bond is represented by one line segment and the double bond is represented by two parallel line segments.
[0047] In the method for extracting molecular structural formulas of the present invention, the character part and the skeleton part of the molecular structural formula are obtained by performing graphic recognition on the molecular structural formula image. In some embodiments, OpenCV is used to perform graphic recognition on the molecular structural formula image. OpenCV is a computer vision and machine learning software library that can implement many general algorithms in image processing and computer vision.
[0048] In these embodiments, OpenCV is used to perform graphic recognition on the molecular structural formula image segmented in the target detection step. For example, first, the connected regions in the molecular structural formula image are recognized. Here, the connected region refers to an image region composed of pixel points with the same pixel value and adjacent positions in the image. Since the pixel points with the same pixel value (such as the black pixel value 0) between the character part and the skeleton part in the molecular structural formula image are not adjacent, while the pixel points with the same pixel value inside the skeleton part are adjacent to each other, the skeleton part can be recognized as a connected region.
[0049] Next, character recognition is performed on the image part outside the connected regions in the molecular structural formula image, and the position of each group of characters and the position of the endpoints on the skeleton part corresponding to this group of characters are recorded. In some embodiments, the position of each group of characters can be the position of the center point of a square box with a predetermined size that encloses this group of characters. The endpoints on the skeleton part corresponding to each group of characters can be the endpoints in the skeleton part closest to this group of characters. The position of the center point and the position of the endpoints can be represented by coordinates.
[0050] Taking compound 300 as an example. In the image part outside the skeleton part 302 of the molecular structural formula image of compound 300, 3 groups of characters, namely Ph, H, and O, can be recognized. Assume Figure 3 that the size of each square grid is 5*5 (coordinate unit), and assume that a square box of 8*8 (coordinate unit) can enclose each group of the 3 groups of characters respectively. The center point coordinates of the square box enclosing Ph are (5,5), the center point coordinates of the square box enclosing H are (53,5), and the center point coordinates of the square box enclosing O are (41,30). The endpoints in the skeleton part 302 closest to Ph, H, and O are the lower left endpoint (coordinates (9,7)) of the skeleton part 302, the lower right endpoint (coordinates (48,8)), and the upper endpoints of two parallel line segments (coordinates (40,25) and (42,25)) respectively. The position of each group of characters and the position of the endpoints on the skeleton part corresponding to this group of characters can be recorded in Table 1 below.
[0051]
Table 1
[0052] Character Position Corresponding point on the skeleton Ph (5,5) (9,7) H (53,5) (48,8) O (41,30) (40,25),(42,25)
[0053] Note that OpenCV is only an example of a graphic recognition tool and is not restrictive. Other graphic recognition tools can also be used to perform the graphic recognition step of the method for extracting molecular structural formulas according to the embodiments of the present invention.
[0054] In the method for extracting a molecular structural formula according to an embodiment of the present invention, the character part and the skeleton part of the molecular structural formula are recognized by using graphic recognition technology, which can simplify the calculation model in the subsequent connection relationship inference step and improve the calculation efficiency.
[0055] (Feature point marking)
[0056] Figure 4 Shows Figure 3 Two types of feature points on the skeleton part 402 of the molecular structural formula image of compound 300 in
[0057] In Figure 4 , two types of feature points are marked on the skeleton part 402, namely the first type of feature point, represented by the letter G, and the second type of feature point, represented by the letter T. The first type of feature point G is an end point that is not adjacent to the characters in the character part and represents a carbon atom; the second type of feature point T is an end point that is adjacent to the characters in the character part and represents an atom or atomic group other than a carbon atom.
[0058] In Figure 4 , the first type of feature point G includes the common end point of two line segments, such as G 5 , and the end point of another line segment that is connected to a part other than the end point of one line segment, such as G 1 and G 2 . Due to the certain width of the line segments in the molecular structural formula image or other reasons, it is possible to mark the end points corresponding to the same carbon atom as two or more points, such as G 1 , G 2 , G 3 and G 7 actually correspond to the same carbon atom, G 4 and G 6 actually correspond to the same carbon atom.
[0059] In this embodiment, the second type of feature point T includes the end point of a line segment that is not connected to other line segments, Figure 4 The T in 1 ~T 4 are all such second type of feature points. Although the character part of the molecular structural formula is not shown in Figure 4 , referring to Figure 3 it can be found that T 1 is the end point adjacent to the character H, T 2 and T 3 are the end points adjacent to the character O, and T 4 is the end point adjacent to the character Ph.
[0060] In some embodiments, the category, number, and position of each feature point can be recorded, and the position can be represented by coordinates. For example, in Figure 4Among them, the endpoints on the skeleton part corresponding to the character Ph in Figure 3 are marked as T 4 , and its coordinates are (9, 7); the endpoints on the skeleton part corresponding to the character H are marked as T 1 , and its coordinates are (48, 8); the endpoints on the skeleton part corresponding to the character O are marked as T 2 and T 3 , and their coordinates are (42, 25) and (40, 25) respectively. Similarly, in Figure 4 , the coordinates of the first type of feature point G 1 are (42, 12), the coordinates of G 2 are (40, 12), the coordinates of G 3 are (40.5, 13), the coordinates of G 4 are (30.5, 4), the coordinates of G 5 are (18, 13), the coordinates of G 6 are (29.5, 4), the coordinates of G 7 are (41.5, 13). Note that Figure 4 the coordinate system used in Figure 3 is the same as the coordinate system in Figure 4 . Thus, all the feature points on the skeleton part 402 in
[0061]
Table 2
[0062]
[0063] In some embodiments, the characters in the character part and their positions can be associated with the second type of feature points to obtain the character association relationship between the characters and the second type of feature points. For example, for the compound 300, the character Ph and its position can be associated with the feature point T 4 , the character H and its position can be associated with the feature point T 1 , and the character O and its position can be associated with the feature points T 2 and T 3 . According to Table 1 and Table 2, the character association relationship between the characters of the molecular structural formula of the compound 300 and the second type of feature points can be obtained, as shown in Table 3 below.
[0064]
Table 3
[0065] Character Position Number of the corresponding second type of feature points Ph (5,5) 4 H (53,5) 1 O (41,30) 2,3
[0066] In the method for extracting the molecular structural formula according to the embodiments of the present invention, by marking the feature points on the skeleton part of the compound image, the calculation model in the subsequent connection relationship inference step can be further simplified, the calculation amount can be reduced, and the inference accuracy can be improved.
[0067] (Inference of connection relationships)
[0068] The method for extracting a molecular structural formula according to an embodiment of the present invention performs an inference step of connection relationships after the feature point marking step.
[0069] In some embodiments, a Markov logic network (MLR) is used to perform the inference step of connection relationships. A Markov logic network is a statistical relational learning model that combines a Markov network with first-order logic and can be used to efficiently and accurately infer the connection relationships between feature points in a molecular structural formula image.
[0070] The connection relationships include three categories of connection relationships, namely the connection relationships between the first type of feature points, the connection relationships between the second type of feature points, and the connection relationships between the first type of feature points and the second type of feature points.
[0071] The connection relationships between the first type of feature points can be the connection relationships between carbon atoms. For example, it includes that multiple feature points correspond to the same carbon atom, there is a single bond between the carbon atoms corresponding to multiple feature points, there is a double bond between the carbon atoms corresponding to multiple feature points, there is a triple bond between the carbon atoms corresponding to multiple feature points, etc.
[0072] The connection relationships between the second type of feature points can be the connection relationships between atoms or atomic groups other than carbon atoms. For example, it includes that there is a single bond between the atoms or atomic groups corresponding to multiple feature points, there is a double bond between the atoms or atomic groups corresponding to multiple feature points, the atoms or atomic groups corresponding to multiple feature points are chiral structures, etc.
[0073] The connection relationships between the first type of feature points and the second type of feature points can be the connection relationships between a carbon atom and other atoms or atomic groups. For example, it includes that there is a single bond between a carbon atom and other atoms or atomic groups, there is a double bond between a carbon atom and other atoms or atomic groups, etc.
[0074] In some embodiments, a Markov logic network is used to infer the three categories of connection relationships according to a predefined inference formula. The inference formula is defined according to first-order predicate logic and includes "evidence" and "query".
[0075] Evidence is a descriptive set based on the molecular structural formula image. For example, the evidence VeryCloseCarbons(G 1 ,G 2 ) indicates that two first-type feature points G 1 and G 2 are very close, and the evidence CarbonLine(G 1 ,G 2)Indicates two feature points G 1 and G 2 are connected by a line segment.
[0076] The query is defined according to the connection relationship between feature points. For example, the query SameCarbonPoint(G 1 ,G 2 ) indicates that two feature points of the first type G 1 and G 2 represent the same carbon atom. The query AreCarbonsBonded(G 1 ,G 2 ) indicates that there is a single bond between two feature points G 1 and G 2 .
[0077] The conclusion in the query can be inferred from one or more pieces of evidence or the operation results of evidence. For example, the inference formula composed of evidence and query can be the following Inference Formula 1, where the symbol "∧" represents the AND operation, and "!" represents the NOT operation. Formula 1 means that when two feature points of the first type G 1 and G 2 are connected by a line segment and these two feature points G 1 and G 2 are not very close, it is inferred that there is a single bond between these two feature points G 1 and G 2 .
[0078] CarbonLine(G 1 ,G 2 )∧!VeryCloseCarbons(G 1 ,G 2 )→AreCarbonsBonded(G 1 ,G 2 )
[0079] (Formula 1)
[0080] In actual operations, for the sake of computational simplicity, the Markov logic network uses the conjunctive normal form corresponding to the inference formula to infer the connection relationship between feature points. The inference formula and the corresponding conjunctive normal form are logically equivalent. For example, the conjunctive normal form logically equivalent to the above Inference Formula 1 can be shown as the following Formula 2. Among them, the symbol "∨" represents the OR operation.
[0081] !CarbonLine(G 1 ,G 2 )∨VeryCloseCarbons(G 1 ,G 2 )∨AreCarbonsBonded(G1 , G 2 )
[0082] (Formula 2)
[0083] Using the conjunctive normal form equivalent to the inference formula, the probability value of each inference formula can be calculated. The larger the probability value, the higher the probability of inferring the conclusion in the query from the evidence.
[0084] The connection relationship can be determined according to the probability values of the inference formulas. In some embodiments, a judgment threshold can be set. When the probability value of the calculated inference formula is above the judgment threshold, it is determined that the inference result of the inference formula is acceptable, and the connection relationship corresponding to the inference formula is retained. When the probability value of the calculated inference formula is below the judgment threshold, it is determined that the inference result of the inference formula is unacceptable, and the connection relationship corresponding to the inference formula is discarded.
[0085] In some embodiments, a weight value can be set for the inference formula. The higher the weight, the stronger the constraint of the inference formula. For example, an inference formula with a weight of 10 is a strong rule, indicating that it is expected that the Markov logic network will follow this rule during inference. An inference formula with a weight of 5 is a medium rule, indicating that generally the Markov logic network will follow this rule. An inference formula with a weight of 3 is a weak rule, which is a rule that the Markov logic network can violate, and can avoid over-inference. In some embodiments, the set weight value of the inference formula can be multiplied by the calculated probability value of the inference formula to obtain the weighted probability value of the inference formula, and the connection relationship can be determined according to the weighted probability value.
[0086] In some embodiments, the initial weight value of the inference formula can be set according to experience. For example, the weight of the following Inference Formula 3 can be set to 10, that is, it has a strong rule. This inference formula indicates that between the first type of feature points G 1 and G 2 when there is a single bond, it is inferred that between the first type of feature points G 2 and G 1 there is also a single bond.
[0087] AreCarbonsBonded(G 1 , G 2 ) → AreCarbonsBonded(G 2 , G 1 ) (Formula 3)
[0088] In some embodiments, the Markov logic network can adjust the weight of the inference formula according to the comparison between the inferred connection relationship between feature points and the actual connection relationship. According to such a feedback mechanism, the Markov logic network can automatically adjust the weights of the inference formulas and improve the accuracy of inferring the connection relationship.
[0089] In some embodiments, the inference formulas are defined respectively for three types of connection relationships. For the connection relationship between the first type of feature points, the following inference formula can be defined: If the distance between two first type of feature points is less than a predetermined threshold, it is inferred that these two first type of feature points represent the same carbon atom. For the connection relationship between the second type of feature points, the following inference formula can be defined: If the distance between two second type of feature points is less than a predetermined threshold, it is inferred that these two second type of feature points represent the same atom or atomic group, and there is a double bond between this atom or atomic group and another atom or atomic group. For the connection relationship between the first type of feature points and the second type of feature points, the following inference formula can be defined: If a first type of feature point and a second type of feature point are connected by a line segment, it is inferred that there is a single bond between the carbon atom corresponding to this first type of feature point and the atom or atomic group corresponding to this second type of feature point.
[0090] The following describes in detail the inference formula examples of the Markov logic network in combination with the examples in Figure 3 and Figure 4 .
[0091] VeryCloseCarbons(G 1 ,G 2 )→SameCarbonPoint(G 1 ,G 2 ) (Equation 4)
[0092] The above Equation 4 is an inference formula defined for the connection relationship between the first type of feature points G on the skeleton part of compound 300. The part before the arrow, VeryCloseCarbons(G 1 ,G 2 ), is the evidence, indicating that G 1 and G 2 are very close. In some embodiments, the distance between G 1 and G 2 can be calculated based on their coordinates. If this distance is less than a predetermined threshold, it is determined that G 1 and G 2 are very close. The part after the arrow, SameCarbonPoint(G 1 ,G 2 ), is the query, indicating that G 1 and G 2 are the same carbon atom. Thus, the inference formula of Equation 4 is that if G 1 and G 2 are very close, then G 1 and G 2 are the same carbon atom. This carbon atom can be denoted as C 1 , C1 The position of, for example, can be represented by the coordinates of the midpoint of the line connecting G 1 and G 2 .
[0093] According to Equation 4 above, it is possible to Figure 4 infer that the G 1 , G 2 , G 3 and G 7 are the same carbon atom, and Figure 4 infer that the G 4 and G 6 are the same carbon atom.
[0094] An example of the inference formula can also be as shown in Equation 5 below.
[0095] VeryCloseFormula(T 2 , T 3 ) → SameFormulaEndPoint(T 2 , T 3 ) (Equation 5)
[0096] The above Equation 5 is an inference formula defined for the connection relationship between the second type of feature points T on the skeleton part of compound 300. Similar to Equation 4, VeryCloseFormula(T 2 , T 3 ) is evidence indicating that T 2 and T 3 are very close. That T 2 and T 3 are very close can also be judged by the distance between them being less than a predetermined threshold. SameCarbonPoint(T 2 , T 3 ) means that T 2 and T 3 represent the same atom or atomic group. Therefore, the inference formula of Equation 5 is that if T 2 and T 3 are very close, then T 2 and T 3 represent the same atom or atomic group. According to the character association relationship between the characters in Table 3 and the second type of feature points, it can be known that T 2 and T 3 represent the same oxygen atom, which can be denoted as O 1 .
[0097] Since all the second - type feature points T are the endpoints adjacent to the characters, that is, the endpoints not connected to other line segments, even if two second - type feature points are very close, there will be no connection between them. Instead, they can only indicate that these two feature points are the endpoints of two parallel line segments respectively, that is, there is a double bond between the atoms or atomic groups represented by these two feature points and other atoms or atomic groups. In this example, according to the inference formula of Equation 5 and the character association relationship in Table 3, it can be inferred that T 2 and T 3 represent an oxygen atom O 1 and there is a double bond between it and other atoms or atomic groups.
[0098] An example of the inference formula can also be shown as Equation 6 below.
[0099] CarbonFormulaLine(G 1 ,T 1 )→AreCarbonAndFormulaBonded(G 1 ,T 1 ) (Equation 6)
[0100] The above Equation 6 is an inference formula defined for the connection relationship between the first - type feature point G and the second - type feature point T on the skeleton part of compound 300. The part before the arrow, CarbonFormulaLine(G 1 ,T 1 ), is evidence indicating that G 1 and T 1 are connected by a line segment. The part after the arrow, AreCarbonAndFormulaBonded(G 1 ,T 1 ), is a query indicating that there is a single bond between the carbon atom corresponding to G 1 and the atom or atomic group corresponding to T 1 . Therefore, the inference formula of Equation 6 is that if G 1 and T 1 are connected by a line segment, then there is a single bond between the carbon atom corresponding to G 1 and the atom or atomic group corresponding to T 1 . According to the character association relationship between the characters and the second - type feature points in Table 3, it can be known that T 1 represents a hydrogen atom H, which can be denoted as H 1 . Therefore, according to the inference formula of Equation 6 and the character association relationship in Table 3, it can be inferred that there is a single bond between the carbon atom C 1 represented by G 1 and the hydrogen atom H 1 represented by T 1 .
[0101] In addition, operations can be performed between pieces of evidence. For example, in the following formula 7 (similar to the above formula 1), the part before the arrow represents G 5 and G 6 are connected by a line segment and G 5 and G 6 are not very close to each other. The part after the arrow represents that there is a single carbon bond between the corresponding carbon atoms of G 5 and G 6 . The carbon atom corresponding to G 5 can be denoted as C 3 , and the carbon atom corresponding to G 6 can be denoted as C 2 . Then, according to the inference formula of formula 7, it can be inferred that there is a single carbon bond between the carbon atom C 3 corresponding to G 5 6 and the carbon atom C 2 corresponding to G 6 .
[0102] CarbonLine(G 5 ,G 6 )∧!VeryCloseCarbons(G 5 ,G 6 )→AreCarbonsBonded(G 5 ,G 6 )
[0103] (Formula 7)
[0104] Similarly, the Markov logic network infers the connection relationships between all the feature points on the skeleton part of the molecular structural formula image of compound 300 using predefined inference formulas. Based on these connection relationships, the feature point positions in Table 2, and the character association relationships in Table 3, the positions of the atoms or atomic groups corresponding to all the feature points on the skeleton part can be obtained. In some embodiments, if two or more feature points correspond to the same atom or atomic group, the average value of the coordinates of these feature points can be calculated as the position of the atom or atomic group. In this way, the positions of each atom or atomic group in the molecular structural formula of compound 300 are obtained, as shown in Table 4 below.
[0105]
Table 4
[0106] Atom or atomic group Position <![CDATA[C 1 > (41,12.5) <![CDATA[C 2 > (30,4) <![CDATA[C 3 > (18,13) <![CDATA[H 1 > (48,8) <![CDATA[O 1 > (41,25) <![CDATA[Ph 1 > (9,7)
[0107] Based on the connection relationships between all the feature points on the skeleton part of the molecular structural formula image of compound 300 inferred by the Markov logic network, the feature point positions in Table 2, and the character association relationships in Table 3, the chemical bonds between the atoms or atomic groups corresponding to all the feature points on the skeleton part can also be obtained, as shown in Table 5 below.
[0108]
Table 5
[0109] Atom or atomic group 1 Atom or atomic group 2 Chemical bond <![CDATA[C 1 > <![CDATA[C 2 > Single bond <![CDATA[C 1 > <![CDATA[H 1 > Single bond <![CDATA[C 1 > <![CDATA[O 1 > Double bond <![CDATA[C 2 > <![CDATA[C 3 > Single bond <![CDATA[C 3 > <![CDATA[Ph 1 > Single bond
[0110] So far, the method for extracting molecular structural formulas according to the present invention has been described in detail in combination with embodiments. This method can quickly, efficiently, and accurately batch-extract molecular structural formulas from information source files.
[0111] (mol file generation)
[0112] The method for extracting molecular structural formulas may optionally further include a mol file generation step to store the extracted molecular structural formulas in a computer-readable form. The mol file is a commonly used file format for storing molecular structural formulas, and software such as ChemDraw can be used to convert the mol file into a visual molecular structural formula.
[0113] The mol file generation step includes generating a mol file based on the character association relationship between the characters in the molecular structural formula image and the second type of feature points, as well as the connection relationship between the feature points. For example, regarding Figure 3 Compound 300 in, a mol file of the molecular structural formula of this compound can be generated based on the position of each atom or atomic group in Table 4 and the chemical bonds between these atoms or atomic groups in Table 5. Note that Ph in Table 4 and Table 5 1 represents a phenyl group, which actually contains 6 carbon atoms in the form of a benzene ring. Therefore, Compound 300 actually has 11 atoms.
[0114] Figure 5 shows Figure 3 mol file 500 of Compound 300 in.
[0115] As Figure 5 shown, in mol file 500 of Compound 300, in line 501 marked by the dashed box, from left to right, the first number 11 indicates that there are 11 atoms or atomic groups in Compound 300; the second number 11 indicates that there are 11 chemical bonds in Compound 300; 0999 is the default value; V2000 indicates the version of this mol file.
[0116] In mol file 500, the matrix 502 marked by the dashed box is the atomic matrix, where the first three columns represent the coordinates of the atoms or atomic groups, and the fourth column is the chemical symbol of the atoms or atomic groups. For example, the first row of the atomic matrix 502 records a carbon atom C, and its coordinates are (-0.0000, 0.4125, 0.0000). Note that the coordinate values in this mol file are only examples and may not be the same as the specific numerical values of the coordinates in other embodiments.
[0117] In the mol file 500, the matrix 503 marked by the dashed box is the bond matrix. The first column and the second column represent the serial numbers of atoms or atomic groups (in the order from top to bottom in the atomic matrix), the third column represents the chemical bonds, and the fourth column represents the presence or absence of chirality. For example, the first row in the bond matrix 503 indicates that there is a double bond between the 10th atom (i.e., the C atom) and the 11th atom (i.e., the C atom) in the atomic matrix 502, without chirality; the second row indicates that there is a single bond between the 9th atom (i.e., the C atom) and the 10th atom (i.e., the C atom), without chirality. From the bond matrix 503, it can be seen that the 6th to 11th atoms in the atomic matrix 502 are in a benzene ring structure, that is, the phenyl group Ph in the compound 300.
[0118] In addition to Figure 5 the part shown in, the mol file may also include other contents such as an attribute matrix.
[0119] Note that using the mol file to store the molecular structural formula is only an example, and the molecular structural formula can also be stored in other forms.
[0120] According to the present disclosure, by storing the extracted molecular structural formula in a computer-readable form, it is convenient for researchers to read the molecular structural formula as needed, which is helpful for chemical research, new drug discovery, chemical manufacturing, etc.
[0121] (Computing device)
[0122] Figure 6 Shows an exemplary configuration of a computing device 600 that can implement the embodiments according to the present invention. The computing device 600 can be used, for example, to implement the method for extracting the molecular structural formula according to the embodiments of the present invention.
[0123] The computing device 600 is an example of a hardware device that can apply the above aspects of the present invention. The computing device 600 can be any machine configured to perform processing and / or calculations. The computing device 600 can be, but is not limited to, a workstation, a server, a desktop computer, a laptop computer, a tablet computer, a personal data assistant (PDA), a smart phone, an in-vehicle computer, or a combination thereof.
[0124] As Figure 6As shown, computing device 600 may include one or more components that may be connected to or communicate with bus 601 via one or more interfaces. Bus 601 may include, but is not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus, etc. Computing device 600 may include, for example, one or more processors 602, one or more input devices 603, and one or more output devices 604. One or more processors 602 may be any kind of processor, and may include, but is not limited to, one or more general-purpose processors or dedicated processors (such as dedicated processing chips). Input device 603 may be any type of input device capable of inputting information to the computing device, and may include, but is not limited to, a mouse, keyboard, touch screen, microphone, and / or remote controller. Output device 604 may be any type of device capable of presenting information, and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer.
[0125] Computing device 600 may also include or be connected to non-transitory storage device 607, which may be any non-transitory storage device capable of implementing data storage, and may include, but is not limited to, disk drives, optical storage devices, solid-state memories, floppy disks, flexible disks, hard disks, magnetic tapes, or any other magnetic medium, compact disks, or any other optical medium, cache memories, and / or any other storage chips or modules, and / or any other medium from which a computer can read data, instructions, and / or code. Computing device 600 may also include random access memory (RAM) 605 and read-only memory (ROM) 606. ROM 606 may store programs, utilities, or processes to be executed in a non-volatile manner. RAM 605 may provide volatile data storage and store instructions related to the operation of computing device 600. Computing device 600 may also include network / bus interface 608 coupled to data link 609. Network / bus interface 608 may be any kind of device or system capable of enabling communication with external devices and / or networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication devices, and / or chip sets (such as Bluetooth TM devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication facilities, etc.).
[0126] In addition, another embodiment of the present invention further provides a computer-readable storage medium, including computer-executable instructions, which, when executed by one or more processors, cause the one or more processors to execute the method for extracting molecular structural formulas as described in the above embodiments.
[0127] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0128] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate a device for realizing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.
[0129] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.
[0130] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.
[0131] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.
[0132] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for extracting a molecular structural formula, comprising: a target detection step of detecting a molecular structural formula in an information source file and segmenting a molecular structural formula image including the molecular structural formula from the information source file; a graphic recognition step of performing graphic recognition on the molecular structural formula image to obtain a character part and a skeleton part of the molecular structural formula; a feature point marking step of marking feature points in the skeleton part as first-type feature points or second-type feature points, associating characters and their positions in the character part with the second-type feature points to obtain a character association relationship between the characters and the second-type feature points, the character association relationship including the correspondence between the characters, the position coordinates of the characters, and the second-type feature points, the skeleton part including one or more line segments, and the feature points including endpoints of the one or more line segments; a connection relationship inference step of using a Markov logic network to infer connection relationships between the first-type feature points, between the second-type feature points, and between the first-type feature points and the second-type feature points according to a predefined inference formula, wherein the first-type feature points represent endpoints not adjacent to characters in the character part, and the second-type feature points represent endpoints adjacent to characters in the character part; According to the connection relationships between all feature points in the skeleton part, the position coordinates of each feature point among all feature points in the skeleton part, and the character association relationship, the position coordinates of atoms or atomic groups corresponding to each feature point among all feature points in the skeleton part and the chemical bonds between atoms or atomic groups corresponding to each feature point among all feature points in the skeleton part are obtained.
2. The method according to claim 1, wherein, the information source file further includes other types of information different from the molecular structural formula image.
3. The method according to claim 1, wherein, the graphic recognition step further includes: recording the positions of characters in the character part.
4. The method according to claim 3, further comprising: a mol file generation step of generating a mol file according to the connection relationship and the character association relationship.
5. The method according to claim 1, wherein, the connection relationship includes at least one of the following relationships: multiple feature points correspond to the same atom or atomic group; there is a single bond between atoms or atomic groups corresponding to multiple feature points; there is a double bond between atoms or atomic groups corresponding to multiple feature points; there is a triple bond between atoms or atomic groups corresponding to multiple feature points; there is a bridge bond between atoms or atomic groups corresponding to multiple feature points; the atoms or atomic groups corresponding to multiple feature points are chiral structures.
6. The method according to claim 1, wherein, the inference formula is defined according to first-order predicate logic and includes evidence and a query, the evidence is a descriptive set based on the molecular structural formula image, and the query is defined according to the connection relationship.
7. The method according to claim 1, wherein, The Markov logic network uses a conjunctive normal form corresponding to the inference formula to infer the connection relationship.
8. The method according to claim 1, wherein, the inference formula is set with weights, and the Markov logic network adjusts the weights according to the comparison between the inferred connection relationship and the actual connection relationship.
9. The method according to claim 1, wherein, the first type of feature points represent carbon atoms, the inference formula includes that when the distance between two first type of feature points is less than a predetermined threshold, it is inferred that the two first type of feature points represent the same carbon atom.
10. The method according to claim 1, wherein, the second type of feature points represent atoms or atomic groups other than carbon atoms, the inference formula includes that when the distance between two second type of feature points is less than a predetermined threshold, it is inferred that the two second type of feature points represent the same atom or atomic group, and there is a double bond between this atom or atomic group and another atom or atomic group.
11. The method according to claim 1, wherein, the first type of feature points represent carbon atoms and the second type of feature points represent atoms or atomic groups other than carbon atoms, the inference formula includes that when a first type of feature point and a second type of feature point are connected by a line segment, it is inferred that there is a single bond between the carbon atom corresponding to the first type of feature point and the atom or atomic group corresponding to the second type of feature point.
12. A non-transitory computer-readable medium storing program instructions that, when executed by a processor, perform the method according to any one of claims 1 to 11.
13. A computing device, comprising: a memory storing program instructions; and one or more processors configured to execute the program instructions, so that the computing device performs the method according to any one of claims 1 to 11.