Molecular structure extraction method, device and equipment

CN116453112BActive Publication Date: 2026-09-15SHENZHEN JINGTAI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211101282.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-09
Publication Date
2026-09-15
Estimated Expiration
2042-09-09

AI Technical Summary

Technical Problem

[0003]然而,大量的分子结构信息是隐藏在文献当中,需要专业人员阅读文献,手动绘制分子结构,导致耗费大量的人力和时间进行汇集,效率低下

Benefits of technology

[0035] The technical solution of this application allows for the identification of original images of molecular structures using multiple different preset optical structure recognition tools. After obtaining a corresponding number of candidate molecular structures, these structures are compared to determine their consistency. In cases of inconsistency, the candidate structures are compared with the molecular structures in the original images to select the final identification result. This design eliminates the need for manual drawing by reviewing literature, improving information collection efficiency and reducing labor costs. Furthermore, by using multiple optical structure recognition tools and comparing and selecting candidate molecular structures with more similar structures as the final identification result even when output structures are inconsistent, the accuracy and reliability of molecular structure identification are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116453112B_ABST
    Figure CN116453112B_ABST
Patent Text Reader

Abstract

The application relates to a molecular structure extraction method and device, a molecular structure dataset equipment, an electronic equipment and a computer readable storage medium. The method comprises the following steps: acquiring an original image of a to-be-recognized molecular structure; recognizing the to-be-recognized molecular structure in the original image according to a plurality of preset optical structure recognition tools, and obtaining corresponding candidate molecular structures; comparing the candidate molecular structures with each other, and when at least part of the candidate molecular structures are inconsistent, screening out a candidate molecular structure close to the to-be-recognized molecular structure in the original image as a recognition result. The scheme provided by the application can quickly and accurately extract the molecular structure in the literature, saves human resources, and improves the information collection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of chemical image recognition technology, and in particular to a method, apparatus and equipment for extracting molecular structures. Background Technology

[0002] Scientific research results are typically published in the form of articles or patents. In synthetic chemistry, natural product research, drug discovery, and many other fields, reading literature is the most common way to obtain information for pharmaceutical research and development. It is estimated that nearly 10,000 academic journals publish articles in the field of chemistry, and more than 20,000 new chemical structures are disclosed annually. Drug developers can analyze the molecular structures and activity data published in the literature to advance the next stage of drug development.

[0003] However, a large amount of molecular structural information is hidden in the literature, requiring professionals to read the literature and manually draw molecular structures, which results in a lot of manpower and time being spent on collecting information, leading to low efficiency. Summary of the Invention

[0004] To address or partially address the problems existing in related technologies, this application provides a molecular structure extraction method and apparatus, a molecular structure dataset device, an electronic device, and a computer-readable storage medium, which can quickly and accurately extract molecular structures from literature, save human resources, and improve information collection efficiency.

[0005] The first aspect of this application provides a method for extracting molecular structures, comprising:

[0006] Acquire the original image of the molecular structure to be identified;

[0007] The molecular structures to be identified in the original image are identified by multiple preset optical structure recognition tools to obtain the corresponding candidate molecular structures.

[0008] The candidate molecular structures are compared with each other. When at least some of the candidate molecular structures are inconsistent, the candidate molecular structure that is close to the molecular structure to be identified in the original image is selected as the identification result.

[0009] In the molecular structure extraction method of this application, the comparison of candidate molecular structures includes:

[0010] The candidate molecular structures are identified by identity discrimination, and discrimination labels are set for the molecular structures to be identified based on whether the discrimination results are consistent.

[0011] In the molecular structure extraction method of this application, the step of selecting candidate molecular structures that are close to the molecular structure to be identified in the original image as the identification result when at least some of the candidate molecular structures are inconsistent includes:

[0012] When the discrimination marker is a non-consistent marker indicating structural inconsistency, the similarity between each candidate molecular structure and the corresponding molecular structure in the original image is evaluated. Based on the similarity between each candidate molecular structure and the molecular structure in the original image, the candidate molecular structure with the highest similarity is taken as the identification result of the molecular structure to be identified.

[0013] The molecular structure extraction method of this application further includes:

[0014] When all candidate molecular structures are identical, the candidate molecular structure is determined to be the identification result of the molecular structure to be identified.

[0015] The molecular structure extraction method of this application further includes:

[0016] According to the preset structural format, the recognition results of each molecular structure to be identified are stored to obtain a molecular structure dataset.

[0017] In the molecular structure extraction method of this application, obtaining the original image of the molecular structure to be identified includes:

[0018] The original file in the preset format is paginated to obtain the corresponding paginated file; each molecular structure to be identified in the paginated file is segmented into an independent image, and the original image corresponding to the molecular structure to be identified is generated respectively.

[0019] In the molecular structure extraction method of this application, the step of segmenting each molecular structure to be identified in the paginated file into independent images and generating the original image of the molecular structure to be identified includes:

[0020] Mask the molecular structure to be identified in the paginated file to generate the corresponding mask image;

[0021] The masked region in the masked image is segmented to generate the original image of the molecular structure to be identified.

[0022] In the molecular structure extraction method of this application, the step of identifying the molecular structure to be identified in the original image using multiple preset optical structure recognition tools to obtain the corresponding candidate molecular structure includes:

[0023] The original images are vectorized to obtain corresponding vector graphics. Molecular structure information in the vector graphics is identified using a preset optical structure recognition tool to obtain corresponding candidate molecular structures. The molecular structure information includes chemical bond information, atom type information, charge state, and atom connection information.

[0024] A second aspect of this application provides a molecular structure dataset device, which stores the identification results of molecular structures to be identified obtained by the molecular structure extraction method described above; wherein:

[0025] Each recognition result is mapped and stored to the original image and the identity discrimination result, respectively.

[0026] A third aspect of this application provides a molecular structure extraction device, comprising:

[0027] The raw image acquisition module is used to acquire raw images of the molecular structure to be identified.

[0028] The chemical structure recognition module is used to identify the molecular structure to be identified in the original image according to multiple preset optical structure recognition tools, and obtain the corresponding candidate molecular structure.

[0029] The structure screening module is used to compare the candidate molecular structures with each other. When at least some of the candidate molecular structures are inconsistent, the candidate molecular structure that is close to the molecular structure to be identified in the original image is selected as the identification result.

[0030] A fourth aspect of this application provides an electronic device, comprising:

[0031] Processor; and

[0032] A memory that stores executable code, which, when executed by the processor, causes the processor to perform the method described above.

[0033] A fifth aspect of this application provides a computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method described above.

[0034] The technical solution provided in this application may include the following beneficial effects:

[0035] The technical solution of this application allows for the identification of original images of molecular structures using multiple different preset optical structure recognition tools. After obtaining a corresponding number of candidate molecular structures, these structures are compared to determine their consistency. In cases of inconsistency, the candidate structures are compared with the molecular structures in the original images to select the final identification result. This design eliminates the need for manual drawing by reviewing literature, improving information collection efficiency and reducing labor costs. Furthermore, by using multiple optical structure recognition tools and comparing and selecting candidate molecular structures with more similar structures as the final identification result even when output structures are inconsistent, the accuracy and reliability of molecular structure identification are improved.

[0036] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0037] The above and other objects, features and advantages of this application will become more apparent from the following description of exemplary embodiments of this application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of this application.

[0038] Figure 1 This is a schematic flowchart of the molecular structure extraction method shown in this application;

[0039] Figure 2 This is a schematic flowchart illustrating a specific example of a molecular structure extraction method according to this application;

[0040] Figure 3 yes Figure 2 A schematic diagram of another molecular structure extraction method is shown;

[0041] Figure 4 yes Figure 2 A partial flowchart of the molecular structure extraction method is shown;

[0042] Figure 5 This is a schematic diagram of the molecular structure extraction device shown in this application;

[0043] Figure 6 This is a schematic diagram of the molecular structure extraction device shown in a specific example of this application;

[0044] Figure 7 This is a schematic diagram of the structure of the electronic device shown in this application. Detailed Implementation

[0045] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.

[0046] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0047] It should be understood that although the terms "first," "third," "third," etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as third information, and similarly, third information may also be referred to as first information. Thus, a feature defined as "first" or "third" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0048] In related technologies, the molecular structure information disclosed in the literature is not stored in a computer-readable form. Professionals need to read the literature and manually draw the molecular structure, which results in a lot of manpower and time being spent on the collection, and is inefficient.

[0049] To address the aforementioned issues, this application provides a molecular structure extraction method that can quickly and accurately extract molecular structures from literature, saving human resources and improving information collection efficiency.

[0050] The technical solution of this application is described in detail below with reference to the accompanying drawings.

[0051] Figure 1 This is a schematic flowchart of the molecular structure extraction method shown in this application.

[0052] See Figure 1 This application discloses a method for extracting molecular structures, which includes:

[0053] S110: Obtain the original image of the molecular structure to be identified.

[0054] The original images can come from publicly available patent documents, scientific literature, and other original documents. It is understood that the original documents may contain one or more molecular structures to be identified, and the original image corresponding to each molecular structure to be identified is obtained separately.

[0055] S120 identifies the molecular structure to be identified in the original image using multiple preset optical structure recognition tools, and obtains the corresponding candidate molecular structure.

[0056] The preset optical structure recognition tool can be a known OCSR (Optical Chemical Structure Recognition) tool, such as OSRA (Optical Structure Recognition Application), Imago (an image digital forensics tool written in Python), MolVec, or tools like Image2SMILES, ChemGrapher, and Imag2Mol. Alternatively, it can be a self-developed OCSR tool; there are no restrictions. Optionally, a more accurate OCSR tool can be selected as the preset optical structure recognition tool. The number of preset optical structure recognition tools can be two, three, or more. By using multiple different preset optical structure recognition tools to independently identify the molecular structure to be identified, different dimensions of recognition methods can be used to cross-verify the recognition results, avoiding biases caused by using a single recognition tool.

[0057] It is understandable that each optical structure recognition tool may generate a series of possible structures for the same molecular structure to be identified. To improve the processing efficiency of subsequent steps, each optical structure recognition tool can output one of the possible structures as a candidate molecular structure. In other words, based on the preset number of optical structure recognition tools, a corresponding number of candidate molecular structures are obtained.

[0058] Furthermore, the original structural representation of the molecular structure to be identified may be a structural formula, a simplified structural formula, etc., and the candidate molecular structure adopts the corresponding representation.

[0059] S130, compare the candidate molecular structures with each other, and when at least some of the candidate molecular structures are inconsistent, select the candidate molecular structure that is close to the molecular structure to be identified in the original image as the identification result.

[0060] When there are two or more candidate molecular structures, they can be compared with each other using relevant techniques to determine whether all candidate molecular structures are consistent. When at least some or all of the candidate molecular structures are inconsistent, one of them needs to be selected as the final identification result.

[0061] Based on this, according to relevant technologies, each candidate molecular structure can be compared with the molecular structure to be identified in the original image, and the candidate molecular structure that is closer to the molecular structure shown in the original image can be selected as the final identification result.

[0062] It is understandable that if all candidate molecular structures are identical, then the candidate molecular structure is the identification result of the molecular structure to be identified.

[0063] As this example illustrates, the molecular structure extraction method of this application can identify the original image of the molecular structure to be identified using multiple different preset optical structure recognition tools. After obtaining a corresponding number of candidate molecular structures, they are compared to determine whether they are consistent. In the case of inconsistency, they are compared with the molecular structure in the original image to select the final identification result. This design eliminates the need for manual drawing by reviewing literature one by one, improving information collection efficiency and reducing labor costs. At the same time, by using multiple optical structure recognition tools and comparing and selecting the candidate molecular structure with the closest structure as the final identification result under the premise of inconsistent output structures, the accuracy of molecular structure identification is improved.

[0064] Figure 2 This is a schematic flowchart illustrating a specific example of a molecular structure extraction method according to this application; Figure 3 yes Figure 2 The diagram shows another flowchart of the molecular structure extraction method.

[0065] See Figure 2 and Figure 3 This application discloses a method for extracting molecular structures, which includes:

[0066] S210: Perform pagination on the original file in the preset format to obtain the corresponding paginated file.

[0067] In this step, the original documents can be, for example, patent documents, journal articles, etc. Various original documents can be represented in a preset format, such as PDF. Of course, other file formats are also possible. The original documents can be converted to the preset format for subsequent steps. The file format of the paginated file can be consistent with the preset format of the original documents to standardize data management. It is understood that the original document may have more than one page. Based on this, the original document can be paginated using relevant technologies, such as known pagination software (e.g., QPDF toolkit) or self-developed pagination software, so that each page of the original document exists independently, forming a paginated file with the corresponding format.

[0068] Optionally, to facilitate data statistics and management, each paginated file can be associated with the original file, for example, by linking the naming of the paginated file to the original file. Taking a patent document as an example, different patents each have unique identifiers such as their own patent publication number and application number. For instance, an original document might be a patent document with publication number WO2021262596A1, named WO2021262596A1.pdf. Correspondingly, each paginated version of this original document can be named sequentially according to its corresponding page number: WO2021262596A1-1.pdf, WO2021262596A1-2.pdf, WO2021262596A1-3.pdf, and so on. This allows each paginated file to be distinguished by its unique file name while also being associated with the original file. Of course, other naming methods can also be used to associate each pagination file with its corresponding original file. This is just an example and is not a limitation.

[0069] S220, each molecular structure to be identified in the paginated file is segmented into an independent image, and the original image corresponding to the molecular structure to be identified is generated respectively.

[0070] It is understood that each paginated file may contain no molecular structure or may contain one or more molecular structures to be identified, which can be detected and determined using software in related technologies. When a paginated file contains one or more molecular structures to be identified, in one embodiment, the molecular structures to be identified in the paginated file are masked to generate a corresponding mask image; the masked regions in the mask image are segmented to generate the original image of the molecular structure to be identified.

[0071] Optionally, the paginated file in the preset format can be converted into paginated images in image format, such as JPG, PNG, or GIF, using relevant technologies. These paginated images can then be used for molecular structure recognition and image segmentation. Furthermore, each paginated image can be associated with the original file for data management. For example, file names related to the original file can be used to associate the paginated images with their corresponding original files. For instance, based on the image format, the file names of each paginated image could be WO2021262596A1-1.png, WO2021262596A1-2.png, WO2021262596A1-3.png, and so on. The naming method for each paginated image is only illustrative and is not intended to be limiting.

[0072] Furthermore, molecular structures are detected and masked based on each paginated image. Within the same frame of paginated images, there may be one or more molecular structures to be identified. Correlation techniques can be used to detect and determine the number of molecular structures, and further, the pixels in the regions where the molecular structures are located are masked using correlation techniques. This results in a masked image with one or more masked regions. Finally, the masked regions at different locations in the masked image are segmented using correlation techniques, allowing each masked region to be extracted as an independent original image. Furthermore, the original images can be labeled to associate the original images of each molecular to be identified with the original file. For example, the original file to which the original image belongs, the specific page in the original file, and the specific location within that page can be labeled for data management and querying.

[0073] To facilitate understanding, we will use a current chemical structure image recognition software as an example. For instance, using the pre-loaded DECIMER (Deep Learning for Chemical Image Recognition) software, each page of a PDF file can be converted into a paged image. Then, based on relevant algorithms such as Mask R-CNNnetwork (a general object instance segmentation framework), one or more molecular structures in each frame of the paged image can be automatically detected, and the distribution area of ​​the molecular structure to be identified within the paged image can be located. For each detected distribution area of ​​the molecular structure to be identified, the main skeleton of the molecular structure can be automatically masked, and the mask range can be further refined and expanded to completely cover the molecular structure, thus obtaining a complete molecular structure mask. When the same frame of the image includes multiple chemical structures, the same masking method is used to mask each molecular structure to be identified separately, thus forming a masked image from the paged image. Further, within the same frame of the masked image, segmentation is performed according to the coverage area of ​​each masked region, thereby obtaining multiple independent original images. Each original image corresponds to one molecular structure to be identified, thus obtaining the original image set.

[0074] Furthermore, each original image of the molecular structure to be identified can be associated with a tag linked to the original file. This tagging links each frame of the original image of the molecular structure to be identified to the original file, allowing for tracing back to the original file based on the tagging of the original image. For example, the original images of the three molecules to be identified segmented from the first frame of the masked image can be named sequentially as WO2021262596A1-1_1_bnw.png, WO2021262596A1-1_2_bnw.png, WO2021262596A1-1_3_bnw.png, etc., so that each original image has its own unique name and is associated with the original file. Of course, the above tagging method is merely illustrative and not intended to be limiting.

[0075] S230 identifies the molecular structure to be identified in the original image using multiple preset optical structure recognition tools, and obtains the corresponding candidate molecular structure.

[0076] In one specific implementation, the original images are vectorized to obtain corresponding vector graphics. Molecular structure information in the vector graphics is then identified using a preset optical structure recognition tool to obtain corresponding candidate molecular structures. The molecular structure information includes chemical bond information, atom type information, charge state, and atom connection information. In this step, the candidate molecular structures are represented in the form of a graph structure.

[0077] Specifically, such as Figure 4 As shown, taking three preset optical structure recognition tools—OSRA, Imago, and MolVec—as examples, each frame of the original image of the molecular structure to be identified is vectorized, forming a vector image of the molecular structure in the original image, thus accurately detecting the molecular structure information in the image. Specifically, by identifying single bonds, double bonds, triple bonds, cyclic bonds, aromatic bonds, and bonds on bridged rings in the compound, all chemical bond information is obtained; by identifying element types, all atom type information is obtained; by identifying atom types, charge states, and chemical bond connections, atomic connection information is obtained; and the tool's built-in dictionary is used to interpret atom labels.

[0078] Furthermore, each optical structure recognition tool can output a series of possible structures based on the molecular structure information it obtains, all represented in graph form. To facilitate comparison and selection in subsequent steps, this step calculates the confidence level for each possible structure output by each optical structure recognition tool. This allows each tool to select its own unique candidate molecular structure based on the confidence level of its respective possible structures. For example, by performing linear regression analysis on the similarity between the real molecular structure and the possible structures generated by the optical structure recognition tools at different resolution levels (e.g., the Tanimoto similarity coefficient), each tool retains only the possible structure with the highest confidence level as a candidate molecular structure. In other words, a total of three candidate molecular structures can be obtained based on the three optical structure recognition tools described above.

[0079] S240, perform identity discrimination on each candidate molecular structure; based on whether the discrimination results are consistent, set the discrimination label corresponding to the molecular structure to be identified.

[0080] In this step, based on relevant technologies, the identity of each candidate molecular structure to be identified can be determined by whether the structures are consistent. That is, all candidate molecular structures are compared to see if they are consistent. If all candidate molecular structures are consistent, it means that the candidate molecular structures are the same molecular structure; otherwise, they are different molecular structures.

[0081] When all candidate molecular structures are identical, the discriminant marker for the molecule to be identified is set to a consistency marker, such as "S" (representing sureness). When any two candidate molecular structures are inconsistent, the discriminant marker for the molecule to be identified is set to a non-consistency marker, such as "U" (representing unsureness). This is just an example. Based on the discriminant markers, it is possible to visually reflect whether the output results of different optical structure identification tools are all consistent. Furthermore, by associating the molecular structure to be identified with the discriminant markers, it is convenient for reviewers to focus on verifying the identification results of the corresponding molecular structures based on the discriminant markers.

[0082] S250, when the discrimination label is a non-consistent label indicating structural inconsistency, evaluate the similarity between each candidate molecular structure and the corresponding molecular structure in the original image; based on the similarity between each candidate molecular structure and the molecular structure in the original image, take the candidate molecular structure with the highest similarity as the identification result of the molecular structure to be identified.

[0083] In other words, when the discrimination marker is a non-consistent marker indicating structural inconsistency, it means that at least some candidate molecular structures are inconsistent. When multiple candidate molecular structures with different structures are identified by different optical structure recognition tools for the same molecular structure to be identified, the similarity between each candidate molecular structure and the molecular structure in the segmented original image in step S220 can be calculated. For example, a Python script can be used to calculate the similarity between the input original image and the graph structures of the three candidate molecular structures output by the OCSR tool, thereby obtaining the similarity between each candidate molecular structure and the molecular structure in the original image. Furthermore, the values ​​of each similarity can be compared; the larger the value, the higher the similarity, indicating that the corresponding candidate molecular structure is closer to the molecular structure in the original image. Preferably, the candidate molecular structure with the highest similarity can be used as the final identification result. That is, one candidate molecular structure is selected from the three candidate molecular structures as the identification result.

[0084] S260, when the discrimination marker is a consistency marker indicating structural consistency, the candidate molecular structure is determined as the identification result of the molecular structure to be identified.

[0085] Either step S260 or step S250 is performed. When all candidate molecular structures are identical, the candidate molecular structure is determined as the identification result of the molecular structure to be identified. Clearly, when all candidate molecular structures are identical, step S240 allows the molecule to be identified to obtain a consistency marker indicating structural consistency. Based on the consistency marker, it can be known that the output results of each optical structure identification tool are consistent, thus the consistent candidate molecular structure can be taken as the identification result of the molecular structure to be identified.

[0086] S270: According to the preset structural format, the recognition results of each molecular structure to be identified are stored to obtain a molecular structure dataset.

[0087] It is understood that the recognition results of steps S250 and S260 can be presented in the same format as the original file, such as the structural formula of the molecular structure to be identified. To facilitate computer readability, the current format of the recognition results can be converted into a preset structural format for storage. For example, the recognition results of the graph structure can be converted into a common SMILES (Simplified molecular input line entry system) string or other formats; this is merely an example.

[0088] It is understandable that by identifying the molecular structures to be identified in each original file, converting the identification results into a predefined structural format, and then aggregating them, a molecular structure dataset can be constructed. Optionally, the identification results can also be stored simultaneously in the form shown in the original images.

[0089] Furthermore, in one embodiment, the recognition results of the preset structural format can be associated and stored with the original file and the discrimination marker, respectively, to obtain a more complete molecular structure dataset. As shown in Table 1 below, the first column, the original image name, is marked with the unique identifier of the original file (i.e., the patent publication number), the corresponding page number, and the sorting within the page. The second column represents the recognition result of the molecular structure at the corresponding position represented by the preset structural format. The third column represents the identity discrimination result of multiple preset optical structure recognition tools during the recognition process. If the discrimination marker is "U", it indicates that the recognition results of each preset optical structure recognition tool are different, and the user can trace and verify the original file based on the original image name. This design can better standardize the management of the molecular structure dataset and facilitate users to trace the original data. For molecular structures with inconsistent markers, users can find the original file for verification as needed.

[0090] Table 1

[0091]

[0092] As can be seen from this example, the molecular structure extraction method of this application can identify the same molecular structure on each page of the original file using multiple preset optical structure recognition tools, obtain a limited number of candidate molecular structures, and perform identity discrimination. By comparing the candidate molecular structures with inconsistent structures with the molecular structures in the original image, the best recognition result can be efficiently selected. In addition, by associating the recognition results with the corresponding discrimination markers and the original file and displaying them in the molecular dataset, users can quickly determine the reliability of the recognition results and further verify them according to their needs. Furthermore, by separately associating and storing the output files of different stages with the original files, complete association information is formed, which facilitates data management and retrieval.

[0093] Compared to traditional methods using a single OCSR tool, this design achieves higher recognition accuracy. Furthermore, for original files with different fonts, drawing styles, or resolutions, this application can reduce the quality fluctuations of input data caused by varying display effects, resulting in higher reliability.

[0094] In one embodiment, this application also provides a molecular structure dataset device, which stores the identification results of the molecular structures to be identified obtained by the above-described molecular structure extraction method; wherein: each identification result is mapped and stored to the original image and the identity discrimination result respectively. Optionally, the molecular structure dataset device can be a computer device capable of providing storage functions, such as a terminal or server.

[0095] As shown in Table 1 above, the recognition results can be stored according to a preset structural format, and the tags of the original images of the corresponding molecular structures can also be stored simultaneously. Optionally, the original images have tags associated with the original files. For example, the filenames of the original images can be named using the unique ID of the original file, such as the patent publication number, to facilitate the differentiation of the original images and the tracing of their origin, thereby constructing a complete and clear dataset.

[0096] Corresponding to the aforementioned application function implementation method embodiments, this application also provides a molecular structure extraction device and corresponding embodiments.

[0097] Figure 5 This is a schematic diagram of the molecular structure extraction device shown in this application.

[0098] See Figure 5 The molecular structure extraction device shown in this application includes an original image acquisition module 310, a chemical structure recognition module 320, and a structure screening module 330.

[0099] The original image acquisition module 310 is used to acquire the original image of the molecular structure to be identified.

[0100] The chemical structure recognition module 320 is used to identify the molecular structure to be identified in the original image according to multiple preset optical structure recognition tools, and obtain the corresponding candidate molecular structure.

[0101] The structure screening module 330 is used to compare the candidate molecular structures with each other. When at least some of the candidate molecular structures are inconsistent, the candidate molecular structure that is close to the molecular structure to be identified in the original image is selected as the identification result.

[0102] Figure 6 This is a schematic diagram of a molecular structure extraction device shown in a specific example of this application.

[0103] See Figure 6The original image acquisition module 310 includes a file processing module 311 and an image segmentation module 312. The file processing module 311 performs pagination on the original file in a preset format to obtain corresponding paginated files. The image segmentation module 312 segments each molecular structure to be identified in the paginated files into independent images, generating original images corresponding to each molecular structure. Specifically, the image segmentation module masks the molecular structures to be identified in the paginated files to generate corresponding mask images; it then segments the masked regions in the mask images to generate original images of the molecular structures to be identified.

[0104] The chemical structure recognition module 320 is used to perform vectorization processing on the original image to obtain the corresponding vector image; and to identify the molecular structure information in the vector image according to the preset optical structure recognition tool to obtain the corresponding candidate molecular structure; wherein, the molecular structure information includes chemical bond information, atom type information, charge state and atom connection information.

[0105] The molecular structure extraction device of this application further includes a discrimination labeling module 340, which is used to perform identity discrimination on each candidate molecular structure and set a discrimination label corresponding to the molecule to be identified based on whether the discrimination results are consistent. Specifically, for discrimination results where the candidate molecular structures are consistent, a consistency label is set on the molecular structure to be identified; for discrimination results where the candidate molecular structures are inconsistent, a non-consistency label is set on the molecular structure to be identified.

[0106] The structure screening module 330 is used to evaluate the similarity between each candidate molecular structure and the corresponding molecular structure in the original image when at least some or all candidate molecular structures are inconsistent, i.e., when the discrimination marker is a non-consistent marker indicating structural inconsistency; based on the similarity between each candidate molecular structure and the molecular structure in the original image, the candidate molecular structure with the highest similarity is selected as the identification result of the molecular structure to be identified. Additionally, the structure screening module 330 is also used to determine the candidate molecular structure as the identification result of the molecular structure to be identified when all candidate molecular structures are identical.

[0107] The molecular structure extraction device of this application further includes a storage module 350, which is used to store the recognition results of each molecular structure to be identified according to a preset structural format, thereby obtaining a molecular structure dataset. Optionally, the storage module 350 is used to map and store the recognized structures in the preset structural format with the original image and the identity discrimination results.

[0108] As can be seen from this example, the molecular structure extraction device of this application can extract the molecular structure to be identified from the original file through different functional modules, thereby obtaining the identification results efficiently, accurately and reliably; in addition, it can collect and store the identification results, build and update a rich molecular structure dataset, which is beneficial to the research and development needs of researchers.

[0109] The specific manner in which each module performs its operation in the above embodiments has been described in detail in the embodiments related to the method, and will not be elaborated further here.

[0110] Figure 7 This is a schematic diagram of the structure of the electronic device shown in this application.

[0111] See Figure 7 The electronic device 1000 includes a memory 1010 and a processor 1020.

[0112] The processor 1020 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0113] Memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by processor 1020 or other modules of the computer. Permanent storage devices may be read-write storage devices. Permanent storage devices may be non-volatile storage devices that retain stored instructions and data even when the computer is powered off. In some embodiments, permanent storage devices use mass storage devices (e.g., magnetic or optical disks, flash memory) as permanent storage devices. In other embodiments, permanent storage devices may be removable storage devices (e.g., floppy disks, optical drives). System memory may be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. System memory may store some or all of the instructions and data required by the processor during operation. Furthermore, memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, the memory 1010 may include a removable storage device that is readable and / or writable, such as a laser disc (CD), a read-only digital multifunction optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, a high-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not contain carrier waves or transient electronic signals transmitted wirelessly or via wired connections.

[0114] The memory 1010 stores executable code, which, when processed by the processor 1020, can cause the processor 1020 to execute part or all of the methods described above.

[0115] Furthermore, the method according to this application can also be implemented as a computer program or computer program product, which includes computer program code instructions for performing some or all of the steps in the method described above.

[0116] Alternatively, this application may be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium) storing executable code (or computer program or computer instruction code) thereon, which, when executed by a processor of an electronic device (or an electronic device, etc.), causes the processor to perform part or all of the steps of the above-described method according to this application.

[0117] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for extracting molecular structures, characterized in that, include: The process involves acquiring the original image of the molecular structure to be identified; wherein the original file in a preset format is paginated to obtain the corresponding paginated file; and each molecular structure to be identified in the paginated file is segmented into an independent image to generate the original image corresponding to the molecular structure to be identified. The molecular structures to be identified in the original image are identified by multiple preset optical structure recognition tools to obtain the corresponding candidate molecular structures. The candidate molecular structures are compared with each other, wherein an identity judgment is performed on each candidate molecular structure, and a judgment label is set for the molecular structure to be identified based on whether the judgment result is consistent; the judgment label is a non-consistent label indicating that at least some of the candidate molecular structures are inconsistent or a consistent label indicating that all the candidate molecular structures are consistent; when at least some of the candidate molecular structures are inconsistent, the similarity between the original image and the graph structure of each candidate molecular structure is calculated, and the similarity between each candidate molecular structure and the molecular structure in the corresponding original image is evaluated; the candidate molecular structure with the highest similarity is taken as the identification result of the molecular structure to be identified. When all candidate molecular structures are identical, the candidate molecular structure is determined to be the identification result of the molecular structure to be identified.

2. The method according to claim 1, characterized in that, The method further includes: The recognition results of the preset structure format are associated and stored with the original file and the discrimination label respectively to obtain the molecular structure dataset.

3. The method according to claim 1 or 2, characterized in that, The method further includes: According to the preset structural format, the recognition results of each molecular structure to be identified are stored to obtain a molecular structure dataset.

4. The method according to claim 1, characterized in that, The step of segmenting each molecular structure to be identified in the paginated file into independent images and generating the original image of the molecular structure to be identified includes: Mask the molecular structure to be identified in the paginated file to generate the corresponding mask image; The masked region in the masked image is segmented to generate the original image of the molecular structure to be identified.

5. The method according to claim 1, characterized in that, The step of identifying the molecular structure to be identified in the original image using multiple preset optical structure recognition tools to obtain the corresponding candidate molecular structure includes: The original images are vectorized to obtain the corresponding vector graphics. The molecular structure information in the vector image is identified using a preset optical structure recognition tool to obtain the corresponding candidate molecular structure; wherein, the molecular structure information includes chemical bond information, atom type information, charge state and atom connection information.

6. A molecular structure dataset device, characterized in that, It stores the identification results of the molecular structure to be identified obtained by the molecular structure extraction method according to any one of claims 1-5; wherein: Each recognition result is mapped and stored to the original image and the identity discrimination result, respectively.

7. A molecular structure extraction device, characterized in that, include: The original image acquisition module is used to acquire the original image of the molecular structure to be identified; wherein, the original file in a preset format is paginated to obtain the corresponding paginated file; each molecular structure to be identified in the paginated file is segmented into an independent image, and the original image corresponding to the molecular structure to be identified is generated respectively. The chemical structure recognition module is used to identify the molecular structure to be identified in the original image according to multiple preset optical structure recognition tools, and obtain the corresponding candidate molecular structure. The structure screening module is used to compare candidate molecular structures with each other. Specifically, it performs identity discrimination on each candidate molecular structure and sets a discrimination label corresponding to the molecular structure to be identified based on whether the discrimination results are consistent. The discrimination label can be a non-consistent label indicating that at least some of the candidate molecular structures are inconsistent, or a consistent label indicating that all candidate molecular structures are consistent. When at least some of the candidate molecular structures are inconsistent, the similarity between the original image and the graph structure of each candidate molecular structure is calculated to evaluate the similarity between each candidate molecular structure and the molecular structure in the corresponding original image. The candidate molecular structure with the highest similarity is taken as the identification result of the molecular structure to be identified. When all the candidate molecular structures are consistent, the candidate molecular structure is determined as the identification result of the molecular structure to be identified.

8. An electronic device, characterized in that, include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the molecular structure extraction method as described in any one of claims 1-5.

9. A computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the molecular structure extraction method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Methods and compounds for restoring mutant p53 function

    WO2021262596A1

  • Image recognition method and device, electronic equipment and readable storage medium

    CN113963197A

  • Systems, methods, and apparatus for processing documents to identify structures

    US20110276589A1