A method and system for chemical structure recognition
By using image computation rules and model recognition technology, chemical structure images are converted into computer-readable target chemical molecule data, solving the problem that chemical molecule images in electronic documents cannot be converted into machine-readable formats, thus improving the efficiency of scientific research.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INFINITE INTELLIGENCE PHARMACEUTICAL TECHNOLOGY CO LTD
- Filing Date
- 2023-02-21
- Publication Date
- 2026-05-12
AI Technical Summary
Existing electronic document management systems cannot convert chemical molecule images into machine-readable formats, resulting in low efficiency in scientific research.
By using image calculation rules, parameter lookup tables, image preprocessing models, and image recognition models, chemical structure images are converted into computer-readable target chemical molecule data, including the recognition and combination of element labels, chemical bond labels, and hypertext images.
It improves the efficiency of chemical molecule retrieval, reduces retrieval time, and makes it easier to store and search chemical molecule images in electronic documents.
Smart Images

Figure CN116071554B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method and system for chemical structure recognition. Background Technology
[0002] Currently, in the fields of chemistry and drug discovery, there is a vast amount of literature, including journals and patents. Most of these documents contain images of chemical molecules. When researchers search for relevant information, they need to consult paper documents or electronic documents to analyze and study the content of the relevant chemical molecules. However, paper documents are not easy to store, and the large number of paper documents makes searching difficult. In contrast, electronic documents are more convenient to store and search.
[0003] However, current systems for managing literature cannot convert chemical molecule images in electronic documents into machine-readable formats for storage. This prevents electronic documents from being filtered and retrieved based on chemical molecules, which to some extent affects the efficiency of scientific research. Summary of the Invention
[0004] To address the issue that chemical molecule images in electronic documents cannot be converted into machine-readable formats for storage, this application provides a chemical structure identification method and system.
[0005] In a first aspect of this application, a method for identifying chemical structures is provided. The method includes:
[0006] Acquire chemical structure images;
[0007] The target value is determined based on the preset image calculation rules and the chemical structure image;
[0008] The model parameters are determined based on the preset parameter lookup table and the target value;
[0009] The preprocessed image is determined based on the image preprocessing model, the chemical structure image, and the model parameters;
[0010] Based on the image recognition model and the preprocessed image, a target image is determined, which includes element labels, chemical bond labels, and hypertext images;
[0011] Based on the chemical molecule construction rules and the target image, the target chemical molecule data is determined.
[0012] As can be seen from the above technical solution, based on the preset image calculation rules, the chemical structure image is analyzed and calculated to determine the target value. According to the parameter lookup table and the target value, the model parameters corresponding to the target value are determined. Based on the image preprocessing model, the chemical structure image, and the model parameters, the chemical structure image is preprocessed to determine the preprocessed image. Based on the image recognition model and the preprocessed image, the target image is determined. Finally, based on the chemical molecule construction rules and the target image, the target chemical molecule data is determined. Converting the chemical structure image into target chemical molecule data in a computer-readable format improves the problem that chemical molecule images in electronic documents cannot be converted into machine-readable formats for storage. This reduces the retrieval time for chemical molecules in scientific research and improves the efficiency of scientific research to a certain extent.
[0013] In one possible implementation, acquiring the chemical structure image includes:
[0014] Obtain electronic documents;
[0015] Based on the image conversion rules and the electronic document, determine the document image;
[0016] The document image is segmented according to the semantic segmentation model to determine the chemical structure image.
[0017] In one possible implementation, determining the target value based on preset image calculation rules and the chemical structure image includes:
[0018] The chemical structure image includes multiple pixels, and each pixel corresponds to a pixel value;
[0019] Retrieve the target pixels whose pixel values are preset values from the pixel values corresponding to each row and each column of pixels;
[0020] Based on the target pixels, determine the maximum number of consecutive target pixels in each row and each column;
[0021] According to the outlier removal rules, the maximum number is removed to determine the target number sequence;
[0022] Calculate the median of the target number sequence, where the median is the target value.
[0023] In one possible implementation, determining the preprocessed image based on the image preprocessing model, the chemical structure image, and the model parameters includes:
[0024] Obtain image erosion and image dilation models;
[0025] Based on the model parameters and the image erosion model, determine the target erosion model;
[0026] Based on the model parameters and the image dilation model, determine the target dilation model;
[0027] The corrosion image is determined based on the chemical structure image and the target corrosion model;
[0028] Based on the erosion image and the target dilation model, a preprocessed image is determined.
[0029] As can be seen from the above technical solutions, in the process of compound image recognition, it is easy to confuse coarse single bonds and chiral bonds. By determining the target expansion model and target corrosion model corresponding to the model parameters, the chemical structure image can be preprocessed using the above target expansion model and target corrosion model. According to different chemical structure images and model parameters, the corresponding preprocessing model can be used, which can improve the recognition accuracy of the subsequent image recognition model on the preprocessed image to a certain extent.
[0030] In one possible implementation, determining the target chemical molecule data based on the chemical molecule construction rules and the target image includes:
[0031] The target image is segmented according to the image recognition model to determine the hypertext image;
[0032] Based on the OCR text recognition model and the hypertext image, the hypertext string is determined;
[0033] According to preset combination rules, the element tags, chemical bond tags, and hypertext strings are combined to form target chemical molecule data.
[0034] In one possible implementation, prior to acquiring the chemical structure image, the semantic segmentation model is further trained.
[0035] Obtain a training dataset, which includes multiple training images, including text labels and chemical molecule location labels;
[0036] The training dataset is input into a preset segmentation model to determine the semantic segmentation model.
[0037] In a second aspect of this application, a model training method is provided. The method includes:
[0038] Obtain a standard image;
[0039] The standard image is input into a preset generative adversarial network model to determine the feature image;
[0040] The standard image and the feature image constitute the first model training set;
[0041] The first model training set is input into the preset target model training framework to determine the target model.
[0042] In one possible implementation, the method further includes, when the target model is an image recognition model, the training process of the generative adversarial network model:
[0043] Obtain the raw compound molecule image dataset;
[0044] The original compound molecule image dataset is input into a preset tool to determine the target compound molecule image dataset;
[0045] The original compound molecule image dataset and the target compound molecule image dataset constitute the second model training set;
[0046] The second model training set is input into a preset adversarial network model training framework to determine the generation of the adversarial network model.
[0047] In one possible implementation, the method further includes, when the target model is a semantic segmentation model, the training process of the generative adversarial network model:
[0048] Obtain the original document image dataset;
[0049] The original document image dataset is input into a preset model to determine the target document image dataset;
[0050] The original document image dataset and the target document image dataset constitute the third model training set;
[0051] The training set of the third model is input into the preset adversarial network model training framework to determine the generation of the adversarial network model.
[0052] In a third aspect of this application, a chemical structure recognition system is provided. The system includes:
[0053] The image acquisition module is used to acquire images of chemical structures;
[0054] The image calculation module is used to determine the target value based on preset image calculation rules and the chemical structure image;
[0055] The parameter determination module is used to determine the model parameters based on a preset parameter lookup table and the target value.
[0056] An image preprocessing module is used to determine a preprocessed image based on an image preprocessing model, the chemical structure image, and the model parameters;
[0057] An image recognition module is used to determine a target image based on an image recognition model and the preprocessed image, wherein the target image includes element labels and chemical bond labels;
[0058] The molecular construction module is used to determine target chemical molecule data based on chemical molecule construction rules and the target image.
[0059] In summary, this application includes at least one of the following beneficial technical effects:
[0060] 1. By employing image calculation rules, parameter lookup tables, image preprocessing models, image recognition models, and chemical molecule construction rules, chemical structure images are segmented and analyzed to determine target chemical molecule data. This process converts chemical structure images into target chemical molecule data in a computer-readable format, addressing the issue that chemical molecule images in electronic documents cannot be converted into machine-readable formats for storage.
[0061] 2. In the process of compound image recognition, it is easy to confuse different chemical bonds. By calculating the target value corresponding to the chemical structure image, and retrieving the corresponding model parameters according to the target value, and determining the image preprocessing model that is suitable for the above chemical structure image according to the model parameters, the chemical structure image is preprocessed by using the suitable image preprocessing model. By using the corresponding preprocessing model for different chemical structure images, the recognition accuracy of the subsequent image recognition model on the preprocessed image can be improved to a certain extent. Attached Figure Description
[0062] Figure 1 This is a flowchart illustrating a chemical structure identification method provided in this application.
[0063] Figure 2 This is a flowchart illustrating a model training method provided in this application.
[0064] Figure 3 This is a system block diagram of a chemical structure recognition system provided in this application.
[0065] In the figure, 200 is the chemical structure recognition system; 201 is the image acquisition module; 202 is the image calculation module; 203 is the parameter determination module; 204 is the image preprocessing module; 205 is the image recognition module; and 206 is the molecular construction module. Detailed Implementation
[0066] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0067] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0068] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.
[0069] This application provides a chemical structure identification method, the main process of which is described below.
[0070] like Figure 1 As shown:
[0071] Step S101: Obtain chemical structure image.
[0072] Specifically, electronic documents are acquired, and document images are determined based on image conversion rules and the aforementioned electronic documents. The document images are then segmented using a semantic segmentation model to determine chemical structure images. The aforementioned electronic documents represent electronic materials from different magazines, journals, and patent documents. These electronic materials can be in Word, PDF, or other formats, and are not limited here. In this embodiment, the aforementioned electronic documents are in PDF format. Converting PDF documents into images is a well-known technique among those skilled in the art and will not be elaborated upon here. The aforementioned document images include images containing text, images containing chemical molecules, and images containing both text and chemical molecules. The aforementioned document images are input into a preset semantic segmentation model, and the images containing chemical molecules are segmented to determine the chemical structure images. These chemical structure images are images containing only chemical structures.
[0073] Before step S101, the process includes training a semantic segmentation model: acquiring a training dataset, which includes multiple training images, each containing text labels and chemical molecule position labels; inputting the training dataset into a preset segmentation model to determine the semantic segmentation model. Different literature sources and journals have different placement and display styles for chemical structures. The text labels include the text layout labels and text font labels of the training images, and the chemical molecule position labels include the insertion position of the compound image and the insertion position of the compound reaction arrow. The training images are then input into the preset segmentation model to determine the semantic segmentation model. The trained semantic segmentation model can cut out the chemical molecule portion from images containing text and chemical molecules; the cut-out portion is the aforementioned chemical structure image. Acquiring the training dataset requires inputting the images converted from the aforementioned PDF documents into a generative adversarial network (GAN) model. The GAN model converts these images into images of different styles, where the positions, bond lengths, and bond thicknesses of chemical molecules within the image differ. These images of different styles constitute the training dataset.
[0074] Step S102: Determine the target value based on the preset image calculation rules and chemical structure image.
[0075] Specifically, the chemical structure image includes multiple pixels, each corresponding to a pixel value. Target pixels with preset values are obtained from the pixel values corresponding to each row of pixels. These target pixels are black pixels in the chemical structure image. Based on the target pixels, the maximum number of consecutive target pixels in each row is determined, i.e., the maximum number of consecutive black pixels is determined, and the maximum value of this consecutive number is obtained. For each column of pixels, black pixels also need to be determined, and the maximum value of consecutive black pixels is obtained. Since there are multiple maximum values, these maximum values are grouped into a first sequence. Outliers in the first sequence are identified and removed according to outlier removal rules. The sequence after outlier removal is the target number sequence. The median of the target number sequence is calculated, and this median is the target value. For example, if the chemical structure image is a rectangular image containing multiple pixels, the maximum number of consecutive black pixels in each row or column of the rectangular image is obtained, resulting in N+M maximum consecutive black pixel count values, where M is the number of columns and N is the number of rows. Then, outliers are identified and removed from these N+M values. After removing outliers, the median of the remaining values is taken as the target value.
[0076] In this embodiment, the chemical structure image consists of 768*465 pixels, i.e., 768 rows and 465 columns. For a given row, the number of consecutive black pixels is counted. This count is either two or four, meaning that there are two adjacent black pixels and four adjacent black pixels in that row. At least one non-black pixel is included between the set of two black pixels and the set of four black pixels. The maximum value for this row is 4. The same operation is performed on other rows and columns, resulting in 768 + 465 = 1233 maximum values, forming a first sequence. Outliers in this first sequence are calculated. The calculation of outliers is a technique known to those skilled in the art and will not be elaborated here. After deleting the outliers, a target number sequence is formed. The median of the target number sequence is calculated, and this median is the target value.
[0077] Step S103: Determine the model parameters according to the preset parameter comparison table and target values.
[0078] Specifically, the above parameter comparison table includes the correspondence between target values and model parameters. The parameter comparison table is manually set and stored in the database in advance. After the target value is determined, the corresponding model parameters can be retrieved from the database.
[0079] Step S104: Determine the preprocessed image based on the image preprocessing model, chemical structure image, and model parameters.
[0080] Specifically, the process involves acquiring image erosion and image dilation models; determining the target erosion model based on the model parameters and the image erosion model; determining the target dilation model based on the model parameters and the image dilation model; determining the eroded image based on the chemical structure image and the target erosion model; and determining the preprocessed image based on the eroded image and the target dilation model. The chemical structure image is first dilated, followed by image erosion. Both image dilation and erosion are techniques well-known to those skilled in the art and will not be elaborated upon here. Both the image dilation and image erosion models include a kernel-size parameter; different values of the kernel-size parameter result in different effects on image processing. By retrieving the corresponding model parameters based on the target values and setting the kernel-size parameter in the image dilation and image erosion models to the corresponding values, the accuracy of subsequent image recognition models in recognizing the preprocessed image can be improved to some extent.
[0081] In the process of compound image recognition, it is easy to confuse relatively thick single bonds with chiral bonds. The aforementioned chiral bonds refer to bonds connected to chiral carbons. These chiral bonds include solid wedges and dashed wedges. To improve the problem of confusion between single bonds and chiral bonds, the thickness of bond lines in chemical images is determined by obtaining the number of consecutive black pixels. Different kernel-size parameters are designed for different bond line thicknesses, and image dilation and erosion processing of chemical structure images are implemented through different kernel-size parameters to determine the preprocessed image.
[0082] Step S105: Determine the target image based on the image recognition model and the preprocessed image.
[0083] Specifically, inputting the preprocessed image into the image recognition model yields the target image, which includes element labels, chemical bond labels, and hypertext images. The element labels include Si, N, Br, S, I, Cl, H, P, O, C, B, F, and Text, among others. Text represents atoms other than those listed above, along with hypertext. The hypertext represents the atomic groups that make up the chemical molecule. The chemical bond labels include Single, Double, Solid Wedge, DashWedge, and Wavy, among others. In this embodiment, the image recognition model is the YOLOv5 model. In other embodiments, object detection models such as Mask R-CNN, Fast R-CNN, EfficientDet, and Swin Transformer can be used instead; no limitation is imposed here.
[0084] Step S106: Determine the target chemical molecule data based on the chemical molecule construction rules and the target image.
[0085] Specifically, the target image is segmented based on the image recognition model to determine the hypertext image; the hypertext string is determined based on the OCR text recognition model and the hypertext image; and the element labels, chemical bond labels, and hypertext string are combined according to the preset combination rules to form the target chemical molecule data.
[0086] When the target image contains hypertext, the hypertext portion is extracted using an image recognition model to form the hypertext image. This hypertext image is then input into an OCR text recognition model, converting it into a computer-readable hypertext string. This hypertext string is then combined with element tags and chemical bond tags from the target image to form the target chemical molecule data. When the target image does not contain hypertext, only element tags and chemical bond tags need to be combined according to preset rules to form the target chemical molecule data.
[0087] The image recognition model described above is used to segment hypertext images from target images. Image recognition models are well-known techniques to those skilled in the art and will not be elaborated upon here. In this embodiment, the OCR text recognition model is the EasyOCR model; however, other OCR models can also be used, and will not be elaborated upon here. The process of combining hypertext strings, element tags, and chemical bond tags is implemented using the rdkit toolkit. The rdkit toolkit and its usage are well-known to those skilled in the art and will not be elaborated upon here.
[0088] It also includes evaluation methods for chemical structure identification methods:
[0089] A test dataset is obtained, comprising multiple test structure images and corresponding target structure data. The target structure data includes target strings and target structure diagrams. The test structure images are processed according to steps S101 to S106 to obtain target chemical molecule data corresponding to the test structure images. The target chemical molecule data includes molecular diagram data and string data. Using the Tanimoto coefficient similarity evaluation algorithm, the target strings and string data are compared to obtain a similarity coefficient. When the similarity coefficient is greater than a preset similarity value, it indicates that the computer-readable string data corresponding to the test structure image can be obtained through the chemical structure recognition method. Using the MCS (Maximum Common Substructure) algorithm, the target structure diagram and molecular diagram data are compared to obtain the maximum common substructure. When the maximum common substructure is identical to the chemical molecule in the molecular diagram data, it indicates that the computer-readable molecular diagram data corresponding to the test structure image can be obtained through the chemical structure recognition method. The MCS (Maximum Common Substructure) algorithm and the Tanimoto coefficient similarity evaluation algorithm are well-known techniques to those skilled in the art and will not be elaborated upon here. The above method can be used to determine whether each test structure image can yield corresponding computer-readable data through chemical structure recognition. The number of test structure images in the test dataset is denoted as the total number of images, and the number of test structure images from which corresponding data can be obtained through chemical structure recognition is denoted as the pass value. The evaluation coefficient of the chemical structure recognition method is the pass value divided by the total number of images.
[0090] This application provides a model training method, the main process of which is described below.
[0091] like Figure 2 As shown:
[0092] Step S201: Obtain a standard image.
[0093] Specifically, standard images of compounds are obtained. These standard images refer to images generated using preset tools or models based on relevant data of chemical molecules.
[0094] When the target model is an image recognition model, the standard image represents a compound image generated using the rdkit tool based on the SMILE string. When the target model is a semantic segmentation model, the standard image represents an image generated using a preset model based on the molecular image and text data of the compound. This preset model can combine text, molecular images, or other images to generate different standard images. The standard image represents the aforementioned molecular image.
[0095] Step S202: Input the standard image into the preset generative adversarial network model to determine the feature image.
[0096] Specifically, the aforementioned standard images are input into a preset generative adversarial network model, which then converts them into feature images of different styles. These feature images refer to compound images with different bond lengths and bond thicknesses. The aforementioned standard images and feature images constitute the first model training set.
[0097] Step S203: Input the first model training set into the preset target model training framework to determine the target model.
[0098] Specifically, the first model training set is used as the input dataset and input into the preset target model training framework to determine the target model.
[0099] When the target model is an image recognition model, the training process of the generative adversarial network model is as follows:
[0100] Specifically, the process involves acquiring a dataset of original compound molecular images, which represents compound images and their corresponding SMILE strings obtained from journal articles of different styles. This dataset is then input into a pre-defined tool to determine the target compound molecular image dataset. Specifically, the SMILE strings corresponding to the compound images are used to generate the target compound molecular image dataset using the rdkit tool. The rdkit tool can be used to adjust parameters such as bond length, bond width, number of dashed lines, line width, bond-character distance, character font, character size, left-aligned or right-aligned superatomic text, and superscript or subscript text. The original and target compound molecular image datasets together form the second model training set. This second model training set is then input into a pre-defined adversarial network model training framework to generate the adversarial network model.
[0101] When the target model is a semantic segmentation model, the training process of the generative adversarial network model is as follows:
[0102] Specifically, the process involves acquiring original document image datasets, which refer to image, text, and molecular image data obtained from the conversion of electronic documents of different journal articles; inputting the text and molecular image data into a pre-defined model to determine the target document image dataset; the pre-defined model can combine text, molecular images, or other images to generate different target document image datasets. The target document image dataset represents PDF images containing molecular image location tags. The pre-defined model is a technique well-known to those skilled in the art and will not be elaborated upon here. The original document image dataset and the target document image dataset constitute the training set for the third model; the training set for the third model is then input into a pre-defined adversarial network model training framework to determine the generation of the adversarial network model.
[0103] This application provides a chemical structure recognition system 200, with reference to... Figure 3 The chemical structure recognition system 200 includes:
[0104] Image acquisition module 201 is used to acquire chemical structure images;
[0105] Image calculation module 202 is used to determine the target value according to preset image calculation rules and the chemical structure image;
[0106] The parameter determination module 203 is used to determine the model parameters according to the preset parameter lookup table and the target value;
[0107] Image preprocessing module 204 is used to determine a preprocessed image based on the image preprocessing model, the chemical structure image, and the model parameters;
[0108] Image recognition module 205 is used to determine a target image based on an image recognition model and the preprocessed image, wherein the target image includes element labels and chemical bond labels;
[0109] The molecular construction module 206 is used to determine target chemical molecule data based on chemical molecule construction rules and the target image.
[0110] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0111] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing application concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions claimed in this application.
Claims
1. A method for identifying chemical structures, characterized in that, include: Acquire chemical structure images; The target value is determined based on the preset image calculation rules and the chemical structure image; Determining the target value includes: the chemical structure image includes multiple pixels, each pixel corresponding to a pixel value; obtaining target pixels with preset values among the pixel values corresponding to each row and each column; determining the maximum number of consecutive target pixels in each row and each column based on the target pixels; removing the maximum number according to outlier removal rules to determine the target number sequence; and calculating the median of the target number sequence, where the median is the target value. The model parameters are determined according to the preset parameter lookup table and the target value; the preprocessed image is determined according to the image preprocessing model, the chemical structure image, and the model parameters. The process of determining the preprocessed image includes: acquiring an image erosion model and an image dilation model; determining a target erosion model based on the model parameters and the image erosion model; determining a target dilation model based on the model parameters and the image dilation model; determining an eroded image based on the chemical structure image and the target erosion model; and determining a preprocessed image based on the eroded image and the target dilation model. Both the image dilation model and the image erosion model include a kernel-size parameter; different values of the kernel-size parameter result in different effects on image processing. The corresponding model parameters are retrieved based on the target values, and the kernel-size parameter in the image dilation model and the image erosion model is set to the corresponding values of the model parameters. Based on the image recognition model and the preprocessed image, a target image is determined, which includes element labels, chemical bond labels, and hypertext images; based on the chemical molecule construction rules and the target image, target chemical molecule data is determined.
2. The chemical structure identification method according to claim 1, characterized in that, The acquisition of chemical structure images includes: Obtain electronic documents; Based on the image conversion rules and the electronic document, determine the document image; The document image is segmented according to the semantic segmentation model to determine the chemical structure image.
3. The chemical structure identification method according to claim 1, characterized in that, The step of determining the target chemical molecule data based on the chemical molecule construction rules and the target image includes: The target image is identified using an image recognition model to determine the hypertext image; Based on the OCR text recognition model and the hypertext image, the hypertext string is determined; According to preset combination rules, the element tags, chemical bond tags, and hypertext strings are combined to form target chemical molecule data.
4. The chemical structure identification method according to claim 2, characterized in that, Prior to acquiring the chemical structure image, the semantic segmentation model is trained: Obtain a training dataset, which includes multiple training images, including text labels and chemical molecule location labels; The training dataset is input into a preset segmentation model to determine the semantic segmentation model.
5. A chemical structure recognition system, characterized in that, include: Image acquisition module (201) is used to acquire chemical structure images; The image calculation module (202) is used to determine the target value according to the preset image calculation rules and the chemical structure image; Determining the target value includes: the chemical structure image includes multiple pixels, each pixel corresponding to a pixel value; obtaining target pixels with preset values among the pixel values corresponding to each row and each column; determining the maximum number of consecutive target pixels in each row and each column based on the target pixels; removing the maximum number according to outlier removal rules to determine the target number sequence; and calculating the median of the target number sequence, where the median is the target value. The parameter determination module (203) is used to determine the model parameters according to the preset parameter reference table and the target value; The image preprocessing module (204) is used to determine the preprocessed image based on the image preprocessing model, the chemical structure image, and the model parameters; The process of determining the preprocessed image includes: acquiring an image erosion model and an image dilation model; determining a target erosion model based on the model parameters and the image erosion model; determining a target dilation model based on the model parameters and the image dilation model; determining an eroded image based on the chemical structure image and the target erosion model; and determining a preprocessed image based on the eroded image and the target dilation model. Both the image dilation model and the image erosion model include a kernel-size parameter; different values of the kernel-size parameter result in different effects on image processing. The corresponding model parameters are retrieved based on the target values, and the kernel-size parameter in the image dilation model and the image erosion model is set to the corresponding values of the model parameters. An image recognition module (205) is used to determine a target image based on an image recognition model and the preprocessed image, wherein the target image includes element labels and chemical bond labels; The molecular construction module (206) is used to determine target chemical molecule data based on chemical molecule construction rules and the target image.