Structured decomposition and information identification method of multi-modal data, medium and equipment

By combining the domain-fine-tuned DETR model with OCSR tools, the problem of information identification in multimodal data in scientific literature in the fields of chemistry, biology and pharmaceuticals has been solved. It has achieved accurate division of logical regions and accurate identification of chemical structure objects, thus improving the accuracy and quality of information identification.

CN121033879APending Publication Date: 2025-11-28HANGZHOU LIWU YINGJI TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511187849.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify information from complexly formatted multimodal data in scientific literature in the fields of chemistry, biology, and pharmaceuticals, particularly in the identification of chemical structure images in illustrations and tables.

Method used

The domain-fine-tuned DETR model is used to identify logical region categories, and OCSR tools are used to extract differentiated information, including text recognition and SMILES string recognition of chemical structure objects. Multi-level verification is used to ensure accuracy.

Benefits of technology

It has achieved accurate identification of information in scientific literature and accurate identification of chemical formula structure images, thus improving the accuracy and quality of information identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033879A_ABST
    Figure CN121033879A_ABST
Patent Text Reader

Abstract

The invention provides a structural decomposition and information identification method of multi-modal data, a medium and equipment, and the method comprises the steps: obtaining to-be-processed multi-modal literature data, the multi-modal literature data being documents or pictures, the types of the documents including word documents and PDF (Portable Document Format) documents; the method comprises the following steps: converting to-be-processed multi-modal literature data into to-be-processed literature data in an image form, preprocessing the to-be-processed literature data to obtain input image data, inputting the input image data into a field fine-tuning DETR model, identifying a logic region category of the input image data through the field fine-tuning DETR model, obtaining a region category identification result, and outputting the region category identification result. The logic region category comprises a title region, an author region, an abstract region, a text region, an illustration region, a table region, a formula region, a footer region and a reference region; and carrying out differential information extraction on each region category identification result to obtain information corresponding to each region category identification result so as to realize accurate identification of information in the multi-modal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and more specifically, to a method, medium, and device for structured decomposition and information recognition of multimodal data. Background Technology

[0002] In the fields of chemistry and biopharmaceuticals, current information processing suffers from problems such as modal fragmentation and siloed knowledge representation. Overcoming these issues necessitates the processing of multimodal data. Existing data sources include scientific literature, patent documents, and professional reports, with scientific literature (such as academic papers and journal articles) exhibiting the most complex formatting. Information recognition from this type of complexly formatted multimodal data is fundamental to building a professional knowledge base. Accurately identifying information from unstructured images or documents, particularly in the fields of chemistry, biology, and pharmaceuticals, presents significant challenges. Various illustrations and tables often contain complex chemical structure images. How to accurately identify information from this type of data is a problem that needs to be solved in this field. Summary of the Invention

[0003] The purpose of this application is to provide a method, medium, and device for structured decomposition and information recognition of multimodal data, so as to achieve accurate recognition of information in multimodal data (especially scientific literature).

[0004] To achieve the above objectives, the embodiments of this application are implemented in the following manner: In a first aspect, embodiments of this application provide a method for structured decomposition and information recognition of multimodal data, comprising: acquiring multimodal literature data to be processed, wherein the multimodal literature data is a document or an image, and the document type includes Word documents and PDF documents; converting the multimodal literature data to be processed into image-based literature data and preprocessing it to obtain input image data; inputting the input image data into a domain fine-tuning DETR model, identifying the logical region categories of the input image data through the domain fine-tuning DETR model, obtaining and outputting region category recognition results, wherein the logical region categories include title region, author region, abstract region, main text region, illustration region, table region, formula region, footer region, and reference region; and extracting differential information from each region category recognition result to obtain information corresponding to each region category recognition result.

[0005] In conjunction with the first aspect, in the first possible implementation of the first aspect, the multimodal literature data belongs to the fields of chemistry, biology, and pharmaceuticals. Differential information extraction is performed on the identification results of each region category to obtain the information corresponding to each region category identification result, including: For the title region, author region, abstract region, footer region, and reference region: text recognition methods are used to identify the text information corresponding to each region; For the illustration region: OCSR is used to identify the SMILES string corresponding to each chemical structure object within the illustration region, and post-processing verification is performed to obtain the SMILES string information corresponding to the illustration region; For the main text region: when there are no chemical structure images within the main text region, text recognition methods are used to identify the text information corresponding to the main text region. When chemical structure images are present within the text area, OCSR is used to identify the information in the chemical structure images, and text recognition methods are used to identify the remaining parts within the text area to determine the comprehensive information corresponding to the text area. This comprehensive information includes text information and the string "SMILES". For table areas: logical structure recognition is performed on the table area to identify the table skeleton containing several cells. Then, information recognition is performed on the data objects within each cell to obtain the comprehensive information corresponding to the table area. When the data object within a cell is a chemical structure image, OCSR is used to identify the information in the data object within the cell to obtain the corresponding string "SMILES". When the data object within a cell is text data, text recognition methods are used to identify the text information corresponding to the cell.

[0006] In conjunction with the first possible implementation of the first aspect, in the second possible implementation of the first aspect, OCSR is used for information recognition to determine the SMILES string corresponding to each chemical structure object within the illustration area, and post-processing verification is performed to obtain the SMILES string information corresponding to the illustration area. This includes: using OCSR for information recognition to determine the SMILES string corresponding to each chemical structure object within the illustration area; verifying the atomic valence state for the SMILES string corresponding to each chemical structure object; verifying the aromaticity perception for the SMILES string corresponding to each chemical structure object; and verifying the atomic valence state for the SMILES string corresponding to each chemical structure object. The system performs stereochemical verification on each chemical structure object. Based on the verification results of atomic valence state, aroma perception, and stereochemical verification, the recognition confidence level for each chemical structure object is determined. When the recognition confidence level of a chemical structure object reaches the confidence level threshold, the identified SMILES string is determined to be the SMILES string corresponding to that chemical structure object. If the recognition confidence level of a chemical structure object does not reach the confidence level threshold, a manual review or feedback loop is triggered. The feedback loop means that a second recognition is performed after adjusting the image preprocessing parameters of the illustration area. If the second recognition still fails to reach the confidence level threshold for the recognition of the chemical structure object, a manual review is triggered.

[0007] In conjunction with the second possible implementation of the first aspect, in the third possible implementation of the first aspect, atomic valence state verification is performed on the SMILES string corresponding to each chemical structure object. This includes: for each SMILES string corresponding to a chemical structure object: determining the explicit valence state of each atom in the chemical structure object based on the number of bonds in the SMILES string; converting the SMILES string into an RDKit molecule object and traversing the atoms in the RDKit molecule object to determine the implicit valence state of each atom in the chemical structure object; calculating the actual valence state based on the explicit and implicit valence states of each atom in the chemical structure object: actual valence state = explicit valence state + implicit valence state; obtaining predefined valence state rules, where the predefined valence state rules include fixed-value valence states and list-value valence states, where fixed-value valence states indicate that the atom can only be one agreed-upon valence state, and list-value valence states indicate that the atom needs to be any valence state in the list; verifying the actual valence state of each atom in the chemical structure object based on the predefined valence state rules to determine the atomic valence state verification result of the chemical structure object.

[0008] In conjunction with the second possible implementation of the first aspect, in the fourth possible implementation of the first aspect, aroma perception verification is performed for the SMILES string corresponding to each chemical structure object, including: for the SMILES string corresponding to each chemical structure object: converting the SMILES string into an RDKit molecular object; detecting each ring system in the RDKit molecular object to determine whether each ring system conforms to the Hückel rule; if the ring system conforms to the Hückel rule, marking this ring system as an aromatic ring; if the ring system does not conform to the Hückel rule, determining that this ring system is not an aromatic ring; judging whether the ring system marked with an aromatic ring is consistent with the aromatic ring represented in the SMILES string; if they are consistent, determining that the aroma perception verification corresponding to this chemical structure object has passed.

[0009] In conjunction with the second possible implementation of the first aspect, in the fifth possible implementation of the first aspect, stereochemical verification is performed on the SMILES string corresponding to each chemical structure object, including: for the SMILES string corresponding to each chemical structure object: converting the SMILES string into an RDKit molecule object; detecting whether a chiral center exists in the RDKit molecule object; if a chiral center exists: detecting whether the number of neighbors of this chiral center is 4; if the number of neighbors of this chiral center is not 4, determining that this chiral center is invalid; if the number of neighbors of this chiral center is 4, determining whether the neighbors of this chiral center are all different; if this chiral center has the same neighbors, determining that this chiral center is invalid; detecting the wedge bonds or dashed bonds of this chiral center; if this chiral center has multiple wedge bonds or dashed bonds pointing in the same direction, determining that this chiral center is invalid; if no invalid chiral center exists, determining that the stereochemical verification corresponding to this chemical structure object passes.

[0010] In conjunction with the second possible implementation of the first aspect, in the sixth possible implementation of the first aspect, the recognition confidence level corresponding to each chemical structure object is determined based on the atomic valence state verification result, aromaticity perception verification result, and stereochemistry verification result corresponding to each chemical structure object. This includes: for each chemical structure object: if the atomic valence state verification result, aromaticity perception verification result, and stereochemistry verification result corresponding to this chemical structure object are all verified as passed, the recognition confidence level corresponding to this chemical structure object is determined to be 1; if at least one of the atomic valence state verification result, aromaticity perception verification result, and stereochemistry verification result corresponding to this chemical structure object fails verification, the recognition confidence level corresponding to this chemical structure object is determined to be x, where x is less than the confidence level threshold.

[0011] In conjunction with the first possible implementation of the first aspect, in the seventh possible implementation of the first aspect, a text recognition method is used to identify the text information corresponding to the main text area, including: using a blank area analysis algorithm to identify the vertical and horizontal blank areas in the main text area to identify the column boundaries in the main text area; for each column in the main text area: identifying all independent characters or symbols through connected component analysis, then applying the K-nearest neighbor clustering algorithm to aggregate characters into words, and then aggregating words into text lines; integrating the text lines of each column in the main text area to obtain the text information corresponding to the main text area.

[0012] Secondly, embodiments of this application provide a storage medium disposed within an electronic device, comprising a stored program, wherein, when the program is executed, it controls the electronic device containing the storage medium to perform the structured decomposition and information recognition method for multimodal data as described in the first aspect or any possible implementation thereof.

[0013] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, they implement the steps of the structured decomposition and information recognition method for multimodal data as described in the first aspect or any possible implementation of the first aspect.

[0014] Beneficial effects: This solution acquires multimodal literature data (documents or images, including Word and PDF documents); transforms the multimodal literature data into image format and preprocesses it to obtain input image data; then, using a domain-adjusted DETR model (domain: chemistry, biology, and pharmaceuticals), it identifies the logical region categories of the input image data (title region, author region, abstract region, body text region, illustration region, table region, formula region, footer region, and reference region); the region category identification results are obtained and output; and then, differential information is extracted for each region category identification result (for title region, author region, abstract region, footer region, and reference region). Reference area: Text recognition methods are used to identify the text information corresponding to each area. Illustration area: OCSR is used to identify the SMILES string corresponding to each chemical structure object within the illustration area, and post-processing verification is performed to obtain the SMILES string information corresponding to the illustration area. Main text area: A combination of text recognition methods and OCSR is used for identification to obtain comprehensive information. Table area: Logical structure recognition is performed on the table area to identify the table skeleton containing several cells, and then information recognition is performed on the data objects within each cell to obtain the comprehensive information corresponding to the table area. This yields the information corresponding to the identification results for each area category. This approach enables precise segmentation of complex scientific literature layouts, adopts differentiated identification schemes for different areas, and accurately identifies chemical formula structure images involved in chemistry, biology, and pharmaceutical fields. Post-processing verification is performed on the SMILES strings identified from the images, thereby achieving accurate identification of information in multimodal data (especially scientific literature).

[0015] Due to the complexity of chemical formulas, to ensure accuracy in identification, the process of identifying chemical formula structure images (mainly in image and table areas, but some simple chemical formula structure images may also exist in the text area) involves performing atomic valence state verification, aromaticity perception verification, and stereochemistry verification after identifying the SMILES string using OCSR tools. Based on the results of these verifications, the identification confidence level for each chemical structure object is determined. When necessary (when the identification confidence level does not reach the confidence threshold), manual review or feedback loops are triggered. By performing multi-level verification, the accuracy of identifying complex chemical structures can be improved, ensuring the quality of information identification.

[0016] Atomic valence verification determines the explicit valence state of each atom by counting the bonds in the SMILES string. After converting the SMILES string into an RDKit molecular object, the implicit valence state of each atom is determined. The actual valence state (actual valence state = explicit valence state + implicit valence state) is then calculated. Combined with predefined valence state rules, this verifies the actual valence state of each atom in the chemical structure object, yielding the atomic valence state verification result. Aromaticity sensing verification uses the RDKit molecular object converted from the SMILES string to check if each ring system conforms to Hückel's rule, determining if the ring system is an aromatic ring. It then checks if the ring system marking the aromatic ring matches the aromatic ring represented in the SMILES string to determine if the aromaticity sensing verification for this chemical structure object passes. For stereochemistry verification, it checks if a chiral center exists in the RDKit molecular object converted from the SMILES string. It determines if the chiral center is invalid by checking if the number of its neighbors is 4, whether the neighbors are all different, and whether multiple wedge or dashed bonds point in the same direction, thus obtaining the stereochemistry verification result. Based on the characteristics of complex chemical structures, this solution employs multi-level post-processing verification. After the SMILES string is identified using OCSR tools, verification can be performed to determine whether the identification results of the chemical structure conform to the objective laws of the chemical field, thereby improving the accuracy of information identification.

[0017] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating the structured decomposition and information recognition method for multimodal data provided in this application embodiment.

[0020] Figure 2 This is a schematic diagram of the area to be illustrated.

[0021] Figure 3 This is a schematic diagram of the table area. Detailed Implementation

[0022] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0023] To achieve information recognition from multimodal data, this embodiment provides a method for structured decomposition and information recognition of multimodal data, applicable to electronic devices (servers or smart terminals). Please refer to [link to relevant documentation]. Figure 1 The structured decomposition and information recognition method for multimodal data may include steps S10, S20, S30 and S40.

[0024] First, the electronic device can run step S10.

[0025] Step S10: Obtain the multimodal literature data to be processed, wherein the multimodal literature data is a document or an image, and the document type includes Word document and PDF document.

[0026] In this embodiment, the electronic device can acquire multimodal literature data to be processed. The multimodal literature data consists of scientific literature in the fields of chemistry, biology and pharmaceuticals. It is in image format (if there are multiple images, there needs to be a document identifier and page number order associated with the same document) or document format, such as Word documents and PDF documents, mainly scientific literature (journals, papers, etc.).

[0027] In order to achieve structured decomposition of multimodal literature data so as to identify differentiated information in each region, the electronic device can run step S20.

[0028] Step S20: Convert the multimodal literature data to be processed into image data and preprocess it to obtain input image data.

[0029] In this embodiment, the electronic device can convert the multimodal literature data to be processed into image-based literature data and perform preprocessing to obtain input image data.

[0030] For example, electronic devices can call APIs to convert PDFs into image formats (for multi-page documents, they are converted into multiple images in sequence, with each image associated with the same document's document identifier and page number order). For Word documents, they can be converted into PDFs and then into image formats (again, with each image associated with the same document's document identifier and page number order).

[0031] As input to the model, the obtained document data needs to be preprocessed to improve the quality of the input image data. Preprocessing includes skew correction, image binarization, and noise removal. Skew correction is used to detect and correct the overall rotation angle of each image in the document data to ensure that the text lines are horizontal, which is a basic prerequisite for the accuracy of subsequent layout analysis. Image binarization can use adaptive thresholding algorithms (such as the Otsu method) to convert grayscale images into black and white binary images to enhance the contrast between the foreground (text, lines) and the background. Noise removal can apply techniques such as Gaussian blur or median filtering to remove random noise (such as salt and pepper noise) generated during the scanning process of PDFs and images (mainly scanned versions).

[0032] After preprocessing is complete, the input image data is obtained. At this point, the electronic device can proceed to step S30.

[0033] Step S30: Input the input image data into the domain fine-tuning DETR model, identify the logical region categories of the input image data through the domain fine-tuning DETR model, obtain the region category identification results and output them. The logical region categories include title region, author region, abstract region, body text region, illustration region, table region, formula region, footer region and reference region.

[0034] In this embodiment, the domain-fine-tuned DETR model employs a Transformer-based DETR model (DEtection TRansformer) to perform object detection and prediction on document page images, outputting the bounding box and category label for each predicted object. Considering the complexity of document layouts in the chemistry, biology, and pharmaceutical fields, after basic training on large-scale public document layout datasets such as PubLayNet and DocLayNet (containing millions of annotated pages), fine-tuning is performed on an internally constructed proprietary dataset containing tens of thousands of journal articles in the chemistry, biology, and pharmaceutical fields. Through this fine-tuning process, the model can identify the following key logical region categories with extremely high accuracy: title, authors, abstract, text body, figure, table, equation, footer, and references. Of course, if it is necessary to expand the document types, such as to expand to the identification of logical region categories of patents, it is necessary to add a dataset formed by patent annotation data in related fields for domain fine-tuning to identify the corresponding logical region categories (such as basic patent information, name, abstract, abstract drawings, claims, technical field, background technology, invention content, description of drawings, detailed embodiments, specification drawings, etc.). After fine-tuning and training, the domain-fine-tuned DETR model can be obtained.

[0035] Accordingly, electronic devices can input image data into a neighborhood fine-tuning DETR model, identify the logical region category of the input image data through the neighborhood fine-tuning DETR model, obtain the region category identification result, and output it.

[0036] Taking journal and academic paper input image data as an example, the domain-fine-tuned DETR model can identify and output the logical region categories of the input image data. These logical region categories include title region, author region, abstract region, body text region, figure region, table region, formula region, footer region, and reference region. The output of this step is a set of bounding boxes with category labels, dividing the document page into several functional regions for subsequent differentiated information extraction.

[0037] After obtaining the region category identification result, the electronic device can run step S40.

[0038] Step S40: Extract differential information from the identification results of each region category to obtain the information corresponding to each region category identification result.

[0039] In this embodiment, the electronic device can extract differentiated information from the identification results of each region category to obtain the information corresponding to each region category identification result.

[0040] For the title, author, abstract, footer, and reference areas: text recognition methods are used to identify the text information corresponding to each area. These areas are primarily text-based, with almost no complex content such as images or tables. Therefore, for these areas, the electronic device can use connected component analysis to identify all individual characters or symbols, then apply the K-nearest neighbor clustering algorithm to aggregate characters into words, and finally aggregate words into lines of text. This process not only accurately segments the text but, more importantly, reconstructs the correct reading order, which is crucial for subsequent natural language processing tasks requiring contextual understanding.

[0041] For the illustration area: the electronic device can use OCSR to identify information, determine the SMILES string corresponding to each chemical structure object in the illustration area, and perform post-processing verification to obtain the SMILES string information corresponding to the illustration area.

[0042] For example, electronic devices can use OCSR (Optical Character Recognition) to identify information and determine the SMILES string corresponding to each chemical structure object within the illustration area. Relying solely on the output of a deep learning model has a certain error rate, especially when dealing with novel or poorly drawn structures. The innovation of this embodiment lies in introducing a post-processing verification engine to perform a series of post-processing verifications.

[0043] For each chemical structure object, the electronic device can perform atomic valence state verification using the corresponding SMILES string. Atomic valence state verification primarily checks whether each atom in the molecule (especially C, N, O, S, etc.) conforms to its allowed valence state according to standard chemical rules. For example, a pentavalent carbon atom would be immediately marked as incorrect, and the verification would fail.

[0044] Specifically, the electronic device can determine the explicit valence state of each atom in the chemical structure object based on the number of bonds in the SMILES string. The explicit valence state is determined by the number of bonds in the SMILES (e.g., C(=O) represents a double bond). Then, the electronic device can convert the SMILES string into an RDKit molecular object (if the conversion fails, it can be directly determined that the verification failed), and iterate through the atoms in the RDKit molecular object to determine the implicit valence state of each atom in the chemical structure object. The implicit valence state is determined by the number and charge of hydrogen atoms inferred by RDKit.

[0045] Then, the actual valence state can be calculated based on the explicit and implicit valence states of each atom in the chemical structure object: Actual valence state = Explicit valence state + Implicit valence state. Specifically, the GetExplicitValence and GetImplicitValence functions of RDKit can be used to calculate the total valence state, avoiding the omission of the influence of implicit hydrogen or charge.

[0046] After calculating the actual valence state of each atom in the chemical structure object, the electronic device can obtain predefined valence state rules, including fixed value valence states and list value valence states. Fixed value valence states mean that the atom can only be a certain valence state (e.g., C: 4, the C atom must strictly conform to this valence state), while list value valence states mean that the atom needs to be any valence state in the list (e.g., N: [3,5], the valence state of the N atom can be any value in the list).

[0047] Therefore, electronic devices can verify the actual valence state of each atom in a chemical structure based on predefined valence state rules, thus determining the verification result. The main point is to check whether the actual valence state of each atom conforms to the predefined valence state rules. If the actual valence state of any atom does not conform to the predefined valence state rules, the verification of the atomic valence state of the current chemical structure fails. If the actual valence state of every atom in the chemical structure conforms to the predefined valence state rules, the verification of the atomic valence state of the current chemical structure passes.

[0048] By checking whether the sum of the explicit and implicit valence states (actual valence states) of each atom in a chemical structure object satisfies predefined chemical rules, the rationality of the molecular structure can be determined. Through predefined valence state rules, common chemical knowledge is encoded into the verification logic, which can be used to filter out invalid molecular structures.

[0049] The pseudocode for verifying atomic valence states is as follows: def check_atom_valency(smiles): mol = Chem.MolFromSmiles(smiles) if mol is None: return False, "Invalid SMILES" valency_rules = { 'C': 4, 'N': [3, 5], 'O': 2, 'S': [2, 4, 6], 'P': [3, 5], 'F': 1, 'Cl': 1, 'Br': 1, 'I': 1} for atom in mol.GetAtoms(): symbol = atom.GetSymbol() if symbol not in valency_rules: continue # Calculate the actual price state explicit_valence = atom.GetExplicitValence() implicit_valence = atom.GetImplicitValence() actual_valence = explicit_valence + implicit_valence # Check if it meets the rules allowed = validity_rules[symbol] if isinstance(allowed, int): if actual_valence != allowed: return False, f"{symbol} atomic valence error: expected {allowed}, actual {actual_valence}" else: # Multiple possible valence states in list form if actual_valence not in allowed: return False, f"{symbol} atomic valence error: expected {allowed} one, actual {actual_valence}" The function returns True, indicating that the price check has passed. Secondly, for each chemical structure object, the electronic device can perform aroma perception verification based on the SMILES string.

[0050] During the information extraction process, aromatic rings are sometimes mistakenly identified as a series of alternating single and double bonds. Aromaticity perception verification can be performed based on chemical principles such as Hückel's rule to check whether the SMILES string corresponding to each identified chemical structure object is accurate.

[0051] For each chemical structure object, the corresponding SMILES string is converted into an RDKit molecular object. Then, each ring system in the RDKit molecular object is examined to determine whether each ring system conforms to Hückel's rule. If the ring system conforms to Hückel's rule, it is marked as an aromatic ring; if the ring system does not conform to Hückel's rule, it is determined that the ring system is not an aromatic ring.

[0052] For example, the main conditions for a ring system to possess aromaticity are: 1. Conjugation: Atoms must form a conjugated system (bonds alternate between single / double bonds or aromatic bonds).

[0053] 2. It conforms to Hückel's rule: the total number of π electrons in a ring system must satisfy 4n+2.

[0054] 3. Atomic geometry: Ring systems must be planar (non-planar rings will not be labeled aromatic even if they satisfy the electron rule).

[0055] Therefore, the electronic device can detect each ring system in the RDKit molecular object and then verify these three conditions: the detection of conjugation and planar configuration is relatively simple, by determining whether the bonds in the ring system are alternating single / double bonds or aromatic bonds, to determine whether a conjugated system is formed between the atoms in the ring system. However, relying on RDKit's two-dimensional processing based on topological structure and electronic system, it defaults to assuming that ring systems that meet the rules are aromatic. Therefore, the key lies in the judgment of Hückel's rule, that is, whether the total number of π electrons in the ring system satisfies 4n+2.

[0056] Specifically, electronic devices need to calculate the π-electron contribution: traversing the ring system, the π-electron contribution of each atom is calculated separately (as shown in Table 1): Table 1. Contribution of π electrons Atom type hybrid state Number of connection keys π electron contribution Carbon atom (C) <![CDATA[sp 2 ]]> any +1 Neutral nitrogen (N) <![CDATA[sp 2 ]]> 2 keys +2 <![CDATA[sp 2 ]]> 3 keys +1 <![CDATA[Positively charged nitrogen (N + )]]> <![CDATA[sp 2 ]]> any 0 Neutral oxygen (O) <![CDATA[sp 2 ]]> 2 keys +2 <![CDATA[Positively charged sulfur (S + )]]> <![CDATA[sp 2 ]]> any +2 The total number of π electrons is calculated by summing them up (if there are special groups such as carbonyl (C=O) or imine (C=NR), the total number of π electrons needs to be corrected. In carbonyl, only the carbon atom provides 1 π electron, and in imine, the carbon atom and nitrogen atom each provide 1 π electron). The result is used to determine whether the 4n+2 condition is met, thereby determining whether the ring system has aromaticity.

[0057] After identifying the ring system with the labeled aromatic ring using RDKit, it can be compared with the aromatic ring represented in the SMILES string (where aromatic rings are represented by lowercase letters, such as c1ccccc1). If they match, the aromaticity sensing verification for this chemical structure object is considered successful. If they do not match, the aromaticity sensing verification for this chemical structure object is considered unsuccessful.

[0058] After completing the aroma perception verification, the electronic device can perform stereochemical verification for the SMILES string corresponding to each chemical structure object.

[0059] For example, for each chemical structure object corresponding to the SMILES string: the electronic device can convert the SMILES string into an RDKit molecular object, and then detect whether a chiral center exists in the RDKit molecular object.

[0060] If a chiral center exists: the electronic device needs to detect whether the number of neighbors of this chiral center is 4. If the number of neighbors of this chiral center is not 4, the chiral center is determined to be invalid. If the number of neighbors of this chiral center is 4, it is determined whether the neighbors of this chiral center are all different. If the chiral center has the same neighbors, the chiral center is determined to be invalid. In addition, it needs to detect the wedge key or dashed key of this chiral center. If the chiral center has multiple wedge keys or dashed keys pointing in the same direction, the chiral center is determined to be invalid.

[0061] If there are no invalid chiral centers, the electronic device can determine that the stereochemical verification corresponding to this chemical structure object is successful; otherwise, it can determine that the stereochemical verification corresponding to this chemical structure object is unsuccessful.

[0062] Additionally, it can be checked whether the double bond cis-trans isomers have been correctly labeled. If there are double bond cis-trans isomers that do not correspond to the label (e.g., double bond cis-trans isomers are not labeled, but double bond cis-trans isomers exist; or double bond cis-trans isomers are labeled, but double bond cis-trans isomers do not exist), it can also be determined that the stereochemical verification has failed.

[0063] Accordingly, electronic devices can perform post-processing verification of each chemical structure object to obtain corresponding atomic valence state verification results, aromaticity perception verification results, and stereochemical verification results.

[0064] Then, the electronic device can determine the recognition confidence level for each chemical structure object based on the verification results of the atomic valence state, the aromaticity perception verification results, and the stereochemistry verification results. If the verification results for the atomic valence state, aroma perception, and stereochemistry of this chemical structure are all passed, the recognition confidence level for this chemical structure can be determined to be 1. If at least one of the verification results for the atomic valence state, aroma perception, and stereochemistry of this chemical structure fails, the recognition confidence level for this chemical structure is determined to be x (for example, the recognition confidence level is 0.67 when one verification fails, 0.33 when two verifications fail, and 0 when three verifications fail), and x is less than the confidence threshold (the threshold is designed to be greater than the recognition confidence level when one verification fails, for example, 0.9).

[0065] Therefore, when the recognition confidence level of the chemical structure object reaches the confidence level threshold, the electronic device can determine that the recognized SMILES string is the SMILES string corresponding to the chemical structure object.

[0066] When the recognition confidence level of a chemical structure object fails to reach the confidence threshold, manual review or a feedback loop can be triggered. For the first recognition where the recognition confidence level of the chemical structure object fails to reach the confidence threshold and the recognition confidence level is 0.67 (i.e., only one post-processing verification fails), a feedback loop is triggered (the feedback loop means adjusting the image preprocessing parameters of the inset area before performing a second recognition); if the second recognition still fails to reach the confidence threshold for the recognition of the chemical structure object, manual review is triggered.

[0067] Finally, the SMILES strings corresponding to each chemical structure object in the illustration area are integrated to obtain the SMILES string information corresponding to the illustration area.

[0068] like Figure 2 As shown, Figure 2 A schematic diagram of the inset area is provided, and the extracted information is as follows: Ellipticine: C1=CC=C(C=C1)C2=Nc3ccccc3C(=O)c3cccc4c3N=C(c3ccccc32)N4 7-Hydroxyellipticine: C1=CC=C(O)C=C1C2=Nc3ccccc3C(=O)c3cccc4c3N=C(c3ccccc32)N4 9-Hydroxyellipticine: C1=CC=C(C=C1)C2=Nc3cc(O)ccc3C(=O)c3cccc4c3N=C(c3ccccc32)N4 13-I-hydroxyellipticine: C1=CC=C(C=C1)C2=Nc3ccccc3C(=O)c3cccc4c3N=C(c3ccccc32)N4.O Ellipticine N2 oxide: C1=CC=C(C=C1)C2=N[N+](=O)c3ccccc3C(=O)c3cccc4c3N=C(c3ccccc32)N4 N7-deoxyguanosine adduct: C1=CC=C(C=C1)C2=Nc3ccccc3C(=O)c3cccc4c3N=C(c3ccccc32)N4.N[C@H]1[C@@H](CO)[C@H](O)[C@@H](O)[C@H]1O "Please note that the SMILES values ​​for N7-deoxyguanosine adduct may need to be adjusted depending on the specific adduct structure." For the main text region: When there is no chemical structure image in the main text region, text recognition methods can be used to identify the text information corresponding to the main text region.

[0069] For example, since the main text area typically contains multiple paragraphs and columns, a blank area analysis algorithm is needed to identify the vertical and horizontal blank areas in the main text area to identify the column boundaries. Then, for each column in the main text area (i.e., the column section; scientific documents such as papers and journals are usually divided into two vertical columns): all independent characters or symbols are identified through connected component analysis, and then the K-nearest neighbor clustering algorithm is applied to aggregate the characters into words, and then the words into text lines. Finally, the text lines of each column in the main text area are integrated to obtain the text information corresponding to the main text area.

[0070] When chemical structure images are present within the text area, OCSR (Optical Character Recognition) is used to identify the information in the chemical structure image, and text recognition methods are used to recognize the text in the remaining parts of the text area to determine the comprehensive information corresponding to the text area. This comprehensive information includes text information and the string "SMILES". The method for identifying information in the chemical structure image is described in the previous section on image area processing; the method for text recognition in the remaining parts is also described in the previous section on text areas without chemical structure images. Generally, chemical structure images do not appear in the text area; therefore, the processing method for cases where chemical structure images are not present in the text area is usually the primary approach.

[0071] For table areas: Electronic devices can perform logical structure recognition on table areas, identify the table skeleton containing several cells, and then perform information recognition on the data objects in each cell to obtain comprehensive information corresponding to the table area. When the data object in the cell is a chemical structure image, OCSR is used to perform information recognition on the data object in the cell to obtain the SMILES string information corresponding to the cell. When the data object in the cell is text data, text recognition methods are used to perform text recognition to determine the text information corresponding to the cell.

[0072] Tables in scientific literature are a treasure trove of structured information, but their layouts are complex and varied, including merged cells, multi-layered nested headers, and unclear boundary lines, which pose a huge challenge to data extraction.

[0073] For table areas, this embodiment first understands the "skeleton" of the table, rather than attempting to read the text. A structure recognition model based on graph neural networks or Transformers can be used to interpret the table image (e.g., ...). Figure 3 As shown in the figure, the output is the detection results of all logical components of the table, including the precise bounding box of each row, each column and each individual cell. The model can also identify and label merged cells (i.e., a cell that spans multiple rows or columns), thereby accurately identifying the table skeleton containing several cells corresponding to the table area.

[0074] Then, after the table skeleton is parsed and recognized, the bounding boxes of each recognized cell are traversed. Then, an OCR engine (such as Tesseract) fine-tuned for scientific literature characters (including Greek letters, superscripts and subscripts) is used to recognize the image area of ​​each cell to extract its text content. In addition, OCR can also be used to recognize information from chemical structure images within cells (see the previous text for the recognition process).

[0075] Ultimately, the information recognition results within each cell can be integrated to obtain comprehensive information corresponding to the table area.

[0076] like Figure 3 As shown, Figure 3 A schematic diagram of a table area is provided, and the extracted information is as follows (the table layout is omitted): CYP107W1 - Streptomyces avermitilis Biological Process and Enzymatic Reaction: Oligomycin biosynthesis C12-hydroxylation reaction of oligomycin C to form oligomycin A The Biological Significance of the Product: Antibiotic Reference:

[35] SMILES String (Oligomycin C): CC[C@H]1C[C@@H](O)[C@H](O)C(=O)[C@H]([C@H]2OC[C@H](O)[C@H](O)[C@H]2O)[C@@H]1O SMILES String (Oligomycin A): CC[C@H]1C[C@@H](O)[C@H](O)C(=O)[C@H]([C@H]2OC[C@H](O)[C@H](O)[C@H]2O)[C@@H]1O / C=C\C=C\CCC CYP107FH5 (CYP TamI) - Streptomyces sp. 307-9 Biological Process and Enzymatic Reaction: Tirandamycin biosynthesis C10 hydroxylation, oxidative conversion of C10 hydroxyl to carbonyl,C11 / C12 epoxidation, C18 hydroxylation The Biological Significance of the Product: Antibiotic Reference:

[36] SMILES String (Tirandamycin precursor): CC[C@H]1C[C@@H](O)[C@H](O)C(=O)[C@H]([C@H]2OC[C@H](O)[C@H](O)[C@H]2O)[C@@H]1O SMILES String (Tirandamycin product): CC[C@H]1C[C@@H](O)[C@H](O)C(=O)[C@H]([C@H]2OC[C@H](O)[C@H](O)[C@H]2O)[C@@H]1O / C=C\C=C\CCC CYP107Z14 (P450-sb21) - Sebekia benihana Biological Process and Enzymatic Reaction: Cyclosporine A pathway Hydroxylating at the 4th N-methyl leucine (MeLeu4) The Biological Significance of the Product: Immunosuppressant Reference:

[37] SMILES String (Cyclosporine A precursor): CC[C@H]1C[C@@H](O)[C@H](O)C(=O)[C@H]([C@H]2OC[C@H](O)[C@H](O)[C@H]2O)[C@@H]1O SMILES String (Cyclosporine A product): CC[C@H]1C[C@@H](O)[C@H](O)C(=O)[C@H]([C@H]2OC[C@H](O)[C@H](O)[C@H]2O)[C@@H]1O / C=C\C=C\CCC CYP107X1 - Streptomyces avermitilis Biological Process and Enzymatic Reaction: Progesterone pathway The 16α-Hydroxylation of progesterone The Biological Significance of the Product: Steroid Reference:

[38] SMILES String (Progesterone): CC[C@H]1C[C@@H](O)[C@H](O)C(=O)[C@H]([C@H]2OC[C@H](O)[C@H](O)[C@H]2O)[C@@H]1O SMILES String (16α-Hydroxyprogesterone): CC[C@H]1C[C@@H](O)[C@H](O)C(=O)[C@H]([C@H]2OC[C@H](O)[C@H](O)[C@H]2O)[C@@H]1O / C=C\C=C\CCC In this embodiment, for information recognition of multimodal literature data, the information extracted from each region after its structured decomposition can generate highly structured data objects (e.g., data in JSON format), thereby realizing information recognition and structured processing of multimodal literature data.

[0077] This application provides a storage medium disposed within an electronic device, comprising a stored program, wherein the program, when running, controls the electronic device containing the storage medium to execute the structured decomposition and information recognition method for multimodal data of this embodiment.

[0078] Furthermore, this application embodiment also provides an electronic device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, the steps of the structured decomposition and information recognition method for multimodal data of this embodiment are implemented.

[0079] In summary, this application provides a method, medium, and device for structured decomposition and information recognition of multimodal data. It acquires multimodal literature data (documents or images, including Word and PDF documents); converts the multimodal literature data into image format and preprocesses it to obtain input image data; uses a domain-adjusted DETR model (domain: chemistry, biology, and pharmaceuticals) to identify logical region categories (title region, author region, abstract region, text region, illustration region, table region, formula region, footer region, and reference region) of the input image data; obtains and outputs the region category recognition results; and then extracts differentiated information from each region category recognition result (for example, for the standard...). The text is divided into several sections: title area, author area, abstract area, footer area, and reference area. Text recognition methods are used to identify the text information corresponding to each area. For illustration areas, OCSR is used to identify the SMILES strings corresponding to each chemical structure object within the illustration area, and post-processing verification is performed to obtain the SMILES string information for the illustration area. For the main text area, a combination of text recognition methods and OCSR is used to obtain comprehensive information. For table areas, logical structure recognition is performed to identify the table skeleton containing several cells, and then information recognition is performed on the data objects within each cell to obtain the comprehensive information corresponding to the table area. This approach allows for precise segmentation of complex scientific literature layouts, employing differentiated recognition schemes for different areas. It also enables accurate recognition of chemical structure images in the fields of chemistry, biology, and pharmaceuticals, and post-processing verification of the SMILES strings from image recognition, thereby achieving accurate identification of information in multimodal data (especially scientific literature).

[0080] Due to the complexity of chemical formulas, to ensure accuracy in identification, the process of identifying chemical formula structure images (mainly in image and table areas, but some simple chemical formula structure images may also exist in the text area) involves performing atomic valence state verification, aromaticity perception verification, and stereochemistry verification after identifying the SMILES string using OCSR tools. Based on the results of these verifications, the identification confidence level for each chemical structure object is determined. When necessary (when the identification confidence level does not reach the confidence threshold), manual review or feedback loops are triggered. By performing multi-level verification, the accuracy of identifying complex chemical structures can be improved, ensuring the quality of information identification.

[0081] Atomic valence verification determines the explicit valence state of each atom by counting the bonds in the SMILES string. After converting the SMILES string into an RDKit molecular object, the implicit valence state of each atom is determined. The actual valence state (actual valence state = explicit valence state + implicit valence state) is then calculated. Combined with predefined valence state rules, this verifies the actual valence state of each atom in the chemical structure object, yielding the atomic valence state verification result. Aromaticity sensing verification uses the RDKit molecular object converted from the SMILES string to check if each ring system conforms to Hückel's rule, determining if the ring system is an aromatic ring. It then checks if the ring system marking the aromatic ring matches the aromatic ring represented in the SMILES string to determine if the aromaticity sensing verification for this chemical structure object passes. For stereochemistry verification, it checks if a chiral center exists in the RDKit molecular object converted from the SMILES string. It determines if the chiral center is invalid by checking if the number of its neighbors is 4, whether the neighbors are all different, and whether multiple wedge or dashed bonds point in the same direction, thus obtaining the stereochemistry verification result. Based on the characteristics of complex chemical structures, this solution employs multi-level post-processing verification. After the SMILES string is identified using OCSR tools, verification can be performed to determine whether the identification results of the chemical structure conform to the objective laws of the chemical field, thereby improving the accuracy of information identification.

[0082] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.

[0083] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for structured decomposition and information recognition of multimodal data, characterized in that, include: Acquire the multimodal literature data to be processed, where the multimodal literature data is either a document or an image, and the document types include Word documents and PDF documents; The multimodal literature data to be processed is transformed into image data and preprocessed to obtain input image data. The input image data is fed into the domain fine-tuning DETR model. The domain fine-tuning DETR model identifies the logical region categories of the input image data, and the region category identification results are obtained and output. The logical region categories include title region, author region, abstract region, body text region, illustration region, table region, formula region, footer region, and reference region. Differential information is extracted from the identification results of each region category to obtain the information corresponding to each region category identification result.

2. The method for structured decomposition and information recognition of multimodal data according to claim 1, characterized in that, The multimodal literature data belongs to the fields of chemistry, biology, and pharmaceuticals. Differential information is extracted from the category identification results for each region to obtain the information corresponding to each region category identification result, including: For the title area, author area, abstract area, footer area, and reference area: text recognition methods are used to identify the text information corresponding to each area; For the illustration area: OCSR is used for information recognition to determine the SMILES string corresponding to each chemical structure object in the illustration area, and post-processing verification is performed to obtain the SMILES string information corresponding to the illustration area; For the main text area: when there is no chemical structure image in the main text area, text recognition method is used to identify the text information corresponding to the main text area; when there is a chemical structure image in the main text area, OCSR is used to identify the information of the chemical structure image, and text recognition method is used to identify the text of the remaining part in the main text area to determine the comprehensive information corresponding to the main text area. The comprehensive information includes text information and SMILES string information. For table areas: Logical structure recognition is performed on the table area to identify the table skeleton containing several cells. Then, information recognition is performed on the data objects in each cell to obtain the comprehensive information corresponding to the table area. When the data object in the cell is a chemical structure image, OCSR is used to recognize the information of the data object in the cell to obtain the SMILES string information corresponding to the cell. When the data object in the cell is text data, text recognition methods are used to recognize the text to determine the text information corresponding to the cell.

3. The method for structured decomposition and information recognition of multimodal data according to claim 2, characterized in that, OCSR was used for information recognition to determine the SMILES string corresponding to each chemical structure object within the illustration area. Post-processing verification was then performed to obtain the SMILES string information corresponding to the illustration area, including: OCSR was used for information recognition to determine the SMILES string corresponding to each chemical structure object within the illustration area; For each chemical structure object, the atomic valence state is verified using the SMILES string. For each chemical structure object, perform aroma perception verification on the corresponding SMILES string; Stereochemical verification is performed on the SMILES string corresponding to each chemical structure object; Based on the verification results of atomic valence state, aroma perception, and stereochemistry for each chemical structure object, the recognition confidence level for each chemical structure object is determined. When the recognition confidence level for a chemical structure object reaches the confidence level threshold, the identified SMILES string is determined to be the SMILES string corresponding to that chemical structure object. When the recognition confidence level for a chemical structure object does not reach the confidence level threshold, manual review or a feedback loop is triggered. The feedback loop means that a second recognition is performed after adjusting the image preprocessing parameters of the illustration area. If the second recognition still fails to reach the confidence level threshold for the recognition of the chemical structure object, manual review is triggered.

4. The method for structured decomposition and information recognition of multimodal data according to claim 3, characterized in that, For each chemical structure object, the corresponding SMILES string is used to verify atomic valence states, including: For each chemical structure object, the corresponding SMILES string is: Based on the number of bonds in the SMILES string, determine the explicit valence state of each atom in the chemical structure object; The SMILES string is converted into an RDKit molecule object, and the atoms in the RDKit molecule object are traversed to determine the implicit valence state of each atom in the chemical structure object. The actual valence state is calculated based on the explicit and implicit valence states of each atom in the chemical structure object: actual valence state = explicit valence state + implicit valence state; Obtain predefined valence rules, which include fixed-value valence rules and list-value valence rules. Fixed-value valence rules mean that an atom can only be one agreed-upon valence state, while list-value valence rules mean that an atom can be any valence state in the list. The actual valence state of each atom in the chemical structure object is verified based on predefined valence state rules, and the atomic valence state verification results of the chemical structure object are determined.

5. The method for structured decomposition and information recognition of multimodal data according to claim 3, characterized in that, For each chemical structure object, perform aroma perception verification on the corresponding SMILES string, including: For each chemical structure object, the corresponding SMILES string is: Convert the SMILES string to an RDKit molecule object; Detect each ring system in the RDKit molecular object to determine whether each ring system conforms to Hückel's rule; If a ring system conforms to Hückel's rule, it is labeled as an aromatic ring; if a ring system does not conform to Hückel's rule, it is determined that the ring system is not an aromatic ring. Determine whether the ring system of the labeled aromatic ring is consistent with the aromatic ring represented in the SMILES string. If they are consistent, the aromaticity perception verification corresponding to this chemical structure object is confirmed to be successful.

6. The method for structured decomposition and information recognition of multimodal data according to claim 3, characterized in that, For each chemical structure object, perform stereochemical verification of the corresponding SMILES string, including: For each chemical structure object, the corresponding SMILES string is: Convert the SMILES string to an RDKit molecule object; Detecting the presence of chiral centers in RDKit molecular objects; If a chiral center exists: Check if the number of neighbors of this chiral center is 4. If the number of neighbors of this chiral center is not 4, determine that this chiral center is invalid. If the number of neighbors of this chiral center is 4, determine whether the neighbors of this chiral center are all different. If the chiral center has the same neighbors, determine that this chiral center is invalid. Detect the wedge-shaped or dashed bonds in this chiral center. If multiple wedge-shaped or dashed bonds in this chiral center point in the same direction, then this chiral center is invalid. If no invalid chiral center exists, the stereochemical verification corresponding to this chemical structure object is confirmed to be successful.

7. The method for structured decomposition and information recognition of multimodal data according to claim 3, characterized in that, Based on the verification results of atomic valence state, aromaticity perception, and stereochemistry for each chemical structure object, the recognition confidence level for each chemical structure object is determined, including: For each chemical structure object: If the verification results of atomic valence state, aromaticity perception and stereochemistry of this chemical structure are all passed, the recognition confidence level of this chemical structure is determined to be 1. If at least one of the verification results for the atomic valence state, the aromaticity perception, and the stereochemistry of this chemical structure fails, the recognition confidence level for this chemical structure is determined to be x, where x is less than the confidence threshold.

8. The method for structured decomposition and information recognition of multimodal data according to claim 2, characterized in that, Text recognition methods are used to identify the text information corresponding to the main text area, including: Use blank area analysis algorithms to identify vertical and horizontal blank areas in the text area in order to identify column boundaries in the text area; For each column in the main text area: all independent characters or symbols are identified through connected component analysis, and then the K-nearest neighbor clustering algorithm is applied to aggregate characters into words, and then words into text lines; By integrating the text lines of each column in the main text area, the corresponding text information of the main text area can be obtained.

9. A storage medium, characterized in that, The storage medium is disposed within an electronic device and includes a stored program, wherein, when the program is executed, it controls the electronic device containing the storage medium to perform the structured decomposition and information recognition method for multimodal data as described in any one of claims 1 to 8.

10. An electronic device comprising a memory and a processor, the memory for storing information including program instructions, and the processor for controlling the execution of the program instructions, characterized in that: When the program instructions are loaded and executed by the processor, they implement the steps of the structured decomposition and information recognition method for multimodal data as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Chemical structure identification method based on visual transformer

    CN121884069A