Drilling engineering document identification method and device based on multi-modal information fusion
By using a multimodal information fusion method, text, image, and structural features in drilling engineering documents are extracted and corrected to generate structured output, which solves the problem of low information utilization in drilling documents and improves recognition accuracy and automated processing efficiency.
Patent Information
- Application Number
- CN202511013207.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-14
AI Technical Summary
In existing technologies, the multimodal data of drilling engineering documents has not been fully explored, resulting in low information utilization, poor retrieval efficiency, and the inability to achieve cross-modal data correlation analysis. Traditional optical character recognition and image recognition algorithms cannot understand professional symbols and contextual information.
By using a multimodal information fusion method, text, image, and structural modal features are extracted, and multi-task recognition, error correction, and semantic association processing are performed to generate structured output results, including text sequences, graphic categories, tabular data, and the position coordinates of symbol labels.
It improved the accuracy of semantic understanding of complex graphics in drilling documents, reduced manual data entry costs, and improved the efficiency of automated document processing and the accuracy of engineering decisions.
Smart Images

Figure CN120954013A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of engineering document processing, specifically an image recognition method and apparatus for drilling engineering documents based on multimodal information fusion. Background Technology
[0002] In the field of drilling engineering, technical documents (such as drilling design reports, well logging curves, and geological profiles) typically contain a large amount of images, charts, and text information. This multimodal data collectively records key information such as drilling parameters, formation characteristics, and equipment status. However, traditional document management methods mainly rely on manual review and text retrieval, making it difficult to efficiently extract and utilize unstructured information from images. This results in low information utilization, poor retrieval efficiency, and an inability to perform cross-modal data correlation analysis.
[0003] The relevant existing technologies include: [1] Li Hongliang; Liu Yuliang; Liao Wenhui; Huang Mingxin; Zhang Shuo; Jin Lianwen. Optical character recognition in the era of large models: current status and prospect [J]. Journal of Image and Graphics, 2025, 30(06): 2023-2050; [2] Lin Botao; Zhu Haitao; Jin Yan; Zhang Jiahao; Han Xueyin. Construction method and application case of digital twin model of oil and gas drilling and production [J]. Petroleum Science Bulletin, 2024, 9(02): 282-296; [3] Zhang Feifei; Wang Qian; Wang Xueying; Yu Yibing; Lou Wenqiang; Peng Fengjia. Multi-source multimodal data fusion technology and prospect of oil and gas well engineering [J]. Natural Gas Industry, 2024, 44(09): 152-166.
[0004] At least the following defects exist in the existing technology: recognition technology based on a single modality (such as plain text or image) has obvious limitations. For example, in reference [1], OCR (Optical Character Recognition) technology can extract text in a document, but cannot understand the engineering semantics in an image; while in reference [2], although conventional image recognition algorithms can detect graphic elements, they lack the ability to analyze professional symbols, charts and contextual information in the drilling field. In addition, in reference [3], images and text in drilling documents often complement each other (such as well structure diagrams with annotation text), and existing methods have not fully explored the correlation between multimodal data, resulting in the loss or misjudgment of key information.
[0005] Therefore, there is an urgent need for an image recognition method that integrates multimodal information. By combining visual features, textual semantics, and domain knowledge, it can achieve intelligent recognition and structured extraction of image content in drilling engineering documents, thereby improving the efficiency of automated document processing and the accuracy of engineering decisions.
[0006] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section. Summary of the Invention
[0007] To address the problems in the existing technology, this application provides a drilling engineering document recognition method and apparatus based on multimodal information fusion. By fusing multimodal information such as text, images, and domain knowledge, it can enhance the semantic understanding of complex graphics in drilling documents and improve the recognition accuracy of professional symbols, charts, and annotation information.
[0008] To solve the above-mentioned technical problems, this application provides the following technical solution:
[0009] Firstly, this application provides a drilling engineering document recognition method based on multimodal information fusion, including:
[0010] Multimodal feature extraction was performed on the preprocessed original drilling document images to obtain text modal features, image modal features, structural modal features, and semantic modal features;
[0011] Multimodal information fusion is performed on the text modal features, image modal features, structural modal features and semantic modal features to obtain high-level fusion features and preliminary error correction identifiers;
[0012] Multi-task recognition and error correction are performed based on the high-level fusion features and preliminary error correction identifiers to obtain the final recognition result; wherein, the final recognition result includes text sequence, graphic category, tabular data, symbol label and corresponding position coordinates;
[0013] The final identification result is subjected to semantic association processing to obtain a structured output result.
[0014] Further, the step of preprocessing the original drilling document image includes:
[0015] The original drilling document image is enhanced to obtain an enhanced drilling document image;
[0016] Region detection is performed on the enhanced image of the drilling document to obtain the corresponding document layout information; wherein, the document layout information includes content categories and corresponding location coordinates; the content categories include text, graphics, tables and symbols;
[0017] The drilling document augmented image is segmented into independent regions based on the document layout information. The independent regions include a set of text blocks, a set of graphic blocks, a set of table blocks, and a set of symbol blocks.
[0018] Furthermore, the multimodal feature extraction of the preprocessed original drilling document image yields text modal features, image modal features, structural modal features, and semantic modal features, including:
[0019] Visual encoding extraction is performed on the set of text blocks to obtain the text modal features;
[0020] Visual label recognition is performed on the set of graphic blocks to obtain the image modal features;
[0021] Graph structures are constructed from the independent region blocks to obtain the structural modal features;
[0022] The semantic modality features are obtained by performing semantic matching on the independent region blocks.
[0023] Furthermore, the multimodal information fusion of the text modal features, image modal features, structural modal features, and semantic modal features to obtain high-level fusion features and preliminary error correction identifiers includes:
[0024] The text modal features, image modal features, structural modal features, and semantic modal features are concatenated at the lowest level to obtain a low-level fusion vector;
[0025] Construct a text-graphic attention model and a table symbol attention model based on the independent region blocks;
[0026] Generate a text-graphic association vector based on the text block set, the graphic block set, and the text-graphic attention model;
[0027] Generate a table symbol association vector based on the table block set, the symbol block set, and the table symbol attention model;
[0028] The text-graphic association vector, the table symbol association vector, and the bottom-level fusion vector are subjected to residual connection processing to obtain the middle-level fusion feature;
[0029] The mid-layer fusion features are validated using a rule engine to obtain the preliminary error correction identifier;
[0030] The high-level fusion features are obtained by performing statistical model correction on the mid-level fusion features.
[0031] Furthermore, the step of performing multi-task recognition and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain the final recognition result includes:
[0032] The constructed joint loss function is used to perform multi-task learning on the high-level fusion features to obtain multi-task learning results;
[0033] Cross-modal consistency verification is performed on the multi-task learning results to obtain the final recognition result and secondary error correction identifier.
[0034] Furthermore, the structured output results include structured data files and document-level semantic graphs; the semantic association processing of the final recognition results to obtain the structured output results includes:
[0035] The final recognition results are then structured to obtain text structure organization results, tabular data organization results, graphic data organization results, and symbol annotation organization results.
[0036] Based on the results of text structure analysis, tabular data analysis, graphical data analysis, and symbol annotation analysis, semantic association modeling is performed to obtain structured data files and document-level semantic graphs.
[0037] Secondly, this application provides a drilling engineering document recognition device based on multimodal information fusion, comprising:
[0038] The modal feature extraction unit is used to extract multimodal features from the preprocessed original drilling document images to obtain text modal features, image modal features, structural modal features, and semantic modal features;
[0039] The modal information fusion unit is used to perform multimodal information fusion on the text modal features, image modal features, structural modal features and semantic modal features to obtain high-level fusion features and preliminary error correction identifiers;
[0040] The final identification unit is used to perform multi-task identification and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain the final identification result; wherein, the final identification result includes text sequence, graphic category, tabular data, symbol label and corresponding position coordinates;
[0041] The structure output unit is used to perform semantic association processing on the final recognition result to obtain a structured output result.
[0042] Furthermore, the modal feature extraction unit includes:
[0043] The document enhancement module is used to enhance the original drilling document image to obtain an enhanced drilling document image;
[0044] The document layout module is used to perform region detection on the enhanced image of the drilling document to obtain corresponding document layout information; wherein, the document layout information includes content categories and corresponding position coordinates; the content categories include text, graphics, tables and symbols;
[0045] The data segmentation module is used to perform multimodal data segmentation on the enhanced image of the drilling document based on the document layout information to obtain independent region blocks; wherein, the independent region blocks include a set of text blocks, a set of graphic blocks, a set of table blocks, and a set of symbol blocks.
[0046] Furthermore, the modal feature extraction unit includes:
[0047] The text modality extraction module is used to perform visual encoding extraction on the set of text blocks to obtain the text modality features;
[0048] The image modality extraction module is used to perform visual label recognition on the set of graphic blocks to obtain the image modality features;
[0049] The structural modality extraction module is used to construct a graph structure for the independent region blocks to obtain the structural modality features;
[0050] The semantic modality extraction module is used to perform semantic matching processing on the independent region blocks to obtain the semantic modality features.
[0051] Furthermore, the modal information fusion unit includes:
[0052] The low-level fusion module is used to concatenate the low-level features of the text modal features, image modal features, structural modal features and semantic modal features to obtain the low-level fusion vector;
[0053] The attention model building module is used to build text graphics attention models and table symbol attention models based on the independent region blocks;
[0054] The text-image association vector generation module is used to generate text-image association vectors based on the text block set, the image block set, and the text-image attention model.
[0055] The table symbol association vector generation module is used to generate a table symbol association vector based on the table block set, the symbol block set, and the table symbol attention model.
[0056] The middle-layer fusion module is used to perform residual connection processing on the text-graphic association vector, the table symbol association vector and the bottom-layer fusion vector to obtain the middle-layer fusion feature;
[0057] The preliminary error correction module is used to perform rule engine verification on the mid-layer fusion features to obtain the preliminary error correction identifier;
[0058] The high-level fusion module is used to perform statistical model error correction on the mid-level fusion features to obtain the high-level fusion features.
[0059] Furthermore, the final identification unit includes:
[0060] The task learning module is used to perform multi-task learning on the high-level fusion features using the constructed joint loss function to obtain multi-task learning results.
[0061] The identification and error correction module is used to perform cross-modal consistency verification on the multi-task learning results to obtain the final identification result and secondary error correction identifier.
[0062] Furthermore, the structured output results include structured data files and document-level semantic graphs; the structured output unit includes:
[0063] The structure sorting module is used to sort the final recognition results into structured forms, resulting in text structure sorting results, tabular data sorting results, graphic data sorting results, and symbol annotation sorting results.
[0064] The semantic association module is used to perform semantic association modeling based on the results of text structure sorting, table data sorting, graphic data sorting, and symbol annotation sorting, to obtain structured data files and document-level semantic graphs.
[0065] Thirdly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the drilling engineering document recognition method based on multimodal information fusion.
[0066] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the drilling engineering document recognition method based on multimodal information fusion.
[0067] Fifthly, this application provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the drilling engineering document recognition method based on multimodal information fusion.
[0068] To address the problems in existing technologies, this application provides a drilling engineering document recognition method and apparatus based on multimodal information fusion. This method and apparatus can solve the problem of misrecognition of technical terms and complex symbols in traditional optical symbol recognition technology through multimodal information complementarity, improving character recognition accuracy on drilling document test sets. It utilizes structured parsing capabilities to achieve semantic association parsing of text, images, and tables in documents, outputting structured results that can be directly used for engineering data analysis, reducing manual data entry costs. Through domain knowledge base and customized model training, it can quickly adapt to different types of drilling engineering documents, including drilling reports and well completion reports. Attached Figure Description
[0069] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 This is a flowchart of the drilling engineering document recognition method based on multimodal information fusion in the embodiments of this application;
[0071] Figure 2 This is a flowchart illustrating the preprocessing of raw drilling document images in an embodiment of this application;
[0072] Figure 3 This is a flowchart illustrating the multimodal feature extraction process in this application embodiment;
[0073] Figure 4 This is a flowchart illustrating the multimodal information fusion process in this application embodiment;
[0074] Figure 5 This is a flowchart illustrating the multi-task identification and error correction process in this application embodiment;
[0075] Figure 6 This is a flowchart illustrating the semantic association processing in the embodiments of this application;
[0076] Figure 7 This is a structural diagram of the drilling engineering document recognition device based on multimodal information fusion in the embodiments of this application;
[0077] Figure 8 This is a structural diagram of the modal feature extraction unit in the embodiments of this application;
[0078] Figure 9 This is a structural diagram of the modal feature extraction unit in the embodiments of this application;
[0079] Figure 10 This is a structural diagram of the modal information fusion unit in the embodiments of this application;
[0080] Figure 11 This is a structural diagram of the final identification unit in the embodiments of this application;
[0081] Figure 12 This is a structural diagram of the output unit in the embodiments of this application;
[0082] Figure 13 This is a schematic diagram of the structure of the electronic device in the embodiments of this application;
[0083] Figure 14 This is an overall flowchart of the embodiments of this application. Detailed Implementation
[0084] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0085] The information collected in the technical solution of this application is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding operation portals are provided for users to choose to authorize or refuse.
[0086] Provide users with corresponding operation entry points, allowing them to choose to agree to or reject the automated decision results; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0087] In one embodiment, see Figure 1 To enhance semantic understanding of complex graphics in drilling documents and improve the accuracy of recognizing technical symbols, charts, and annotations by fusing multimodal information such as text, images, and domain knowledge, this application provides a drilling engineering document recognition method based on multimodal information fusion, including:
[0088] S101: Perform multimodal feature extraction on the preprocessed original drilling document images to obtain text modal features, image modal features, structural modal features, and semantic modal features;
[0089] S102: Perform multimodal information fusion on the text modal features, image modal features, structural modal features and semantic modal features to obtain high-level fusion features and preliminary error correction identifiers;
[0090] S103: Perform multi-task recognition and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain the final recognition result (also known as the document recognition result); wherein, the final recognition result includes text sequence, graphic category, tabular data, symbol label and corresponding position coordinates;
[0091] S104: Perform semantic association processing on the final recognition result to obtain a structured output result.
[0092] It is understood that the method provided in this application sequentially performs preprocessing on the original drilling document image, multimodal feature extraction, multimodal information fusion, multi-task recognition and error correction, and semantic association processing to finally obtain a structured output result.
[0093] In the preprocessing stage, the original drilling document images can be color scans of the corresponding documents. Image tilt correction is performed, table areas are segmented into rows and columns, and coordinate axes and data point contours are extracted from the graph areas. In the feature extraction stage, text modality, image modality, structural modality, and semantic modality features are extracted from the preprocessed original drilling document images. In the fusion and recognition stage, the results of the previous processing are sequentially fused at the bottom layer, fused at the middle layer, and validated at the top layer. In the semantic association processing and structured output stage, the recognition results are transformed into structured data. Specifically, the table data is stored as a two-dimensional matrix, and key coordinate points of the graph data are extracted and associated with the measured values in the table, ultimately generating a JSON file that can be imported into an engineering database.
[0094] As described above, the drilling engineering document recognition method based on multimodal information fusion provided in this application can solve the problem of misrecognition of professional terms and complex symbols by traditional optical symbol recognition technology through the complementarity of multimodal information, thereby improving the character recognition accuracy on the drilling document test set; it can realize the semantic association parsing of text, images and tables in the document by using structured parsing capabilities, and output structured results that can be directly used for engineering data analysis, reducing manual data entry costs; through domain knowledge base and customized model training, it can quickly adapt to different types of drilling engineering documents, including drilling reports and well completion reports.
[0095] In one embodiment, see Figure 2 The step of preprocessing the original drilling document image includes:
[0096] S201: The original drilling document image is enhanced to obtain an enhanced drilling document image;
[0097] S202: Perform region detection on the enhanced image of the drilling document to obtain the corresponding document layout information; wherein, the document layout information includes content categories and corresponding location coordinates; the content categories include text, graphics, tables, and symbols;
[0098] S203: Perform multimodal data segmentation on the enhanced image of the drilling document according to the document layout information to obtain independent region blocks; wherein, the independent region blocks include a set of text blocks, a set of graphic blocks, a set of table blocks, and a set of symbol blocks.
[0099] It is understandable that steps S201 to S203 can be understood as the process of multimodal data preprocessing.
[0100] Input: Raw drilling document image (color / grayscale, resolution ≥150dpi).
[0101] (1) Image enhancement
[0102] A 3×3 median filter is used to remove salt-and-pepper noise, and the denoised image I is output. denoised ;
[0103] I is analyzed using the Otsu algorithm. denoised Perform binarization processing to generate a black and white binary image I. bin ;
[0104] For I bin Based on the Hough transform, the baseline of the text lines is detected, and then an affine transform is used to output the tilt-corrected image I. corrected (Angle error ≤ 0.5°).
[0105] (2) Document layout analysis
[0106] Using the PaddleLayout model to analyze I corrected Perform region detection and output a set of four types of bounding boxes (ROIs) containing coordinate information. i |i=1,2,...,N}. Each ROI is labeled with its category (text / graphic / table / symbol) and coordinates (x1,y1,x2,y2).
[0107] (3) Multimodal data segmentation
[0108] Based on the layout detection results, from I corrected Cut out independent regions from the middle:
[0109] The set of text blocks ROI-T = {T1,T2,...,T} M (Preserve original resolution, 300dpi);
[0110] The set of graphic blocks ROI-G = {G1, G2, ..., G...} K (Scaling to 224x224 while maintaining aspect ratio);
[0111] The set of table blocks ROI-Tab = {Tab1, Tab2, ..., Tab...} L (Preserve original dimensions, add row and column line detection results);
[0112] The symbol block set ROI-S = {S1,S2,...,S} P (Crop to 32×32 pixels, center positioned).
[0113] Output: A set of multimodal region blocks {ROI-T, ROI-G, ROI-Tab, ROI-S} and a global coordinate mapping table (used to record the position of each block in the original document).
[0114] As can be seen from the above description, the drilling engineering document recognition method based on multimodal information fusion provided in this application can preprocess the original drilling document image.
[0115] In one embodiment, see Figure 3 The step of extracting multimodal features from the preprocessed original drilling document images to obtain text modal features, image modal features, structural modal features, and semantic modal features includes:
[0116] S301: Visual encoding extraction is performed on the set of text blocks to obtain the text modal features;
[0117] S302: Perform visual label recognition on the set of graphic blocks to obtain the image modal features;
[0118] S303: Construct a graph structure for the independent region blocks to obtain the structural modal features;
[0119] S304: Perform semantic matching processing on the independent region blocks to obtain the semantic modality features.
[0120] It is understandable that steps S301 to S304 can be understood as the process of multimodal feature extraction.
[0121] Input: {ROI-T,ROI-G,ROI-Tab,ROI-S} and the global coordinate mapping table generated in the previous steps.
[0122] Processing procedure:
[0123] (1) Text Modal Feature Extraction
[0124] Each text block T i Input an improved CRNN+Transformer model, combined with Beam Search decoding (loading a domain dictionary of 20,000+ terms).
[0125] ① Perform CRNN visual encoding
[0126] Character visual features are extracted using a 3-layer CNN: each layer performs a 3×3 convolution, using the ReLU activation function + 2×2 max pooling.
[0127] ② Sequence context is captured through 2-layer Bi-LSTM modeling
[0128] A two-layer bidirectional LSTM is used for processing, with a hidden layer dimension of 512.
[0129] ③ A 512-dimensional text semantic vector V is generated by a 6-layer Transformer Encoder with 8 self-attention heads. The self-attention is calculated using the Softmax function.text,i ;
[0130] ④ Combining Beam Search decoding, load the domain dictionary D containing N=20000 terms and calculate the sequence score: Where L is the sequence length, and w ∈ D is forced. Output the character sequence s. i and character-level coordinate matrix C text,i (Precision ±1 pixel). Each position in the sequence is represented by t (or each sample); p is the conditional probability, representing the probability in the model subparameter W. q Under the dictionary domain D constraint, a sequence of t elements W t The probability of occurrence; W t : Local elements of a sequence; W q : The model's "sub-parameter set"; D: Domain dictionary.
[0131] (2) Image modal feature extraction
[0132] Each graphic block G j Scaled to 224×224, input the ResNet-50 model through the first 4 residual blocks;
[0133] The feature map of the 4th residual block is extracted and then subjected to global average pooling to obtain a 1024-dimensional visual feature vector V. graph,j :
[0134] The EAST detector is used to locate the coordinate axis label regions, and the label content is recognized by the OCR model in the previous step to generate a 512-dimensional label semantic vector V. label,j .
[0135] (3) Structural modal feature extraction
[0136] Construct all region blocks (ROI-T / G / Tab / S) into a graph structure.
[0137] Where node v k The ∈V attribute includes: geometric features (center coordinates, aspect ratio, area percentage) + 4-dimensional class encoding; edge e kl ∈ε is defined by Euclidean distance (≤50 pixels) and orientation relationship (4-dimensional binary, up, down, left, and right);
[0138] A 64-dimensional structural feature vector v is extracted using a 2-layer GAT network (8-head attention). struct,k Describes the spatial dependencies between nodes:
[0139] Single-head attention weights: Where W is the weight matrix, and N(i) are the neighbors of node i. α ijAttention weights represent the attention coefficients of the i-th query to the j-th key, ranging from [0,1]; LeakyReLU: Linear Unit Activation Function; W: Weight Matrix; h i : The i-th query vector; h j : The j-th key vector; h k : The k-th key vector in set N(i).
[0140] (4) Semantic modality feature extraction
[0141] For each region block, semantic matching is performed based on the domain knowledge base.
[0142] Text block: Extracting 300-dimensional Word2Vec embedding vector v through terminology matching term,i ;
[0143] Symbol blocks: Generate symbol semantic labels by matching HOG features (360-dimensional) through a symbol mapping table. s and the corresponding embedding vector v sym,p ;
[0144] Table / Graph Block: Matches typical chart templates to generate a 256-dimensional structured prior vector v. temp,m .
[0145] Output: Set of feature vectors for each modality {v text ,v graph ,v struct ,v semantic} and coordinate association table (used to record the region location corresponding to the feature).
[0146] As can be seen from the above description, the drilling engineering document recognition method based on multimodal information fusion provided in this application can extract multimodal features from the preprocessed original drilling document image to obtain text modal features, image modal features, structural modal features and semantic modal features.
[0147] In one embodiment, see Figure 4 The process of fusing multimodal information from the text modal features, image modal features, structural modal features, and semantic modal features to obtain high-level fusion features and preliminary error correction identifiers includes:
[0148] S401: The text modal features, image modal features, structural modal features and semantic modal features are concatenated at the lowest level to obtain the lowest level fusion vector;
[0149] S402: Construct a text-graphic attention model and a table symbol attention model based on the independent region blocks;
[0150] S403: Generate a text-graphic association vector based on the text block set, the graphic block set, and the text-graphic attention model;
[0151] S404: Generate a table symbol association vector based on the table block set, the symbol block set, and the table symbol attention model;
[0152] S405: Perform residual connection processing on the text-graphic association vector, the table symbol association vector and the bottom-level fusion vector to obtain the middle-level fusion feature;
[0153] S406: Perform rule engine verification on the mid-layer fusion features to obtain the preliminary error correction identifier;
[0154] S407: Perform statistical model correction on the mid-layer fusion features to obtain the high-layer fusion features.
[0155] It is understandable that steps S401 to S407 can be understood as a process of multimodal information fusion.
[0156] Input: Text features v generated in the preceding steps text Graphic features v graph Structural features v struct Semantic features v semantic (All aligned according to the region block index.)
[0157] Processing procedure:
[0158] (1) Low-level feature splicing
[0159] For each region block k, the four types of feature vectors are concatenated according to their dimensions to form a 2688-dimensional bottom-level fusion vector: f low,k =[v text,k ;v graph,k ;v struct,k ;v semantic,k ].
[0160] (2) Mid-level cross-modal attention fusion
[0161] Constructing a bidirectional attention mechanism for text-graphics and tables-symbols:
[0162] Text-graphic association: using text blocks as v text,k For Query, the v of the graph block graph,j For key / value pairs, the association vector a is calculated using Scaled Dot-Product Attention. tg,kj Establish a semantic mapping between "data description and graphical representation".
[0163] Attention mechanism:
[0164] Where Query is a text semantic vector, and Key / Value is a graphical visual feature.
[0165] Table-Symbol Association: Using the text vector of a table cell as the query and the HOG features of a symbol block as the key / value pair, a symbol semantic correction vector 'a' is generated through 8-head attention. ts,mp .
[0166] The attention-related vector and the underlying fusion vector are merged through a residual connection and input into a two-layer fully connected network to generate a 1024-dimensional mid-layer fusion feature f. mid,k .
[0167] (3) High-level domain knowledge verification
[0168] Rule engine validation: For text blocks containing numerical values, check unit consistency based on the terminology knowledge base (e.g., "well depth" must be followed by "m"), and correct illegal units by editing distance (threshold ≤ 2); perform format compliance checks on table fields (e.g., hash symbols match regular expressions), and mark illegal results.
[0169] Statistical model error correction: Using a 5-gram language model trained on a corpus of 100,000 text samples, the perplexity of the text sequence is calculated, erroneous segments with a perplexity > 200 are replaced, and a validated high-level fusion feature f is generated. high,k :
[0170] Sequence perplexity PP(S) calculation:
[0171] Where PP(S): perplexity of sequence S; L: length of sequence S; t: position of sequence element; L: position of sequence element; p(W t |W 1-t+1 ): Conditional probability, representing the probability of obtaining the first t-1 elements W1, W2, ..., Wn. t-1 In the context of the t-th element W t The probability of occurrence; W t : The t-th element in sequence S; W 1-t+1 : The context sequence consisting of the first to the (t-1)th elements in sequence S.
[0172] Output: The set of feature vectors after fusion and verification {f high,k} and preliminary identification results (including tags to be corrected).
[0173] As can be seen from the above description, the drilling engineering document recognition method based on multimodal information fusion provided in this application can perform multimodal information fusion on the text modal features, image modal features, structural modal features and semantic modal features to obtain high-level fusion features and preliminary error correction identifiers.
[0174] In one embodiment, see Figure 5 The step of performing multi-task recognition and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain the final recognition result includes:
[0175] S501: The constructed joint loss function is used to perform multi-task learning on the high-level fusion features to obtain multi-task learning results;
[0176] S502: Perform cross-modal consistency verification on the multi-task learning results to obtain the final recognition result and secondary error correction identifier.
[0177] It is understandable that steps S501 to S502 can be understood as the process of performing multi-task recognition and error correction.
[0178] Input: fusion features {f} generated in the preceding steps high,k} and preliminary identification results (including tags to be corrected).
[0179] Processing procedure:
[0180] (1) Multi-task learning model
[0181] Text recognition task: using text block images and {f high,k} is the input, and the output is the character sequence s. i The loss function is CTCLoss + dictionary-constrained cross-entropy;
[0182] Graph classification task: using block images and graphs data Input: Graphic category labels (10 categories, such as stratigraphic profiles and drilling trajectory maps);
[0183] Table parsing task: Taking table block images and row and column line detection results as input, detect cells using Faster R-CNN, combine fused features to parse cell content, and generate a two-dimensional table matrix.
[0184] Symbol recognition task: using symbol block images and {f high,p} is the input, and the output is the symbol semantic label. The loss function is focal loss.
[0185] Joint loss function: L = λ1L ocr +λ2L class+λ3L table +λ4L sym .
[0186] Wherein, λ1, λ2, λ3, and λ4 are the task loss weights, which are the "importance" coefficients of the loss of each subtask.
[0187] L ocr CTC Loss + Dictionary-Constrained Cross Entropy
[0188] L class Cross-entropy loss for image classification
[0189] L table Faster R-CNN loss for table cell detection
[0190] L sym Symbol recognition focus loss
[0191] (2) Domain Error Correction Module
[0192] Perform cross-modal consistency verification on the output results of multiple tasks: for example, check the difference between the value of "outer diameter of sleeve" in the table and the value marked "Φ" in the corresponding equipment diagram. If the difference is greater than 10%, it will be corrected with the standard value of the knowledge base (median of historical data).
[0193] Based on the rule verification results of the previous steps, secondary error correction (such as unit replacement and format completion) is triggered for the identification results marked as illegal.
[0194] Output: The final recognition results after error correction (text sequence, graphic category, tabular data, symbol label) and the coordinate information of each modal object.
[0195] As can be seen from the above description, the drilling engineering document recognition method based on multimodal information fusion provided in this application can perform multi-task recognition and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain the final recognition result.
[0196] In one embodiment, see Figure 6 The structured output results include structured data files and document-level semantic graphs; the semantic association processing of the final recognition results to obtain the structured output results includes:
[0197] S601: The final recognition result is structured and sorted to obtain text structure sorting results, tabular data sorting results, graphic data sorting results, and symbol annotation sorting results;
[0198] S602: Based on the results of text structure analysis, table data analysis, graphic data analysis, and symbol annotation analysis, semantic association modeling is performed to obtain structured data files and document-level semantic graphs.
[0199] It is understandable that steps S501 to S502 can be understood as the process of performing structured output and application.
[0200] Input: The error-corrected recognition results generated in the previous steps (text / graphics / tables / symbol content and coordinates).
[0201] Processing procedure:
[0202] (1) Data structuring
[0203] Text document: Arrange text blocks in coordinate order, retain font size (inferred from text block height, with an error of ±2px) and alignment (left / center / right), and generate formatted paragraph text. struct ;
[0204] Table data: Parsing results of the table Convert to a two-dimensional array to record cell content, coordinates, and merge markers (rowspan / colspan), and support direct import from Excel / database;
[0205] Graphical data: Extract axis labels and data point coordinates from the graphic blocks (using Hough transform detection, error ≤ 5 pixels), and generate a coordinate sequence Graph in JSON format. data ;
[0206] Symbol annotation: Record the coordinates, semantic labels, and associated objects (nearest text / graphics within 30 pixels) of the symbol, and generate an annotation list.
[0207] (2) Semantic association modeling
[0208] Construct a document-level semantic graph: Each modal object is a node (including ID, type, content, and coordinates), and relationships such as "description", "correspondence", and "label" are edges (e.g., "text T5 describes graph G1"). It is stored in JSON-LD format and supports semantic retrieval in engineering software.
[0209] Output: Structured data files (JSON / XML format) and document-level semantic graphs, which can be directly imported into petroleum engineering data analysis systems for further processing.
[0210] As can be seen from the above description, the drilling engineering document recognition method based on multimodal information fusion provided in this application can perform semantic association processing on the final recognition result to obtain a structured output result.
[0211] Next, see Figure 14 An embodiment is given based on the method provided in this application.
[0212] 1. Preprocessing stage
[0213] The input image is a 300dpi color scan with slight wrinkles and stains;
[0214] Image denoising was performed using the U-Net model, and text regions (70%), table regions (20%), and line graph regions (10%) were detected using PaddleLayout.
[0215] The text area is skewed, the table area is divided into rows and columns, and the coordinate axes and data point outlines are extracted from the graph area.
[0216] 2. Feature Extraction Stage
[0217] Text modality: An improved CRNN model is used, which takes a text block image as input and outputs a character sequence and position coordinates. At the same time, a drilling domain dictionary (containing 5000+ technical terms) is loaded to optimize beamsearch decoding.
[0218] Image modality: Visual features were extracted from the curve region using ResNet-50, with a focus on axis labels (such as "depth / m" and "pressure / MPa") and curve trend characteristics;
[0219] Structural modality: The coordinate positions of text blocks, table blocks, and curve blocks are transformed into graph structures, with nodes representing regions and edges representing spatial distances and relative positional relationships. GAT (Graph Attention Network) is used to model structural features.
[0220] 3. Fusion and Recognition Stage
[0221] Bottom-level fusion: Text features (768-dimensional), image features (512-dimensional), and structural features (64-dimensional) are concatenated into a 1344-dimensional vector;
[0222] Mid-layer fusion: Through a cross-modal attention mechanism, a semantic association is established between the "well depth" data in the table and the depth coordinate axis in the curve, correcting the OCR error of misidentifying "3500m" as "3500n";
[0223] High-level verification: Illegal identification results are filtered using knowledge base rules (such as "well depth unit must be m" and "pressure unit must be MPa"), and the N-gram model is used to perform secondary error correction on the text sequence.
[0224] 4. Output Stage
[0225] The recognition results are transformed into structured data, where tabular data is stored as a two-dimensional matrix, key coordinate points of the curve data are extracted and associated with the measured values in the table, and finally a JSON file that can be imported into the engineering database is generated.
[0226] The results of testing on the drilling documentation dataset are as follows:
[0227] index Traditional OCR The method provided in this application Character recognition accuracy 82.3% 97.5% Table parsing accuracy 75.0% 92.0% Symbol recognition accuracy 68.0% 89.5% Structured output completeness 60.0% 95.0%
[0228] Based on the same inventive concept, this application also provides a drilling engineering document recognition device based on multimodal information fusion, which can be used to implement the method described in the above embodiments, as described in the following embodiments. Since the principle of solving the problem by the drilling engineering document recognition device based on multimodal information fusion is similar to that of the drilling engineering document recognition method based on multimodal information fusion, the implementation of the drilling engineering document recognition device based on multimodal information fusion can refer to the implementation of the software performance benchmark determination method, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the system described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0229] In one embodiment, see Figure 7 In order to enhance the semantic understanding of complex graphics in drilling documents and improve the recognition accuracy of professional symbols, charts and annotations by integrating multimodal information such as text, images and domain knowledge, this application provides a drilling engineering document recognition device based on multimodal information fusion, including: a modal feature extraction unit 701, a modal information fusion unit 702, a final recognition unit 703 and a structure output unit 704.
[0230] The modal feature extraction unit 701 is used to extract multimodal features from the preprocessed original drilling document image to obtain text modal features, image modal features, structural modal features and semantic modal features;
[0231] The modal information fusion unit 702 is used to perform multimodal information fusion on the text modal features, image modal features, structural modal features and semantic modal features to obtain high-level fusion features and preliminary error correction identifiers;
[0232] The final identification unit 703 is used to perform multi-task identification and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain the final identification result; wherein, the final identification result includes text sequence, graphic category, tabular data, symbol label and corresponding position coordinates;
[0233] The structure output unit 704 is used to perform semantic association processing on the final recognition result to obtain a structured output result.
[0234] In one embodiment, see Figure 8The modal feature extraction unit 701 includes: a document enhancement module 801, a document layout module 802, and a data segmentation module 803.
[0235] Document enhancement module 801 is used to enhance the original drilling document image to obtain an enhanced drilling document image;
[0236] The document layout module 802 is used to perform region detection on the enhanced image of the drilling document to obtain corresponding document layout information; wherein, the document layout information includes content categories and corresponding position coordinates; the content categories include text, graphics, tables and symbols;
[0237] The data segmentation module 803 is used to perform multimodal data segmentation on the enhanced image of the drilling document according to the document layout information to obtain independent region blocks; wherein, the independent region blocks include a set of text blocks, a set of graphic blocks, a set of table blocks, and a set of symbol blocks.
[0238] In one embodiment, see Figure 9 The modal feature extraction unit 701 includes: a text modal extraction module 901, an image modal extraction module 902, a structural modal extraction module 903, and a semantic modal extraction module 904.
[0239] The text modality extraction module 901 is used to perform visual encoding extraction on the text block set to obtain the text modality features;
[0240] The image modality extraction module 902 is used to perform visual label recognition on the graphic block set to obtain the image modality features;
[0241] The structural modality extraction module 903 is used to construct a graph structure for the independent region blocks to obtain the structural modality features;
[0242] The semantic modality extraction module 904 is used to perform semantic matching processing on the independent region blocks to obtain the semantic modality features.
[0243] In one embodiment, see Figure 10 The modal information fusion unit 702 includes: a bottom-level fusion module 1001, an attention model construction module 1002, a text-image association vector generation module 1003, a table number association vector generation module 1004, a middle-level fusion module 1005, a preliminary error correction module 1006, and a high-level fusion module 1007.
[0244] The low-level fusion module 1001 is used to perform low-level feature concatenation on the text modal features, image modal features, structural modal features and semantic modal features to obtain a low-level fusion vector;
[0245] Attention model building module 1002 is used to build a text graphics attention model and a table symbol attention model based on the independent region blocks;
[0246] The text-image association vector generation module 1003 is used to generate a text-image association vector based on the text block set, the image block set, and the text-image attention model.
[0247] Table number association vector generation module 1004 is used to generate table symbol association vectors based on the table block set, the symbol block set and the table symbol attention model;
[0248] The middle-layer fusion module 1005 is used to perform residual connection processing on the text-graphic association vector, the table symbol association vector and the bottom-layer fusion vector to obtain the middle-layer fusion feature;
[0249] The preliminary error correction module 1006 is used to perform rule engine verification on the mid-layer fusion features to obtain the preliminary error correction identifier;
[0250] The high-level fusion module 1007 is used to perform statistical model error correction on the mid-level fusion features to obtain the high-level fusion features.
[0251] In one embodiment, see Figure 11 The final identification unit 703 includes a task learning module 1101 and an identification and error correction module 1102.
[0252] The task learning module 1101 is used to perform multi-task learning on the high-level fusion features using the constructed joint loss function to obtain multi-task learning results.
[0253] The identification and error correction module 1102 is used to perform cross-modal consistency verification on the multi-task learning results to obtain the final identification result and secondary error correction identifier.
[0254] In one embodiment, see Figure 12 The structured output results include structured data files and document-level semantic graphs; the structured output unit 704 includes:
[0255] The structure sorting module 1201 is used to sort the final recognition result in a structure, and obtain text structure sorting result, table data sorting result, graphic data sorting result and symbol annotation sorting result;
[0256] The semantic association module 1202 is used to perform semantic association modeling based on the text structure sorting results, table data sorting results, graphic data sorting results and symbol annotation sorting results to obtain structured data files and document-level semantic graphs.
[0257] From a hardware perspective, in order to enhance the semantic understanding of complex graphics in drilling documents and improve the recognition accuracy of professional symbols, charts, and annotations by fusing multimodal information such as text, images, and domain knowledge, this application provides an embodiment of an electronic device for implementing all or part of the aforementioned drilling engineering document recognition method based on multimodal information fusion. The electronic device specifically includes the following components:
[0258] The system comprises a processor, a memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to realize information transmission between the drilling engineering document recognition device based on multimodal information fusion and core business systems, user terminals, and related databases and other related devices; the logic controller can be a desktop computer, tablet computer, or mobile terminal, etc., and this embodiment is not limited to these. In this embodiment, the logic controller can be implemented with reference to the embodiments of the drilling engineering document recognition method based on multimodal information fusion and the embodiments of the drilling engineering document recognition device based on multimodal information fusion in the embodiments, the content of which is incorporated herein, and repeated details will not be described again.
[0259] It is understood that the user terminal may include smartphones, tablet computers, network set-top boxes, portable computers, desktop computers, personal digital assistants (PDAs), in-vehicle devices, smart wearable devices, etc. Among these, the smart wearable devices may include smart glasses, smartwatches, smart bracelets, etc.
[0260] In practical applications, parts of the drilling engineering document recognition method based on multimodal information fusion can be executed on the electronic device side as described above, or all operations can be completed in the client device. The choice can be made based on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations on this. If all operations are completed in the client device, the client device may further include a processor.
[0261] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.
[0262] Figure 13 This is a schematic block diagram illustrating the system configuration of the electronic device 9600 according to an embodiment of this application. Figure 13 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that... Figure 13 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.
[0263] In one embodiment, the drilling engineering document recognition method based on multimodal information fusion can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control:
[0264] S101: Perform multimodal feature extraction on the preprocessed original drilling document images to obtain text modal features, image modal features, structural modal features, and semantic modal features;
[0265] S102: Perform multimodal information fusion on the text modal features, image modal features, structural modal features and semantic modal features to obtain high-level fusion features and preliminary error correction identifiers;
[0266] S103: Perform multi-task recognition and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain the final recognition result; wherein, the final recognition result includes text sequence, graphic category, tabular data, symbol label and corresponding position coordinates;
[0267] S104: Perform semantic association processing on the final recognition result to obtain a structured output result.
[0268] As described above, the drilling engineering document recognition method based on multimodal information fusion provided in this application can solve the problem of misrecognition of professional terms and complex symbols by traditional optical symbol recognition technology through the complementarity of multimodal information, thereby improving the character recognition accuracy on the drilling document test set; it can realize the semantic association parsing of text, images and tables in the document by using structured parsing capabilities, and output structured results that can be directly used for engineering data analysis, reducing manual data entry costs; through domain knowledge base and customized model training, it can quickly adapt to different types of drilling engineering documents, including drilling reports and well completion reports.
[0269] In another embodiment, the drilling engineering document recognition device based on multimodal information fusion can be configured separately from the central processing unit 9100. For example, the data composite transmission device based on multimodal information fusion can be configured as a chip connected to the central processing unit 9100, and the function of the drilling engineering document recognition method based on multimodal information fusion can be realized through the control of the central processing unit.
[0270] like Figure 13 As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily need to include these components. Figure 13 All components shown; in addition, the electronic device 9600 may also include Figure 13 For components not shown, please refer to existing technologies.
[0271] like Figure 13 As shown, the central processing unit 9100, sometimes also referred to as a controller or operating control, may include a microprocessor or other processor device and / or logic device, which receives inputs and controls the operation of various components of the electronic device 9600.
[0272] The memory 9140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information, and also store a program for executing that information. The central processing unit 9100 may execute the program stored in the memory 9140 to perform information storage or processing, etc.
[0273] Input unit 9120 provides input to central processing unit 9100. Input unit 9120 may be, for example, a keypad or touch input device. Power supply 9170 provides power to electronic device 9600. Display 9160 displays images and text. Display may be, for example, an LCD display, but is not limited thereto.
[0274] The memory 9140 can be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), a SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROMs. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 via the central processing unit 9100.
[0275] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for the electronic device's communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0276] The communication module 9110 is a transmitter / receiver that sends and receives signals via the antenna 9111. The communication module (transmitter / receiver) 9110 is coupled to the central processing unit 9100 to provide input signals and receive output signals, which is the same as in a conventional mobile communication terminal.
[0277] Based on different communication technologies, multiple communication modules 9110 can be configured in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication module (transmitter / receiver) 9110 is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby realizing typical telecommunications functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Additionally, the audio processor 9130 is also coupled to a central processing unit 9100, enabling on-device recording via the microphone 9132 and on-device playback of stored sound via the speaker 9131.
[0278] Embodiments of this application also provide a computer-readable storage medium capable of implementing all steps of the drilling engineering document recognition method based on multimodal information fusion, where the execution subject is a server or client, as described in the above embodiments. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all steps of the drilling engineering document recognition method based on multimodal information fusion, where the execution subject is a server or client, as described in the above embodiments. For example, when the processor executes the computer program, it implements the following steps:
[0279] S101: Perform multimodal feature extraction on the preprocessed original drilling document images to obtain text modal features, image modal features, structural modal features, and semantic modal features;
[0280] S102: Perform multimodal information fusion on the text modal features, image modal features, structural modal features and semantic modal features to obtain high-level fusion features and preliminary error correction identifiers;
[0281] S103: Perform multi-task recognition and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain the final recognition result; wherein, the final recognition result includes text sequence, graphic category, tabular data, symbol label and corresponding position coordinates;
[0282] S104: Perform semantic association processing on the final recognition result to obtain a structured output result.
[0283] As described above, the drilling engineering document recognition method based on multimodal information fusion provided in this application can solve the problem of misrecognition of professional terms and complex symbols by traditional optical symbol recognition technology through the complementarity of multimodal information, thereby improving the character recognition accuracy on the drilling document test set; it can realize the semantic association parsing of text, images and tables in the document by using structured parsing capabilities, and output structured results that can be directly used for engineering data analysis, reducing manual data entry costs; through domain knowledge base and customized model training, it can quickly adapt to different types of drilling engineering documents, including drilling reports and well completion reports.
[0284] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0285] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0286] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.
[0287] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0288] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A drilling engineering document recognition method based on multimodal information fusion, characterized in that, include: Multimodal features are extracted from the original drilling document images to obtain modal features; the modal features include text modal features, image modal features, structural modal features, and semantic modal features; Multimodal information fusion is performed on the modal features to obtain high-level fusion features and preliminary error correction identifiers; Based on the high-level fusion features and preliminary error correction identifiers, multi-task recognition and error correction are performed to obtain document recognition results; wherein, the document recognition results include text sequences, graphic categories, tabular data, symbol labels and corresponding position coordinates; The document recognition results are subjected to semantic association processing to obtain structured output results.
2. The drilling engineering document recognition method based on multimodal information fusion according to claim 1, characterized in that, Also includes: The original drilling document image is enhanced to obtain an enhanced drilling document image; Region detection is performed on the enhanced image of the drilling document to obtain the corresponding document layout information; wherein, the document layout information includes content categories and corresponding location coordinates; the content categories include text, graphics, tables and symbols; The drilling document augmented image is segmented into independent regions based on the document layout information. The independent regions include a set of text blocks, a set of graphic blocks, a set of table blocks, and a set of symbol blocks.
3. The drilling engineering document recognition method based on multimodal information fusion according to claim 2, characterized in that, Multimodal feature extraction is performed on the original drilling document images to obtain modal features, including: Visual encoding extraction is performed on the set of text blocks to obtain the text modal features; Visual label recognition is performed on the set of graphic blocks to obtain the image modal features; Graph structures are constructed from the independent region blocks to obtain the structural modal features; The semantic modality features are obtained by performing semantic matching on the independent region blocks.
4. The drilling engineering document recognition method based on multimodal information fusion according to claim 2, characterized in that, The process of fusing multimodal information on the modal features to obtain high-level fusion features and preliminary error correction identifiers includes: The text modal features, image modal features, structural modal features, and semantic modal features are concatenated at the lowest level to obtain a low-level fusion vector; Construct a text-graphic attention model and a table symbol attention model based on the independent region blocks; Generate a text-graphic association vector based on the text block set, the graphic block set, and the text-graphic attention model; Generate a table symbol association vector based on the table block set, the symbol block set, and the table symbol attention model; The text-graphic association vector, the table symbol association vector, and the bottom-level fusion vector are subjected to residual connection processing to obtain the middle-level fusion feature; The mid-layer fusion features are validated using a rule engine to obtain the preliminary error correction identifier; The high-level fusion features are obtained by performing statistical model correction on the mid-level fusion features.
5. The drilling engineering document recognition method based on multimodal information fusion according to claim 1, characterized in that, The step of performing multi-task recognition and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain document recognition results includes: The high-level fusion features are learned through multi-task learning using a pre-constructed joint loss function to obtain multi-task learning results; Cross-modal consistency verification is performed on the multi-task learning results to obtain the document recognition results.
6. The drilling engineering document recognition method based on multimodal information fusion according to claim 1, characterized in that, The structured output results include structured data files and document-level semantic graphs; the semantic association processing of the document recognition results to obtain the structured output results includes: The document recognition results are then structured to obtain text structure organization results, tabular data organization results, graphic data organization results, and symbol annotation organization results. Based on the results of text structure analysis, tabular data analysis, graphical data analysis, and symbol annotation analysis, semantic association modeling is performed to obtain structured data files and document-level semantic graphs.
7. A drilling engineering document recognition device based on multimodal information fusion, characterized in that, include: The modal feature extraction unit is used to extract multimodal features from the original drilling document images to obtain modal features; the modal features include text modal features, image modal features, structural modal features, and semantic modal features; The modal information fusion unit is used to perform multimodal information fusion on the modal features to obtain high-level fusion features and preliminary error correction identifiers; The final recognition unit is used to perform multi-task recognition and error correction based on the high-level fusion features and preliminary error correction identifiers to obtain document recognition results; wherein, the document recognition results include text sequences, graphic categories, tabular data, symbol labels and corresponding position coordinates; The structure output unit is used to perform semantic association processing on the document recognition results to obtain structured output results.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the drilling engineering document recognition method based on multimodal information fusion as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the drilling engineering document recognition method based on multimodal information fusion as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the drilling engineering document recognition method based on multimodal information fusion as described in any one of claims 1 to 6.