An engineering drawing title block recognition method and system based on multi-modal information extraction

Through multimodal information extraction technology, combined with text and image information, the accuracy and universality of drawing information extraction are solved, and the end-to-end information extraction from image format drawings is realized, and a unified extraction framework is adapted to different drawings.

CN119964171BActive Publication Date: 2025-07-11ZHEJIANG HUADONG ENG DIGITAL TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510450366.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

When extracting drawing information of drawing files in image formats, the prior art has problems that the OCR processing results do not have complete natural semantics, lack of unified standards, and traditional methods cannot handle complex information relationships.

Method used

Multimodal information extraction technology is adopted, combining text semantics, images and Layout layout information, and cross-modal information extraction models are used to realize cross-modal information extraction, build a unified extraction framework, and support unified extraction of quadrants and triples.

Benefits of technology

It improves the accuracy and universality of drawing information extraction, and can accurately extract information when there are differences in the location, format and content of drawing drawings, and adapt to various relationship types and scenario requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964171B_ABST
    Figure CN119964171B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for engineering drawing title block recognition based on multi-modal information extraction. The method includes obtaining an engineering drawing image and preprocessing the image, and setting a data structure Schema for controlling extraction targets and relationship patterns; detecting the title block area in the drawing through a trained title block detection model to obtain the bounding box coordinates of each title block, mapping the detected bounding box coordinates back to the coordinate system of the original image, and cropping out the corresponding title block image from the original image; performing text recognition on the cropped title block image to extract the text content and the corresponding text box coordinate information; inputting the cropped title block image and the text recognition result into a trained multi-modal information extraction model, and performing information extraction according to the set Schema to output a structured extraction result. This method can flexibly define extraction targets, support the unified extraction of binary tuples and triple tuples, and has high extraction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence deep learning and image processing, and particularly relates to a method and system for recognizing a drawing title block based on multi-modal information extraction. Background Art

[0002] In industries such as construction engineering, mechanical manufacturing, and electronic engineering, design drawings, as an important carrier for recording and communicating design schemes, play an important role in the processes of project management, construction execution, and post-maintenance. With the increase in the quantity and variety of drawings, the retrieval, management, and archiving of these drawings have become key tasks in the informatization construction of relevant enterprises and institutions. Generally, design drawings are drawn using CAD (Computer-Aided Design) software and contain rich graphic information and text descriptions. A corresponding title block is generally designed on the drawing to record key information such as the project name, drawing name, drawing number, designer, reviewer, and formation date of the drawing. How to structurally extract these key information from the title block is an important part of drawing informatization management.

[0003] There are mainly two directions for the recognition and extraction of drawing title blocks. One is to recognize the vector information (such as line elements and text elements) of CAD vector drawings, and the other is to use technologies in the fields of vision and NLP (Natural Language Processing) to recognize drawings converted into image formats. Considering that the original CAD files cannot be obtained in all cases, and in some scenarios, only picture format files of design drawings can be collected and obtained, the applicability of extracting title block information from drawings in image format is more extensive.

[0004] Currently, there are still some difficulties in extracting title block information from drawing files in image format. First, the title block itself is mostly in the form of a table or other forms with layout information. The results directly obtained by OCR (Optical Character Recognition) processing do not have complete natural semantics, and the method of extracting NLP information through OCR processing usually cannot obtain accurate output. Second, there is no unified standard among different drawing units and departments, and there are differences in the position, format, content, etc. of drawing title blocks. Therefore, the method of rule or template matching cannot achieve a high universality on a wide variety of title blocks. Third, the information relationships in drawing title blocks include not only binary groups in the form of key-value pairs but also more complex ternary groups. Traditional extraction methods cannot simultaneously complete the information extraction of these two forms. Summary of the Invention

[0005] Based on the above background, the purpose of the present invention is to provide a method and system for structurally extracting the title block information in engineering drawings based on multi-modal information extraction technology, which combines text semantics, images, and Layout layout information for cross-modal information extraction to improve the accuracy of information extraction.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions:

[0007] In the first aspect, the present invention provides a method for identifying the title block of an engineering drawing based on multi-modal information extraction, including:

[0008] Obtain the engineering drawing image and preprocess the image, and set a data structure Schema for controlling the extraction target and relationship mode, including a binary data structure and / or a triple data structure;

[0009] Detect the title block area in the drawing through a trained title block detection model to obtain the bounding box coordinates of each title block, and map the detected bounding box coordinates from the coordinate system of the preprocessed image back to the coordinate system of the original image to crop the corresponding title block area image from the original image;

[0010] Perform text recognition on the cropped title block image to extract the text content and the corresponding text box coordinate information;

[0011] Construct a multi-modal information extraction model and input the cropped title block image and the text recognition result into the trained multi-modal information extraction model, and perform information extraction according to the set data structure Schema for controlling the extraction target and relationship mode, and output a structured extraction result.

[0012] In the second aspect, the present invention provides a system for identifying the title block of an engineering drawing based on multi-modal information extraction, including:

[0013] An original picture acquisition module for obtaining the engineering drawing image and preprocessing the image;

[0014] A data structure definition module for obtaining the set data structure Schema for controlling the extraction target and relationship mode, including a binary data structure and / or a triple data structure;

[0015] A title block detection module for detecting the title block area in the drawing through a trained title block detection model to obtain the bounding box coordinates of each title block, and mapping the detected bounding box coordinates from the coordinate system of the preprocessed image back to the coordinate system of the original image to crop the corresponding title block area image from the original image;

[0016] A text recognition module for performing text recognition on the cropped title block image to extract all the text content and the corresponding text box coordinate information;

[0017] A multimodal information extraction module, which is used to construct a multimodal information extraction model, input the cropped label images and text recognition results into the trained multimodal information extraction model, and perform information extraction according to the set control extraction target and the data structure Schema of the relationship mode, and output the structured extraction results;

[0018] A post-processing module, which is used to post-process the extraction results, including format verification of specific fields and similarity detection and filtering of redundant results.

[0019] The beneficial effects of the present invention are as follows:

[0020] The engineering drawing label recognition method based on multimodal information extraction provided by the embodiment of the present invention combines text semantics, images and Layout layout information for cross-modal information extraction, has the ability to directly perform end-to-end information extraction from pictures, and has strong universality. Even if there are differences in the position, format, content, etc. of the drawing labels, accurate information can be extracted.

[0021] At the same time, the present invention provides a unified extraction framework, which realizes flexible definition of extraction targets through Prompt construction, supports unified extraction of binary and ternary groups, adapts to various relationship types and scenario requirements, without additional adjustment of the model.

[0022] The embodiment of the present invention further proposes a complete process from image preprocessing, object detection, OCR recognition to cross-modal information extraction and post-processing, provides a complete cross-modal information extraction system, and improves the accuracy of information extraction. Description of the Drawings

[0023] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0024] Figure 1 It is a flowchart of the engineering drawing label recognition method based on multimodal information extraction provided by the embodiment of the present invention;

[0025] Figure 2 It is a schematic diagram of the training process of the label detection model provided by the embodiment of the present invention;

[0026] Figure 3 It is a schematic diagram of the training process of the multimodal information extraction model provided by the embodiment of the present invention;

[0027] Figure 4Schematic diagram of the engineering drawing title block recognition system based on multi-modal information extraction provided by the embodiments of the present invention;

[0028] Figure 5 It is an example diagram of binary annotation provided by the embodiments of the present invention;

[0029] Figure 6 It is an example diagram of triple annotation provided by the embodiments of the present invention. Detailed implementation manners

[0030] To further understand the present invention, the preferred implementation manners of the present invention will be described below in conjunction with embodiments. However, it should be understood that these descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention.

[0031] Term explanation:

[0032] Schema refers to the data structure that controls the extraction target and relationship pattern;

[0033] Prompt is a prompt word used to guide the model to perform a specific task or generate specific information;

[0034] SSI (Structural Schema Instructor) is a prompt mechanism used to guide the model to complete a specific information extraction task; by adding a specific prompt (Prompt) before the input data, key information such as task requirements and data types is embedded into the input of the model to help the model better understand and process multi-modal data;

[0035] Transformer is a neural network architecture based on the self-attention mechanism;

[0036] LayoutLMv3 is a multi-modal Transformer model that combines text and layout information;

[0037] YOLOv8 (You Only Look Once v8) is a deep learning model for object detection;

[0038] LabelImg is an annotation tool for image object detection, supporting multiple annotation formats;

[0039] Label Studio is an open-source data annotation tool, supporting annotation of multiple data types;

[0040] Batch size is the number of samples input at one time during training, which affects training efficiency and stability.

[0041] Embodiment 1

[0042] See Figure 1 , this embodiment provides a method for identifying the title block of engineering drawings based on multi-modal information extraction, including

[0043] S1: Obtain the engineering drawing image and preprocess the image, and set the data structure Schema that controls the extraction target and relationship mode, including the binary data structure and / or the triple data structure;

[0044] S2: Detect the title block area in the drawing through the trained title block detection model, obtain the bounding box coordinates of each title block, and map the detected bounding box back to the coordinate system of the original image from the coordinate system of the preprocessed image using the scaling ratio, so as to crop the corresponding title block area image from the original image;

[0045] S3: Perform text recognition on the cropped title block image, and extract the text content and the corresponding text box coordinate information;

[0046] S4: Construct a multi-modal information extraction model and input the cropped title block image and the recognized text content into the trained multi-modal information extraction model, and perform information extraction according to the set data structure that controls the extraction target and relationship mode, and output the structured extraction result.

[0047] S5: Post-process the extraction result to ensure the integrity and accuracy of the output information.

[0048] Specifically:

[0049] Step S1 specifically includes

[0050] S101: Obtain the engineering drawing image and preprocess the image.

[0051] The obtained engineering drawing image is in the format of TIFF or PNG, and it should be ensured that the clarity of the drawing meets the requirements of subsequent processing. Perform preprocessing operations on the input engineering drawing, including decoding the image into an array format, resizing the image, and normalizing it, so as to facilitate subsequent model processing.

[0052] S102: Obtain the predefined data structure Schema that controls the extraction target and relationship mode.

[0053] In the binary relationship, Schema is a list of keys. For example Figure 5 if you want to extract the binary information of review and check, Schema is ["review", "check"], Figure 6If you want to extract triple information, the Schema is ["drawing name": ["drawing number", "drawing issue date"]]. The Schema defines the types of information to be extracted and associated, and this information is usually parsed into tree nodes, each node representing an entity or relationship type, which is used to construct the SSI Prompt and guide subsequent information extraction.

[0054] Step S2 is to detect the title block area in the drawing through a trained title block detection model, obtain the class label (label) and bounding box coordinates (bbox) of each title block, and use the scaling ratio to map the detected bounding box from the coordinate system of the preprocessed image back to the coordinate system of the original image, so as to crop the corresponding title block area image from the original image.

[0055] Preferably, geometric correction such as orientation and super-resolution technology can also be used to perform image enhancement processing on the title block area image. There may be problems such as incorrect orientation and low resolution in the cropped title block area. Geometric correction can use image processing techniques such as Hough transform or deep learning models to detect the orientation of the image, and then correct the image to the correct orientation through rotation operations; super-resolution can use a pre-trained super-resolution model, such as the ESRGAN (Enhanced Super-Resolution Generative Adversarial Network) model, to convert a low-resolution image into a high-resolution image while maintaining the details and realism of the image.

[0056] See Figure 2 , the training process of the title block detection model specifically includes,

[0057] S201: Drawing data preparation and annotation. Collect a large number of real engineering drawing images to ensure that different types of title blocks are covered, and formulate labels according to these different types. For example, they can be classified as "architectural engineering title blocks", "hydropower engineering title blocks", "transportation engineering title blocks", etc. according to the industry, or classified as "review type title blocks", "specification type title blocks", "catalogue type title blocks", etc. according to the usage; then use the open-source tool LabelImg to annotate the images, draw a rectangular box for each title block area and assign a label. After annotation, export it as a label file in YOLO format, which contains the position coordinates (center point coordinates x, y, width w, and height h) and class ID of each object.

[0058] S202: Detection dataset construction. Divide the annotated dataset into a training set, a validation set, and a test set according to the ratio of 70%:15%:15%. Since YOLO provides online data augmentation, there is no need to perform offline augmentation processing on the dataset.

[0059] S203: Training environment setup. Equip the hardware environment, using a server equipped with NVIDIA GPUs, such as NVIDIA Tesla V100, RTX 3090, etc.; build the software environment, and install necessary libraries and frameworks, such as Python, PyTorch, and Ultralytics.

[0060] S204: Parameter configuration and training. For this part, it is sufficient to use a relatively new version of YOLO (such as above v8) for training. The training process is as follows: Modify model training parameters such as batch size and learning rate. The batch size can be set to 16, the initial learning rate can be set to 0.001, use cosine annealing decay, and use the Adam optimizer. The loss function part does not need to be modified; start training after completing the above configurations.

[0061] S205: Model evaluation and saving. After training, if the amount of training data is large, for the task of label detection, the accuracy can reach over 97%; save the model weights obtained from training.

[0062] Step S3 is to use an OCR tool to perform text recognition on the label image and extract all text content and its corresponding text box coordinate information, that is, the bounding box coordinates of the text sequence, specifically including the x and y coordinates of the center, width w, and height h.

[0063] Step S4 is to input the cropped label image and text recognition results into the trained multi-modal information extraction model, and perform information extraction according to the set control extraction target and the data structure of the relationship pattern, and output the structured extraction result.

[0064] See Figure 3 , the training process for constructing the multi-modal information extraction model specifically includes:

[0065] S401: Data annotation and dataset construction. For engineering drawing images, after performing label detection using the label detection model, a large number of drawings of label regions with unlabeled data are obtained. Use Label Studio to annotate the label information. For Figure 5 the binary tuple labels shown, for each pair of information, the key-value pair of key-value needs to be annotated. For Figure 6 the triple tuple labels shown, the key-value pair (key-value) must be annotated and associated (Link) to the subject (Subject). Divide the annotated dataset into a training set, a validation set, and a test set according to the ratio of 80%:10%:10%.

[0066] S402: Document data preprocessing. Obtain the document image and use an OCR tool to extract the text content and the corresponding two-dimensional layout information, i.e., the bounding box coordinates of the text sequence, specifically including the x and y coordinates of the center, width w, and height h.

[0067] S403: Construct a structured pattern guidance. The core idea of the structured pattern instructor SSI is to utilize the structured Schema information of the task to construct a prefix prompt word Prompt to guide the model to perform information extraction of a specific type.

[0068] Specifically, all information extraction IE (Information Extraction) tasks can be regarded as the conversion from text to structure, and this conversion can be decomposed into two operations: one is spotting, which is used to locate the segments of specific semantic types; the other is associating, which is used to associate these segments according to the predefined Schema.

[0069] In the information extraction task, the Schema is used to define the information structure to be extracted. Through the predefined Schema, the model can more accurately extract relevant entities, relationships, and events from the text. Different information extraction IE tasks have different structured patterns Schema. The structured pattern instructor (Structural Schema Instructor, SSI) is used to guide the model to perform information extraction according to a specific Schema. Specifically, the information about spotting and associating in the Schema is embedded into the input sequence of the model in a specific format, which is equivalent to providing a clear instruction to the model, telling the model what type of information should be extracted and how to organize this information.

[0070] The implementation method of the structured pattern instructor SSI mainly constructs a Prompt prefix containing specific types of Tokens. Tokens are the basic units that make up the Prompt prefix and are used to convey key information such as task intent, data type, location information, and relationship type to the model, guiding the behavior of the model to better complete the information extraction task. In this embodiment, SSI includes the following three types of content:

[0071] Spotting name SpotName: Represents the key in a binary or ternary tuple.

[0072] Association name AssoName: Represents the subject in a ternary tuple.

[0073] Special Symbols: The three special symbols used to separate and identify SpotName, AssoName, and the original text sequence are: [spot], [asso], and [text].

[0074] Add the corresponding special symbols [spot] and [asso] as prefixes to the two parts of SpotName and AssoName respectively, and then connect them with the [text] special symbol to form the SSI Prompt.

[0075] For example Figure 5 If you want to extract the binary tuple information of review and verification in, the Prompt is "[spot] review [spot] verification [text]". Figure 6 If you want to extract the triple information in, the Prompt is "[asso] drawing name [spot] drawing number [spot] drawing date [text]".

[0076] S404: Generate multimodal input. The multimodal input specifically includes the following text vectors and image vectors, specifically:

[0077] The text vector is the fusion of the word vector and the position vector. The word vector is the text content extracted by the OCR tool in step S402, and the SSI Prompt has been incorporated in the previous step S403, that is, the SSI Prompt is directly concatenated in front of the text content extracted by the OCR tool; the position vector includes the 1D position vector and the 2D layout information extracted by the OCR tool. The 1D position vector is the position index of each word in the text sequence; the 2D layout information includes the x coordinate, y coordinate, width, and height of the center coordinates, and they are normalized. The fusion method is to add the word vector, the one-dimensional position vector of the word, and the two-dimensional position vector element by element (add).

[0078] The image vector follows the processing principle of LayoutLMv3. The document image is segmented into image patches of a fixed size, a linear projection vector is generated for each image patch, and it is mapped to the same dimension as the text vector; and a 1D position vector is added to each image patch to represent its order in the image.

[0079] Finally, concatenate the text vector and the image vector to form a multimodal input. Specifically, use the cat operation to concatenate along the sequence dimension (dim=1). For example, if the shape of the text embedding is (2, 10, 768), it means 2 samples, each sample has 10 tokens, and the embedding dimension of each token is 768; the shape of the image embedding is (2, 197, 768), which means 2 samples, each sample has 196 patches + 1 [CLS] token, and the embedding dimension of each is 768. Then the shape of the concatenated embedding is (2, 207, 768), where 207 is the sum of the text token number (10) and the image patch number (197). In this way, the embedding vectors of text and image are integrated into a unified multimodal input, which can be further passed to the multimodal Transformer model for processing.

[0080] S405: Building the multimodal Transformer model. Use the multi-layer Transformer architecture of LayoutLMv3 to capture complex features. Each layer contains:

[0081] Multi-Head Self-Attention: Capture the interaction information among text, image, and SSI Prompt.

[0082] Position Bias: Introduce semantic 1D relative position and spatial 2D relative position to enhance the alignment ability between modalities.

[0083] Feed-Forward Neural Network: Perform a non-linear transformation on the output of the multi-head self-attention mechanism to further extract and transform features.

[0084] Then stack multiple layers of Transformer. Specifically, use 12 layers of Transformer encoders, set the hidden layer size to 768, and use 12 heads (12-head) in the multi-head self-attention mechanism. Each layer of Transformer gradually extracts cross-modal context representations through self-attention and feed-forward networks, deepening the model's understanding of multimodal data.

[0085] Perform joint multimodal encoding, that is, concatenate the vector representations of text, image, and modality prompt words along the sequence dimension to form a unified multimodal input sequence and input it into the Transformer model for processing.

[0086] S406: By adding two independent fully connected layers on the basis of the Transformer output, predict the start position and end position of the target value respectively.

[0087] Fully connected layer prediction and loss function definition.

[0088] Specifically: taking the output of the Transformer as the input, connecting two independent feed-forward neural networks, that is, fully connected layers, which are respectively used to output the start position and end position of the predicted value. For each key to be extracted, forward inference can be performed to obtain the corresponding value.

[0089] During training, the true labels of the start position and end position are obtained by fusing the annotation information of S401 and the OCR text recognition information of S402. The specific fusion process is as follows: matching the value in the annotation information to the corresponding text in the OCR recognition result, so as to obtain the start and end position indexes of the target value in the OCR text recognition result.

[0090] The cross-entropy loss function is used to measure the difference between the predicted position and the true position, and the model parameters are updated by the optimizer. In this embodiment, the binary cross-entropy is used to calculate the loss function, and the formula is as follows: ;

[0091] Among them, is the true label (ground truth), is the predicted value (probability value) of the model, that is, to measure the difference between the probability distribution predicted by the model and the true distribution, and it is expected that the model predicts a high probability for positive samples and a low probability for negative samples.

[0092] S407: Model fine-tuning. In this embodiment, the main model uses the pre-trained model of LayoutLMv3 as the initial weight, sets hyperparameters such as batch size and learning rate, starts training and adjusts the hyperparameters according to the evaluation of the model structure on the validation set, and saves the model after training. Among them, determining whether the hyperparameters are appropriate is mainly evaluated by monitoring indicators such as the loss, accuracy, and learning curve of the training set and the validation set. For example, if the training loss drops too fast while the validation loss rises, it may indicate that the learning rate is too high or the model is overfitting. At this time, the learning rate can be adjusted smaller (such as from 0.01 to 0.001). A feasible example is to use the Adam optimizer, with a batch size of 16 and a learning rate of 1e-5. Observing the training time, gradient changes, and early stopping method (such as setting to stop training when the validation loss does not decrease within 5 epochs) can avoid overfitting and improve training efficiency.

[0093] Step S5 is to post-process the extraction results, including format verification of specific fields (such as date format parsing) and similarity detection and filtering of redundant results, to ensure the integrity and accuracy of the output information.

[0094] The engineering drawing title block recognition method based on multi-modal information extraction provided in this embodiment combines cross-modal information extraction of text semantics, images, and Layout layout information, has the ability to directly perform end-to-end information extraction from pictures, and has high universality. Even if there are differences in the position, format, content, etc. of the drawing title block, it also has high accuracy in extracting information.

[0095] Embodiment 2

[0096] See Figure 4 , this embodiment provides an engineering drawing title block recognition system based on multi-modal information extraction, which is used to implement the engineering drawing title block recognition method based on multi-modal information extraction described in Embodiment 1, including,

[0097] An original picture acquisition module, which is used to acquire an engineering drawing image and preprocess the image;

[0098] A Schema definition module, which is used to acquire a data structure Schema that sets the control extraction target and relationship mode, including a binary data structure and / or a triple data structure;

[0099] A title block detection module, which is used to obtain the bounding box coordinates of each title block, and map the detected bounding box coordinates from the coordinate system of the preprocessed image back to the coordinate system of the original image, so as to crop out the corresponding title block area image from the original image;

[0100] A text recognition module, which is used to perform text recognition on the cropped title block image and extract all text contents and corresponding text box coordinate information;

[0101] A multi-modal information extraction module, which is used to input the cropped title block image and the text recognition result into a trained multi-modal information extraction model, and perform information extraction according to the data structure Schema that sets the control extraction target and relationship mode, and output a structured extraction result;

[0102] A post-processing module, which is used to post-process the extraction result, including format verification of specific fields (such as date format parsing) and similarity detection and filtering of redundant results, to ensure the integrity and accuracy of the output information.

[0103] The above complete cross-modal information extraction system improves the accuracy of information extraction from the complete process of image preprocessing, object detection, OCR recognition to cross-modal information extraction and post-processing.

[0104] The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

[0105] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0106] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the steps as sequential processes, many of the steps can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operations are completed, but there can also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.

Claims

1. An engineering drawing title block recognition method based on multi-modal information extraction, characterized in that including, obtaining an engineering drawing image and preprocessing the image, and setting a data structure Schema for controlling extraction targets and relationship patterns, including a binary tuple data structure and / or a triple tuple data structure; detecting the title block area in the drawing through a trained title block detection model to obtain the bounding box coordinates of each title block, and mapping the detected bounding box coordinates from the coordinate system of the preprocessed image back to the coordinate system of the original image, so as to crop the corresponding title block area image from the original image; performing text recognition on the cropped title block image to extract the text content and the corresponding text box coordinate information; constructing a multi-modal information extraction model and inputting the cropped title block image and the text recognition result into the trained multi-modal information extraction model, and performing information extraction according to the set data structure Schema for controlling extraction targets and relationship patterns, and outputting a structured extraction result; the constructing of the multi-modal information extraction model includes, annotating the drawing detected by the title block detection model and constructing a training data set, and the annotation includes binary tuple title block annotation and triple tuple title block annotation; obtaining a drawing image and performing recognition to extract the text content and two-dimensional layout information of the image, and the two-dimensional layout information is the bounding box coordinate information of the text sequence; constructing a structured pattern guide prompt based on the pre-set structured Schema information to guide the multi-modal information extraction model to perform specific type of information extraction; fusing the text vector and the image vector to construct a multi-modal input vector; building a multi-modal model architecture, and taking the output of the multi-modal model as input and connecting it to two independent feed-forward neural networks, and the two independent feed-forward neural networks are respectively used to predict the start position and the end position of the target value; wherein, building the multi-modal model architecture includes using the multi-layer Transformer architecture of LayoutLMv3 to capture complex features and performing multi-layer Transformer stacking, and each layer of Transformer gradually extracts cross-modal context representations through self-attention and a feed-forward network; each layer of Transformer includes a multi-head self-attention mechanism, a position bias, and a fully-connected feed-forward network; the multi-head self-attention mechanism is used to capture the interaction information among the text, the image, and the structured pattern guide prompt; the position bias is used to introduce the semantic one-dimensional relative position and the spatial two-dimensional relative position to enhance the alignment ability between modalities; the fully-connected feed-forward network is used to perform non-linear transformation on the output of the multi-head self-attention mechanism to extract and transform features; training the model based on the constructed data set, and evaluating the model structure according to the validation set divided in the training data set and adjusting the model hyperparameters; performing post-processing on the extraction result, including format verification of specific fields and similarity detection and filtering of redundant results.

2. The method for identifying the title block of an engineering drawing based on multi-modal information extraction according to claim 1, wherein, the preprocessing of the image includes decoding the image into an array format, adjusting the image size, and normalizing.

3. The method for identifying the title block of an engineering drawing based on multimodal information extraction according to claim 1, wherein, the training process of the title block detection model includes, Collect real engineering drawings, create labels according to different types, and annotate the images. Draw rectangular boxes for each drawing label area and assign labels. Each annotated file after annotation contains the position coordinates and class ID of each object; Divide the annotated dataset into a training set, a validation set, and a test set according to a preset ratio to train an open-source detection model, adjust the model parameters, and build a drawing label detection model.

4. The engineering drawing title block recognition method based on multi-modal information extraction according to claim 1, characterized in that After cropping the corresponding drawing label area image from the original image, it further includes Performing directional geometric correction and super-resolution image enhancement processing on the drawing label area image.

5. The engineering drawing title block recognition method based on multi-modal information extraction according to claim 1, wherein The binary drawing label annotation includes annotating the key-value pairs key-value in each pair of binary drawing label information. The triple drawing label annotation includes annotating the key-value pairs key-value in the triple drawing label information and associating Link to the Subject.

6. The engineering drawing title block recognition method based on multi-modal information extraction according to claim 1, wherein The constructing a structured schema-guided prompt to guide the multi-modal information extraction model to perform specific type of information extraction based on the preset structured Schema information includes Determining the content type of the prompt prefix and locating the fragments of specific semantic types; Including, SpotName for location name: representing the key in the binary or triple; AssoName for association name: representing the subject in the triple; Special Symbols: three special symbols used to separate and identify SpotName, AssoName, and the original text sequence, which are: [spot], [asso], and [text]; Associating the fragments according to the predefined Schema; Including, adding the corresponding special symbols [spot] and [asso] as prefixes to the two parts of SpotName and AssoName respectively, and then connecting them with the [text] special symbol to form the structured schema-guided prompt.

7. The engineering drawing title block recognition method based on multi-modal information extraction according to claim 1, characterized in that The fusing the text vector and the image vector to construct a multi-modal input vector includes Obtaining the text vector, which is the fusion of the word vector and the position vector. The word vector is the extracted text content, and the position vector includes a one-dimensional position vector and the extracted two-dimensional layout information; the one-dimensional position vector refers to the position index of each word in the text sequence, and the two-dimensional layout information includes the x coordinate, y coordinate, width, and height of the center coordinate, and is normalized; the fusion method for obtaining the text vector is to add the word vector, the one-dimensional position vector of the word, and the two-dimensional position vector element by element; Obtaining the image vector, dividing the document image into image patches of a fixed size, and linearly projecting each patch into the same dimension as the text vector; at the same time, adding a one-dimensional position vector to each image patch to retain its order information in the image; Unifying the dimensions of the text vector and the image vector and performing splicing fusion to form a multi-modal input.

8. An engineering drawing title block recognition system based on multi-modal information extraction, characterized in that, For implementing the engineering drawing label recognition method based on multi-modal information extraction described in any one of the above claims 1-7, it includes An original picture acquisition module, which is used to acquire the engineering drawing image and preprocess the image; A data structure definition module, which is used to obtain a data structure Schema for setting control extraction targets and relationship patterns, including a binary tuple data structure and / or a triple tuple data structure; A drawing label detection module, which is used to detect the drawing label area in the drawing through a trained drawing label detection model, obtain the bounding box coordinates of each drawing label, and map the detected bounding box coordinates from the coordinate system of the preprocessed image back to the coordinate system of the original image, so as to crop out the corresponding drawing label area image from the original image; A text recognition module, which is used to perform text recognition on the cropped drawing label image and extract all text contents and corresponding text box coordinate information; A multi-modal information extraction module, which is used to input the cropped drawing label image and the text recognition result into a trained multi-modal information extraction model, and perform information extraction according to the data structure Schema for setting control extraction targets and relationship patterns, and output a structured extraction result; A post-processing module, which is used to perform post-processing on the extraction result, including format verification of specific fields and similarity detection and filtering of redundant results.

Citation Information

Patent Citations

  • Drawing text information extraction method, device and equipment

    CN115995092A

  • Deep neural network-based system for detection and classification of construction elements in construction engineering drawings

    US20240428351A1