Engineering drawing label identification method and system based on multi-modal information extraction
Through a multimodal information extraction method, combined with text semantics, images and Layout layout information, the difficulty of extracting image format drawing files and drawing information is solved, and the information extraction effect with high accuracy and universality is achieved.
Patent Information
- Application Number
- CN202510450366.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The prior art has difficulties in extracting drawing information on drawing files in image formats, including the lack of natural semantics in OCR processing results, the lack of unified standards, the poor matching of rules or templates, and the inability of traditional methods to complete information extraction in the form of bi-tuple and triple at the same time.
A method based on multimodal information extraction is adopted, combining text semantics, images and Layout layout information, cross-modal information extraction models are used to extract cross-modal information, and the structured extraction results are output.
It improves the accuracy of information extraction and is highly universal. Even if there are differences in the position, format, content, etc. of the drawings, accurate information can be extracted, and unified extraction of quadruple and triple.
Smart Images

Figure CN119964171A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence deep learning and image processing technology, and in particular to a method and system for identifying engineering drawing labels based on multimodal information extraction. Background Art
[0002] In the fields of construction engineering, mechanical manufacturing, electronic engineering and other industries, design drawings, as an important carrier for recording and communicating design plans, play an important role in project management, construction execution, and post-maintenance. With the increase in the number and types of drawings, the retrieval, management and archiving of these drawings have become the key tasks of the information construction of relevant enterprises and institutions. Usually, design drawings are drawn using CAD (Computer-Aided Design) software, which contains rich graphic information and text descriptions. There are generally corresponding drawing labels designed on the drawings to record the project name, drawing name, drawing number, designer, reviewer, date of formation and other key information of the drawing. How to extract these key information from the drawing labels in a structured manner is an important part of the information management of drawings.
[0003] There are two main directions for the recognition and extraction of drawing labels: one is to recognize the vector information (such as line primitives and text primitives) of CAD vector drawings, and the other is to recognize drawings converted into image formats using technologies in the fields of vision and NLP (Natural Language Processing). Considering that the original CAD files cannot be obtained in all cases, in some scenarios only the image format files of the design drawings can be collected and obtained, the applicability of extracting label information from drawings in image format is more extensive.
[0004] At present, there are still some difficulties in extracting label information from drawing files in image format. First, the label itself is mostly in the form of tables and other forms with layout information. The results directly obtained by OCR (Optical Character Recognition) processing do not have complete natural semantics. The method of OCR processing and NLP information extraction usually cannot obtain accurate output. Second, there is a lack of unified standards among different mapping units and departments. The location, format, content, etc. of drawing labels are different. Therefore, the use of rule or template matching methods cannot achieve high universality on a wide variety of label types. Third, the information relationship in the drawing label has not only the binary form of key-value pairs but also the more complex triple form. Traditional extraction methods cannot complete these two forms of information extraction at the same time. Summary of the invention
[0005] Based on the above background, the purpose of the present invention is to provide a method and system for structured extraction of drawing label information in engineering drawings based on multimodal information extraction technology, combining text semantics, images and Layout information for cross-modal information extraction, and improving the accuracy of information extraction.
[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a method for identifying an engineering drawing label based on multimodal information extraction, comprising: Acquire engineering drawing images and pre-process the images, and set a data structure Schema for controlling extraction targets and relationship patterns, including a binary data structure and / or a triple data structure; The label area in the drawing is detected by the trained label detection model to obtain the bounding box coordinates of each label, and the detected bounding box coordinates are mapped from the coordinate system of the preprocessed image back to the coordinate system of the original image to crop the corresponding label area image from the original image; Perform text recognition on the cropped label image to extract the text content and the corresponding text box coordinate information; A multimodal information extraction model is constructed and the cropped label image and text recognition results are input into the trained multimodal information extraction model. Information is extracted according to the set data structure Schema that controls the extraction target and the relationship pattern, and structured extraction results are output.
[0007] In a second aspect, the present invention provides an engineering drawing label recognition system based on multimodal information extraction, comprising: The original image acquisition module is used to acquire engineering drawing images and pre-process the images; A data structure definition module is used to obtain the data structure Schema of the set control extraction target and relationship model, including a binary data structure and / or a triple data structure; The label detection module is used to detect the label area in the drawing through the trained label detection model, obtain the bounding box coordinates of each label, and map the detected bounding box coordinates from the coordinate system of the preprocessed image back to the coordinate system of the original image, so as to crop the corresponding label area image from the original image; The text recognition module is used to perform text recognition on the cropped label image and extract all text content and corresponding text box coordinate information; The multimodal information extraction module is used to build a multimodal information extraction model and input the cropped label image and text recognition results into the trained multimodal information extraction model, and extract information according to the set control extraction target and the data structure Schema of the relationship model, and output the structured extraction results; The post-processing module is used to post-process the extraction results, including format verification of specific fields and similarity detection and filtering of redundant results.
[0008] The beneficial effects of the present invention are as follows: The engineering drawing label recognition method based on multimodal information extraction provided by the embodiment of the present invention combines text semantics, images and Layout information to perform cross-modal information extraction. It has the ability to extract information directly from pictures end-to-end and has strong universality. Even if there are differences in the position, format, content, etc. of the drawing labels, accurate information can be extracted.
[0009] At the same time, the present invention provides a unified extraction framework, which can flexibly define extraction targets through Prompt construction, support unified extraction of tuples and triples, and adapt to various relationship types and scenario requirements without additional adjustments to the model.
[0010] The embodiment of the present invention further proposes a complete process from image preprocessing, target detection, OCR recognition to cross-modal information extraction and post-processing, providing a complete cross-modal information extraction system and improving the accuracy of information extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0012] Figure 1 A flowchart of a method for identifying an engineering drawing signature based on multimodal information extraction provided by an embodiment of the present invention; Figure 2 A schematic diagram of the training process of the image signature detection model provided in an embodiment of the present invention; Figure 3 A schematic diagram of the training process of the multimodal information extraction model provided by an embodiment of the present invention; Figure 4 A schematic diagram of an engineering drawing label recognition system based on multimodal information extraction provided by an embodiment of the present invention; Figure 5 is an example diagram of a two-tuple annotation provided by an embodiment of the present invention; Figure 6 This is a triplet annotation example diagram provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0013] In order to further understand the present invention, preferred embodiments of the present invention are described below in conjunction with examples. However, it should be understood that these descriptions are only for further illustrating the features and advantages of the present invention, rather than limiting the claims of the present invention.
[0014] Terminology explanation: Schema refers to the data structure that controls the extraction target and relationship model; Prompt, a prompt word used to guide the model to perform specific tasks or generate specific information; SSI (Structural Schema Instructor) is a prompt mechanism used to guide the model to complete specific information extraction tasks. By adding specific prompts before inputting data, key information such as task requirements and data types are embedded into the model input, helping the model to better understand and process multimodal data. Transformer, a neural network architecture based on self-attention mechanism; LayoutLMv3, a multimodal Transformer model that combines text and layout information; YOLOv8 (You Only Look Once v8), a deep learning model for object detection; LabelImg, an image object detection labeling tool that supports multiple labeling formats; Label Studio, an open source data labeling tool that supports labeling of multiple data types; Batch size is the number of samples input at one time during training, which affects the training efficiency and stability.
[0015] Example 1
[0016] See also Figure 1 This embodiment provides a method for identifying engineering drawing labels based on multimodal information extraction, including: S1: Acquire an engineering drawing image and pre-process the image, and set a data structure Schema for controlling the extraction target and relationship pattern, including a binary data structure and / or a triple data structure; S2: Detect the label area in the drawing using the trained label detection model to obtain the bounding box coordinates of each label, and map the detected bounding box from the coordinate system of the preprocessed image back to the coordinate system of the original image using a scaling ratio, so as to crop the corresponding label area image from the original image; S3: Perform text recognition on the cropped label image to extract the text content and the corresponding text box coordinate information; S4: Construct a multimodal information extraction model and input the cropped label image and the recognized text content into the trained multimodal information extraction model, and extract information according to the set control extraction target and the data structure of the relationship model, and output the structured extraction result.
[0017] S5: Post-process the extraction results to ensure the integrity and accuracy of the output information.
[0018] Specifically: Step S1 specifically includes: S101: Acquire an engineering drawing image and pre-process the image.
[0019] The obtained engineering drawing images, such as TIFF or PNG format images, should ensure that the drawing clarity meets the requirements of subsequent processing. Perform preprocessing operations on the input engineering drawings, including decoding the image into array format, adjusting the image size (resize), and normalizing (normalize) to facilitate subsequent model processing.
[0020] S102: Obtain a predefined data structure Schema that controls the extraction target and relationship model.
[0021] In a two-tuple relationship, the schema is a list of keys, for example Figure 5 If you want to extract the tuple information of review and verification, the schema is ["review", "verification"], Figure 6 If you want to extract triple information, the schema is ["Figure Name": ["Figure Number", "Date of Production"]]. The schema defines the types of information that need to be extracted and associated, and this information is usually parsed into tree nodes, each node represents an entity or relationship type, which is used to construct SSI prompts and guide subsequent information extraction.
[0022] Step S2 is to detect the label area in the drawing through the trained label detection model to obtain the category label (label) and bounding box coordinates (bbox) of each label, and map the detected bounding box from the coordinate system of the preprocessed image back to the coordinate system of the original image using a scaling ratio, thereby cropping the corresponding label area image from the original image.
[0023] Preferably, the image of the label area can also be subjected to geometric correction such as orientation and image enhancement processing such as super-resolution technology. The cropped label area may have problems such as incorrect orientation and low resolution. Geometric correction can use image processing technology such as Hough transform or deep learning model to detect the orientation of the image, and then correct the image to the correct orientation through rotation operation; super-resolution can use pre-trained super-resolution models, such as ESRGAN (Enhanced Super-Resolution Generative Adversarial Network) model, to convert low-resolution images into high-resolution images while maintaining the details and realism of the image.
[0024] See also Figure 2 ,The training process of the image signature detection model specifically includes, S201: Drawing data preparation and labeling. Collect a large number of real engineering drawing images to ensure that different types of labels are covered, and formulate labels according to these different types, such as "construction engineering drawing labels", "water and power engineering drawing labels", "transportation engineering drawing labels", etc., or classify them according to their uses into "review type drawing labels", "standard type drawing labels", "catalog type drawing labels", etc.; then use the LabelImg open source tool to label the images, draw a rectangular box for each label area and assign a label. After the labeling is completed, export it in the YOLO format label file, which contains the location coordinates of each object (center point coordinates x, y and width w, height h) and category ID.
[0025] S202: Construction of detection data set. The labeled data set is divided into training set, validation set and test set in a ratio of 70%:15%:15%. Since YOLO provides online data enhancement, there is no need to perform offline enhancement processing on the data set.
[0026] S203: Set up the training environment. Prepare the hardware environment and use servers equipped with NVIDIA GPUs, such as NVIDIA Tesla V100, RTX 3090, etc.; build the software environment and install necessary libraries and frameworks, such as Python, PyTorch, and Ultralytics.
[0027] S204: Parameter configuration and training. This part can be trained with a newer version of YOLO (such as v8 or above). The training process is as follows: modify the model training parameters such as batch size and learning rate. The batch size can be set to 16, the initial learning rate can be set to 0.001, cosine annealing decay is used, and the Adam optimizer is used. The loss function does not need to be modified. After completing the above configuration, start training.
[0028] S205: Model evaluation and saving. After the training is completed, if the amount of training data is large, the accuracy of the image label detection task can reach more than 97%; save the model weight obtained by training.
[0029] Step S3 is to use an OCR tool to perform text recognition on the label image, extract all text content and its corresponding text box coordinate information, that is, the bounding box coordinates of the text sequence, specifically including the center coordinates x, y, width w and height h.
[0030] Step S4 is based on inputting the cropped label image and text recognition results into the trained multimodal information extraction model, extracting information according to the set control extraction target and the data structure of the relationship model, and outputting a structured extraction result.
[0031] See also Figure 3 ,The training process of building a multimodal information extraction model specifically includes: S401: Data labeling and dataset construction. For engineering drawing images, after label detection by the label detection model, a large number of drawings with label areas without labeled data are obtained. Label Studio is used to label the label information. Figure 5 The two-tuple graph shown in the figure has a key-value pair for each pair of information. Figure 6 The triplet labels shown must be labeled with key-value pairs and linked to the subject. The labeled data set is divided into training set, validation set and test set in a ratio of 80%:10%:10%.
[0032] S402: Document data preprocessing: Obtain a document image and use an OCR tool to extract text content and corresponding two-dimensional layout information, namely, the bounding box coordinates of the text sequence, including the center coordinates x, y, width w, and height h.
[0033] S403: Constructing a structured schema guide. The core idea of the structured schema guide SSI is to use the structured Schema information of the task to construct a prefix prompt to guide the model to extract specific types of information.
[0034] Specifically, all information extraction (IE) tasks can be regarded as the transformation from text to structure, and this transformation can be decomposed into two operations: one is Spotting, which is used to locate fragments of specific semantic types; the other is Associating, which is used to associate these fragments according to a predefined Schema.
[0035] In the information extraction task, Schema is used to define the information structure that needs to be extracted. Through the predefined Schema, the model can more accurately extract relevant entities, relationships, and events from the text. Different information extraction IE tasks have different structural schemas. The Structural Schema Instructor (SSI) is used to guide the model to extract information according to a specific Schema. Specifically, the information about the need to locate (Spotting) and associate (Associating) in the Schema is embedded in the input sequence of the model in a specific format, which is equivalent to providing a clear instruction to the model, telling the model what type of information should be extracted and how to organize this information.
[0036] The implementation of the structured pattern guide SSI is mainly through constructing a prompt prefix containing a specific type of token. Token is the basic unit of the prompt prefix, which is used to convey key information such as task intent, data type, location information and relationship type to the model, guide the behavior of the model, and enable it to better complete the information extraction task. In this embodiment, SSI includes the following three types of content: SpotName: represents the key in a binary or triplet.
[0037] Association name AssoName: represents the subject in the triple.
[0038] Special Symbols: Three special symbols used to separate and identify SpotName, AssoName, and original text sequences: [spot], [asso], and [text].
[0039] Add the corresponding special symbols [spot] and [asso] as prefixes to SpotName and AssoName respectively, and then connect them with the special symbol [text] to form the SSI Prompt.
[0040] For example Figure 5 If you want to extract the two-tuple information of review and verification, the prompt word is "[spot] review [spot] verification [text]", Figure 6 If you want to extract triple information, the prompt is "[asso]Figure name[spot]Figure number[spot]Date of drawing[text]".
[0041] S404: Generate multimodal input, which specifically includes the following text vectors and image vectors, specifically: The text vector is a fusion of the word vector and the position vector, where the word vector is the text content extracted by the OCR tool in step S402, and the SSI Prompt has been integrated in the previous step S403, that is, the SSI Prompt is directly spliced in front of the text content extracted by the OCR tool; the position vector includes the 1D position vector and the 2D layout information extracted by the OCR tool. The 1D position vector is the position index of each word in the text sequence; the 2D layout information includes the x-coordinate, y-coordinate, width and height of the center coordinate, and it is normalized. The fusion method is to add the word vector, the one-dimensional position vector of the word, and the two-dimensional position vector element by element.
[0042] The image vector follows the processing principle of LayoutLMv3. The document image is divided into fixed-size image blocks. A linear projection vector is generated for each image block and mapped to the same dimension as the text vector. A 1D position vector is added to each image block to indicate its order in the image.
[0043] Finally, the text vector and the image vector are concatenated to form a multimodal input. Specifically, the cat operation is used to concatenate on the sequence dimension (dim=1). For example, the shape of the text embedding is (2, 10, 768), which means 2 samples, each with 10 tokens, and the embedding dimension of each token is 768; the shape of the image embedding is (2, 197, 768), which means 2 samples, each with 196 small blocks + 1 [CLS] token, and the dimension of each embedding is 768. The shape of the concatenated embedding is (2, 207, 768), where 207 is the sum of the number of text tokens (10) and the number of image small blocks (197). In this way, the embedding vectors of text and image are integrated into a unified multimodal input, which can be further passed to the multimodal Transformer model for processing.
[0044] S405: Multimodal Transformer model construction. Use the multi-layer Transformer architecture of LayoutLMv3 to capture complex features. Each layer contains: Multi-Head Self-Attention: Captures the interactive information between text, images, and SSI Prompt.
[0045] Position Bias: Introduces semantic 1D relative position and spatial 2D relative position to enhance inter-modal alignment capabilities.
[0046] Fully connected feed-forward neural network: performs nonlinear transformation on the output of the multi-head self-attention mechanism to further extract and transform features.
[0047] Then, multiple layers of Transformers are stacked. Specifically, a 12-layer Transformer encoder is used, the hidden layer size is set to 768, and the number of heads in the multi-head self-attention mechanism is 12 (12-head). Each layer of Transformer gradually extracts cross-modal context representations through self-attention and feedforward networks, deepening the model's understanding of multimodal data.
[0048] Perform joint multimodal encoding, that is, concatenate the vector representations of text, images, and pattern prompt words according to the sequence dimension to form a unified multimodal input sequence that is input into the Transformer model for processing.
[0049] S406: By adding two independent fully connected layers based on the Transformer output, the starting position and the ending position of the target value are predicted respectively.
[0050] Fully connected layer prediction and loss function definition.
[0051] Specifically, the output of the Transformer is used as input, and two independent feedforward neural networks, namely the fully connected layers, are connected to output the start and end positions of the predicted value respectively. For each key to be extracted, forward reasoning is performed to obtain the corresponding value.
[0052] During training, the real labels of the start position and the end position are obtained by fusing the annotation information of S401 and the OCR text recognition information of S402. The specific fusion process is: matching the value in the annotation information to the corresponding text in the OCR recognition result, thereby obtaining the start and end position indexes of the target value in the OCR text recognition result.
[0053] The cross entropy loss function is used to measure the difference between the predicted position and the actual position, and to update the model parameters through the optimizer. In this embodiment, the loss function is calculated using binary cross entropy, and the formula is as follows: ;
[0054] in, is the true label (ground truth), is the predicted value (probability value) of the model, which measures the difference between the probability distribution predicted by the model and the true distribution. It is expected that the model predicts a high probability for positive samples and a low probability for negative samples.
[0055] S407: Model fine-tuning. In this embodiment, the main model uses the pre-trained model of LayoutLMv3 as the initial weight, sets hyperparameters such as batch size and learning rate, starts training, and adjusts the hyperparameters based on the model structure evaluated by the validation set. After the training, the model is saved. The suitability of the hyperparameters is mainly evaluated by monitoring indicators such as the loss, accuracy, and learning curve of the training set and validation set. For example, if the training loss drops too fast and the validation loss rises, it may indicate that the learning rate is too high or the model is overfitting. At this time, you can try to reduce the learning rate (such as from 0.01 to 0.001). A feasible example is to use the Adam optimizer with a batch size of 16 and a learning rate of 1e-5. Observing the training time, gradient changes, and early stopping (such as stopping training when the validation loss no longer decreases within 5 epochs) can avoid overfitting and improve training efficiency.
[0056] Step S5 is to post-process the extraction results, including format verification of specific fields (such as date format parsing) and similarity detection and filtering of redundant results to ensure the integrity and accuracy of the output information.
[0057] The engineering drawing label recognition method based on multimodal information extraction provided in this embodiment combines the cross-modal information extraction of text semantics, images and Layout information, and has the ability to extract information directly from pictures end-to-end. It has high universality and can extract information with high accuracy even if there are differences in the position, format, content, etc. of the drawing labels.
[0058] Example 2 See also Figure 4 This embodiment provides an engineering drawing label recognition system based on multimodal information extraction, which is used to implement the engineering drawing label recognition method based on multimodal information extraction described in Example 1, including: The original image acquisition module is used to acquire engineering drawing images and pre-process the images; Schema definition module, used to obtain the data structure Schema of the set control extraction target and relationship mode, including a binary data structure and / or a triple data structure; A label detection module is used to obtain the bounding box coordinates of each label, and map the detected bounding box coordinates from the coordinate system of the preprocessed image back to the coordinate system of the original image, so as to crop the corresponding label area image from the original image; The text recognition module is used to perform text recognition on the cropped label image and extract all text content and corresponding text box coordinate information; The multimodal information extraction module is used to input the cropped label image and text recognition results into the trained multimodal information extraction model, extract information according to the set control extraction target and relational model data structure Schema, and output structured extraction results; The post-processing module is used to post-process the extraction results, including format verification of specific fields (such as date format parsing) and similarity detection and filtering of redundant results to ensure the integrity and accuracy of the output information.
[0059] The above complete cross-modal information extraction system has a complete process from image preprocessing, target detection, OCR recognition to cross-modal information extraction and post-processing, which improves the accuracy of information extraction.
[0060] The above embodiments are only used to help understand the method and core idea of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0062] It should be noted that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
Claims
1. A method for identifying engineering drawing labels based on multimodal information extraction, characterized in that: include, Acquire engineering drawing images and pre-process the images, and set a data structure Schema for controlling extraction targets and relationship patterns, including a binary data structure and / or a triple data structure; The label area in the drawing is detected by the trained label detection model to obtain the bounding box coordinates of each label, and the detected bounding box coordinates are mapped from the coordinate system of the preprocessed image back to the coordinate system of the original image to crop the corresponding label area image from the original image; Perform text recognition on the cropped label image to extract the text content and the corresponding text box coordinate information; Construct a multimodal information extraction model and input the cropped label image and text recognition results into the trained multimodal information extraction model, extract information according to the set control extraction target and relational model data structure Schema, and output structured extraction results; Post-process the extraction results, including format verification of specific fields and similarity detection and filtering of redundant results.
2. The method for identifying engineering drawing signatures based on multimodal information extraction according to claim 1 is characterized in that: The preprocessing of the image includes decoding the image into an array format, adjusting the image size and normalizing the image.
3. The method for identifying engineering drawing labels based on multimodal information extraction according to claim 1 is characterized in that: The training process of the image signature detection model includes: Collect real engineering drawings, create labels according to different types, and annotate the images. Draw a rectangular box for each label area and assign a label. After the annotation is completed, each annotation file contains the location coordinates and category ID of each object. The labeled data set is divided into training set, validation set and test set according to preset ratios to train the open source detection model, adjust the model parameters, and build a label detection model.
4. The method for identifying engineering drawing labels based on multimodal information extraction according to claim 1 is characterized in that: After cutting out the corresponding image label area image from the original image, the method further includes: The image in the label area is oriented and geometrically corrected, and the image is enhanced with super-resolution.
5. The method for identifying engineering drawing labels based on multimodal information extraction according to claim 1 is characterized in that: The construction of the multimodal information extraction model includes: Annotating drawings detected by the image label detection model and constructing a training data set, wherein the annotations include two-tuple image label annotations and three-tuple image label annotations; Acquire the drawing image and perform recognition to extract the text content and two-dimensional layout information of the image, where the two-dimensional layout information is the bounding box coordinate information of the text sequence; Based on the pre-set structured Schema information, a structured mode guide prompt word is constructed to guide the multimodal information extraction model to extract specific types of information; Fuse text vectors and image vectors to construct multimodal input vectors; Build a multimodal model architecture and use the output of the multimodal model as input to two independent feedforward neural networks, which are used to predict the start and end positions of the target value respectively. The model is trained based on the constructed dataset, and the model structure is evaluated and model hyperparameters are adjusted based on the validation set divided from the training dataset.
6. The method for identifying engineering drawing labels based on multimodal information extraction according to claim 5 is characterized in that: The two-tuple image labeling includes labeling the key-value pair key-value in each pair of two-tuple image label information, and the three-tuple image labeling includes labeling the key-value pair key-value in the three-tuple image label information and associating Link to the subject Subject.
7. The method for identifying engineering drawing labels based on multimodal information extraction according to claim 5 is characterized in that: The method of constructing a structured mode guide prompt word based on the preset structured Schema information to guide the multimodal information extraction model to extract a specific type of information includes: Determine the content type of the cue word prefix and locate the fragments of specific semantic type; Includes, positioning name SpotName: represents the key in a binary or triple; association name AssoName: represents the subject in a triple; special symbols SpecialSymbols: three special symbols used to separate and identify SpotName, AssoName and the original text sequence, namely: [spot], [asso] and [text]; Associate the fragments according to the predefined Schema; Including adding the corresponding special symbols [spot] and [asso] as prefixes to SpotName and AssoName respectively, and then connecting them with the special symbol [text] to form a structured pattern guide prompt word.
8. The method for identifying engineering drawing labels based on multimodal information extraction according to claim 5 is characterized in that: The fusing of text vectors and image vectors to construct a multimodal input vector includes: Get the text vector. The text vector is a fusion of the word vector and the position vector. The word vector is the extracted text content, and the position vector includes a one-dimensional position vector and extracted two-dimensional layout information. The one-dimensional position vector refers to the position index of each word in the text sequence. The two-dimensional layout information includes the x-coordinate, y-coordinate, width and height of the center coordinate, and is normalized. The fusion method of obtaining the text vector is to add the word vector, the one-dimensional position vector of the word, and the two-dimensional position vector element by element. Get the image vector, split the document image into fixed-size image blocks, and linearly project each block to the same dimension as the text vector; at the same time, add a one-dimensional position vector to each image block to preserve its order information in the image; Unify the dimensions of text vectors and image vectors and concatenate and fuse them to form multimodal input.
9. The method for identifying engineering drawing labels based on multimodal information extraction according to claim 5 is characterized in that: The multimodal model architecture is constructed including: Use LayoutLMv3’s multi-layer Transformer architecture to capture complex features and stack multiple layers of Transformers, where each layer of Transformer gradually extracts cross-modal context representations through self-attention and feed-forward networks; Each layer of Transformer contains a multi-head self-attention mechanism, position bias, and a fully connected feedforward network; Multi-head self-attention mechanism to capture the interactive information between text, image and structured pattern guide prompt words; Position bias is used to introduce semantic one-dimensional relative position and spatial two-dimensional relative position to enhance the inter-modal alignment capability; the fully connected feedforward network is used to perform nonlinear transformation on the output of the multi-head self-attention mechanism to extract and convert features.
10. An engineering drawing label recognition system based on multimodal information extraction, characterized in that: include, The original image acquisition module is used to acquire engineering drawing images and pre-process the images; A data structure definition module is used to obtain the data structure Schema of the set control extraction target and relationship model, including a binary data structure and / or a triple data structure; The label detection module is used to detect the label area in the drawing through the trained label detection model, obtain the bounding box coordinates of each label, and map the detected bounding box coordinates from the coordinate system of the preprocessed image back to the coordinate system of the original image, so as to crop the corresponding label area image from the original image; The text recognition module is used to perform text recognition on the cropped label image and extract all text content and corresponding text box coordinate information; The multimodal information extraction module is used to input the cropped label image and text recognition results into the trained multimodal information extraction model, extract information according to the set control extraction target and relational model data structure Schema, and output structured extraction results; The post-processing module is used to post-process the extraction results, including format verification of specific fields and similarity detection and filtering of redundant results.
Citation Information
Patent Citations
Drawing text information extraction method, device and equipment
CN115995092A
Deep neural network-based system for detection and classification of construction elements in construction engineering drawings
US20240428351A1
Cited By
Document content matching method and system based on multiple modes
CN120182990A
Engineering drawing intelligent identification method and system based on deep learning
CN120748003A
A Deep Learning-Based Intelligent Recognition Method and System for Engineering Drawings
CN120748003B
Method and device for automatically generating equipment data report based on drawing, and storage medium
CN120766305A
CAD drawing processing method and device, equipment and storage medium
CN120783351A