Engineering drawing legend automatic matching method and system based on block detection and feature memory bank
Patent Information
- Application Number
- CN202610827673.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-25
AI Technical Summary
这种方法对图像噪声、旋转或缩放非常敏感,并且在复杂背景下(如密集的线条或文字重叠)容易出现错误
[0040]1、高召回与高精度兼顾:本发明有效解决了传统方法在工程图纸中难以兼顾召回率与分类精度的问题。既显著缓解了因输入分辨率受限导致的小目标漏检问题,又能够精准区分外观高度相似但语义不同的图例类别,避免了单一检测模型因特征表达不足造成的误分类,真正实现高效、可靠的图例实例定位。
Smart Images

Figure CN122821583A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent recognition of engineering drawings and computer vision technology, and in particular to an automatic matching method and system for engineering drawing legends based on block detection and feature memory. Background Technology
[0002] In engineering drawing processing, legends are standardized symbols used to represent different parts of a building, such as doors, windows, columns, and various types of walls. In practical applications, the implementation of scenarios such as automatically calculating quantities, reviewing design compliance with specifications, and intelligently searching for specific components typically requires a standard legend diagram to centrally identify and classify corresponding component instances across the entire engineering drawing.
[0003] However, existing methods encounter several challenges in addressing this problem. First, template-matching methods treat each legend as a fixed template, matching it by constructing a sliding window on the drawing. This approach is highly sensitive to image noise, rotation, or scaling, and is prone to errors in complex backgrounds (such as dense lines or overlapping text). Second, single-stage deep learning testing... Figure 1 The traditional method of identifying all components on a drawing at once is insufficient. Most detection models, limited by computational resources, must reduce image resolution during detection. This can lead to some small components in the original high-resolution drawing being overlooked, and similar components becoming difficult to distinguish, resulting in inaccurate classification. Finally, pure feature retrieval methods extract features from the legend and then search for the best match across the entire drawing. This approach is computationally very expensive on high-resolution drawings, and severe background interference can significantly impact accuracy.
[0004] The information disclosed in this background section is intended only to enhance the understanding of the general background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide an automatic matching method and system for engineering drawing legends based on block detection and feature memory, so as to solve the technical problems existing in the prior art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] An automatic matching method for engineering drawing legends based on block detection and feature memory includes the following steps:
[0008] S1. Modal data of a feature library synthesized based on legend images and their annotations;
[0009] S2. Construct training data based on drawings and train the detector;
[0010] S3. Based on engineering drawings, block-based high-recall detection generates candidate region modal data;
[0011] S4. Based on the detection candidate region and feature memory, fine-grained feature matching and fusion decision-making generate the final matching result.
[0012] Preferably, step S1 includes:
[0013] S11. Parse and standardize multi-format legend annotation data;
[0014] S12. Add contextual margins around the annotation area;
[0015] S13. Scale the legend image proportionally to the model input size;
[0016] S14. Load the DINOv2 model and extract the visual features of the legend;
[0017] S15. Perform L2 normalization on the feature vectors and construct a structured feature library;
[0018] S16. Save the preprocessed legend image and visualization results.
[0019] Preferably, step S2 includes:
[0020] S21. Cut high-resolution engineering drawings into fixed-size sub-drawings;
[0021] S22. Generate candidate regions for the legend using a heuristic algorithm;
[0022] S23. Manually correct the candidate boxes and complete the annotation;
[0023] S24. Perform engineering drawing-specific data augmentation on the training samples;
[0024] S25. Unify the annotation format to the YOLO standard format;
[0025] S26. Fine-tune the YOLO detection model and export weights in multiple formats.
[0026] Preferably, step S3 includes:
[0027] S31. Divide the engineering drawings into local sub-blocks according to the grid;
[0028] S32. Perform YOLO object detection inference in parallel on each sub-block;
[0029] S33. Convert the local detection box coordinates to the original image global coordinates;
[0030] S34. Perform cross-grid nonmaximum suppression and deduplication;
[0031] S35. Output the set of high-recall candidate regions and generate visualization results.
[0032] Preferably, step S4 includes:
[0033] S41. YOLO test results are directly adopted for structural stability categories;
[0034] S42. Perform fine-grained feature matching on easily confused categories;
[0035] S43. Apply a class-level weight penalty mechanism to highly obfuscated subclasses;
[0036] S44. Integrate decision-making and generate structured matching results;
[0037] S45. Generate global statistics and visualization verification files.
[0038] Preferably, a system is provided for the aforementioned automatic matching method of engineering drawing legends based on block detection and feature memory.
[0039] By adopting the above technical solution, the present invention has the following beneficial effects:
[0040] 1. Balancing High Recall and High Precision: This invention effectively solves the problem of traditional methods struggling to balance recall and classification accuracy in engineering drawings. It significantly alleviates the problem of missed detection of small targets due to limited input resolution, and can accurately distinguish between image categories that are highly similar in appearance but different in meaning. This avoids misclassification caused by insufficient feature representation in a single detection model, truly achieving efficient and reliable image instance localization.
[0041] 2. Strong Robustness and Stability: Addressing the common issues in engineering drawings such as numerous symbol variations, inconsistent drawing standards, and varying scanning quality, this invention designs a robustness enhancement strategy for real-world business scenarios. The adopted binary fusion decision logic not only improves the overall matching accuracy but also significantly enhances its generalization ability and operational stability under different projects, drawing specifications, and image quality conditions, ensuring long-term reliable operation in real-world engineering environments.
[0042] 3. High feasibility for engineering implementation: This invention fully considers the actual deployment environment, supports automatic switching between CPU and GPU to adapt to different hardware conditions; prioritizes loading local model weights, and automatically reverts to online download if missing, ensuring process continuity.
[0043] 4. Explainability and Traceability: This invention can automatically generate visual images and structured log files for each stage of the entire process, which facilitates development and debugging, result acceptance and subsequent iterative optimization, significantly reduces engineering integration and operation and maintenance costs, and ensures that the entire process is transparent, visible and easy to track and verify.
[0044] This invention not only solves the problems of poor recognition stability, low reasoning efficiency, and insufficient fine-grained discrimination ability in the prior art, but also provides an efficient and reliable solution for the automated processing of engineering drawings, and has broad application prospects. Attached Figure Description
[0045] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0046] Figure 1 This is a schematic diagram of a structured feature memory provided in an embodiment of the present invention;
[0047] Figure 2 This is an example image of a candidate region for generating a gridded detection map, provided in an embodiment of the present invention.
[0048] Figure 3 An example diagram of the final legend matching and recognition result provided in the embodiments of the present invention. Detailed Implementation
[0049] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0051] Combination Figures 1 to 3As shown, this embodiment provides a three-stage collaborative automatic matching method for engineering drawing legends. The overall process adopts an optimized workflow of "feature memory construction – block-based high-recall detection – fine-grained feature matching and fusion decision-making," combining the YOLO object detection model and DINOv2 visual feature extraction capabilities. While ensuring processing efficiency, it significantly improves the localization accuracy and fine-grained classification ability of legend instances in complex engineering drawings. The following detailed explanation, combined with actual input / output and key technical details, is provided.
[0052] S1 Modal data of a feature library synthesized based on legend images and their annotations;
[0053] This step primarily involves extracting standardized visual features from the user-provided legend images and annotations, and constructing a structured feature memory library. This provides a high-quality, traceable feature modality data foundation for subsequent legend matching and recognition tasks. See the appendix for details. Figure 1 .
[0054] This step uses standard legend images and their corresponding component annotation files as core inputs, and is compatible with multiple mainstream annotation formats such as Pascal VOCXML, LabelMe JSON, and YOLO TXT. First, all types of annotation files are uniformly parsed and converted into an internal general data structure containing category labels, bounding box coordinates, and original image size information. Then, a 20-pixel context margin is extended around each annotation area to preserve the semantic meaning of key symbols in the engineering drawings that depend on the neighborhood structure. The cropped legend image is scaled to 224×224 pixels proportionally to the original aspect ratio, while the minimum side is forced to be no less than 16 pixels to prevent texture degradation. Next, the pre-trained DINOv2-large model is loaded, and after normalization preprocessing by AutoImageProcessor, the image is input into the encoder to obtain a token sequence of [1,257,1024]. The class token is discarded, and global mean pooling is performed on 256 patch tokens to generate a 1024-dimensional compact feature vector. L2 normalization is then performed to ensure the stability of cosine similarity calculation. Finally, a hierarchical and structured feature memory is constructed to fully retain metadata for traceability, and all preprocessed legend images and visualization results with annotation boxes are saved to a temporary directory for manual review and debugging verification.
[0055] S11. Parse and standardize multi-format legend annotation data;
[0056] This step mainly involves uniformly parsing and structurally converting the various annotation formats input by the user, providing a consistent data interface for subsequent image cropping and feature extraction.
[0057] First, the annotation file provided by the user is read, and the corresponding parser is called according to its format type to extract the category name, bounding box coordinates and original image size of each legend instance. All annotation information is then mapped to a unified internal data structure to ensure that downstream modules can access key information by the same field.
[0058] S12. Add contextual margins around the annotation area;
[0059] This step mainly involves extending the margin of each legend bounding box by 20 pixels in four directions, providing support for preserving the structural information around the symbol and enhancing semantic integrity.
[0060] Considering that the recognition of some symbols in engineering drawings is highly dependent on neighboring graphic elements, a new cropping area is generated by extending the original bounding box outward by 20 pixels (limited by the image boundary), which effectively incorporates the necessary context and avoids the loss of key visual cues due to overly strict cropping.
[0061] S13. Scale the legend image proportionally to the model input size;
[0062] This step primarily involves scaling the legend area with margins to 224×224 pixels while maintaining the original aspect ratio, thus preventing excessive compression of extremely small legends and providing compliant and faithful input for the DINOv2 model.
[0063] This step systematically calculates the scaling ratio, aligning the long side of the image to 224 pixels and scaling the short side proportionally. If any side is smaller than 16 pixels after scaling, it is forcibly enlarged to 16 pixels while maintaining the aspect ratio, and the remaining areas are filled with grayscale or edges.
[0064] S14. Load the DINOv2 model and extract the visual features of the legend;
[0065] This step mainly involves loading the DINOv2-large model, image preprocessing, and feature encoding to generate a high-dimensional semantic representation.
[0066] Pre-trained weights are loaded from the local path first. The input image is standardized by AutoImageProcessor and then sent to the encoder. The output is a token sequence of [1,257,1024]. The 0th class token is discarded. Global average pooling is performed on the remaining 256 patch tokens in the channel dimension to obtain a 1024-dimensional feature vector.
[0067] S15. Perform L2 normalization on the feature vectors and construct a structured feature library;
[0068] This step mainly completes the standardization and organization of feature vectors, providing a numerically stable and structurally clear data foundation for efficient retrieval and similarity calculation.
[0069] For each 1024-dimensional feature vector, L2 normalization is performed to make it lie on the unit hypersphere. The feature vector is then integrated with its corresponding class label, original image path, bounding box coordinates, margin parameters and other metadata to build a hierarchical feature library.
[0070] S16. Save the preprocessed legend image and visualization results;
[0071] This step mainly completes the persistent storage of intermediate products, providing a visual basis for manual review, quality verification and debugging.
[0072] All cropped and scaled legend images are saved to a temporary directory, and additional visualizations with bounding boxes and category labels are generated to facilitate checking annotation accuracy, evaluating margin effects, or assisting in model iterative analysis.
[0073] S2 builds training data and trains the detector based on a small number of drawings;
[0074] This step mainly involves constructing a target detection training set adapted to the characteristics of the drawings from a small number of engineering drawing samples, and fine-tuning the YOLO detection model to provide high recall and robust detection capabilities for subsequent large-format block detection.
[0075] After the feature memory is built, a small number of representative engineering drawings are introduced. A complete process of "sample cropping—candidate region generation—manual correction—data augmentation—annotation standardization—model fine-tuning" is used to construct a target detection model specifically for engineering drawing scenarios. First, the original drawings are cut into multiple sub-drawings at fixed sizes. Then, heuristic methods such as weak detectors, local connected component analysis, and line clustering rules are used to automatically generate legend candidate regions. Annotators only need to correct the suggested boxes. Based on this, data augmentation strategies tailored to the characteristics of engineering drawings are implemented, including rotation, random scaling, local cropping, overlaying interfering text, and simulating scanning stains and wire mesh occlusion. The augmented samples are uniformly converted to YOLO format. Then, the YOLO model is fine-tuned based on this training set, employing strategies such as multi-scale training, class-balanced sampling, and small target focusing loss, focusing on optimizing the recall rate of small-scale legends. After training, ONNX and PyTorch weight and class configuration files are exported to ensure cross-platform deployment compatibility.
[0076] S21. Cut high-resolution engineering drawings into fixed-size sub-drawings;
[0077] This step mainly involves dividing the original large-format drawing into blocks to reduce the computational load on a single drawing and adapt to the input limitations of the detection model.
[0078] The high-resolution engineering drawings provided by the user are cropped according to a configurable a×a grid division strategy to generate multiple sub-image blocks, while retaining the original coordinate mapping relationship, which facilitates the stitching and positioning of subsequent detection results on the original image.
[0079] S22. Generate candidate regions for the legend using a heuristic algorithm;
[0080] This step mainly completes the automatic prediction of possible legend locations based on weakly supervised methods, providing efficient initial suggestions for manual annotation.
[0081] By combining lightweight weak detectors, local connectivity analysis, and line clustering patterns commonly found in engineering drawings, preliminary candidate boxes are generated on each sub-graph to cover potential legend locations, significantly reducing the workload of manual annotation.
[0082] S23. Manually correct the candidate boxes and complete the annotation;
[0083] This step mainly involves manual review and refinement of the candidate regions generated by the system to ensure the accuracy of the training labels.
[0084] Annotators can add, delete, adjust boundaries, or correct category labels on candidate boxes output by heuristic algorithms using a visual interface to generate high-quality annotation results.
[0085] S24. Perform engineering drawing-specific data augmentation on the training samples;
[0086] This step mainly involves expanding the data to simulate the diversity and degradation factors of real drawings, thereby improving the model's generalization ability.
[0087] A series of targeted enhancement operations are applied to all labeled sub-images, including random rotation, scaling, local cropping, adding simulated text watermarks, and partial occlusion.
[0088] S25. Unify the annotation format to the YOLO standard format;
[0089] This step mainly involves standardizing various annotation results into the normalized coordinate format required for YOLO training, ensuring training consistency.
[0090] The manually corrected annotation files were uniformly converted into YOLO TXT format, which means that each line contains the category ID, the normalized center point x / y coordinates, the normalized width and height, and all coordinates are normalized relative to the subplot size.
[0091] S26. Fine-tune the YOLO detection model and export weights in multiple formats;
[0092] This step mainly involves training the YOLO model on the engineering drawing training set and outputting a multi-platform compatible model file.
[0093] Load the pre-trained YOLO backbone network and fine-tune it on the constructed training set. Use optimization strategies such as multi-scale input, focus loss or EIoU to improve the recall rate of small targets. Enable class balanced sampling during training. After training, export the ONNX model, PyTorch native weight and class name configuration file.
[0094] S3 generates candidate region modal data based on block-based high-recall detection using engineering drawings.
[0095] This step primarily involves performing gridded block detection on high-resolution engineering drawings. Through parallel inference and cross-block deduplication, it generates a set of candidate regions with high recall, low redundancy, and accurate coordinates, providing high-quality input for subsequent fine-grained feature matching. See the appendix for details. Figure 2 .
[0096] Faced with ultra-high resolution engineering drawings with large target scales and complex background interference, a configurable a×a grid partitioning strategy is adopted to uniformly divide the entire image into multiple local sub-blocks, and the global offset of the top-left corner of each sub-block in the original image is recorded. All sub-images are fed in parallel into a pre-trained YOLO detection model. During inference, a lower confidence threshold is used to improve the recall rate of small targets, and the maximum number of detection boxes output per block is limited to n. After detection, all local bounding boxes are uniformly transformed back to the global coordinate system of the original image based on the sub-block offset. Since adjacent sub-blocks overlap, the same physical target may be detected multiple times. Therefore, cross-grid non-maximum suppression (Cross-grid NMS) is introduced: all detection boxes are first sorted in descending order of confidence, and then duplicates with an IoU exceeding the threshold of the retained boxes are removed one by one. The final output candidate region set has the characteristics of high recall, low redundancy and accurate localization, and a visualization result with category color, label and confidence is generated and saved to a temporary directory.
[0097] S31. Divide the engineering drawings into local sub-blocks according to the grid;
[0098] This step mainly completes the regular block processing of high-resolution drawings, providing sub-graph units that are adapted to the model input size for subsequent parallel detection.
[0099] Based on the preset a×a grid parameters, the original drawing is evenly divided into a² sub-image blocks, each with a size of (H / a)×(W / a), and the offset of the upper left corner of each sub-block in the original drawing coordinate system (x0, y0) is recorded synchronously.
[0100] S32. Perform YOLO object detection inference in parallel on each sub-block;
[0101] This step primarily involves using the trained YOLO model to perform efficient and high-recall legend detection on all subgraphs.
[0102] Load locally trained YOLO weights, perform forward inference independently for each sub-block, set the confidence threshold to a low value (e.g., 0.2), and limit the output of a single sub-block to a maximum of n detection boxes (e.g., 50), balancing efficiency and detection capability.
[0103] S33. Convert the local detection box coordinates to the original image global coordinates;
[0104] This step mainly completes the spatial alignment of the detection results, ensuring that all candidate boxes are accurately positioned in the original image coordinate system.
[0105] Add the global offset (x0, y0) of the corresponding sub-block to the detected bounding box within each sub-block, and map it back to the absolute coordinate system of the original drawing to generate a list of candidate regions under a unified spatial reference.
[0106] S34. Perform cross-grid nonmaximum suppression and deduplication;
[0107] This step primarily aims to eliminate duplicate detections caused by overlapping blocks, thereby improving the simplicity and reliability of the candidate region set.
[0108] Sort all detection boxes in global coordinates from highest to lowest confidence and iterate through them sequentially: if the IoU between the current box and any of the retained boxes exceeds a preset threshold (e.g., 0.5), it is determined to be a duplicate and is removed.
[0109] S35. Output the set of high-recall candidate regions and generate visualization results;
[0110] This step mainly completes the structured output of the test results and provides visualization assistance for quality verification and debugging.
[0111] The deduplicated candidate regions are output in a structured format, including fields such as category, confidence level, and global coordinates. At the same time, all detection boxes are overlaid on the original drawing, with different categories marked with different colors, and the visualization image is saved to a temporary directory.
[0112] S4 generates the final matching result based on fine-grained feature matching and fusion decision-making using detected candidate regions and feature memory.
[0113] This step primarily involves classifying high-recall candidate regions. A binary fusion decision strategy is used to directly adopt YOLO results for structurally stable categories, while DINOv2 fine-grained feature matching is performed on easily confused categories with a class-level penalty mechanism. Finally, high-precision, interpretable final legend matching results are generated. See the appendix for details. Figure 3 .
[0114] After obtaining the high-recall candidate regions output in step S3, the optimal discrimination path is dynamically selected based on the semantic and geometric characteristics of the legend category: for categories with stable geometric structures and significant appearance differences (such as doors, windows, and columns), the category and confidence score output by YOLO are directly adopted as the final judgment; while for fine-grained, highly similar categories (especially walls and their subcategories), a refined feature matching process is initiated: the corresponding detection region is cropped from the original drawing, and the preprocessing logic used in building the feature memory in step S1 is strictly reused; subsequently, the same DINOv2 as in step S1 is used. The -large model extracts 1024-dimensional L2 normalized feature vectors and retrieves only all subclass instances under the same major class from the feature memory to calculate cosine similarity. For some subclasses that are prone to mismatch due to local texture similarity, a class-level weight penalty mechanism is introduced to multiply their initial similarity by a penalty coefficient less than 1. Finally, YOLO confidence and feature similarity are fused to generate the final class label and comprehensive confidence for each candidate region. The output includes a structured JSON file, a global statistics file stats.json, and a visualization image with colored annotation boxes.
[0115] S41. YOLO test results are directly adopted for structural stability categories;
[0116] This step primarily involves skipping feature matching for geometrically clear and visually distinct legend categories (such as doors, windows, and columns), and directly outputting YOLO predictions as the final decision, thereby improving efficiency and avoiding redundant calculations.
[0117] The candidate region is determined to belong to the category based on the preset "high confidence category list". If a match is successful and the YOLO confidence score is higher than the threshold (e.g., 0.5), its category and confidence score are directly used as the final result.
[0118] S42. Perform fine-grained feature matching on easily confused categories;
[0119] This step mainly completes the high-precision identification of fine-grained subclasses such as walls. By reusing the preprocessing and feature extraction process during the construction of the legend library, the consistency matching between the query and the library is achieved.
[0120] For candidate regions not processed by S41, their detection regions are cropped from the original drawings, the preprocessing logic of the S1 stage is strictly applied, and a 1024-dimensional L2 normalized feature vector is extracted using the DINOv2-large model. Cosine similarity is then calculated within the same class range in the feature memory.
[0121] S43. Apply a class-level weight penalty mechanism to highly obfuscated subclasses;
[0122] This step mainly completes the similarity correction of the mismatched subclasses and suppresses overfitting and misjudgment caused by local texture similarity.
[0123] This step maintains a "high-confusion subclass list". If the current matching target belongs to this list, its original similarity is multiplied by a preset penalty coefficient (such as 0.85). This coefficient is based on the statistical analysis of mismatch frequency in real engineering data.
[0124] S44. Integrate decision-making and generate structured matching results;
[0125] This step mainly integrates the YOLO results with the feature matching results to output standardized and parsable final matching data.
[0126] The results are processed according to the category paths: the S41 path directly retains the YOLO output; the S42–S43 paths use the maximum similarity after penalty as the confidence level and assign corresponding subclass labels; all results are organized into a structured JSON format according to the drawing, which is fully compatible with the LabelMe standard.
[0127] S45. Generate global statistics and visualization verification files;
[0128] This step mainly completes the output quality assessment indicators and visualization results, supporting manual review and system debugging.
[0129] Summarize the matching results of all drawings and generate a stats.json file, which records key indicators such as total number of detections, distribution of each main / subclass, matching success rate, and average confidence. At the same time, the final category label is drawn on the original drawing with a colored bounding box and saved to a temporary directory.
[0130] In summary, existing technologies in the field of automatic parsing of engineering drawings mainly rely on two types of methods: traditional template matching or rule engines, which are sensitive to legend deformation, scaling, and occlusion, and have poor generalization ability; and end-to-end deep learning detection models (such as Faster R-CNN and YOLO series), which, although able to locate targets, are insufficient in fine-grained category differentiation and are limited by input resolution, resulting in low recall rates for small-scale components. Furthermore, existing solutions generally do not consider the unique characteristics of engineering drawings, such as high resolution, dense layout, and symbol context dependencies, leading to high false negative rates and frequent misclassifications in practical applications.
[0131] To address the aforementioned problems, this invention integrates a three-stage collaborative architecture with several key technologies, achieving the following significant advantages and technical effects:
[0132] 1. Significantly improves the detection recall rate of small target components. The grid-based detection mechanism avoids the loss of small target information caused by downsampling of the entire drawing. Real-world testing shows that, on actual engineering drawings, the recall rate for small components such as wall segments and micro-devices with a width of less than 32 pixels is significantly improved compared to traditional whole-drawing YOLO detection. This solves a core problem in the automated processing of CAD and scanned drawings—the problem of missed detection of small targets—providing a complete data foundation for subsequent BIM modeling or quantity surveying.
[0133] 2. Achieve high-precision, fine-grained legend classification, effectively suppressing confusion and misjudgment. By introducing DINOv2 feature matching and a class-level weight penalty mechanism, the classification accuracy of wall subclasses is significantly improved. This overcomes the limitation of traditional detection models that can only identify coarse-grained categories, achieving stable and reliable subclass differentiation for the first time in an engineering scenario, meeting the requirements of professional design specifications for the semantic accuracy of legends.
[0134] 3. Balancing processing efficiency and recognition accuracy, it possesses strong engineering practicality. Employing a binary fusion decision strategy, it directly adopts YOLO results for targets with clear structures, saving feature computation overhead; for the remaining complex targets, it uses a high-cost matching process. The average processing time per image across the entire process is controlled within 15 seconds, far superior to pure feature matching schemes. It significantly improves throughput efficiency without sacrificing accuracy, meeting industrial-grade batch processing needs and avoiding the dilemma of "high accuracy but unusable" in practical applications.
[0135] 4. The system is highly robust, supporting offline private deployment and full-process traceability. Technical benefits: The system requires no internet connection and supports adaptive CPU / GPU operation; output results include the judgment method, confidence score, and visual evidence, facilitating manual review and error attribution. It meets the stringent data security and process auditing requirements of industries such as construction, power, and rail transportation, addressing the pain point that existing tools cannot be applied to classified projects.
[0136] In summary, this invention not only outperforms existing technologies in key technical indicators such as small target recall rate, fine-grained classification accuracy, and processing efficiency, but also achieves a leap from "laboratory algorithm" to "industrial-grade tool" through modular decoupling design and engineering-friendly output. It has outstanding substantive features and significant progress, and is applicable to a wide range of scenarios such as architectural design, construction drawing review, and BIM forward design.
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatic matching of engineering drawing legends based on block detection and feature memory, characterized in that, Includes the following steps: S1. Modal data of a feature library synthesized based on legend images and their annotations; S2. Construct training data based on drawings and train the detector; S3. Based on engineering drawings, block-based high-recall detection generates candidate region modal data; S4. Based on the detection candidate region and feature memory, fine-grained feature matching and fusion decision-making generate the final matching result.
2. The automatic matching method for engineering drawing legends based on block detection and feature memory as described in claim 1, characterized in that, Step S1 includes: S11. Parse and standardize multi-format legend annotation data; S12. Add contextual margins around the annotation area; S13. Scale the legend image proportionally to the model input size; S14. Load the DINOv2 model and extract the visual features of the legend; S15. Perform L2 normalization on the feature vectors and construct a structured feature library; S16. Save the preprocessed legend image and visualization results.
3. The automatic matching method for engineering drawing legends based on block detection and feature memory as described in claim 1, characterized in that, Step S2 includes: S21. Cut high-resolution engineering drawings into fixed-size sub-drawings; S22. Generate candidate regions for the legend using a heuristic algorithm; S23. Manually correct the candidate boxes and complete the annotation; S24. Perform engineering drawing-specific data augmentation on the training samples; S25. Unify the annotation format to the YOLO standard format; S26. Fine-tune the YOLO detection model and export weights in multiple formats.
4. The automatic matching method for engineering drawing legends based on block detection and feature memory as described in claim 1, characterized in that, Step S3 includes: S31. Divide the engineering drawings into local sub-blocks according to the grid; S32. Perform YOLO object detection inference in parallel on each sub-block; S33. Convert the local detection box coordinates to the original image global coordinates; S34. Perform cross-grid nonmaximum suppression and deduplication; S35. Output the set of high-recall candidate regions and generate visualization results.
5. The automatic matching method for engineering drawing legends based on block detection and feature memory as described in claim 1, characterized in that, Step S4 includes: S41. YOLO test results are directly adopted for structural stability categories; S42. Perform fine-grained feature matching on easily confused categories; S43. Apply a class-level weight penalty mechanism to highly obfuscated subclasses; S44. Integrate decision-making and generate structured matching results; S45. Generate global statistics and visualization verification files.
6. A system, characterized in that, The method for automatically matching engineering drawing legends based on block detection and feature memory as described in any one of claims 1 to 5.