A kind of segmented guide carbon fiber reinforcement CAD drawing BoM generation method

CN122821584APending Publication Date: 2026-09-25SHANDONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611049039.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0006]为解决现有技术存在的CFRP加固工程材料统计中人工识图效率低、栅格化图纸难以直接解析、细长加固目标识别不稳定、工程文字与构件关联困难以及材料清单生成缺乏自动化支撑等问题,本发明提供一种分割引导的碳纤维加固CAD图纸BoM生成方法,旨在以栅格化CAD图纸为输入,不依赖DWG/DXF矢量图层、块属性或BIM语义信息,直接从图像像素层面识别CFRP相关构件和工程文字区域,并进一步结合OCR识别、几何属性提取、文本—实例关联和工程规则计算,自动生成结构化材料清单

Benefits of technology

1.不依赖结构化 CAD/BIM 语义。本发明直接从栅格化 CAD 图纸图像中提取目标实例和尺寸文字,即使原始 DWG/DXF 图层、块属性或 BIM 构件语义不可获得,也可以执行CFRP 加固区域识别与材料清单生成,适用范围更广。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821584A_ABST
    Figure CN122821584A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of bill of materials generation, and particularly relates to a kind of segmented guide's carbon fiber reinforcement CAD drawing BoM generation method.Method includes: the pretreatment and equal proportion overlap slice of carbon fiber reinforcement CAD drawing are obtained grid CAD drawing slice image;Grid CAD drawing slice image is input into the CFRP-TGNet network constructed, and instance mask and text detection frame are obtained;Text detection frame and material specification area are executed OCR text analysis and specification field extraction, and OCR recognition text is obtained;Instance mask, text detection frame and OCR recognition text are input into segmented guide's BoM generation process, and structured bill of materials is formed.The present application can directly identify CFRP related components and engineering text area from image pixel level, and further combine OCR recognition, geometric attribute extraction, text-instance association and engineering rule calculation to automatically generate structured bill of materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of bill of materials generation technology, specifically relating to a method for generating a carbon fiber reinforcement CAD drawing BoM guided by segmentation. Background Technology

[0002] The Bill of Materials (BoM) and quantity calculations are crucial for material procurement, construction organization, cost accounting, and quality verification in CFRP reinforcement projects. Their accuracy directly impacts project implementation efficiency and project management reliability. In practical engineering, carbon fiber reinforced polymer (CFRP) composite materials are widely used in the reinforcement and repair of concrete beams, slabs, columns, and other structural components due to their high strength, light weight, corrosion resistance, and ease of construction. The layout, statistics, procurement, and construction verification of components such as carbon fiber cloth, carbon fiber plates, strips, U-shaped hoops, anchors, and cement-based reinforcement areas typically require CAD design drawings as a basis.

[0003] The generation of the Bill of Materials (BoM) for CFRP reinforcement projects relies on geometric and semantic information from the drawings, including the location, length, width, area, number of layers, spacing, anchoring method, material specifications, and quantity rules of the reinforcement area. However, in current engineering practice, this information is mainly read, annotated, and summarized manually by engineers. This process is easily affected by dense annotations on the drawings, elongated reinforcement areas, overlapping text and lines, and interference from elements with similar colors. This results in problems such as long processing times, omissions, errors, and inconsistent statistical standards among different personnel, thus affecting the accuracy of material statistics and construction management.

[0004] Compared to general architectural floor plans, CFRP reinforcement drawings have stronger domain-specific characteristics. Target objects in these drawings are typically elongated strips, locally curved shapes, or small-scale discrete components, and overlap with dimension text, leader lines, color-coded lines, and structural outlines. Relying solely on manual drawing interpretation or general drawing recognition algorithms makes it difficult to simultaneously perform tasks such as fine-grained instance segmentation, dimension text localization, text-geometric instance association, and material quantity calculation according to engineering rules.

[0005] Therefore, it is necessary to study a segmentation-guided method for generating BoM (BoM) drawings for carbon fiber reinforcement CAD drawings, so as to provide a reliable data foundation for automatic BoM generation and subsequent consistency verification. Summary of the Invention

[0006] To address the problems of low efficiency in manual drawing interpretation, difficulty in directly parsing rasterized drawings, unstable identification of slender reinforcement targets, difficulty in associating engineering text with components, and lack of automated support for bill of materials generation in existing technologies for CFRP reinforcement engineering material statistics, this invention provides a segmentation-guided method for generating a Bill of Materials (BoM) from CAD drawings for carbon fiber reinforcement. This method aims to use rasterized CAD drawings as input, without relying on DWG / DXF vector layers, block attributes, or BIM semantic information, to directly identify CFRP-related components and engineering text areas at the image pixel level. Furthermore, it combines OCR recognition, geometric attribute extraction, text-instance association, and engineering rule calculation to automatically generate a structured bill of materials.

[0007] To achieve the above objectives, the present invention provides the following solution: A method for generating a BoM (Board of Name) of carbon fiber reinforcement CAD drawings guided by segmentation, the method comprising: Preprocessing and proportionally overlapping slicing of carbon fiber reinforced CAD drawings yields rasterized CAD drawing slice images. The rasterized CAD drawing slice image is input into the constructed CFRP-TGNet network to obtain instance masks and text detection boxes; OCR text parsing and specification field extraction are performed on the text detection box and material description area to obtain OCR-recognized text; The BoM generation process, guided by the instance mask, the text detection box, and the OCR-recognized text input segmentation, forms a structured bill of materials.

[0008] Preferably, the method for preprocessing and proportionally overlapping slices of carbon fiber reinforced CAD drawings to obtain rasterized CAD drawing slice images includes: Convert the carbon fiber reinforced CAD drawings into raster images; The raster image is preprocessed by performing white background normalization, noise removal, color space unification, and resolution adjustment to obtain a preprocessed image; The preprocessed image is sliced ​​with equal-proportion overlapping slices to obtain rasterized CAD drawing slice images.

[0009] Preferably, the CFRP-TGNet network includes: a high-resolution feature extraction backbone, a text-color decoupling module, a deformable strip geometry thinning module, and a detection / segmentation output head; The high-resolution feature extraction backbone is used to perform high-resolution feature extraction on the rasterized CAD drawing slice image to obtain multi-scale features. The text-color decoupling module is used to suppress interference from text, lines, and non-target primitives with similar colors in the multi-scale features to obtain clean features. The deformable strip geometry refinement module is used to enhance the purification features to obtain geometrically enhanced features. The detection / segmentation output head is used to generate instance masks and text detection boxes based on the geometric enhancement features.

[0010] Preferably, the text-color decoupling module performs interference suppression on the multi-scale features by text, leaders, and non-target primitives with similar colors to obtain clean features, including the following methods: Based on the multi-scale features, the color embedding branch of the text-color decoupling module is used to extract color perception features; Based on the multi-scale features, the text prediction branch of the text-color decoupling module is used to output a text / label probability map; The multi-scale features, the color perception features, and the text / label probability map are concatenated and fused through convolution to obtain a text-color interference map; Based on the multi-scale features and the text-color interference map, the text-color suppression branch and feature reweighting operation of the text-color decoupling module are used to obtain cleaned features.

[0011] Preferably, the deformable strip geometry refinement module enhances the purification features to obtain geometrically enhanced features using the following methods: Based on the purification features, geometric adaptive features are obtained by deformable convolution of the deformable strip geometry refinement module. Based on the aforementioned geometric adaptive features, the horizontal stripe context branch and the vertical stripe context branch of the deformable stripe geometry refinement module are used to obtain the horizontal stripe context features and the vertical stripe context features. The geometric adaptive features, the horizontal strip context features, and the vertical strip context features are fused to obtain a geometric attention map; Based on the geometric adaptive features and the geometric attention map, geometric enhancement features are obtained.

[0012] Preferably, the method for forming a structured bill of materials (BOM) through the instance mask, the text detection box, and the OCR-recognized text input segmentation-guided BoM generation process includes: Using the instance mask, the text detection box, and the OCR-recognized text as inputs to the segmentation-guided BoM generation process, the process sequentially performs geometric attribute extraction, OCR field parsing, image-text association, engineering rule calculation, and merging of similar items. The corresponding preset engineering rules are called according to the material category to form a structured material list.

[0013] Preferably, the method for extracting geometric attributes includes: The object candidate box, instance mask, and text detection box on each rasterized CAD drawing slice image are mapped from the slice local coordinates back to the scaled drawing coordinates, and then mapped back to the original drawing coordinates according to the scaling ratio. For the same target across slice boundaries or overlapping regions after coordinate mapping, perform merging based on mask cross-union ratio, boundary distance, category consistency, and centerline continuity; For each merged graphic instance, extract geometric attributes based on its mask contour.

[0014] Preferably, the method for associating images and text includes: Multi-clue text-instance association is achieved by comprehensively using three types of cues: lead connectivity, spatial proximity, and semantic consistency.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Independent of structured CAD / BIM semantics. This invention directly extracts target instances and dimension text from rasterized CAD drawings. Even if the original DWG / DXF layers, block attributes, or BIM component semantics are unavailable, it can still perform CFRP reinforcement area identification and bill of materials generation, thus having a wider range of applications.

[0016] 2. Improve the ability to identify the boundaries of slender CFRP targets. This invention employs high-resolution feature extraction and DSGRM geometric refinement structure. Through deformable sampling, horizontal / vertical strip context, geometric attention, and boundary point correction, it improves the segmentation ability of slender, locally curved, and boundary-sensitive targets, and reduces CFRP region fractures, anchorage omissions, and boundary offsets.

[0017] 3. Reduce interference from text and color-similar annotations. The TCDM of this invention generates a text-color interference map using a color embedding branch and a text prediction branch, and performs soft suppression on the multi-scale features extracted by the high-resolution feature extraction backbone. This effectively weakens the response of large-sized text, annotation lines, and non-target lines with similar colors, thereby improving the purity of target recognition.

[0018] 4. Implement an end-to-end engineering process from recognition to inventory list. Compared with general vision methods that only output detection boxes or masks, this invention further combines OCR field parsing, lead connectivity, spatial proximity, semantic consistency, and engineering rule calculation to convert visual recognition results into BoM entries that can be directly used for material estimation.

[0019] 5. Improve the accuracy of image-text association and the traceability of results. This invention not only outputs material names, specifications, and quantities, but also saves instance masks, text detection boxes, OCR text, association scores, and calculation rules corresponding to BoM entries, facilitating review by engineers.

[0020] 6. Improve the efficiency and consistency of manual quantity calculation. This invention can automatically output material items, specifications, rules, quantities, and units, reducing the workload of manually reading drawings item by item and repetitive calculations, and reducing statistical discrepancies caused by different personnel interpreting the drawings.

[0021] 7. It has good module replaceability and engineering deployment feasibility. The input and output between the modules of this invention are clear. The slicing strategy, recognition network, OCR engine, association rules and quantity calculation rules can all be replaced or extended according to project needs, which is convenient for integration into existing engineering drawing management, cost estimation and material procurement systems. Attached Figure Description

[0022] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of the process for generating a BoM (Board of Name) for carbon fiber reinforcement CAD drawings guided by segmentation, according to an embodiment of the present invention. Figure 2 This is a schematic diagram of proportionally overlapping slices in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating high-resolution feature extraction and multi-scale fusion in an embodiment of the present invention; Figure 4 This is a schematic diagram of the text-color decoupling module in an embodiment of the present invention; Figure 5 This is a schematic diagram of the deformable strip geometry refinement module according to an embodiment of the present invention; Figure 6 This is a flowchart illustrating the process of generating a bill of materials guided by segmentation in an embodiment of the present invention. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0026] Existing implementations similar to this invention: 1. Engineering quantity statistics method based on CAD / BIM structured data Existing quantity surveying methods typically utilize structured information already present in CAD or BIM files for analysis. For example, CAD solutions can read vector entities, layers, blocks, linetypes, and annotations from DWG / DXF files; BIM solutions can use categories, dimensions, materials, and spatial relationships in IFC models or BIM components to generate quantities or bills of materials according to preset rules. These methods are highly efficient when layer specifications are standardized, object attributes are complete, and vector entities are accessible. However, in projects involving scanned PDFs, rasterized CAD drawings, non-standard drafting documents, or missing layer semantics, it is difficult to directly obtain reliable object boundaries and material properties.

[0027] 2. A method for understanding architectural or engineering drawings based on general image recognition For rasterized drawings, existing technologies include visual understanding methods for architectural floor plans or engineering drawings, including the detection and segmentation of walls, rooms, doors and windows, symbols, pipelines, text areas, engineering symbols, and graphic element lines. These methods can extract certain graphic elements from pixel-level input, but their main targets are usually general building components, symbols, or room boundaries, and they do not specifically model slender reinforcement areas, anchors, U-shaped hoops, color codes, and dense dimension specifications in CFRP reinforcement drawings.

[0028] 3. Component recognition method based on instance segmentation network Instance segmentation networks such as Mask R-CNN, Cascade Mask R-CNN, and PointRend have been used to detect and segment target objects from images and can be extended to some construction drawing component recognition tasks. These methods output object positions and boundaries through target detection boxes and instance masks, providing a visual basis for subsequent material statistics. However, general instance segmentation networks are easily affected by factors such as slender target boundaries, local bending, small-scale anchors, and interference from dimension text and similar-colored annotations in CFRP reinforcement drawings, leading to instance breakage, edge offset, false detections, and missed detections.

[0029] 4. A method for generating a bill of materials based on OCR and rule matching Existing OCR technology can recognize dimensions, specifications, and text descriptions in drawings. Combined with manual input or rule templates, it can extract fields such as material width, thickness, number of layers, spacing, and units. However, if a list is generated solely based on OCR results, it's crucial to accurately determine the specific geometric instance corresponding to the text description and apply appropriate calculation rules based on different material categories. In CFRP drawings, the relationship between text and targets can be determined by leader connections, spatial proximity, and contextual semantics. Using OCR or simple proximity rules alone can easily lead to incorrect associations in dense annotations, multiple adjacent components, and text annotations across regions, resulting in errors in BoM entries, specification fields, or quantity calculations.

[0030] The existing technology has the following drawbacks: 1. High reliance on structured CAD / BIM information. Traditional CAD / BIM quantity take-off methods rely on accessible vector entities, layer definitions, block attributes, and component parameters. When drawings exist in PDF, image, or non-standard CAD file formats, this structured information is often unavailable or incomplete, making it difficult for existing methods to directly identify and calculate the quantity of carbon fiber reinforced components.

[0031] 2. General image segmentation models are difficult to adapt to the characteristics of CFRP reinforcement drawings. CFRP fabric, sheets, strips, U-shaped hoops, and anchoring components often appear in drawings as thin strips, local bends, small-scale discrete parts, or color-coded areas. They are small in area, have a large aspect ratio, and sensitive boundaries, making them easy to confuse with dimension lines, text, leader lines, and structural lines.

[0032] 3. Severe interference from similar-colored text, leader lines, and other elements. CFRP drawings often contain densely packed dimension text, leader lines, structural lines, and color-coded elements. Some non-target elements have similar colors or spatial overlap with the target area. Existing methods, without explicit modeling of text areas and color interference, can easily misidentify dimension text or leader lines as reinforced areas.

[0033] 4. Difficulty in instance-level text association. Bill of materials generation requires not only recognizing text in drawings but also identifying the specific reinforcement components corresponding to the text descriptions. For example, "two layers of 300mm wide carbon fiber cloth bonded to the bottom of the beam" needs to be matched with the corresponding carbon fiber area instance. Existing methods, if relying solely on spatial distance, are prone to incorrect associations when annotations are dense or multiple targets are adjacent.

[0034] 5. Lack of a complete process for target identification, specification parsing, and engineering rule calculation. Existing solutions often only address one part of target detection, text recognition, or rule calculation, lacking a closed-loop process from drawing input to component identification, text association, specification parsing, quantity calculation, and structured BoM output.

[0035] Example 1 like Figure 1 As shown, this invention provides a method for generating a BoM (Board of Name) for carbon fiber reinforcement CAD drawings guided by segmentation, comprising: Preprocessing and proportionally overlapping slicing of carbon fiber reinforced CAD drawings yields rasterized CAD drawing slice images. The rasterized CAD drawing slice image is input into the constructed CFRP-TGNet network to obtain instance masks and text detection boxes; OCR text parsing and specification field extraction are performed on the text detection box and material description area to obtain OCR-recognized text; The BoM generation process, guided by the instance mask, the text detection box, and the OCR-recognized text input segmentation, forms a structured bill of materials.

[0036] The specific implementation process of this invention is as follows: To address the shortcomings of existing technologies, this invention provides a segmentation-guided method for generating a Bill of Materials (BoM) from CAD drawings for carbon fiber reinforcement. This method takes rasterized CAD drawings as input and does not rely on DWG / DXF vector layers, block attributes, or BIM semantic information. Instead, it directly identifies CFRP-related components and engineering text areas at the image pixel level. Furthermore, it combines OCR recognition, geometric attribute extraction, text-instance association, and engineering rule calculations to automatically generate a structured bill of materials.

[0037] like Figure 1 As shown, the overall process of this invention includes: drawing input and preprocessing, proportionally overlapping slicing, CFRP-TGNet instance segmentation and text detection, slicing result back-mapping and instance merging, OCR recognition and specification parsing, text-instance association, engineering rule calculation, and BoM table output. Specifically, it includes: 1. Drawing input, standardization, and proportional overlapping slicing Given an input carbon fiber reinforcement CAD drawing I, the PDF, scanned image, or exported CAD drawing is first converted into a raster image. Preprocessing steps include white background normalization, noise removal, color space unification, and resolution adjustment to obtain a preprocessed image. For drawings containing scale bars or dimensions, the conversion relationship between pixel dimensions and actual engineering dimensions can be established through scale bar recognition or manual input of scale parameters.

[0038] The preprocessed image is sliced ​​with equal-scale overlap to obtain a rasterized CAD drawing slice image. For example... Figure 2 As shown, proportional scaling is used instead of direct stretching to a fixed size to avoid geometric distortion of slender CFRP targets. If the original image size is... (Where H represents the original image height and W represents the original image width), and the image scaling factor is s, then the scaled image is denoted as... Its dimensions are: Where s represents the image scaling factor, This represents the floor function, used to ensure that the height and width of the scaled image are integer pixel values; This indicates the height of the scaled image. This indicates the width of the image after scaling.

[0039] Then, adjust the scaled drawing to the window size. Slicing is performed using the overlap ratio ρ, with horizontal and vertical step sizes as follows: in, This represents the floor function.

[0040] The (m, n)th rasterized CAD drawing slice image can be represented as: in, This represents the slice function, where m represents the slice number in the horizontal direction and n represents the slice number in the vertical direction. This represents the x-coordinate of the top-left corner of the m-th slice window in the scaled image. The vertical coordinates of the top-left corner of the nth slice window in the scaled image are defined as follows: in, This function takes the smaller value to ensure that the slice window does not exceed the boundaries of the scaled image.

[0041] When the slice window is close to the right or bottom boundary of the drawing, the above constraints ensure that the last slice still falls completely within the image area. This proportionally overlapping slicing method allows for localized high-resolution input while maintaining the drawing's geometric proportions, and reduces recognition errors caused by scaling distortion or slice boundary truncation of slender reinforced targets.

[0042] In the data preparation phase of model training, the original drawings and their annotation information are simultaneously scaled proportionally and overlapped into slices. The annotation boxes and polygon masks are then transformed into the local coordinate system of the slices, forming slice-level training samples. During the training phase, the slice images and corresponding local annotations are input into the CFRP-TGNet network for parameter learning. In the inference phase, the slice-level prediction results are backmapped to the coordinates of the original drawings, and a complete instance is obtained through an overlapping region merging strategy. This design can reduce the breakage problem of slender targets caused by scaling deformation or slice boundaries.

[0043] 2. CFRP-TGNet Network Structure CFRP-TGNet is the core recognition network of this invention, used to detect and segment CFRP-related objects from rasterized CAD drawing slice images. This network includes a high-resolution feature extraction backbone, a Text-Color Disentanglement Module (TCDM), a Deformable StripGeometry Refinement Module (DSGRM), and a detection / segmentation output head. The inputs, outputs, and functions of each module in the CFRP-TGNet network are shown in Table 1.

[0044] In terms of output categories, CFRP-TGNet outputs instance masks for CFRP Area, Anchor, and Cement Area annotation categories, and text detection boxes for Dimension Text annotation categories. Instance masks are used for subsequent geometric attribute extraction, and text detection boxes are used for subsequent OCR and specification parsing.

[0045] Table 1 2.1 High-resolution feature extraction backbone like Figure 3 As shown, the high-resolution feature extraction backbone employs a multi-resolution parallel feature extraction structure. Given a rasterized CAD drawing slice image... First, basic features are obtained through an initial convolutional layer. Then, L parallel branches with different resolutions are constructed, and the t-th stage is denoted as... The characteristics of each resolution branch are: Where l = 1, 2, ..., L. The high-resolution branch is used to preserve thin lines, text edges, and stripe boundaries, while the low-resolution branch is used to extract stronger semantic context information. In each stage, each resolution branch first performs convolutional feature extraction separately, and then achieves cross-scale fusion through upsampling, downsampling, and channel transformation.

[0046] The fusion feature of the l-th resolution branch in the (t+1)-th stage can be represented as: in, Indicates the first Convolutional feature extraction operation with each resolution branch Indicates from the first The resolution branch to the first Scaling transformation operation for each resolution branch Indicates the t-th stage. Features of each resolution branch. When hour, Downsampling is performed using stride convolution; when hour, Resolution alignment is achieved through upsampling and 1×1 convolution; when hour, It is an identity mapping.

[0047] Finally, the features from each resolution branch are uniformly upsampled to a high-resolution scale and fused to obtain the fused multi-scale features: in, Indicates an upsampling operation. Indicates channel splicing. This indicates the convolution fusion operation.

[0048] Compared to ordinary layer-by-layer downsampling backbone networks, this structure can continuously retain high-resolution spatial information, making it more suitable for handling slender strips, small-scale anchors, dimensional text, and complex boundaries in CFRP reinforcement drawings, and helping to reduce detail loss and boundary offset.

[0049] 2.2 Text-Color Decoupling Module TCDM like Figure 4 As shown, the TCDM module is used to suppress interference from non-target primitives such as text, leaders, and similar colors on the multi-scale features, reducing the interference of text, leaders, and similarly colored annotations on target recognition. Given the multi-scale features F extracted by the high-resolution feature extraction backbone, the TCDM module includes a color embedding branch, a text prediction branch, a text-color suppression branch, and a feature reweighting operation.

[0050] Color embedding branch extracts color perception features This is used to characterize the color distribution in a drawing that is related to a target or annotation. F represents the feature transformation function corresponding to the color embedding branch, which is usually composed of convolutional layers, normalization layers and nonlinear activation layers. It is used to extract information such as color distribution, color boundaries and color similar regions from F; C represents the obtained color perception features.

[0051] The text prediction branch outputs a text / label probability graph. Its monitoring signal is obtained by converting the DimensionText text detection box into a binary text region map. Among them, This represents the feature transformation function corresponding to the text prediction branch, used to predict the response of text and annotation areas in the drawing based on multi-scale features F; This represents the Sigmoid activation function, used to... The output is normalized to between 0 and 1 to obtain the probability that each pixel belongs to the text / label region; This represents the predicted text / label probability graph.

[0052] Then, F, C, and T are concatenated and fused using convolution to obtain a text-color interference map: Where [F,C,T] represents the concatenation of multi-scale feature F, color perception feature C, and text / annotation probability map T along the channel dimension; This represents the convolutional fusion function corresponding to the text-color suppression branch, which is used to combine color features and text probability information to generate an interference response; represents the Sigmoid activation function, used to normalize the interference response to between 0 and 1; S represents the text-color interference map, the larger the value, the more likely the location is to be interfered with by text, leaders, or non-target primitives with similar colors.

[0053] Finally, the multi-scale features extracted by the high-resolution feature extraction backbone are softly suppressed through a text-color suppression branch and a feature reweighting operation. Specifically, the text-color suppression branch generates a text-color interference map S based on the multi-scale fusion feature F, color-aware feature C, and text / annotation probability map T output by the high-resolution feature extraction backbone; subsequently, the feature reweighting operation suppresses the interference region responses in the multi-scale fusion feature F output by the high-resolution feature extraction backbone based on the text-color interference map S, resulting in cleaned features: in, The suppression coefficient is used to control the multi-scale features extracted from the high-resolution feature extraction backbone by the interference map. The intensity of inhibition; This represents the cleaned-up features after text and color interference suppression. In other words, this soft suppression process is achieved through two parts: first, the text-color suppression branch generates the interference map S; second, the feature reweighting operation utilizes… The multi-scale fusion feature F is weighted positionally to reduce the feature response of non-target regions such as text, lines, and similar colors, while preserving the geometric and semantic features of the real CFRP-reinforced target.

[0054] This invention addresses the problem that Dimension Text, construction instructions, leader lines, and non-target primitives with similar colors can easily interfere with CFRP target recognition. It designs a text-color decoupling module, which explicitly weakens the interference of text and similar-colored primitives on target features through color embedding, text region prediction, and interference map suppression mechanisms, thereby improving the robustness of recognition in complex annotation environments.

[0055] 2.3 Deformable Strip Geometry Refinement Module DSGRM like Figure 5 As shown, the DSGRM module is used to enhance cleanroom features, obtaining geometrically enhanced features to improve the representation of slender, locally curved, and boundary-sensitive engineering targets. The DSGRM module includes deformable convolution, horizontal stripe context branches, and vertical stripe context branches. Given the cleanroom features output by TCDM... First, geometrically adaptive features are obtained through deformable convolution: in, This refers to deformable convolution operations, which learn additional spatial offsets on top of the fixed sampling positions of ordinary convolution, enabling the convolution sampling points to adaptively adjust according to the local morphology of the target. Specifically, deformable convolution can dynamically adjust the feature sampling positions according to the local shape of CFRP strips, U-shaped hoops, and anchor boundaries, thereby making the sampling region more closely fit the target structure that is slender, curved, and has irregular boundaries. This represents the geometrically adaptive features obtained after deformable convolution.

[0056] Subsequently, horizontal stripe context branches and vertical stripe context branches were introduced to adapt the geometric features. Perform directional structural modeling: in, It represents the contextual features of horizontal stripes, used to capture information about slender structures that extend laterally in drawings; The feature transformation function corresponding to the horizontal strip context branch is represented by strip convolution or strip pooling operation of size 1×k, which expands the receptive field in the horizontal direction; This represents the contextual features of vertical stripes, used to capture information about slender structures that extend longitudinally; This represents the feature transformation function corresponding to the vertical strip context branch, which uses a strip convolution or strip pooling operation of size k×1 to expand the receptive field in the vertical direction. k represents the length of the strip convolution kernel or strip pooling window, used to control the range of directional context modeling. The larger k is, the wider the range of long strip structures that the model can perceive.

[0057] Next, geometric adaptive features will be used. Horizontal strip context features and vertical strip context features The data is then stitched together and merged, and a geometric attention graph is generated using a geometric attention enhancement mechanism. in, This indicates that the three types of features are concatenated along the channel dimension; The geometric attention fusion function, typically composed of convolutional layers, normalization layers, and nonlinear activation layers, is used to fuse deformable local features with horizontal and vertical strip contextual information. This represents the Sigmoid activation function, used to normalize the attention response to a range of 0 to 1; This represents the generated geometric attention map; a larger value indicates that the geometric structure response at that location needs to be enhanced.

[0058] Finally, the geometrically adaptive features are enhanced based on the geometric attention map to obtain the geometrically enhanced features: in, This represents geometric enhancement features.

[0059] This operation enhances the feature response of slender strips, locally bent regions, and boundary-sensitive regions, thereby improving the integrity and boundary accuracy of the CFRP target instance mask.

[0060] Furthermore, based on geometric enhancement features The detection / segmentation output head further generates object candidate boxes, text detection boxes, and instance masks. Object candidate boxes refer to candidate regions for entities such as CFRP Areas, Anchors, and Cement Areas, used for subsequent instance mask prediction and geometric attribute extraction. Text detection boxes refer to the detection results of Dimension Text or other engineering text regions, used for subsequent OCR text recognition and specification field extraction. Both object candidate boxes and text detection boxes are predicted by the detection / segmentation output head within the same drawing coordinate system. Object candidate boxes serve for target segmentation and quantity calculation, while text detection boxes serve for text parsing. In the subsequent BoM generation stage, the two establish a relationship through spatial proximity, leader connectivity, and semantic consistency to determine the specific graphic instance corresponding to a given text description.

[0061] For uncertain regions such as CFRP strip boundaries, the segmentation branch first predicts a coarse mask; then a point-based boundary refinement head can be used to sample and correct the uncertain points at the coarse mask boundary, thereby improving the boundary accuracy of the instance mask.

[0062] This invention addresses the problem that targets such as CFRP cloth, CFRP plate, and U-shaped hoop often have slender strips, local bends, and boundary-sensitive shapes. It introduces deformable convolution, horizontal / vertical strip context, and geometric attention enhancement mechanisms to enable the model to adaptively capture strip structures and local geometric changes, thereby improving mask integrity and boundary accuracy.

[0063] 2.4 Training Objectives and Outputs The network training objective L consists of detection loss, instance segmentation loss, and text region loss: Among them, detection loss Used for object classification and bounding box regression, instance segmentation loss. Used for instance mask supervision, text region loss. The binary cross-entropy loss is used, and its abbreviated form is: Where T is the predicted text / label probability graph. This is a supervised image of the text region generated from the text detection box and adjusted to the same spatial resolution. This represents the binary cross-entropy loss function, used to measure the relationship between the predicted probability map T and the text region supervision map. The differences between them.

[0064] Furthermore, if the text probability graph contains N pixel locations, the binary cross-entropy loss can be expanded as follows: in, Let represent the predicted text probability at the i-th pixel position. This represents the text region supervision label at the i-th pixel position. When this pixel is within the corresponding area of ​​the text detection box... ;otherwise The two equations above are the abbreviated form and pixel-level expanded form of the text region loss, respectively. They represent the same loss term and will not be repeatedly included in the total training objective.

[0065] This invention addresses the problem of text areas in drawings being highly mixed with reinforcement targets, making it difficult for models to distinguish between text annotations and real components. It converts Dimension Text annotation boxes into text area supervision maps, which serve as auxiliary training targets for the TCDM branch. This enables the network to learn the spatial distribution of text areas in drawings and use this information to purify target features, reducing text false detections and target confusion.

[0066] In summary, this invention addresses the problem of mixed distribution of multiple object types in CFRP reinforcement drawings, including slender reinforcement areas, anchorages, cement-based reinforcement areas, and dimensional text. It proposes a text-geometry-aware instance segmentation framework for CFRP reinforcement drawings—the CFRP-TGNet instance segmentation framework. This framework enables the model to simultaneously output instance masks and text detection boxes for CFRP-related targets, providing a unified visual foundation for subsequent geometric attribute extraction and BoM generation.

[0067] 3. Slicing result backmapping, instance merging, and geometric attribute extraction During the inference phase, the object candidate bounding boxes, instance masks, and text detection boxes on each rasterized CAD drawing slice image are first mapped from the slice's local coordinates back to the scaled drawing coordinates, and then mapped back to the original drawing coordinates according to the scaling ratio. Let the top-left corner offset of the i-th rasterized CAD drawing slice image in the scaled drawing be denoted as . If the scaling factor is s, and the coordinates of the predicted point within the rasterized CAD drawing slice image are (u, v), then its coordinates in the original drawing can be expressed as: in, This represents the original drawing coordinates after backmapping. This coordinate transformation unifies the detection boxes, instance masks, and text regions in different slices to the original drawing coordinate system, facilitating subsequent instance merging and quantity calculations.

[0068] For the same target across slice boundaries or overlapping regions, merging can be performed based on mask crossover ratio, boundary distance, category consistency, and centerline continuity to avoid duplicate statistics. Specifically, for two candidate instances... and Instance merging scoring can be defined. for: in, This represents the cross-union ratio of two instance masks. Indicates category consistency. Indicates the continuity of the centerline. Indicates the boundary distance. For weighting coefficients. When When the value exceeds a preset threshold, the two candidate instances are determined to belong to the same target and are merged to obtain the merged graphic instance.

[0069] For each merged graphic instance, geometric attributes are extracted based on its mask contour, including instance type, number of instances, contour region, center position, orientation, circumscribed length, circumscribed width, pixel area, actual length, actual area, and boundary confidence. Let the scaling factor from drawing pixels to actual size be k, and the pixel area of ​​the instance mask be... The pixel length is Then the actual area and actual length can be expressed as: in, Indicates the actual area. Indicates the actual length.

[0070] For continuous materials, the main calculations are based on area or length; for discrete components, the main calculations are based on the number of instances.

[0071] This invention addresses the problems of direct scaling of large-format CAD drawings causing deformation of slender targets and direct slicing potentially disrupting target continuity. It designs an overlapping sliding window slicing strategy that maintains the aspect ratio and merges instances after back-mapping the slicing results to the original image coordinates during the inference phase, thereby reducing target breakage, duplicate detection, and geometric distortion.

[0072] 4. OCR text parsing and specification field extraction OCR text parsing and specification field extraction are performed on the detected text detection boxes and material description areas to obtain OCR-recognized text. After obtaining the parsed text content through OCR text parsing, fields such as material type, width, thickness, number of layers, spacing, unit, construction location, and structural description are extracted using preset rule templates, domain dictionaries, or lightweight text parsing models, as shown in Table 2. For example, "two layers of 300mm wide carbon fiber cloth are pasted at the bottom of the beam" can be parsed as: material type is carbon fiber cloth, width is 300mm, number of layers is two, and the application location is the bottom of the beam.

[0073] Table 2 5. Multi-threaded text-instance association To determine the specific graphic instance corresponding to the text description, this invention comprehensively utilizes three types of clues: lead connectivity, spatial proximity, and semantic consistency. A text detection box is defined. With graphic examples The associated score is Then it can be expressed as: in, Indicates spatial proximity score, This indicates the score for the connectivity of the leads. The semantic consistency score is represented by λ1, λ2, and λ3, which are weight coefficients that can be set based on the validation set or empirical rules.

[0074] For each graphic instance, the text description with the highest comprehensive score that exceeds the threshold is selected as the association result; when there is a conflict between the connectivity of the lead line and the spatial proximity relationship, semantic consistency is used for correction. For example, text containing "carbon fiber cloth" or "carbon fiber plate" is preferentially associated with CFRP instances, text containing "anchor" or "anchor fastener" is preferentially associated with Anchor instances, and text containing "cement-based" or "mortar reinforcement" is preferentially associated with Cement Area instances.

[0075] This invention addresses the problem of mismatches between material specifications, dimensions, and target instances in densely labeled scenarios. It establishes a text-instance association mechanism by integrating leader connectivity, spatial proximity, and semantic consistency. When multiple candidate matches conflict, it makes a comprehensive decision, thereby improving the accuracy of the correspondence between text information such as material specifications, number of layers, and spacing and target instances.

[0076] 6. Bill of Materials Generation The BoM generation process, guided by instance masks, text detection boxes, and OCR-recognized text input segmentation, forms a structured bill of materials.

[0077] like Figure 6 As shown, the BoM generation stage uses instance masks, text detection boxes, and OCR-recognized text as inputs to the segmentation-guided BoM generation process. It sequentially performs the above-mentioned geometric attribute extraction, OCR field parsing, and image-text association. Based on the results of geometric attribute extraction, OCR field parsing, and image-text association, it performs engineering rule calculations and merges similar items. Then, it calls the corresponding preset engineering rules according to the material category and converts the visual recognition results into a structured list that can be directly used for material estimation.

[0078] This invention addresses the problem of inconsistent statistical rules for different CFRP material categories, making it difficult to adopt a unified calculation method. Combining the characteristics of components such as CFRP fabric, CFRP sheet, U-shaped hoop, and anchors, it proposes a quantity calculation method based on CFRP engineering rules: using engineering rules such as area calculation, length calculation, instance counting, and spacing derivation, it achieves automated output from drawing recognition results to the material quantity required for engineering applications.

[0079] An example of a structured BoM output is shown in Table 3.

[0080] Table 3 Table 3 shows representative output formats. The actual quantity is determined by the geometric attributes of the drawing instance, the OCR specification field, the scale conversion relationship, and the corresponding engineering rules.

[0081] This invention addresses the problem that general instance segmentation methods can only output target masks and are difficult to directly serve engineering material statistics. It proposes a segmentation-guided BoM generation process, which unifies the modeling of instance mask geometric attributes, OCR specification fields, text-instance association, and engineering rule calculations, and automatically outputs a structured list of materials containing fields such as material type, specifications, calculation rules, quantity, and unit.

[0082] 7. Examples and Optional Experimental Configurations In one embodiment, a dataset of CFRP (Cyber-Reinforced Plastic) reinforced engineering drawings can be constructed as training and validation data for the model. Data sources may include vector PDFs exported from CAD files, rasterized PNG images of PDFs, and scanned drawing images. Annotation categories include CFRP Area, Anchor, Cement Area, and Dimension Text, with the first three categories using polygonal mask annotations and Dimension Text using bounding box annotations. The dataset can be divided into training, validation, and test sets.

[0083] In one embodiment, the slice size during training can be set to 512×512 pixels, and the overlap ratio ρ can be set to 0.25; the optimizer can be SGD or AdamW; the TCDM suppression coefficient beta, the text loss weight λ, and the strip convolution kernel size k can be adjusted according to the validation set performance. The above values ​​are only optional implementations and do not constitute a limitation on the scope of protection.

[0084] In summary, this invention provides a segmentation-guided method for generating a Bill of Materials (BoM) from CAD drawings for carbon fiber reinforcement. The input can be a PDF exported from the CAD drawing, a rasterized PNG image, or a scanned drawing image. First, the drawing is rasterized, normalized with a white background, has a uniform resolution, is scaled proportionally, and overlapped into slices. Then, a CFRP-TGNet network is used to extract high-resolution multi-scale features, and through a text-color decoupling module and a deformable strip geometry refinement module, instance masks and text detection boxes for the CFRP reinforcement area, anchorage, and cement-based reinforcement area are output. Finally, the instance masks, text detection boxes, and OCR-recognized text are input into the segmentation-guided BoM generation process to form a structured bill of materials. Compared to existing technologies, the main technical advantages of this invention are: (1) In view of the problem that traditional CAD / BIM engineering quantity statistics methods rely on standard layers, block attributes and component semantics, this invention establishes a drawing understanding process for rasterized CFRP reinforcement drawings. It can directly identify reinforcement areas, anchors, cement-based reinforcement areas and dimension text areas in the absence of original vector information and BIM models, thereby improving the automatic recognition capability and applicability in unstructured drawing scenarios.

[0085] (2) To address the problem of identifying slender, locally curved, boundary-sensitive, and easily fractured targets such as CFRP fabric, U-shaped hoop, strips, and anchors, this invention introduces high-resolution feature extraction, deformable strip geometry refinement, and boundary point correction mechanisms to enhance the model's ability to express slender strip structures, locally bent areas, and complex boundaries, reduce problems such as CFRP area fractures, boundary offsets, and missed detections of small-scale anchors, and improve the completeness and accuracy of instance segmentation results.

[0086] (3) In view of the problem that dimension text, construction instructions, leader lines, color-coded lines and similar color markings in CFRP reinforcement drawings can easily interfere with target recognition, this invention designs a text-color decoupling mechanism to explicitly learn text areas and color similarity interference information, so that the model can distinguish between real reinforcement targets and non-target marking elements, thereby reducing the situation of misidentifying text, leader lines or similar color elements as reinforcement components and improving the recognition robustness in complex drawing environments.

[0087] (4) To address the issue that general instance segmentation methods can only output target locations or masks and cannot directly serve bill of materials (BOM) generation, this invention further establishes an automatic conversion process from recognition results to a BoM table. This process extracts geometric attributes such as length, width, area, and quantity based on instance masks, combines OCR field parsing to obtain semantic information such as material name, specifications, number of layers, and spacing, and achieves automatic association between text and components through leader connectivity, spatial proximity, and semantic consistency. Finally, it calculates material usage according to engineering rules and generates structured fields such as material type, specifications, calculation rules, quantity, and unit.

[0088] (5) Unlike previous methods that only focused on drawing target detection or single component identification, this invention integrates CFRP reinforcement component identification, engineering text positioning, text-instance association, and engineering rule-based quantity calculation into a complete bill of materials (BOM) generation chain, achieving end-to-end process support from drawing pixels to structured BoM. This method can reduce the workload of manual drawing interpretation and manual quantity calculation, reduce omissions, miscalculations, and inconsistent statistical standards, and improve the consistency, traceability, and reusability of CFRP reinforcement engineering material estimation.

[0089] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A method for generating a BoM (Bill of Materials) of carbon fiber reinforcement CAD drawings guided by segmentation, characterized in that, The method includes: Preprocessing and proportionally overlapping slicing of carbon fiber reinforced CAD drawings yields rasterized CAD drawing slice images. The rasterized CAD drawing slice image is input into the constructed CFRP-TGNet network to obtain instance masks and text detection boxes; OCR text parsing and specification field extraction are performed on the text detection box and material description area to obtain OCR-recognized text; The BoM generation process, guided by the instance mask, the text detection box, and the OCR-recognized text input segmentation, forms a structured bill of materials.

2. The method according to claim 1, characterized in that, Methods for preprocessing and proportionally overlapping slices of carbon fiber reinforced CAD drawings to obtain rasterized CAD drawing slice images include: Convert the carbon fiber reinforced CAD drawings into raster images; The raster image is preprocessed by performing white background normalization, noise removal, color space unification, and resolution adjustment to obtain a preprocessed image; The preprocessed image is sliced ​​with equal-proportion overlapping slices to obtain rasterized CAD drawing slice images.

3. The method according to claim 1, characterized in that, The CFRP-TGNet network includes: a high-resolution feature extraction backbone, a text-color decoupling module, a deformable strip geometry thinning module, and a detection / segmentation output head; The high-resolution feature extraction backbone is used to perform high-resolution feature extraction on the rasterized CAD drawing slice image to obtain multi-scale features. The text-color decoupling module is used to suppress interference from text, lines, and non-target primitives with similar colors in the multi-scale features to obtain clean features. The deformable strip geometry refinement module is used to enhance the purification features to obtain geometrically enhanced features. The detection / segmentation output head is used to generate instance masks and text detection boxes based on the geometric enhancement features.

4. The method according to claim 3, characterized in that, The text-color decoupling module performs interference suppression on the multi-scale features by text, lines, and non-target primitives with similar colors to obtain clean features. The method for obtaining clean features includes: Based on the multi-scale features, the color embedding branch of the text-color decoupling module is used to extract color perception features; Based on the multi-scale features, the text prediction branch of the text-color decoupling module is used to output a text / label probability map; The multi-scale features, the color perception features, and the text / label probability map are concatenated and fused through convolution to obtain a text-color interference map; Based on the multi-scale features and the text-color interference map, the text-color suppression branch and feature reweighting operation of the text-color decoupling module are used to obtain cleaned features.

5. The method according to claim 3, characterized in that, The deformable strip geometry refinement module enhances the purification features to obtain geometrically enhanced features through the following methods: Based on the purification features, geometric adaptive features are obtained by deformable convolution of the deformable strip geometry refinement module. Based on the aforementioned geometric adaptive features, the horizontal stripe context branch and the vertical stripe context branch of the deformable stripe geometry refinement module are used to obtain the horizontal stripe context features and the vertical stripe context features. The geometric adaptive features, the horizontal strip context features, and the vertical strip context features are fused to obtain a geometric attention map; Based on the geometric adaptive features and the geometric attention map, geometric enhancement features are obtained.

6. The method according to claim 1, characterized in that, The method for generating a structured bill of materials (BOM) by using the instance mask, the text detection box, and the OCR-recognized text input segmentation-guided BoM generation process includes: Using the instance mask, the text detection box, and the OCR-recognized text as inputs to the segmentation-guided BoM generation process, the process sequentially performs geometric attribute extraction, OCR field parsing, image-text association, engineering rule calculation, and merging of similar items. The corresponding preset engineering rules are called according to the material category to form a structured material list.

7. The method according to claim 6, characterized in that, The method for extracting geometric attributes includes: The object candidate box, instance mask, and text detection box on each rasterized CAD drawing slice image are mapped from the slice local coordinates back to the scaled drawing coordinates, and then mapped back to the original drawing coordinates according to the scaling ratio. For the same target across slice boundaries or overlapping regions after coordinate mapping, perform merging based on mask cross-union ratio, boundary distance, category consistency, and centerline continuity; For each merged graphic instance, extract geometric attributes based on its mask contour.

8. The method according to claim 6, characterized in that, The method for linking images and text includes: Multi-clue text-instance association is achieved by comprehensively using three types of cues: lead connectivity, spatial proximity, and semantic consistency.