Drawing processing method and device, medium and product

By identifying and cutting functional areas of industrial drawings and combining deep learning models with optical character recognition technology, the problem of low drawing processing efficiency in existing technologies has been solved, efficient drawing management and text information extraction have been achieved, and query accuracy has been significantly improved and costs have been reduced.

CN120808380APending Publication Date: 2025-10-17LANWO (SHANGHAI) INTELLIGENT DIGITAL TECHNOLOGY CO LTD +1

Patent Information

Application Number
CN202510816367.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

When processing industrial drawings, existing technologies lack the ability to accurately identify and cut the functional areas of drawings. They cannot effectively decompose complex industrial drawings into sub-drawings with specific functions. It is difficult to accurately extract and associate text information in the drawings, and lack the general processing capabilities for various types of industrial drawings, resulting in low processing efficiency.

Method used

The detection model is used to parse the drawings, identify functional areas and cut them, and optical character recognition technology is combined to extract text information. The drawings are processed using a deep learning model of feature extraction layer, fusion layer, inspection layer and association layer to achieve automatic parsing and structured storage of the drawings.

Benefits of technology

It has improved the efficiency of drawing management and utilization, increased the query accuracy by more than 60%, significantly shortened the product development cycle by 30%, and reduced manual marking and management costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808380A_ABST
    Figure CN120808380A_ABST
Patent Text Reader

Abstract

The invention discloses a drawing processing method and device, a medium and a product, and belongs to the field of image processing. According to the drawing processing method, a to-be-processed drawing is analyzed, at least one functional area is determined, the to-be-processed drawing is cut based on the functional area, and at least one sub-graph is obtained; and identifying the sub-graph to obtain text information associated with the sub-graph. Compared with a traditional manual retrieval mode, the query accuracy is improved by 60% or above, and the product research and development cycle is remarkably shortened by 30% on average; according to the method, the sub-graphs in the to-be-processed drawing can be automatically and accurately cut, and the text information in the sub-graphs is associated with the sub-graphs, so that the management and utilization rate of the drawing are facilitated, and the manual annotation and management cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image processing, in particular to a drawing processing method and device, medium and product. BACKGROUND

[0002] With the deepening of industrial automation and digital transformation, the demand for drawing processing and management of industrial enterprises is increasing. In modern industrial production, drawings as important technical documents carry a large amount of design information, process parameters and technical requirements, which are of great significance to product design, manufacturing and maintenance. However, the traditional drawing processing method mainly relies on manual operation, which has the problems of low efficiency and high cost.

[0003] The prior art has the following problems in processing industrial drawings: first, the existing drawing processing method lacks accurate recognition and cutting ability of the functional area of the drawing, and cannot effectively decompose complex industrial drawings into sub-drawings with specific functions; second, the existing technology has deficiencies in extracting and associating text information in the drawing, and it is difficult to accurately associate the recognized text information with the corresponding graphic elements; third, the existing method is mainly for specific fields or specific types of images, and lacks general processing ability for multiple types of industrial drawings; finally, the existing technology often cannot fully utilize the structured characteristics of PDF format drawings, resulting in low processing efficiency.

[0004] Therefore, there is an urgent need for a method that can efficiently process industrial drawings, accurately recognize and cut the functional area of the drawing, and accurately extract and associate the text information in the drawing to improve the management and utilization efficiency of industrial drawings. SUMMARY

[0005] In view of the above problems, the present application provides a drawing processing method, device, medium and product which can accurately extract and associate the text information in the drawing to improve the management and utilization efficiency of industrial drawings.

[0006] The present application provides a drawing processing method, comprising:

[0007] analyzing the to-be-processed drawing to determine at least one functional area, cutting the to-be-processed drawing based on the functional area to obtain at least one sub-drawing;

[0008] identifying the sub-drawing to obtain text information associated with the sub-drawing.

[0009] Optionally, the analyzing the to-be-processed drawing to determine at least one functional area, cutting the to-be-processed drawing based on the functional area to obtain at least one sub-drawing comprises:

[0010] obtaining the to-be-processed drawing;

[0011] adopting a detection model to identify an image region in the to-be-processed drawing, determining a function category corresponding to the image region, forming a function region according to the function type and the image region, each function region being associated with a set of region boundary coordinates;

[0012] cutting the to-be-processed drawing according to the region boundary coordinates to obtain the sub-drawing.

[0013] Optionally, the function category includes at least two of the following categories: solid view category, two-dimensional view category, section view category, explosion view category, and technical requirement category.

[0014] Optionally, the detection model includes a feature extraction layer, a fusion layer, a checking layer, and an association layer.

[0015] The adopting a detection model to identify an image region in the to-be-processed drawing, determining a function category corresponding to the image region, forming a function region according to the function type and the image region, each function region being associated with a set of region boundary coordinates, includes:

[0016] The to-be-processed drawing is input to the feature extraction layer, and a multi-scale feature map is extracted through the feature extraction layer;

[0017] The multi-scale feature map is subjected to semantic feature and detail feature fusion through the fusion layer to obtain a fusion feature;

[0018] The fusion feature is processed through the checking layer to obtain a function category and a regression result of each candidate region, the regression result including boundary coordinates of the candidate region and a confidence degree of the candidate region;

[0019] The association layer determines a function region according to the confidence degree of the candidate region, and associates the function region with the boundary coordinates.

[0020] Optionally, the cutting the to-be-processed drawing according to the region boundary coordinates to obtain the sub-drawing includes:

[0021] The function region associated with the region boundary coordinates in the to-be-processed drawing is cut according to the region boundary coordinates to obtain the sub-drawing.

[0022] Optionally, the identifying the sub-drawing to obtain text information associated with the sub-drawing includes:

[0023] The sub-drawing is preprocessed to obtain a candidate image, the candidate image is identified to determine position information of element data, the element data is identified according to the position information, and text information is extracted.

[0024] associate the text information with the subgraph.

[0025] Optionally, the to-be-processed drawing adopts a PDF format.

[0026] The application further provides an electronic device, comprising one or more processors, and a memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method described above.

[0027] The application further provides a computer readable medium, which stores computer program instructions, which can be executed by a processor to implement the method described above.

[0028] The application further provides a computer program product, comprising computer program / instructions, which, when executed by a processor, implement the steps of the method described above.

[0029] The technical scheme has the following beneficial effects:

[0030] In the technical scheme, the drawing processing method analyzes the to-be-processed drawing to determine at least one functional region, cuts the to-be-processed drawing based on the functional region to obtain at least one subgraph, and identifies the subgraph to obtain text information associated with the subgraph. Compared with the traditional manual retrieval method, the query accuracy is improved by more than 60%, and the product development cycle is shortened by an average of 30%. The subgraph in the to-be-processed drawing can be accurately cut automatically, the text information in the subgraph is associated with the subgraph, the management and utilization of the drawing are facilitated, and the manual labeling and management costs are greatly reduced. BRIEF DESCRIPTION OF DRAWINGS

[0031] One or more embodiments are exemplarily illustrated by pictures in the drawings corresponding to the embodiments, which do not constitute a limitation on the embodiments, and elements with the same reference numerals in the drawings represent similar elements, unless otherwise specified, and the drawings do not constitute a proportional limitation.

[0032] Figure 1 A flowchart of an embodiment of the drawing processing method described in the application;

[0033] Figure 2 A flowchart of an embodiment of the drawing processing method described in the application;

[0034] Figure 3 An effect diagram of an embodiment of the drawing processing method described in the application;

[0035] Figure 4A plot view of an embodiment of the detection model of the present application at a certain confidence threshold;

[0036] Figure 5 A plot of performance of the detection model of the present application at different classification thresholds;

[0037] Figure 6 A schematic view of an embodiment of the subgraph of the present application for recognizing text information;

[0038] Figure 7 A schematic view of another embodiment of the subgraph of the present application for recognizing text information;

[0039] Figure 8 An exemplary structural diagram of the electronic device of the present application. DETAILED DESCRIPTION

[0040] The advantages of the present application are further set forth in the detailed description below.

[0041] The exemplary embodiments will be described in detail herein below with reference to the accompanying drawings. In the following description, unless otherwise indicated, like numbers in the drawings indicate contact or similar components. The following exemplary embodiments are described in enough detail to enable those skilled in the art to practice the application. Additionally, the description itself is not intended to limit the scope of the application, but is intended to provide software, tools, and data that enable one to practice the application. Other aspects, features, and embodiments of the application will become apparent to those skilled in the art from the description.

[0042] The terminology used in the disclosure of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used in the description of the application and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0043] It will be understood that, although the terms first, second, third, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used only as a shorthand notation to distinguish one item from another. For example, a first item can be termed a second item, and similarly, a second item can be termed a first item without departing from the scope of the present application. As used herein, the term "if' can be construed to mean "when" or "in response to determining" depending on the context.

[0044] In the description of the present application, it should be understood that the numerical reference numbers before the steps do not indicate the order of execution of the steps before and after, but are only used for the convenience of describing the present application and distinguishing each step, and therefore should not be understood as a limitation of the present application.

[0045] The drawing processing method of the embodiment of the present application can be applied to the fields of manufacturing industry, engineering design, digital factory management, etc., can automatically analyze the structure and process information in the PDF format drawing, improve the drawing management efficiency, reduce the labor cost, and provide data support for intelligent manufacturing. The present application analyzes the to-be-processed drawing, determines at least one functional area, cuts the to-be-processed drawing based on the functional area, and obtains at least one sub-drawing; the sub-drawing is identified to obtain the text information associated with the sub-drawing. Compared with the traditional manual retrieval method, the query accuracy is improved by more than 60%, and the product development cycle is shortened by an average of 30%; the sub-drawing in the to-be-processed drawing can be accurately cut, the text information in the sub-drawing is associated with the sub-drawing, which is convenient for the management and utilization of the drawing, and greatly reduces the labor marking and management cost.

[0046] Embodiment one

[0047] The present application aims at the problem that the existing factory relies on manual management of image-based CAD drawings, which leads to low efficiency and high cost, specifically manifested as inefficient drawing retrieval, unstructured data storage, high labor marking cost, and difficult process information extraction, etc., and provides a drawing processing method. Referring to Figure 1 The drawing processing method of the embodiment can include:

[0048] S1. Analyzing the to-be-processed drawing to determine at least one functional area, cutting the to-be-processed drawing based on the functional area, and obtaining at least one sub-drawing;

[0049] It should be noted that the to-be-processed drawing in the embodiment can be in PDF format, PNG format or DWG format. The sub-drawing can be in JSON format.

[0050] S2. Identifying the sub-drawing to obtain the text information associated with the sub-drawing.

[0051] In the embodiment, the drawing processing method analyzes the to-be-processed drawing to determine at least one functional area, cuts the to-be-processed drawing based on the functional area, and obtains at least one sub-drawing; the sub-drawing is identified to obtain the text information associated with the sub-drawing. Compared with the traditional manual retrieval method, the query accuracy is improved by more than 60%, and the product development cycle is shortened by an average of 30%; the sub-drawing in the to-be-processed drawing can be accurately cut, the text information in the sub-drawing is associated with the sub-drawing, which is convenient for the management and utilization of the drawing, and greatly reduces the labor marking and management cost.

[0052] Embodiment two

[0053] Embodiment two of the present application relates to a drawing processing method. Embodiment two is an improvement based on embodiment one, and the specific improvement is that step S1 can include:

[0054] S11. obtaining the to-be-processed drawing;

[0055] In this embodiment, the to-be-processed drawing can be an engineering drawing in PDF format, which usually contains multiple views and technical information of different types. For example, a mechanical design drawing can contain a three-dimensional entity view, a two-dimensional engineering drawing, a sectional view, and related technical requirement descriptions.

[0056] Before step S12 is performed, the integrity of the to-be-processed drawing can be detected, and rasterization processing (default 300 dpi) can be performed on the to-be-processed drawing.

[0057] S12. identifying an image region in the to-be-processed drawing by using a detection model, determining a function category corresponding to the image region, and forming a function region according to the function type and the image region, each function region being associated with a set of region boundary coordinates.

[0058] It should be noted that the function category can include at least two of an entity view category, a two-dimensional view category, a sectional view category, an exploded view category, and a technical requirement category.

[0059] In this embodiment, the function regions cannot overlap each other, and each function category can be covered by a color, and the corresponding function category can be realized in the function region. The following table is obtained through experimental comparison:

[0060] Table 1

[0061]

[0062]

[0063] As can be seen, the drawing processing method of the present application can reduce the labeling threshold and greatly improve the labeling efficiency and accuracy in the drawing display function region and function category.

[0064] In this embodiment, the detection model can include a feature extraction layer, a fusion layer, a checking layer, and a correlation layer.

[0065] Further, referring to Figure 2 Step S12 can include:

[0066] S121. inputting the to-be-processed drawing into the feature extraction layer, and extracting a multi-scale feature map by using the feature extraction layer;

[0067] In the feature extraction layer, the to-be-processed drawing is input into the model and processed through a multi-layer convolutional neural network to extract multi-scale feature maps. These feature maps contain visual information of different scales in the drawing, from detailed textures to overall structures. For example, for an engineering drawing containing multiple views, the feature extraction layer can capture both small-scale size annotation text and large-scale overall view outlines.

[0068] In this embodiment, the feature extraction layer can use CSPDarknet53+ (an improved version of Cross Stage Partial Darknet53) as the feature extractor. By splitting and processing the feature map channels through cross-stage partial connection (Cross Stage Partial) and then merging them, the calculation redundancy is reduced, the gradient flow is enhanced, and the detection capability of small targets (such as size annotations) is improved. In addition, a spatial pyramid pooling (Spatial Pyramid Pooling) module is used, which performs maximum pooling of different scales (such as 4x4, 8x8, 16x16) in parallel, so that the model can capture both macro layout (such as view arrangement) and micro details (such as small font technical requirement text) in the drawing.

[0069] S122. Perform semantic feature and detail feature fusion on the multi-scale feature maps through the fusion layer to obtain fused features;

[0070] In the fusion layer, the system processes the multi-scale feature maps obtained from the feature extraction layer and fuses semantic features and detail features. Semantic features mainly express the high-level meaning of the image, such as the type of view and the approximate location of the functional area; while detail features contain precise boundary information, text location and other detailed content. Through a feature pyramid network or a feature fusion module, features at different levels are effectively combined to generate fused features. This fusion enables the model to understand both the overall structure and the local details of the drawing.

[0071] In this embodiment, the fusion layer can use a bidirectional architecture that partially fuses FPN (Feature Pyramid Network) and PANet (Path Aggregation Network). FPN uses a top-down propagation path to pass high-level semantic information (such as recognizing the "two-dimensional view" label) to the bottom layer features; while PANet strengthens the geometric detail positioning (such as accurately framing the lead arrow of the thread annotation) through a bottom-up path. This bidirectional design is similar to the cognitive process of engineers first overviewing the overall structure of the drawing and then focusing on the local details, which in testing improves the detection recall rate of the sectional view to 88.4%.

[0072] S123. The fusion features are processed by the inspection layer to obtain the functional categories of each candidate region and regression results, the regression results including boundary coordinates of the candidate region and confidence of the candidate region;

[0073] In the inspection layer, the system further processes the fusion features to generate candidate regions and their corresponding functional categories and regression results. The functional categories include entity view class, two-dimensional view class, cross-sectional view class, explosion view class, and technical requirement class, etc. The regression results contain the boundary coordinates (usually the upper left and lower right coordinates of the rectangular box) of each candidate region and the confidence score. The confidence score represents the degree of certainty of the model's classification result for the region, and the higher the score indicates the higher the possibility of belonging to the predicted category.

[0074] In this embodiment, the inspection layer adopts a Decoupled Head structure to separate the classification task from the regression task, and processes the two tasks through independent fully connected layers respectively, avoiding mutual interference. The classification task mainly predicts the category of the functional region (such as "two-dimensional view", "three-dimensional view", "technical requirement", etc.); while the regression task mainly predicts the boundary box coordinates (x, y, w, h) and the confidence (whether it contains the target).

[0075] S124. The functional region is determined according to the confidence of the candidate region by the association layer, and the functional region is associated with the boundary coordinates.

[0076] In the association layer, the system filters according to the confidence of the candidate region, usually sets a threshold (such as 0.7 or 0.8), and only keeps the candidate regions with confidence higher than the threshold as the final functional region. Then, the system associates the determined functional region with its corresponding boundary coordinates to form the accurate positioning information of the functional region.

[0077] In this embodiment, the detection model can be a YOLO v11 model, which is a further optimized target detection model based on the YOLO series. Its core value lies in achieving a breakthrough in both speed and accuracy. As a one-stage target detection algorithm, it directly maps the input image to the detection result through end-to-end training, eliminating the region proposal step of traditional two-stage methods such as Faster R-CNN, and the actual inference speed is improved by more than 3 times. The detection model adopts a heterogeneous feature fusion strategy to detect the multi-modal characteristics of CAD drawings (including graphics, text, symbols, etc.). For example, it can simultaneously associate geometric figures and tolerance text when detecting size annotations. This capability enables the mAP@0.5 of the model in the industrial drawing analysis task to reach 92.3%, which is more than 35 percentage points higher than traditional template matching methods. Among them, mAP@0.5 is a core indicator of model performance, indicating that the model can effectively detect various targets in the drawing (such as parts and annotations), and the overlap between the detection box and the true annotation meets the high-precision requirements of the industrial scene (IoU ≥ 50%), and the overall performance is excellent.

[0078] The detection model in this embodiment can use an improved YOLOv5 network structure, which has better performance in processing large-size drawings and complex layouts.

[0079] S13. Cutting the to-be-processed drawing according to the region boundary coordinates to obtain the sub-drawing.

[0080] Further, step S13 can include:

[0081] According to the region boundary coordinates, the functional region associated with the region boundary coordinates in the to-be-processed drawing is cut to obtain the sub-drawing.

[0082] After determining the functional region, the system cuts the to-be-processed drawing according to the region boundary coordinates. Specifically, the system crops the corresponding region from the original drawing according to the boundary coordinates associated with each functional region to form an independent sub-drawing. For example, if the boundary coordinates of an entity view region are (100, 200, 500, 600) (indicating the upper left corner coordinates are (100, 200) and the lower right corner coordinates are (500, 600)), the system will crop this rectangular region from the original drawing to generate a sub-drawing containing the entity view.

[0083] In actual application, the to-be-processed drawing input into the detection model can be an unlabeled CAD drawing containing different modules such as two-dimensional views, three-dimensional views, and technical requirements. The data processed by the detection model is a JSON format file, which can contain the bounding box (bbox) of different views, the polygon segmentation layer covering the target view, etc.Figure 3 Different color borders and functional areas can be used on the original drawing to identify different views (such as 2D view, 3D view, technical requirements, etc.) according to the functional area category (for example, the overlapping 2D view in the figure is the union of the two) Figure Three , and the labels (functional categories) and confidence levels of each view are displayed on the drawing. The initial output data is a sub-drawing in png format cut according to the mask (single mask) result of the detection model.

[0084] When training the detection model, 1000 artificially annotated drawings and their labels on the annotation platform are used as the training set to generate the initial detection model. Then, the initial model is used to batch infer and predict new drawings, and after artificial rapid review, the accuracy of the model is checked. If the accuracy meets the standard, it is finished; if the accuracy does not meet the standard, the drawings that the model infers incorrectly are uploaded to the annotation platform for artificial annotation. Whenever the artificially annotated drawings that are incorrectly identified accumulate to 1000, the model is incrementally trained on these drawings to continuously improve the model effect. Such a training process can achieve 95% model performance of full data using only 40% artificial annotation data, thereby greatly reducing the cost of artificial annotation. See Figure 4 and Figure 5 , the F1 curve and PR curve of the initial detection model trained for 1000 artificially annotated drawings on the test set are as follows. In addition, mAP50 can reach about 0.9. After the training set is expanded, these indicators will be improved.

[0085] See Figure 4 and Figure 5 , the comprehensive performance curve of the initial detection model trained for 1000 artificially annotated drawings on the test set at a specific confidence threshold (see Figure 4 ) and the performance curve of the detection model at different classification thresholds (see Figure 5 ). In addition, mAP50 can reach about 0.9. After the training set is expanded, these indicators will be improved.

[0086] In Figure 4In the figure, the horizontal axis represents the model's confidence (probability value) in the prediction result; the vertical axis represents the score F1, which is the harmonic mean of precision (Precision) and recall (Recall), F1 = 2 × (Precision × Recall) / (Precision + Recall). The score F1 reflects the overall performance of the detection model at a specific confidence threshold (balancing false positives and false negatives). In the high-threshold area (right side of the figure): When only high-confidence predictions are retained, the F1 score may be higher (because the predictions are more reliable), but the sample coverage rate is reduced (some low-confidence samples are discarded). In the low-threshold area (left side of the figure): More samples are covered, but the noise increases, and the F1 score may decrease. The curves in the figure represent the different F1 scores corresponding to sub-graphs at different confidence levels.

[0087] exist Figure 5 In the figure, the horizontal axis represents Recall (also known as Recall Rate), which refers to the proportion of positive samples correctly predicted by the detection model to all true positive samples (avoiding missed detections); the vertical axis represents Precision (also known as Precision Rate), which refers to the proportion of samples predicted as positive by the detection model that are actually positive (avoiding false positives). The PR curve in the figure is generated by adjusting the classification threshold: the default threshold (0.5) means that the detection model outputs a probability > 0.5 as a positive class, and otherwise as a negative class. Dynamically adjusting the threshold (e.g., from 0 to 1) calculates Precision and Recall at different thresholds: As the threshold is lowered (Recall↑, Precision↓), more samples are classified as positive, and Recall increases, but false positives (FP) may increase, and Precision decreases. As the threshold is raised (Recall↓, Precision↑), only high-confidence samples are classified as positive, reducing false positives and improving Precision, but missed detections (FN) may increase, and Recall decreases. AP (Average Precision) = area under the PR curve, which is used to comprehensively evaluate the performance of the detection model: the higher the AP, the better the overall performance of the model.

[0088] The accuracy standard of the detection model is:

[0089] Each sub-view on each drawing (including two-dimensional views, three-dimensional views and technical requirements, etc.) is divided correctly, without missing any digital labels, edge bubble cheongsam markings, partial cross-sections, etc.

[0090] Tables, border lines, etc. are not mistakenly identified as views.

[0091] There is no misidentification of different types of sub-views, such as misidentifying a 2D view as a 3D view.

[0092] The accuracy data of the detection model is shown in Table 2:

[0093] Table 2

[0094]

[0095] The detection model processes a 1024x1024 resolution drawing in only about 48 ms, meeting the real-time requirements of the production line.

[0096] Embodiment Three

[0097] Embodiment Three of the present application relates to a drawing processing method. Embodiment Three is an improvement based on Embodiment One, and the specific improvement is that step S2 can include:

[0098] S21. Preprocessing the sub-drawing to obtain a candidate image, identifying the candidate image, determining the position information of the element data, identifying the element data according to the position information, and extracting text information;

[0099] S22. Associating the text information with the sub-drawing.

[0100] In this embodiment, the element data can be data such as annotations, dimensions, text, process parameters, etc. in the sub-drawing.

[0101] In this embodiment, for each sub-drawing obtained by cutting, the system performs further recognition processing to obtain text information associated with the sub-drawing. First, the sub-drawing is preprocessed, including image enhancement, noise removal, binarization, etc. to obtain a candidate image that is more suitable for recognition. Then, the system identifies the candidate image to determine the position information of the element data (such as text, symbols, annotations, etc.) contained therein. According to these position information, the system can accurately locate each element and extract text information through optical character recognition (OCR) and other technologies. Finally, the system associates the extracted text information with the corresponding sub-drawing to form a complete drawing analysis result.

[0102] Further, this embodiment uses an adaptive binarization algorithm in the sub-drawing preprocessing stage, which can better handle drawings of different lighting conditions and qualities, improving the accuracy of text recognition. Specifically, the system dynamically adjusts the binarization threshold according to the local region characteristics of the sub-drawing, rather than using a global fixed threshold. This can better preserve text information and reduce the interference of background noise.

[0103] In this embodiment, the subgraph can be recognized in step S2 using an optical character recognition model. The optical character recognition model converts the text information in the subgraph into structured text data, and the core process can be divided into four key stages: image preprocessing, text detection, text recognition, and post-processing. Image preprocessing mainly improves image quality and enhances the recognizability of text by grayscale, binarization, de-noising, and morphological processing, thereby improving the accuracy of subsequent OCR recognition. Text detection mainly analyzes the preprocessed image to find the location information of the text. Through the location information of the text, the corresponding text image region can be extracted from the image to obtain the preliminary text recognition result. The post-processing of the text is based on the recognition result of the text, and the spelling errors are corrected or the rules are added to improve the rationality of the overall semantics through simple but complete logical judgment.

[0104] The optical character recognition model used in this embodiment is an end-to-end OCR model. This model can complete the text detection and recognition tasks simultaneously in a single architecture, avoiding the error accumulation caused by the two-stage process of traditional methods. In the feature extraction stage, the model uses ResNet-50 as the backbone network to extract deep visual features of the image. Then, a module based on the Transformer encoder-decoder architecture is used to model the image features, fully exploiting their spatial relationships and contextual semantic information, thereby enhancing the modeling capability for complex text structures such as curved text and dense text. In the decoding stage, the model introduces a set of learnable queries, each corresponding to a potential text instance. Each query is decoded through a set of parallel prediction heads (usually linear layers or multi-layer perceptrons, MLPs) to output multiple key elements of the text instance, including: the center point coordinates or center line position of the text, the boundary contour information, the corresponding character transcription result, and the confidence score of the prediction.

[0105] Unlike the previous embodiments, this embodiment uses an OCR model optimized specifically for engineering drawings in the text information extraction stage. This model has been trained on a large number of engineering drawing texts and can more accurately recognize professional text content such as dimension labels and technical parameters. In addition, the system also integrates a professional term dictionary to correct the OCR recognition results, further improving the accuracy of text recognition.

[0106] Through the above processing flow, the drawing processing method of this embodiment can automatically analyze engineering drawings, recognize different functional areas, and extract relevant text information, greatly improving the efficiency and accuracy of drawing processing. This embodiment can be applied to processing complex and comprehensive engineering drawings.

[0107] It should be noted that, unlike the previous embodiment, this embodiment also supports processing of multiple formats of drawing files, including not only PDF format, but also common image formats such as TIFF, JPG, and PNG. The system will automatically select the appropriate parsing method according to the format of the input file. For PDF format, it will first be converted into a high-resolution image before processing; for other image formats, it will be processed directly. This multi-format support greatly enhances the applicability of the system.

[0108] The drawing processing method of the present embodiment supports multiple file formats, improving the versatility and practicality of the system, and meeting the needs of drawing processing in different scenarios. The present application can store subgraphs and associated text information in a database for later query management.

[0109] The present application uses a target detection model based on YOLO v11 to automatically parse and segment drawings, and combines OCR technology to identify key information of subgraphs, achieving intelligent processing of CAD drawings. Compared with traditional manual retrieval methods, the query accuracy is improved by more than 60%; through automatic parsing of drawing structure, multi-modal information fusion, and process knowledge extraction, the system provides key technical support for the digital transformation of manufacturing, significantly shortening the new product development cycle by an average of 30%; as the number of labeled drawings increases, the accuracy of the automatic segmentation model improves from 63.7% (1000 labeled drawings) to 85.1% (3000 labeled drawings); processing a 1024x1024 resolution drawing on an NVIDIA T4 GPU only takes about 48ms, meeting the real-time requirements of the production line; through incremental training, only 40% of the manual annotation data is used to achieve 95% of the model performance of the full amount of data, greatly reducing the cost of manual annotation.

[0110] Figure 6 and Figure 7 The figure shows that the trained optical character recognition model can successfully extract key special symbols and size information (i.e. text information) on the subgraph, without missing or errors.

[0111] The drawing processing method of the present application can address the problem of low efficiency and high cost caused by relying on manual management of image-based CAD drawings in existing factories. Through a deep learning model, the functional areas of the drawing are automatically identified and classified into entity graphs, two-dimensional views, and auxiliary Figure Three Large structure modules, combined with OCR technology to accurately extract process parameter features, and construct a process manufacturing-oriented drawing semantic atlas. The system realizes the structural analysis of drawing content and the digitalization of process knowledge, and compared with traditional manual retrieval methods, the query accuracy is improved by more than 60%, the new product development cycle is shortened by an average of 30%, and provides key technical support for the digital transformation of manufacturing.

[0112] The steps of the various methods above are divided only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this patent.

[0113] In addition, some embodiments of the present application further provide an electronic device. The electronic device may be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, etc. The electronic device may also be various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices.

[0114] The electronic device includes: one or more processors; and a memory storing computer program instructions, wherein the computer program instructions, when executed, enable the processor to perform the steps of the method provided in any one or more of the above embodiments. Figure 8 An exemplary structural diagram of the electronic device is disclosed. Figure 8 As shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Among them, the components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0115] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103 and the output device 1104 may be connected via a bus or other means. Figure 8 The bus connection is taken as an example.

[0116] The input device 1103 can receive input of a number or character information, and generate a key signal input relating to user settings and function controls of the electronic device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 1104 can include a display device, an auxiliary lighting device (e.g., an LED), a haptic feedback device (e.g., a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, and a plasma display. In some embodiments, the display device can be a touch screen.

[0117] To provide for interaction with a user, the electronic device can be a computer. The computer has a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0118] In the embodiments of the present application, the computer program / instruction is stored on the computer readable medium, and the computer program / instruction is executed by the processor to implement the steps of the method provided by any one or more of the embodiments. The computer readable medium can be included in the electronic device described in the above embodiments; or can exist separately and not be assembled into the device. The computer readable medium carries one or more computer readable instructions.

[0119] The memory 1102 can be used as a non-transitory computer readable storage medium for storing non-transitory software programs, non-transitory computer executable programs and modules. The processor 1101 executes various functions and data processing of the server by running the non-transitory software programs, instructions and modules stored in the memory 1102, so as to implement the program instructions / modules corresponding to the method provided by any one or more of the embodiments in the present application.

[0120] The memory 1102 can include a program region that can store an operating system, an application program required for at least one function, and a data region that can store data created according to usage of the electronic device, etc. In addition, the memory 1102 can include a high-speed random access memory, and also can include a non-transitory memory such as at least one of a magnetic disk storage device, a flash memory device, or other non-transitory solid state storage device. In some embodiments, the memory 1102 can optionally include a memory disposed remotely from the processor 1101, which can be connected to the electronic device through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0121] It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus.

[0122] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage device, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0123] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0124] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. For example, the embodiments can be implemented by an application specific integrated circuit (ASIC), a general purpose computer or any other similar hardware device. In some embodiments, the software programs of the embodiments can be executed by a processor to implement the above steps or functions. Similarly, the software programs (including related data structures) of the embodiments can be stored in a computer readable recording medium, such as a RAM memory, a magnetic or optical drive or a floppy disk and the like. In addition, some steps or functions of the embodiments can be implemented by hardware, such as a circuit cooperating with a processor to perform the steps or functions.

[0125] The computer program product provided by the embodiments includes one or more computer programs / instructions, which, when executed by a processor, generate all or part of the processes or functions described in the embodiments. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)) and the like.

[0126] The computer program product of the present application can be a computer program embedded in a computer readable medium such as hard disk, CD, DVD, flash memory, etc. The computer readable medium can be directly readable by a computer or the computer can read the computer readable medium through a drive unit. The computer program product can also be a downloadable computer program transmitted via a network, e.g., the Internet, a local area network, etc. The downloadable computer program can be transmitted via a wired network or a wireless network, e.g., a satellite network, a cellular network, etc. The computer program product can also be a computer program product that is directly loadable into the internal memory of the computer.

[0127] The scope of the present application is defined by the appended claims rather than by the description of the specification. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope. No feature of the application should be considered critical unless the claims expressly identify it as being essential to the application. Furthermore, no element, component, or step in the claims is to be construed as critical, material, or essential to the extent that it is not recited in each of the claims. The citation of references herein does not constitute a concession that such references are prior art.

[0128] The above description is only the specific implementation of the present application, but the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above examples should be regarded as exemplary and non-limiting.

Claims

1. A drawing processing method, characterized in that: include: Parsing the drawing to be processed to determine at least one functional area, and cutting the drawing to be processed based on the functional area to obtain at least one sub-drawing; The sub-image is identified, and text information associated with the sub-image is obtained.

2. The drawing processing method according to claim 1, characterized in that: The step of parsing the drawing to be processed, determining at least one functional area, and cutting the drawing to be processed based on the functional area to obtain at least one sub-drawing includes: Obtaining the drawing to be processed; Using a detection model to identify image regions in the drawing to be processed, determining a function category corresponding to the image region, forming the function regions according to the function type and the image region, and associating each function region with a set of region boundary coordinates; The drawing to be processed is cut according to the region boundary coordinates to obtain the sub-graph.

3. The drawing processing method according to claim 2, characterized in that: The functional categories include at least two categories of: solid view category, two-dimensional view category, section view category, exploded view category and technical requirement category.

4. The drawing processing method according to claim 2, characterized in that: The detection model includes: a feature extraction layer, a fusion layer, a check layer and a correlation layer; The detection model is used to identify the image area in the drawing to be processed, and the function category corresponding to the image area is determined. The function area is formed according to the function type and the image area, and each function area is associated with a set of area boundary coordinates, including: The drawing to be processed is input into the feature extraction layer, and a multi-scale feature map is extracted by the feature extraction layer; fusing semantic features and detail features of the multi-scale feature map through the fusion layer to obtain fused features; Processing the fused features through the inspection layer to obtain a functional category and a regression result of each candidate region, wherein the regression result includes the boundary coordinates of the candidate region and the confidence of the candidate region; A functional area is determined according to the confidence of the candidate area through the association layer, and the functional area is associated with the boundary coordinates.

5. The drawing processing method according to claim 2, characterized in that: The step of cutting the drawing to be processed according to the region boundary coordinates to obtain the sub-drawing includes: The functional area associated with the area boundary coordinates in the drawing to be processed is cut according to the area boundary coordinates to obtain the sub-drawing.

6. The drawing processing method according to claim 1, characterized in that: The identifying the sub-image and obtaining text information associated with the sub-image includes: Preprocessing the sub-image to obtain a candidate image, identifying the candidate image, determining position information of element data, identifying the element data according to the position information, and extracting text information; The text information is associated with the sub-image.

7. The drawing processing method according to claim 1, characterized in that: The drawings to be processed are in PDF format.

8. An electronic device, characterized in that: The electronic device comprises: one or more processors; and A memory storing computer program instructions, which, when executed, cause the processor to perform the steps of the method according to any one of claims 1 to 7.

9. A computer-readable medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Functional region recognition method and device, functional region model construction method and device and storage medium

    CN115424287A

  • Drawing text information extraction method, device and equipment

    CN115995092A

  • Engineering drawing identification method and device and storage medium

    CN117576717A

  • Industrial drawing key symbol semantic recognition method and system

    CN119229469A

Cited By

  • Image processing method and device, equipment, medium and product

    CN121544877A

  • CAD electrical scheme identification method and system

    CN121616894A