Engineering image-text cross-modal automatic reconstruction method based on large model
Through the YOLOv5 model and PDF analysis technology based on deep learning, the gas roadmap of IC manufacturing equipment is identified and reconstructed, and the cross-modal understanding of large models in industrial applications is solved, and efficient multi-source information fusion and process generation are achieved.
Patent Information
- Application Number
- CN202410156091.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-04
- Publication Date
- 2025-08-05
AI Technical Summary
Existing large-modal technology lacks the ability to understand, reason, and content generation in professional fields, and is difficult to directly use multimodal data in industrial applications. The computing power resources requirements for training and inference processes are high.
The YOLOv5 model based on deep learning is used to identify the main area and devices of engineering drawings, combine PDF analysis and LSD algorithm to identify straight lines, generate standard input documents for large language models, and achieve cross-modal automatic reconstruction through data set training and retraining.
It realizes the deep fusion and reconstruction of multi-source information of the gas circuit diagram of IC manufacturing equipment, automatically extracts key information and generates process generation models, improving the recognition rate and model applicability.
Smart Images

Figure CN120430135A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence multi-modal large models, and specifically relates to an engineering graphics and text cross-modal automatic reconstruction method based on a large model. Background Art
[0002] The intelligentization of the equipment manufacturing process includes user intention recognition, text intelligent understanding, process intelligent generation, quality intelligent detection, defect root cause intelligent backtracking, etc. Existing large model technologies lack the ability of cross-modal understanding, reasoning, and content generation in professional fields, and there are problems of high computing power resource requirements in the training and reasoning processes. The input and output of large models are general natural language data, while industrial applications have multi-modal data such as images, videos, texts, and time-series data, which are difficult to be directly used for large model training and reasoning. Therefore, data conversion is required. First, use intelligent recognition algorithms to identify the features of images, videos, and time-series data, study data-efficient generation and data conversion methods based on zero-shot or few-shot samples, and format the input and output of large models to meet the needs of domain-specific large model fine-tuning and industrial applications. Summary of the Invention
[0003] Regarding the characteristics of special gases used in the IC manufacturing process and the uses of gases, corresponding transportation and control systems need to be designed. The manufacturing process of the IC equipment gas cabinet mainly includes links such as drawing recognition and understanding, process planning, pipeline welding, pipeline assembly, and quality inspection. The present invention is an engineering graphics and text cross-modal automatic reconstruction method based on a large model, which adaptively recognizes and understands the user's gas circuit engineering drawings, generates corresponding feature files, and uses them as the standard input documents of the large language model for model fine-tuning, and finally forms an industry large model for automatic understanding of engineering graphics and text and process generation.
[0004] The technical solution adopted by the present invention to achieve the above object is: an engineering graphics and text cross-modal automatic reconstruction method based on a large model, including the following steps:
[0005] Obtain a drawing file for expressing a gas pipeline diagram;
[0006] Construct a deep learning network model;
[0007] Construct a data set according to the drawing file, input it into the deep learning network model for model training, and obtain a main area recognition result file and a device recognition result file;
[0008] After covering the text and devices in the main area of the drawing file, perform straight line recognition to generate a straight line recognition result file;
[0009] Integrate the main area recognition result file, the device recognition result file, and the straight line recognition result file into a description document for describing the pipeline gas circuit engineering drawing, and realize the graphic and text reconstruction of the gas pipeline diagram.
[0010] The drawing file is in PDF format and is used to express the pneumatic pipeline diagram, including legends and annotations.
[0011] To obtain the drawing file for expressing the pneumatic pipeline diagram, all text information in the drawing file is recognized through PDF parsing to form a text recognition description document.
[0012] To construct a dataset based on the drawing file, input it into a deep learning network model for model training, and obtain the main area recognition result and device recognition result of the drawing file, including the following steps:
[0013] Use a marking tool to mark the main pipeline area, legend area, and annotation area of each drawing file, and divide them into a training set and a validation set at a set ratio;
[0014] Input the training set into the YOLOv5 network for training to obtain a trained YOLOv5 main pipeline area target recognition network model;
[0015] Input the drawing file into the trained YOLOv5 main pipeline area target recognition network model to obtain the main pipeline area target detection frame; the target detection frame information is stored in the main area recognition result file.
[0016] To construct a dataset based on the drawing file, input it into a deep learning network model for model training, and obtain the main area recognition result and device recognition result of the drawing file, including the following steps:
[0017] Use a marking tool to mark the pipeline devices of each drawing file, and divide them into a training set and a validation set at a set ratio;
[0018] Input the training set into the YOLOv5 network for training to obtain a trained YOLOv5 device recognition network model;
[0019] Input the drawing file into the trained YOLOv5 device recognition network model to obtain the recognition frames of various devices; the recognition frame information is stored in the device recognition result file.
[0020] In the main area of the drawing file, after covering the text and devices, perform straight line recognition to generate a straight line recognition result document, including the following steps:
[0021] 1) Read the text recognition description document of the drawing file, obtain the positions of the text information, and fill the position boxes of all text with white pixels;
[0022] 2) Read the main area recognition result file, and fill and cover all pixels outside the main pipeline area target detection frame in the result map obtained in step 1) with white;
[0023] 3) Read the device recognition result file, and fill the positions of various recognized devices in the result graph obtained in step 2) with white to obtain the connection line graph after the main area is blocked by non-linear graphics;
[0024] 4) Use the LSD line detection algorithm to identify line segments, obtain the position information of all line segments, and write it into the line recognition result file.
[0025] The described generation of a description document for describing the pipeline gas circuit engineering drawing is used to input the description document of the engineering drawing and the corresponding processing prompt words into a pre-trained process model for retraining to obtain an engineering graphic analysis process model for generating processes.
[0026] An engineering graphic cross-modal automatic reconstruction system based on a large model includes:
[0027] A drawing acquisition module for acquiring a drawing file for expressing a gas circuit diagram;
[0028] A model construction module for constructing a deep learning network model;
[0029] A model training module for constructing a data set according to the drawing file, inputting it into the deep learning network model for model training, and obtaining a main area recognition result file and a device recognition result file;
[0030] A line recognition module for performing line recognition after blocking text and devices in the main area of the drawing file, and generating a line recognition result file;
[0031] A file integration module for integrating the main area recognition result file, the device recognition result file, and the line recognition result file into a description document for describing the pipeline gas circuit engineering drawing, realizing the graphic reconstruction of the gas circuit diagram.
[0032] An engineering graphic cross-modal automatic reconstruction device based on a large model includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement the described engineering graphic cross-modal automatic reconstruction method when executing the computer program.
[0033] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the described engineering graphic cross-modal automatic reconstruction method is implemented.
[0034] The present invention has the following beneficial effects and advantages:
[0035] 1) The present invention uses the deep learning yolov5 model to achieve the recognition of the main area of engineering images and the multi-object recognition of components. The model has a small number of parameters and a high recognition rate, and is applicable to the recognition of engineering graphic targets.
[0036] 2) The present invention comprehensively applies advanced technologies such as artificial intelligence, PDF parsing, and line segment recognition to achieve the perception of the information of the gas circuit diagram of IC manufacturing equipment, data analysis, and the deep integration and reconstruction of multi-source engineering information.
[0037] 3) In terms of graphic and text parsing, the present invention is an automatic extraction technology for image and annotation information in engineering drawings based on large model technology, which can automatically extract key information such as dimensions, geometric features, and processing standards from engineering drawings, use large language models for reasoning, and convert the reasoning results into target data of specific types. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a schematic flow diagram of the present invention;
[0039] Figure 2 is the YOLOv5 model structure of the present invention;
[0040] Figure 3a is a diagram of the training result of the main area recognition of the present invention;
[0041] Figure 3b is a diagram of the training result of the device recognition of the present invention;
[0042] Figure 4a is a diagram of the device recognition result of the present invention;
[0043] Figure 4b is a diagram of the line segment recognition result of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The following further describes the present invention in detail with reference to the drawings and embodiments.
[0045] The present invention is a cross-modal automatic reconstruction method for engineering graphics and texts based on a large model, which adaptively recognizes and understands the gas circuit engineering drawings of users, generates corresponding feature files, fine-tunes the model as the standard input document of the large language model, and finally forms an industry large model for automatic understanding of engineering graphics and texts and process generation.
[0046] For the gas path diagram of integrated circuit manufacturing equipment, this invention first uses the PDF parsing method to parse all the text in the diagram and write it into a text parsing result file. Then, the main area of the image is recognized by the YOLOv5 training model and written into the main area recognition result file. Next, various devices are recognized by the YOLOv5 training model and written into the device recognition result file. The main area of the engineering drawing is obtained through the overlay function, and the text and devices in the main area are filtered out. The straight lines are recognized by the LSD algorithm and written into the straight line recognition result file. All the parsing result files are integrated to form the input parsing file of the multi-modal large model for engineering drawings. The drawing parsing file and the corresponding process processing prompt words are input into the pre-trained general large model for retraining to obtain the engineering drawing text process generation large model.
[0047] The user's drawing file is in PDF format, mainly containing the gas path pipeline diagram, with legends and annotations. The PDF file structure is a directed graph formed by connecting pages, metadata, fonts, and resources. PDFminer is a Python library file that provides a series of functions for parsing PDFs. PDFParser is a library function in PDFminer that can extract data from files. This invention extracts the text in the PDF diagram through the LTTextBox component in the LTPage function of PDFParser and saves it as a CSV format description document, which stores text information, location information, etc.
[0048] Since the PDF parsing method cannot understand the overall structure of device graphics, the deep learning object detection method is used for the recognition of the main area of the engineering drawing and component recognition. Since the deep learning method requires a certain amount of training data, and only 9 engineering drawings were collected in the early stage with a small sample size, data augmentation is needed. Through the PDF editor, first convert the 9 PDF drawings into 9 original images, set the resolution to 37.8 pixels / cm, and obtain an image resolution of 3178*2245 for each image. The 9 original images are sequentially subjected to grayscale transformation, random rotation, horizontal rotation, and vertical rotation based on the torchvision.transoforms library to obtain 9*2 4 = 144 images, where the grayscale transformation adjusts the brightness, contrast, saturation, and hue to 0.5 ratio, and the random rotation angle is 3°.
[0049] The 144 original drawings are labeled with the main pipeline area, legend area, and annotation area of each drawing through the LabelImg annotation tool and divided into a training set and a validation set at a set ratio; the training set is input into the YOLOv5 network for training to obtain a trained YOLOv5 network model. The original drawings are input into the trained network model to obtain the target detection boxes for the main pipeline area, and the recognition information and location information are stored in a txt file as the main area recognition result.
[0050] 144 original drawings are marked with pipeline devices on each drawing using the LabelImg annotation tool, which are respectively manual valves (MV), normally closed pneumatic valves (PV), normally closed three-way pneumatic valves (PV3), flow controllers (MFC), filters (FV), pressure regulators (PR), and pressure switches (PS), and are divided into a training set and a validation set according to a set ratio; the training set is input into the YOLOv5 network for training to obtain a trained YOLOv5 device recognition network model. The original drawings are input into the trained network model to obtain the recognition frames of various devices. The recognition information and position information are stored in a txt file as the device recognition result.
[0051] To identify the connection lines in the drawings, non-linear content in the drawings needs to be removed first, that is, in the main area of the original drawing, after obscuring the text and devices, line recognition is performed. The present invention designs 3 occlusion steps, which are executed in sequence: 1. Read the PDF text recognition description document to obtain the positions of the text information, and fill the positions of all the text with white pixels. 2. Read the main area recognition result txt file, and fill all the pixels outside the detection frame of the main pipeline area with white in the result map of step 1. 3. Read the device recognition result txt file, and fill the positions of the recognized various devices with white in the result map of step 2. After step 3 is executed, the connection line diagram after obscuring non-linear graphics in the main area is obtained.
[0052] The present invention uses the LSD fast line detection algorithm to identify line segments and obtains the position information of all line segments. The LSD algorithm analyzes the local area of the image to obtain the pixel point set of the line, and then verifies and solves it through hypothesis parameters, combines the pixel point set with the error control set, and then adaptively controls the number of false detections. Write the line segment recognition position information into a txt document as the line recognition result.
[0053] Read the text recognition result description document, main area recognition result txt, device recognition txt, and line segment recognition result txt obtained by the present invention, integrate them into a csv file to obtain a pipeline gas circuit engineering drawing description document, and input the engineering drawing description document and the corresponding processing technology prompt words into a pre-trained general large model for retraining to obtain an engineering graphics parsing process generation large model.
[0054] Embodiment
[0055] The experimental environment based on the present invention is as follows: The hardware configuration of the PC is i7-12700 2.1GHz, 32G memory, GPU4060ti, the software environment is based on the Window11 operating system, and the PyCharm2020.3.3 compilation environment manages deep learning libraries such as python3.10 and pytorch2.0.0.
[0056] The present invention proposes an engineering graphic and text cross-modal automatic reconstruction method based on a large model. For the gas circuit diagram of integrated circuit manufacturing equipment, all information such as text, main regions, devices, line segments, etc. in the above figure is parsed to form an input parsing file for the engineering drawing multi-modal large model. The drawing parsing file and the corresponding processing technology prompt words are input into a pre-trained general large model for retraining to obtain an engineering graphic and text process generation large model. The flowchart of the present invention is as Figure 1 shown.
[0057] The present invention uses the YOLOv5 model network structure to identify the main regions and various devices of engineering drawings. The YOLOv5 model network structure consists of four main parts, as Figure 2 shown as Input, Backbone, Neck, and Output respectively. Input adopts the Mosaic data augmentation method and can simultaneously perform adaptive scaling on the image, automatically calculating the optimal anchor box value of the dataset. Backbone is mainly composed of the Focus structure and the CSPNet structure. The main function of the Focus structure is to slice the image, turning the original image with a resolution of 640×640×3 into a feature image of 320×320×32. CSPNet is a cross-stage local fusion network, and its main function is to extract the features of the feature map, obtaining the features of different layers to enrich the image information. In the Neck part, the FPN+PAN structure is used. FPN uses the upsampling method to obtain the spliced feature map, and then aggregates the shallow features from bottom to top through the PAN network structure to fully integrate the image features of each layer. In the Output part of YOLOv5, CIOU_Loss is used as the loss function, and non-maximum suppression is used to filter out redundant bounding boxes, predicting the image features and finding the best detection position.
[0058] The present invention extracts the text in the PDF diagram through the LTTextBox component in the LTPage function of PDFParser and saves it as a description document in csv format, which stores text information, position information, etc.
[0059] The present invention uses the method of deep learning object detection for the recognition of the main regions and components of engineering drawings. Since the deep learning method requires a certain amount of training data, and only 9 engineering drawings were collected in the early stage with a small number of samples, data augmentation is needed. Through the PDF editor, the 9 PDF drawings are first converted into 9 original images, with the resolution set to 37.8 pixels / cm, and the resolution of each image is 3178*2245. The 9 original images are sequentially subjected to grayscale transformation, random rotation, horizontal rotation, and vertical rotation based on the torchvision.transoforms library to obtain 9*2 4= 144 images, where the grayscale transformation adjusts the brightness, contrast, saturation, and hue to a proportion of 0.5, and the random rotation angle is 3°.
[0060] Label each of the 144 original drawings with the main pipeline area, legend area, and annotation area using the LabelImg annotation tool, and divide them into a training set and a validation set at a ratio of 8:2; there are 115 images in the training set and 29 images in the test set. Use yolov5x as the pre-trained model, and at the same time modify the model configuration file, change the number of model classifications to 3, and change the classification categories to the pipeline area, legend area, and text area. Input the training set into the YOLOv5 network for training, with the number of iterations being 100, to obtain a trained YOLOv5 network model. The training results of the training set are as Figure 3a shown. Among them, mAP_0.5 is the average precision of all recognized categories when the intersection-over-union threshold is 0.5, and mAP_0.95 is the average precision of all recognized categories when the intersection-over-union threshold is 0.95. Percision is the recognition accuracy rate, recall is the recall rate. Box_loss is the localization loss, representing the error between the predicted box and the calibrated box. Cls_loss is the classification loss, representing whether the calculated anchor box and the corresponding calibrated classification are correct. Obj_loss is the confidence loss, representing the calculation of the confidence of the network. Input the original drawing into the trained network model to obtain the target detection box for the main pipeline area, and store the recognition information and location information in a txt file, which is the recognition result for the main area.
[0061] Label the pipeline devices of each of the 144 original drawings using the LabelImg annotation tool, and divide them into a training set and a validation set at a ratio of 8:2; there are 115 images in the training set and 29 images in the test set. Use yolov5x as the pre-trained model, and at the same time modify the model configuration file, change the number of model classifications to 7, which are manual valve (MV), normally closed pneumatic valve (PV), normally closed three-way pneumatic valve (PV3), flow controller (MFC), filter (FV), pressure regulator (PR), and pressure switch (PS), input the training set into the YOLOv5 network for training, with the number of iterations being 300, to obtain a trained YOLOv5 device recognition network model. The training results of the training set are as Figure 3b shown. Input the original drawing into the trained network model to obtain the recognition boxes for various devices, as Figure 4a shown, and store the recognition information and location information in a txt file, which is the device recognition result.
[0062] When identifying connecting lines in engineering drawings, non - straight content in the drawing needs to be removed first. That is, in the main area of the original drawing, after covering the occluded text and devices, straight - line recognition is performed. The present invention designs 3 occlusion steps, which are executed in sequence: 1. Read the PDF text recognition description document to obtain the positions of text information, and fill the position frames of all text with white pixels. 2. Read the main - area recognition result txt file, and fill all pixels outside the detection frame of the main pipeline area with white in the result graph of step 1. 3. Read the device recognition result txt file, and fill the positions of various recognized devices with white in the result graph of step 2. After step 3 is executed, the connecting - line graph after occluding non - straight graphics in the main area is obtained.
[0063] The present invention uses the LSD fast line detection algorithm to identify line segments and obtains the position information of all line segments, as Figure 4b shown. The LSD algorithm analyzes the local area of the image to obtain the pixel point set of the line, then verifies and solves it through assumed parameters, combines the pixel point set with the error control set, and then adaptively controls the number of false detections. Write the line - segment recognition position information into a txt document as the straight - line recognition result.
[0064] Read the text recognition result description document, the main - area recognition result txt, the device recognition txt, and the line - segment recognition result txt obtained by the present invention, integrate them into a csv file to obtain the pipeline gas - path engineering drawing description document, and input the engineering drawing description document and the corresponding processing - technology prompt words into a pre - trained general large - model for retraining to obtain an engineering graphic - parsing process - generation large - model.
Claims
1. A large-scale model-based cross-modal automatic reconstruction method for engineering graphics, characterized by: The following steps are involved: Obtain the drawing file used to express the gas line and pipeline diagram; Build deep learning network models; Build a data set based on the drawing file, input it into the deep learning network model for model training, and obtain the main area recognition result file and the device recognition result file; After blocking text and components in the main area of the drawing file, perform straight line recognition and generate a straight line recognition result file; The main area recognition result file, the device recognition result file and the straight line recognition result file are integrated into a description document for describing the pipeline gas path engineering drawing, thereby realizing the graphic and text reconstruction of the gas path pipeline diagram.
2. The method for automatic cross-modal reconstruction of engineering graphics based on a large model according to claim 1, characterized in that: The drawing file is in PDF format and is used to express the gas pipeline diagram, including legends and annotations.
3. The method for automatic cross-modal reconstruction of engineering graphics based on a large model according to claim 1, characterized in that: The method comprises obtaining a drawing file for expressing a gas circuit diagram, identifying all text information in the drawing file through PDF parsing, and forming a text recognition description document.
4. The method for automatic cross-modal reconstruction of engineering graphics based on a large model according to claim 1, characterized in that: The method of constructing a data set based on the drawing file and inputting it into the deep learning network model for model training to obtain the main area recognition result and device recognition result of the drawing file includes the following steps: Use annotation tools to mark the main route area, legend area, and annotation area of each drawing file, and divide it into training set and verification set according to the set ratio; The training set is input into the YOLOv5 network for training to obtain the trained YOLOv5 main road area target recognition network model; The drawing file is input into the trained YOLOv5 main road area target recognition network model to obtain the main road area target detection frame; the target detection frame information is stored in the main area recognition result file.
5. The method for automatic cross-modal reconstruction of engineering graphics based on a large model according to claim 1, characterized in that: The method of constructing a data set based on the drawing file and inputting it into the deep learning network model for model training to obtain the main area recognition result and device recognition result of the drawing file includes the following steps: Use annotation tools to mark the piping components in each drawing file and divide them into training and validation sets according to the set ratio; Input the training set into the YOLOv5 network for training to obtain a trained YOLOv5 device recognition network model; The drawing file is input into the trained YOLOv5 device recognition network model to obtain the recognition boxes of various devices; the recognition box information is stored in the device recognition result file.
6. The method for automatic cross-modal reconstruction of engineering graphics based on a large model according to claim 1, characterized in that: In the main area of the drawing file, after blocking text and components, perform line recognition and generate a line recognition result document, including the following steps: 1) Read the text recognition description document of the drawing file, obtain the position of the text information, and fill the position box of all text with white pixels; 2) Read the main area recognition result file, and fill all pixels outside the main road area target detection frame in the result image obtained in step 1) with white; 3) Reading the device identification result file, filling the positions of various devices identified in the result diagram obtained in step 2) with white, and obtaining a connection line diagram after the main area blocks the non-linear graphics; 4) Use the LSD line detection algorithm to identify line segments, obtain the position information of all line segments, and write the line recognition result file.
7. The method for automatic cross-modal reconstruction of engineering graphics based on a large model according to claim 1, characterized in that: The description document for describing the pipeline gas path engineering drawing is generated, and the description document of the engineering drawing and the corresponding processing prompt words are input into the pre-trained process model for retraining to obtain the engineering drawing and text parsing process model for generating the process.
8. A large-scale model-based cross-modal automatic reconstruction system for engineering graphics, characterized by: include: A drawing acquisition module is used to acquire a drawing file for expressing a gas circuit and pipeline diagram; Model building module, used to build deep learning network models; The model training module is used to construct a data set based on the drawing file, input it into the deep learning network model for model training, and obtain the main area recognition result file and the device recognition result file; The straight line recognition module is used to identify straight lines after blocking text and devices in the main area of the drawing file and generate a straight line recognition result file; The file integration module is used to integrate the main area recognition result file, the device recognition result file and the straight line recognition result file into a description document for describing the pipeline gas path engineering drawing, thereby realizing the graphic and text reconstruction of the gas path pipeline diagram.
9. A large-scale model-based cross-modal automatic reconstruction device for engineering graphics, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement a large model-based cross-modal automatic reconstruction method for engineering graphics as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by the processor, the method for automatic cross-modal reconstruction of engineering drawings and texts based on a large model as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Method and device for automatically designing gas path diagram
CN120974674A
Method and apparatus for automatically designing a pneumatic circuit
CN120974674B