Structured data generation method for engineering field drawings

By classifying and normalizing drawings in the engineering field, combining multi-dimensional reinforcement learning strategies and Vision Transformer, a solution that can automatically output structured data from engineering drawings in multiple formats is generated, solving the problem of low plug-in dependence and recognition accuracy in the existing technology.

CN120564210APending Publication Date: 2025-08-29CHINA HAISUM ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510631944.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

The prior art cannot directly generate structured data from drawings in engineering fields in different formats, and requires the development of specific plug-ins and optical character recognition technology has low accuracy in recognition of fuzzy or complex drawings, resulting in the inability to directly extract structured data.

Method used

By classifying and normalizing drawings in the engineering field, fine-tuning the base model using a multi-dimensional reinforcement learning strategy, combining Vision Transformer and large-model knowledge distillation to generate structured data, including image vector embedding and text vector embedding, performing graphic and text modal alignment, and finally output structured data.

Benefits of technology

It realizes automatic generation of structured data from engineering drawings in multiple formats, solves the problems of plug-in dependence and manual annotation, and improves the recognition accuracy of fuzzy or complex drawings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564210A_ABST
    Figure CN120564210A_ABST
Patent Text Reader

Abstract

The invention discloses a method for generating structured data of engineering field drawings, which comprises the following steps of: extracting data in different formats and drawing field types from the engineering field drawings in different file formats, and normalizing the data in different formats and the engineering field drawings, the method comprises the following steps: performing corresponding field matching according to a drawing field type and a multi-field sample image-text corpus, selecting a sample with high similarity in a matched field, constructing a cue word context for the sample, generating structured data based on large model knowledge distillation, and filtering to obtain fine-tuning corpora of different engineering fields; image vector embedding and text vector embedding are obtained based on fine-tuning corpora in different engineering fields, fine-tuning training is carried out in combination with the base model, an engineering field drawing analysis model is obtained, and structural data are obtained by using the engineering field drawing analysis model. The method solves the problems that the structured data can be generated only when multiple engineering field plug-ins are developed to generate full-design process data, and the structured data need to be manually annotated in paper drawings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of engineering drawing processing, and in particular relates to a method for generating structured data of engineering drawings. Background Art

[0002] Engineering drawings are generally designed using engineering software. By developing corresponding plug-ins for engineering software, the entire design process data of engineering drawing development is recorded and output as PDF documents or printed as paper documents. Finally, the entire design process data of PDF documents or paper documents is processed based on optical character recognition (OCR) to generate structured data.

[0003] Current structured data generation methods have the following shortcomings:

[0004] 1) The full design process data must be generated based on the corresponding plug-in of the engineering software, which makes it impossible to directly generate the full design process data based on existing engineering field drawings, and further makes it impossible to generate the structured data corresponding to the existing engineering field drawings.

[0005] 2) Since engineering drawings belong to different engineering fields, in order to record the full process data of different engineering fields, it is necessary to customize and develop plug-ins for specific engineering fields.

[0006] 3) It is necessary to indirectly convert the data of the entire design process into structured data, and it is impossible to directly extract structured data from engineering drawings.

[0007] 4) For paper drawings without full design process data, optical character recognition technology is generally used to identify engineering drawings and manually annotate them as structured data. However, optical character recognition technology can only recognize relatively clear engineering drawings and paper scans in PDF format with simple line structures. The recognition accuracy is poor for engineering drawings that are blurred and have complex line structures. Summary of the Invention

[0008] The purpose of the present invention is to classify engineering field drawings into engineering field categories, normalize engineering field drawings of different input formats, automatically generate fine-tuning corpus corresponding to the engineering field drawings, fine-tune the base model based on the above-mentioned fine-tuning corpus using a multi-dimensional reinforcement learning strategy, and finally obtain a fine-tuned engineering field drawing parsing model, use the engineering field drawing parsing model to parse the engineering field drawings, and finally output structured data of the engineering field drawings.

[0009] In order to achieve the above-mentioned object of the invention, the technical solution of the present invention provides a method for generating structured data of engineering drawings, comprising the following steps:

[0010] For engineering drawings in CAD file format, the corresponding CAD text information of the engineering drawings is extracted; for engineering drawings in PDF format, the corresponding PDF text information of the engineering drawings is extracted. The CAD text information and PDF text information include title bars, annotation information and tables.

[0011] According to the keyword library of different engineering fields, the CAD text information and PDF text information are matched and extracted with domain keywords. If the match is successful, the CAD keyword information after engineering field extraction and the PDF keyword information after engineering field extraction as well as the corresponding drawing field type are extracted.

[0012] For engineering drawings in hand-drawn and / or scanned formats, as well as engineering drawings that failed to match, Vision Transformer is used in combination with keyword libraries from different engineering fields to parse them. If the confidence reaches the preset confidence level, information extracted in other formats from different engineering fields is obtained. If the preset confidence level is not reached, manual intervention is performed to directly classify and obtain information extracted in other formats from the engineering field and the corresponding drawing field type.

[0013] The engineering drawings are marked according to the corresponding drawing field types, and normalized with the CAD keyword information extracted from the engineering field, the PDF keyword information extracted from the engineering field, and the information extracted from other formats in the engineering field. The normalized text data is formatted and stored using json, and finally the normalized graphic and text data is obtained.

[0014] According to the drawing domain type, the corresponding domain is matched with the multi-domain sample graphic and text corpus. The normalized graphic and text data is used to perform similarity matching of graphic and text feature vectors in the matching domain. Samples with high similarity are selected from the multi-domain sample graphic and text corpus according to the preset proportion. The samples are used as prompt words to construct the prompt word context. Structured data is generated based on large-scale model knowledge distillation. The rejection sampling method based on confidence is used to filter the structured data with low confidence. The drawings in the normalized graphic and text data and the corresponding structured data pairs are obtained as fine-tuning corpora for different engineering fields. The multi-domain sample graphic and text corpus internally stores typical graphic and text corpora for different engineering fields. The typical graphic and text corpora are used to describe typical engineering drawings in the engineering field and the corresponding structured data pairs.

[0015] Drawings in fine-tuned corpora from different engineering fields are encoded to generate image vector embeddings. Texts in fine-tuned corpora from different engineering fields are segmented and embedded to obtain text vector embeddings. For image vector embeddings and text vector embeddings, image-text modality alignment is performed based on image-text embeddings to generate a batch matrix of input data to be trained. Based on the input data matrix, multi-dimensional reinforcement learning is used to fine-tune the weight matrix of the base model, and finally a fine-tuned engineering field drawing parsing model is obtained. The engineering field drawing parsing model is used for reasoning, with the reasoning input being the engineering field drawings to be parsed, and finally the structured data of the engineering field drawings is output.

[0016] Preferably, the CAD text information is extracted from engineering drawings in CAD file format using an API or ezdxf python parsing library.

[0017] Preferably, the CAD text information is extracted from engineering drawings in PDF format using optical character recognition based on keywords.

[0018] Preferably, the confidence level is calculated using a Softmax-based probability method.

[0019] Preferably, the normalization process includes uniformly converting file formats, coordinates and units.

[0020] Preferably, the image-text vector similarity matching is cosine similarity matching.

[0021] Preferably, the fine-tuning of the base model weight matrix using multi-dimensional reinforcement learning based on the input data matrix comprises the following steps:

[0022] Generate multi-dimensional scores through the multi-value head output of the reward strategy and reward model;

[0023] Based on the standardization of the grouped dimension scores, the relative advantages of different outputs are generated according to the advantage calculation formula;

[0024] Generate a KL penalty based on the log-cumulative probability of the output of the initial reference model and the training model;

[0025] Based on the above logarithmic cumulative generation probability, relative advantage, the single output training loss is calculated according to the loss calculation formula and KL penalty;

[0026] Based on the above single-output training loss, the multi-output normalization within the group is performed to obtain the comprehensive loss of the single input;

[0027] Update the model weights based on the input batch and gradient accumulation.

[0028] Preferably, the advantage calculation formula is as follows:

[0029]

[0030] Among them, r i,j represents the reward for the jth answer of sample i, G represents the total G answers generated by sample i, and ε is a small parameter coefficient to prevent division by zero.

[0031] Preferably, the loss calculation formula is as follows:

[0032]

[0033] Among them, i represents the i-th answer of a sample, t represents the t-th token in the i-th answer, A i represents the advantage of sample i, m i,t Indicates the mask of the t-th position of the i-th sample, L i It represents the loss of the i-th sample calculated in combination with the penalty coefficient β of the KL divergence, and B represents the number of samples in a batch.

[0034] The technical solution of the present invention provides a method for generating structured data of engineering field drawings, which includes: extracting different format data and drawing field types from engineering field drawings of different file formats, normalizing the different format data and engineering field drawings, matching the corresponding fields with the multi-field sample graphic and text corpus according to the drawing field type, selecting samples with high similarity in the matching fields, constructing prompt word contexts for the samples and generating structured data based on large model knowledge distillation, filtering to obtain fine-tuning corpora of different engineering fields, obtaining image vector embedding and text vector embedding based on the fine-tuning corpora of different engineering fields, and fine-tuning training with a base model to obtain an engineering field drawing parsing model, and using the engineering field drawing parsing model to obtain structured data. This method solves the problem that multiple engineering field plug-ins must be developed to generate full design process data in order to generate structured data, and that paper drawings require manual annotation of structured data. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A specific flow chart for a typical drawing field classification for drawings input in different formats;

[0036] Figure 2 It is a typical JSON format for normalized drawing text information in the mechanical field;

[0037] Figure 3 A typical multi-domain sample image and text corpus and a flowchart for similarity matching based on the sample image and text corpus to generate a structured corpus of drawings;

[0038] Figure 4 Design drawings for a typical pump in the mechanical field;

[0039] Figure 5 For Figure 4 Parsed structured data pair information;

[0040] Figure 6 Design drawings for pumps, another typical mechanical field;

[0041] Figure 7 For Figure 5 Parsed structured data pair information;

[0042] Figure 8 It is another typical json format for normalized drawing text information in the mechanical field;

[0043] Figure 9 Fine-tune the structured data pair information of a typical mechanical field corpus in JSON format;

[0044] Figure 10 It is a typical JSON format for normalized drawing text information in the electrical field;

[0045] Figure 11 Fine-tune the structured data pair information of a typical electrical field corpus in JSON format;

[0046] Figure 12 Normalize the drawing text information for a typical architectural field in JSON format;

[0047] Figure 13 Fine-tune the structured data pair information of a typical JSON-formatted architectural domain corpus;

[0048] Figure 14 This is a typical batch fine-tuning corpus input. The training process of a batch of fine-tuning corpus is completed through multimodal encoding and modality alignment, and multi-dimensional reinforcement learning.

[0049] Figure 15 Diagram of the architecture for generating multi-dimensional scores from a multi-value head output of a multi-dimensional reward strategy and reward model. DETAILED DESCRIPTION

[0050] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0051] An embodiment of the present invention provides a method for generating structured data of drawings in the engineering field, comprising the following steps:

[0052] Step 1: Domain identification of engineering drawings, such as Figure 1 shown.

[0053] If the engineering drawings are in CAD file format, the corresponding CAD text information is extracted using the API or the ezdxf Python parsing library. If the engineering drawings are in PDF format, keyword optical character recognition is used to extract the corresponding PDF text information. CAD and PDF text information includes title blocks, annotations, and tables.

[0054] Specifically, for specific CAD file types, such as dxf format files, the ezdxf Python library is used to parse and extract text annotations. For PDF files, OCR + keyword matching is used, such as tesseract OCR or EasyOCR to read the text content in the drawings.

[0055] According to the keyword library of different engineering fields, the CAD text information and the PDF text information are matched and extracted with domain keywords. If the match is successful, the CAD keyword information after engineering field extraction and the PDF keyword information after engineering field extraction are obtained.

[0056] Engineering domain classification matching based on domain keyword library, for example:

[0057] Common keywords in mechanical drawings: diameter M (thread), shaft, gear.

[0058] Common keywords in electrical drawings: voltage, current, circuit, terminal.

[0059] Common keywords in architectural drawings: wall, column, beam, floor.

[0060] In order to build an effective engineering keyword library, different engineering fields should have relatively independent and representative keyword sets. The selection of keywords should meet the following conditions:

[0061] 1. Able to accurately reflect the typical structure, symbols or terminology of drawings in this engineering field;

[0062] 2. It has a high frequency of appearance and representativeness in drawings in this field;

[0063] 3. It has good recognition friendliness for image recognition models and is easy to extract automatically;

[0064] 4. It is highly distinguishable from keywords in other engineering fields to support subsequent accurate classification.

[0065] If the matching fails or the format of the engineering drawings is hand-drawn drawings and / or scanned drawings, VisionTransformer is used in combination with keyword libraries of different engineering fields and the confidence level is calculated. If the confidence level reaches the preset confidence level, information extracted in other formats of the engineering field is obtained. If the preset confidence level is not reached, manual intervention is performed to directly classify and obtain information extracted in other formats of the engineering field.

[0066] Specifically, for other formats such as hand-drawn drawings and / or scanned images, Vision Transformer is used in combination with a domain keyword library for domain classification, and parsing confidence is used to measure the accuracy of the classification results to determine whether manual intervention is required.

[0067] The confidence calculation adopts the probability calculation based on Softmax, and the combination with the domain keyword library adopts the fewshots method, that is, the prompt words provide typical hand-drawn drawings in various fields, the actual classification type, and the relevant classification basis, including the domain keywords in the description drawings, the domain-specific drawing components in the drawings, etc., to provide classification examples for the model.

[0068] Step 2: Generate domain-based fine-tuning corpus.

[0069] The engineering field drawings, their corresponding field types, CAD keyword information after engineering field extraction, PDF keyword information after engineering field extraction, and other format extraction information of the engineering field, etc. are normalized in file format, coordinates and units, and the normalized data is stored using json.

[0070] Specifically, due to the inconsistency of formats in engineering drawings, CAD keywords after extraction from different engineering fields, PDF keywords after extraction from different engineering fields, and information extracted in other formats from different engineering fields, it is necessary to normalize the formats before fine-tuning to ensure the consistency of the overall format.

[0071] 1) File format normalization: Multiple file formats (DWG, DXF, PDF, JPG) → unified conversion to a standard format (SVG / PNG)

[0072] 2) Coordinate normalization: (CAD absolute coordinates vs PDF image relative coordinates) → uniformly converted to normalized coordinates (0-1)

[0073] 3) Unit normalization: (mm, cm, inch) → uniformly converted to millimeters (mm).

[0074] Figure 2 This is the normalized result of a drawing in the mechanical field.

[0075] The engineering field drawing information is formatted and stored using JSON to generate formatted text information, including the drawing field type, drawing text information, etc. The drawing field type is generated based on step 1 and is used for similarity matching based on the field sample image and text corpus in the current step.

[0076] like Figure 3 As shown in the figure, a multi-domain sample image-text corpus is constructed by generating a multi-level structured data pair. Each sample corresponds to a image-text feature vector. In the actual corpus generation process, the sample image-text corpus of the corresponding field is matched according to the drawing field type. The normalized drawings and the above-mentioned formatted text data are used to perform image-text vector similarity matching to obtain the top 3 samples. The top 3 samples are used as prompt words to construct the prompt word context, and the structured data pair of the current drawing is generated based on the large model knowledge distillation. The image-text similarity matching uses the image-text feature vector for cosine similarity matching. The confidence-based rejection sampling method is used to filter out low-quality generated structured data pairs, and only retains the structured data pairs with high confidence to avoid low-quality corpus pairs causing noise and interference in the actual fine-tuning process.

[0077] Among them, the construction process of the multi-domain sample graphic and text corpus includes: for different engineering fields (mechanical, electrical, construction, etc.), experts from different engineering fields construct a batch of high-quality typical graphic and text corpus pairs in the format of engineering field drawings + structured data pairs, which are required to cover various typical drawing scenarios in the engineering field, and the structured data pairs are presented in key-value form.

[0078] The key setting rules are as follows: if the engineering drawings contain key descriptions, use them directly; if the engineering drawings do not contain key descriptions, the actual key description of the component must be derived based on the overall information of the engineering drawings. Figure 4 and Figure 6 , corresponding to structured data pairs such as Figure 5 and Figure 7 shown.

[0079] Figure 4 Corresponding structured data pairs Figure 5 The fourth key-value pair corresponds to the core size information of the pump inlet and outlet, but the figure does not indicate that it is the inlet and outlet of the pump body. Therefore, it is necessary to give a reasonable description of the key in combination with the overall information of engineering drawing 1.

[0080] Examples of structured data pairs in different engineering fields (mechanical, electrical and architectural) are as follows: Figures 8-13 shown.

[0081] Step 3: RL-based fine-tuning training.

[0082] like Figure 14As shown in the figure, after generating fine-tuning corpora for each domain in step 2, the corpus is encoded and embedded. The actual corpus is a graphic-text corpus in the format of drawings + structured data pairs. For images, the CNN / ViT image encoder is used to encode and generate image vector embeddings. For structured data, the word segmenter is used to perform text segmentation. The text vector embeddings are then obtained by word segmentation embedding through the model's embedding layer. Based on the graphic-text embeddings, the graphic-text modality is aligned to generate a batch matrix of input data to be trained. Based on this input data matrix, the weight matrix of the base model is fine-tuned using multi-dimensional reinforcement learning. Finally, a fine-tuned engineering domain drawing parsing model is obtained (this model is fine-tuned and trained based on the above-mentioned multi-domain engineering fine-tuning corpus. Compared with the base model, its model weight matrix contains domain information). This domain model is used for reasoning, with the reasoning input being the engineering drawing to be parsed, and the structured data of the drawing is finally output.

[0083] Among them, an algorithm logic that can perform reinforcement learning based on multi-dimensional scoring is designed. The base model is fine-tuned based on the input data matrix after the above modal alignment to fine-tune the base model weight matrix. Finally, a fine-tuned engineering field drawing parsing model is obtained. The specific steps are as follows:

[0084] like Figure 15 As shown, multi-dimensional scores are generated through the multi-value head output of the reward strategy and reward model, and the dimensional scores are standardized based on the grouping. The relative advantages of different outputs are generated according to the advantage calculation formula. The KL penalty is generated based on the logarithmic cumulative probability of the output of the initial reference model and the training model. Based on the above logarithmic cumulative generation probability and relative advantage, the single-output training loss is calculated according to the loss calculation formula and the KL penalty. Based on the above single-output training loss, the multi-output standardization within the group is performed to obtain the comprehensive loss of the single input. The weight of the model is updated based on the input batch and gradient accumulation.

[0085] The advantage calculation formula is as follows:

[0086]

[0087] Among them, r i,j represents the reward for the jth answer of sample i, S z The score generated by the z-th reward strategy, P z The weighted average weight value of the z-th reward strategy.

[0088] This term indicates that the reward for the j-th answer of sample i is obtained by the weighted average of the reward scores of K reward strategies.

[0089]

[0090] Here, G represents the total number of G answers generated by sample i.

[0091] This term represents the average reward of G answers generated by sample i.

[0092]

[0093] This term represents the standard deviation of the G responses generated by sample i.

[0094]

[0095] Here, ε prevents small parameter coefficients from dividing by zero.

[0096] This item expresses the calculated relative advantage of the j-th answer of sample i.

[0097] The loss calculation formula is as follows:

[0098] KL i,t =exp(log p ref,i,t -logp i,t )-(logp ref,i,t -logp i,t )-1

[0099] Where i represents the i-th answer of a sample, t represents the t-th token in the i-th answer, log p represents the logarithmic probability of each token of the current model (per_token_logps), log p ref Represents the log probability of the reference model (ref_model).

[0100] This item uses the second-order expansion approximation of Reverse KL to measure the difference between the token probability distribution output by the current model and the reference model, and penalizes the model for generating outputs that are too different from the reference model, thereby maintaining the stability of training.

[0101]

[0102] A i represents the advantage of sample i, m i,t represents the mask of the t-th position of the i-th sample, and β represents the penalty coefficient of KL divergence.

[0103] This term calculates the loss for each token, amplifies the reward signal through the advantage function, and adds a KL divergence penalty to limit the degree of deviation from the reference policy, thereby balancing exploration and stability during optimization.

[0104]

[0105] L i Represents the loss of the i-th sample.

[0106] This item calculates the average loss of each sample by weighting the loss items of all generated tokens by mask, and only calculates the valid (masked) token part.

[0107]

[0108] B represents the number of samples in a batch.

[0109] This term averages the loss of all samples in the batch to obtain the overall loss for backpropagation.

[0110] The above reward strategy can include dimensions such as parsing accuracy, confidence, stability, and error correction capability. The reward score is obtained through the multi-dimensional reward strategy to calculate the advantage. The specific reward strategy is described as follows:

[0111] 1) Parsing accuracy reward Racc is used to reward high parsing accuracy and ensure parsing accuracy, where Racc = 10 × total number of correctly parsed data items / total number of data items.

[0112] 2) Confidence reward Rconf, which is used to reward the selection of high-confidence parsing strategies and avoid low-confidence predictions, where Rconf = 5×(Cvit-0.6), Cvit is the average confidence of the parsing results, which is in the range of 0-1.

[0113] 3) Parsing stability reward Rstab, used to reward stability over multiple rounds of parsing, where Rstab = 3 × (1-Tacc,n / Uacc,n). Tacc,n is the standard deviation of the parsing accuracy over the last n rounds, and Uacc,n is the mean of the parsing accuracy over the last n rounds.

[0114] 4) Error correction reward Rfix, used to reward error corrections after multiple rounds of parsing, where Rfix = 5×(repaired data items / data items to be repaired).

[0115] For the above-mentioned fine-tuning based on multi-dimensional reinforcement learning, an early stopping strategy based on the fine-tuning loss percentage is designed. This strategy is more flexible than a fixed patience and is applicable to loss values ​​of different scales. The specific strategy is as follows:

[0116] Calculate the loss change rate, monitor the validation set loss (Validation Loss), and calculate the percentage change in loss for the most recent p rounds.

[0117] Set a threshold. If the loss decreases by less than a certain threshold (such as 0.1%) within p rounds, fine-tuning is considered to be close to convergence and training is stopped.

[0118] Patience parameter allows small fluctuations in loss and sets the waiting period for the patience round. Early stopping is triggered only when consecutive patience rounds meet the stopping conditions.

[0119] The beneficial effects of the embodiments of the present invention are as follows:

[0120] 1) Realized the field classification of engineering drawings;

[0121] 2) Achieved high-quality generation of fine-tuned corpora of engineering drawings from multiple fields;

[0122] 3) Implemented model fine-tuning of multi-dimensional reinforcement learning strategies based on engineering drawings.

Claims

1. A method for generating structured data of drawings in the engineering field, characterized in that: The following steps are involved: Extract the CAD text information corresponding to the engineering drawings in CAD file format, and extract the PDF text information corresponding to the engineering drawings in PDF format. The CAD text information and PDF text information include title bars, annotation information, and tables. According to the keyword library of different engineering fields, the CAD text information and PDF text information are matched and extracted with domain keywords. If the match is successful, the CAD keyword information after engineering field extraction and the PDF keyword information after engineering field extraction as well as the corresponding drawing field type are extracted; For engineering drawings in hand-drawn and / or scanned formats, as well as engineering drawings that fail to match, VisionTransformer is used in combination with keyword libraries from different engineering fields to parse them. If the confidence level reaches a preset level, information extracted in other formats from different engineering fields is obtained. If the confidence level does not reach the preset level, manual intervention is performed to directly classify the information extracted in other formats from the engineering field and the corresponding drawing field type. The engineering drawings are marked according to the corresponding drawing field types, and normalized with the CAD keyword information extracted from the engineering field, the PDF keyword information extracted from the engineering field, and the information extracted from other formats in the engineering field. The normalized text data is formatted and stored using JSON, and finally the normalized graphic and text data is obtained; According to the domain type of drawings, the corresponding domains are matched with the multi-domain sample graphic and text corpus. The normalized graphic and text data is used to match the similarity of graphic and text feature vectors in the matching domain. Samples with high similarity are selected from the multi-domain sample graphic and text corpus according to the preset proportion. The samples are used as prompt words to construct the prompt word context. Structured data is generated based on large-scale model knowledge distillation. The rejection sampling method based on confidence is used to filter out the structured data with low confidence. The drawings and corresponding structured data pairs in the normalized graphic and text data are obtained as fine-tuning corpora for different engineering fields. The multi-domain sample graphic and text corpus stores typical graphic and text corpora in different engineering fields. These corpora are used to describe typical engineering drawings and corresponding structured data pairs in the engineering field. Drawings in fine-tuned corpora from different engineering fields are encoded to generate image vector embeddings. Texts in fine-tuned corpora from different engineering fields are segmented and embedded to obtain text vector embeddings. For image vector embeddings and text vector embeddings, image-text modality alignment is performed based on image-text embeddings to generate a batch matrix of input data to be trained. Based on the input data matrix, multi-dimensional reinforcement learning is used to fine-tune the weight matrix of the base model, and finally a fine-tuned engineering field drawing parsing model is obtained. The engineering field drawing parsing model is used for reasoning, with the reasoning input being the engineering field drawings to be parsed, and finally the structured data of the engineering field drawings is output.

2. The method for generating structured data of engineering drawings according to claim 1, wherein: For engineering drawings in CAD file format, the CAD text information is extracted using the API or the ezdxf Python parsing library.

3. The method for generating structured data of engineering drawings according to claim 1, wherein: Optical character recognition is used to extract the CAD text information from engineering drawings in PDF format based on keywords.

4. The method for generating structured data of engineering drawings according to claim 1, wherein: The confidence level is calculated using Softmax-based probability.

5. The method for generating structured data of engineering drawings according to claim 1, wherein: The normalization process includes uniformly converting file formats, coordinates, and units.

6. The method for generating structured data of engineering drawings according to claim 1, wherein: The image-text vector similarity matching is cosine similarity matching.

7. The method for generating structured data of engineering drawings according to claim 1, wherein: The method of fine-tuning the base model weight matrix using multi-dimensional reinforcement learning based on the input data matrix includes the following steps: Generate multi-dimensional scores through the multi-value head output of the reward strategy and reward model; Based on the standardization of the grouped dimension scores, the relative advantages of different outputs are generated according to the advantage calculation formula; Generate a KL penalty based on the log-cumulative probability of the output of the initial reference model and the training model; Based on the above logarithmic cumulative generation probability, relative advantage, the single output training loss is calculated according to the loss calculation formula and KL penalty; Based on the above single-output training loss, the multi-output normalization within the group is performed to obtain the comprehensive loss of the single input; Update the model weights based on the input batch and gradient accumulation.

8. The method for generating structured data of engineering drawings according to claim 7, characterized in that: The advantage calculation formula is as follows: Among them, r i,j represents the reward for the jth answer of sample i, G represents the total G answers generated by sample i, and ε is a small parameter coefficient to prevent division by zero.

9. The method for generating structured data of engineering drawings according to claim 7, wherein: The loss calculation formula is as follows: Among them, i represents the i-th answer of a sample, t represents the t-th token in the i-th answer, A i represents the advantage of sample i, m i,t Indicates the mask of the t-th position of the i-th sample, L i It represents the loss of the i-th sample calculated in combination with the penalty coefficient β of the KL divergence, and B represents the number of samples in a batch.