A method for automatic line annotation based on an AI multi-modal large model
Through the automatic line labeling method based on AI multimodal large model, the problem of traditional labeling dependence on development experience and high customization costs is solved, and the rapid, accurate and automatic labeling of drawings is achieved, which improves design efficiency and quality.
Patent Information
- Application Number
- CN202411253330.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-09
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-09-09
AI Technical Summary
Traditional annotation methods rely on the experience of developers, and customized development scripts label features, resulting in high requirements for developers and high customization costs.
Automatic line labeling method based on AI multimodal large model is adopted to realize automatic labeling of drawings through data set processing, model training and full parameter fine-tuning.
This method can help engineers quickly and accurately mark drawings, improve design efficiency and quality, reduce developer requirements and customization costs.
Smart Images

Figure CN119205975B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of AI multimodal large models, and specifically to an automatic line annotation method based on an AI multimodal large model. Background Art
[0002] In the CAD / CAM field, modeling technology is a key method for transforming physical objects and their attributes into digital representations inside a computer, which not only includes the geometric topology information of parts, but also many non-geometric information such as material properties, process requirements, annotation lines, etc., so it has rich engineering semantics;
[0003] In the field of modern machine learning, Multimodal Large Language Models (MM-LLMs) refer to large language models that combine multiple information modalities. These models can not only process text data, but also process other forms of data such as images, audio, video, models, etc. The emergence of MM-LLMs has promoted the development of multimodal machine learning, enabling machines to more comprehensively understand and process different forms of data, thereby improving the understanding and modeling capabilities of the complex real world;
[0004] Multimodal pre-training technology refers to a method of pre-training on multiple data modalities, aiming to improve the model's understanding and processing capabilities for multiple data types. By performing multimodal pre-training on a large-scale dataset, the model can learn the correlations and semantic information between different data modalities, and thus perform better on various upstream and downstream tasks;
[0005] Full-parameter fine-tuning technology refers to, when performing transfer learning, fine-tuning all the parameters (full parameters) of the model to adapt to a specific task or dataset. In the case of multimodal pre-training, full-parameter fine-tuning technology can help the model better adapt to the data in a specific field or task, and improve the model's performance on specific tasks;
[0006] Explore the situation of the same 3D model in two modalities: the modeling process and the geometric shape. The modeling process and the geometric shape can be regarded as different manifestations of the same entity, that is, the association between the abstract representation of an object in the CAD modeling process and its geometric form in the real world. By studying the relationship between these two modalities, the CAD / CAM modeling process can be better understood and optimized. This research method helps to explore the application of multimodal data in the engineering field, provides new ideas and methods for the development of machine learning technology in the CAD / CAM field, and using an automatic annotation model in this method can improve the efficiency and accuracy of engineering design;
[0007] However, there are still the following disadvantages: traditional annotation often uses customized development scripts based on the developer's experience to annotate features, which places high demands on developers and has high customization costs.
[0008] Therefore, an automatic line labeling method based on AI multimodal large model is proposed. Summary of the invention
[0009] Technical issues solved
[0010] In view of the shortcomings of the prior art, the present invention provides an automatic line labeling method based on an AI multimodal large model, which has the advantages of being versatile and can be widely used in engineering design and manufacturing fields, helping engineers to quickly and accurately label drawings, and improving design efficiency and quality. It solves the problem that traditional labeling often labels features based on customized development scripts based on the developer's experience, which has high requirements for developers and high customization costs.
[0011] Technical Solution
[0012] In order to achieve the above-mentioned universality and be widely applied in the field of engineering design and manufacturing, to help engineers quickly and accurately mark drawings and improve design efficiency and quality, the present invention provides the following technical solutions:
[0013] A method for automatic line annotation based on an AI multimodal large model comprises the following steps:
[0014] Step 1: Dataset and data processing: First, collect 2D and 3D drawing data sets in a targeted manner;
[0015] Step 2: Then convert the three-dimensional structure diagram of the drawing into two-dimensional three-view drawing;
[0016] Step 3: Automatically mark and manually review the linear, radius, diameter, and angle of the 2D three-view drawings and 2D drawings based on the features of the 3D structural drawing;
[0017] Step 4: Perform structured text processing on the 2D drawings and verify the legitimacy of the data, and remove and repair poor quality data;
[0018] Step 5: Segment the 2D drawing plan, convert the entire plan into N blocks of fixed size, and then compress the structured text of all blocks;
[0019] Step 6: Process the data set and divide it into an unsupervised data set for pre-training and a supervised data set for fine-tuning. Then split the supervised data set into a training set and a test set.
[0020] Step 7: Model training: Based on the MM-LLMs model, find the hot words in the data set and expand the original vocabulary;
[0021] Step Eight: Update and expand the original model parameters. Given a set of pre-trained large language model parameters Θ and a dataset D = {(x text , x dimension )} that contains multimodal data with text x text and annotation x dimension ;
[0022] To continue pre-training the model and update the parameters, the following formula can be used:
[0023]
[0024] Step Nine: Fine-tune all parameters of the pre-trained model. Given a supervised dataset D {dimension_task} = {(x i , y i )};
[0025] The goal of fine-tuning is to minimize the loss function on the task dataset, that is, to minimize the loss function L(Θ; D dimension_task ). To perform all-parameter fine-tuning and update the parameters, the following formula is used to adapt to the automatic annotation task:
[0026]
[0027] Step Ten: Test and manually evaluate the results of the model.
[0028] Preferably, in Step Four, the specific objects of structured text processing are lines, types, blocks, rounded corners, arcs, text information, and two-dimensional coordinates.
[0029] Preferably, in Step Five, while compressing the structured text, information irrelevant to annotation is removed.
[0030] Preferably, in Step Ten, for incorrect data, after data augmentation, the Reinforcement Learning Decision Policy Optimization (DPO) algorithm is used to re-train.
[0031] Preferably, in Step Eight, where Θ (t) is the model parameter of the t-th iteration, Θ (t+1) is the updated parameter, η is the learning rate, represents the gradient of the parameter Θ.
[0032] Preferably, by calculating the gradient of the loss function with respect to the parameters and updating the parameters according to the gradient, the model can continue to learn on multimodal data and improve performance.
[0033] Preferably, in Step Nine, x iis the unlabeled input drawing, y i is the output annotation information.
[0034] Compared with the prior art, the present invention provides a method for automatically annotating lines based on an AI multi-modal large model, which has the following beneficial effects:
[0035] The method for automatically annotating lines based on the AI multi-modal large model has universality and can be widely applied to the fields of engineering design and manufacturing, helping engineers quickly and accurately annotate drawings, improving design efficiency and quality, and solving the problems of high requirements for developers and high customization costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the data processing flow chart of the automatic annotation method of the present invention;
[0037] Figure 2 is the model training flow chart of the automatic annotation method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0039] Please refer to Figure 1 - Figure 2 ;
[0040] A method for automatically annotating lines based on an AI multi-modal large model includes the following steps:
[0041] Step 1: Processing of the data set and data: First, collect two-dimensional and three-dimensional drawing data sets specifically;
[0042] Step 2: Then convert the three-dimensional structure diagram of the drawing into two-dimensional three-view drawings;
[0043] Step 3: Automatically annotate and manually review the linear, radius, diameter, and angle of the two-dimensional three-view drawings and two-dimensional drawings based on the features of the three-dimensional structure diagram;
[0044] Step 4: Perform structured text processing on the two-dimensional drawings and verify the data legality, eliminate and repair low-quality data. The specific objects of the structured text processing are lines, types, blocks, fillets, arcs, text information, and two-dimensional coordinates;
[0045] Step 5: Segment the 2D drawing solution, convert the entire solution into N fixed-size blocks, then compress the structured text of all blocks, and remove and annotate irrelevant information while compressing the structured text;
[0046] Step 6: Process the dataset, divide it into an unsupervised dataset for pre-training and a supervised dataset for fine-tuning, and then split the supervised dataset into a training set and a test set;
[0047] Step 7: Model training: Based on the MM-LLMs model, find the hot words in the dataset and expand the original vocabulary;
[0048] Step 8: Update and expand the original model parameters. Given a set of parameters Θ of a pre-trained large language model and a dataset D = {(x text and annotation x dimension )} of multimodal data containing text x text and annotation x dimension ;
[0049] To continue pre-training the model and update the parameters, the following formula can be used:
[0050]
[0051] where Θ (t) are the model parameters at the t-th iteration, Θ (t+1) are the updated parameters, η is the learning rate, represents the gradient of the parameters Θ. By calculating the gradient of the loss function with respect to the parameters and updating the parameters according to the gradient, the model can continue to learn on multimodal data and improve its performance;
[0052] Step 9: Perform full-parameter fine-tuning on the pre-trained model. Given a supervised dataset D {dimension_task} = {(x i , y i )};
[0053] The goal of fine-tuning is to minimize the loss function on the task dataset, that is, to minimize the loss function L(Θ; D dimension_task ). To perform full-parameter fine-tuning and update the parameters, the following formula is used to adapt to the automatic annotation task:
[0054] x i is the unannotated input drawing, and y i is the output annotation information;
[0055] Step 10: Test the results of the model using a test set and manual evaluation. For incorrect data, perform data augmentation and then retrain using the reinforcement learning decision policy optimization algorithm Decision Policy Optimization (DPO).
[0056] The beneficial effects of the present invention are as follows: This method for automatically annotating lines based on an AI multi-modal large model has universality and can be widely applied in the fields of engineering design and manufacturing. It helps engineers quickly and accurately annotate drawings, improves design efficiency and quality, and solves the problems of high requirements for developers and high customization costs.
[0057] During use, parse the historical solutions, extract features, reconstruct semantic features and data structures, and export the dataset. Utilize the reconstructed semantic feature topology vocabulary to pre-train an unsupervised annotation large model, and then fine-tune the annotation large model using the supervised annotation dataset to achieve automatic annotation of the linearity, radius, diameter, angle, etc. of the drawings using the multi-modal large model.
[0058] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to the embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for automatic line annotation based on AI multimodal large model, characterized in that: The following steps are involved: Step 1: Dataset and data processing: First, collect 2D and 3D drawing data sets in a targeted manner; Step 2: Then convert the three-dimensional structure diagram of the drawing into two-dimensional three-view drawing; Step 3: Automatically mark and manually review the linear, radius, diameter, and angle of the 2D three-view drawings and 2D drawings based on the features of the 3D structural drawing; Step 4: Perform structured text processing on the two-dimensional drawings processed in step 3 and verify the legitimacy of the data, and remove and repair data of poor quality; Step 5: Segment the 2D drawing plan, convert the entire plan into N blocks of fixed size, and then compress the structured text of all blocks; Step 6: Process the data set and divide it into an unsupervised data set for pre-training and a supervised data set for fine-tuning. Then split the supervised data set into a training set and a test set. Step 7: Model training: Based on the MM-LLMs model, find the hot words in the data set and expand the original vocabulary; Step 8: Update and expand the original model parameters. Given a pre-trained large language model parameter set Θ, a language model containing text x etxt and label x dimension The multimodal data set D = {(x text ,x dimension )}; To continue pre-training the model and update the parameters, the following formula can be used: Step 9: Fine-tune all parameters of the pre-trained model, given a supervised dataset D {dimension_task} ={(x i ,y i )}; The goal of fine-tuning is to minimize the loss function on the task dataset, that is, to minimize the loss function L(Θ; D dimension_task ), in order to fine-tune all parameters and update the parameters, the following formula is used to adapt to the automatic labeling task: Step 10: Test the model results and conduct manual evaluation.
2. The method for automatic line annotation based on AI multimodal large model according to claim 1 is characterized in that: The specific objects of structured text processing in step 4 are lines, types, blocks, fillets, arcs, text information and two-dimensional coordinates.
3. The method for automatic line annotation based on AI multimodal large model according to claim 1 is characterized in that: In the step 5, irrelevant information is removed and annotated while compressing the structured text.
4. The method for automatic line annotation based on AI multimodal large model according to claim 1 is characterized in that: In the step 10, the erroneous data is enhanced and then retrained using the reinforcement learning decision policy optimization algorithm Decision Policy Optimization (DPO).
5. The method for automatic line annotation based on AI multimodal large model according to claim 1 is characterized in that: In the step eight, θ (t) is the model parameter of the tth iteration, Θ (t+1) is the updated parameter, η is the learning rate, represents the gradient with respect to the parameter Θ.
6. The method for automatic line annotation based on AI multimodal large model according to claim 5 is characterized in that: By calculating the gradient of the loss function with respect to the parameters and updating the parameters according to the gradient, the model can continue to learn on multimodal data and improve performance.
7. The method for automatic line annotation based on AI multimodal large model according to claim 1 is characterized in that: Step nine x i is to input unmarked drawings, y i Output annotation information.
Citation Information
Patent Citations
News event search method and system based on multi-level image-text semantic alignment model
WO2023093574A1
Model training method and apparatus, and storage medium
WO2024174723A1