Engineering drawing auditing ability evaluation method for multi-modal large model

Through the evaluation method of engineering drawing review ability of multimodal large models, its analytical understanding, information recognition and professional common sense judgment capabilities are systematically evaluated, which solves the problem of lack of evaluation methods in existing technologies and realizes efficient and accurate multimodal large model review ability evaluation.

CN120688746APending Publication Date: 2025-09-23BEIJING QDING INTERCONNECTION TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510787646.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing technology lacks testing and evaluation methods for the review capabilities of multi-modal large-model engineering drawings, resulting in low efficiency and prone to errors.

Method used

This paper provides an evaluation method for the engineering drawing review ability of multimodal large models. By obtaining multiple engineering drawings for training and testing of marking and annotation, component recognition, information extraction and common sense judgment tasks, the evaluation of each ability is carried out on the engineering drawing analysis and understanding, information recognition and professional common sense judgment ability of each large model, and a weighted sum is taken to determine its review ability.

Benefits of technology

It has achieved systematic evaluation of large multimodal models, ensuring the objectivity and scientific nature of the evaluation results, improving the robustness of the evaluation, being able to meet the recognition needs of blurred or incomplete drawings, and improving the comprehensiveness and depth of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688746A_ABST
    Figure CN120688746A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an engineering drawing auditing ability evaluation method for a multi-modal large model, and the method comprises the steps: obtaining a plurality of engineering drawings from different construction engineering projects, and the plurality of engineering drawings comprise a construction drawing, a structure drawing, an equipment drawing and an electrical drawing; all the engineering drawings are marked, and the marked content comprises component information, space information, standard information and defect information; training and testing of a component recognition task, an information extraction task and a common sense judgment task are carried out on the multiple to-be-evaluated multi-modal large models through the marked engineering drawings; according to a test result of each task, respectively evaluating a drawing analysis understanding capability, an information identification capability and a professional common sense judgment capability of each large model; and carrying out weighted summation on the evaluation score of each capability, and determining the engineering drawing auditing capability of each large model. According to the embodiment of the invention, objective evaluation of the multi-modal large model can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a method for evaluating engineering drawing review capabilities for multimodal large models. Background Art

[0002] Engineering drawing review is a crucial process in the fields of architecture and engineering. Traditional methods rely heavily on manual review, which is inefficient and prone to errors.

[0003] With the development of artificial intelligence technology, multimodal large models have gradually been introduced into the field of drawing review. However, there is currently no method to test and evaluate the engineering drawing review capabilities of multimodal large models. Summary of the Invention

[0004] An embodiment of the present invention provides a method for evaluating engineering drawing review capabilities for multimodal large models to solve the above technical problems.

[0005] In a first aspect, an embodiment of the present invention provides a method for evaluating engineering drawing review capabilities for multimodal large models, comprising:

[0006] Acquiring a plurality of engineering drawings from different construction projects, the plurality of engineering drawings including architectural drawings, structural drawings, equipment drawings, and electrical drawings;

[0007] Mark each engineering drawing, including component information, space information, specification information and defect information;

[0008] Using the annotated engineering drawings, multiple large multimodal models to be evaluated are trained and tested for component recognition, information extraction, and common sense judgment tasks.

[0009] Based on the test results of each task, the model's ability to interpret and understand drawings, recognize information, and make professional judgments was evaluated.

[0010] The evaluation scores of each capability are weighted and summed to determine the engineering drawing review capabilities of each major model.

[0011] In a second aspect, an embodiment of the present invention provides an electronic device, comprising:

[0012] one or more processors;

[0013] a memory for storing one or more programs,

[0014] When the one or more programs are executed by the one or more processors, the one or more processors implement the method for evaluating engineering drawing review capabilities for multimodal large models described in any embodiment.

[0015] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for evaluating the engineering drawing review capability for multimodal large models described in any embodiment.

[0016] In summary, this embodiment provides a method for evaluating the engineering drawing review ability of multimodal large models, and provides a systematic evaluation system that covers the ability to parse and understand drawings, the accuracy of information recognition, and the ability to judge professional common sense. It also constructs corresponding model training tasks according to the evaluation purpose. Through training and feedback on large models, the large models are given a certain degree of accuracy and stability in engineering drawing review capabilities, and then these stable large models are evaluated. The ability level of the model is quantified through a scoring system to ensure the objectivity and scientific nature of the evaluation results. At the same time, the robustness of the evaluation is improved to meet the recognition needs of fuzzy or incomplete drawings. The evaluation of specification association ability and logical reasoning ability is introduced to improve the comprehensiveness and depth of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is a flow chart of a method for evaluating engineering drawing review capabilities for multi-modal large models provided by an embodiment of the present invention;

[0019] Figure 2 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are also within the scope of protection of the present invention.

[0021] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0022] In the description of the present invention, it should also be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, they may refer to fixed, detachable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0023] Figure 1 This is a flow chart of a method for evaluating the ability of engineering drawing review for multi-modal large models provided by an embodiment of the present invention. The method is executed by an electronic device, such as Figure 1 As shown, the specific steps include:

[0024] S110. Acquire multiple engineering drawings from different construction projects, where the multiple engineering drawings include architectural drawings, structural drawings, equipment drawings, and electrical drawings.

[0025] This example first collects engineering drawings from various construction projects as a dataset for subsequent multimodal large-scale model evaluation. Optionally, the collected engineering drawings include architectural drawings, structural drawings, equipment drawings, electrical diagrams, and more. These drawings include both complete drawings (e.g., a full set of construction drawings) and blurred / defective drawings (e.g., damaged, blurred, or incomplete drawings). The data formats include various drawing formats, such as CAD (DWG, DXF), PDF, JPG, and PNG.

[0026] After data collection, perform data cleaning and preprocessing. Optionally, perform denoising and image enhancement on blurred or incomplete drawings to improve recognition. Then, convert drawings of varying formats to a unified format (such as PNG or JPG) for easier processing.

[0027] S120. Mark each engineering drawing.

[0028] This step labels and annotates the pre-processed engineering drawings. For example, it adds annotation information (such as component type and dimension markings), spatial information, specification information (compliant specifications), defect information (such as design flaws and existing issues), and other information that requires attention during engineering drawing review. This process can be completed manually by experienced engineers, providing reference data for subsequent model training.

[0029] After annotation, the drawings are categorized and stored by type (e.g., architectural drawings, structural drawings) to facilitate subsequent model training and testing. Optionally, a database can be created to record basic information about each drawing (e.g., drawing type, size, source, etc.) for quick retrieval and recall.

[0030] S130. Using the annotated engineering drawings, the multiple multimodal large models to be evaluated are trained and tested for component recognition tasks, information extraction tasks, and common sense judgment tasks.

[0031] This embodiment uses pre-trained models that support dual-modal input of images and text (such as CLIP and LayoutLM), and uses labeled datasets to fine-tune the large model to ensure stable image review capabilities. The fine-tuned models are then tested to compare their image review capabilities.

[0032] In one specific implementation, three training tasks are first constructed for all large language models, including:

[0033] Component recognition task: training the model to identify components and their positional relationships in drawings;

[0034] Information extraction task: training the model to extract dimension, quantity, relationship and other marked information from drawings;

[0035] Common sense judgment task: The model is trained to perform logical reasoning and error judgment based on the drawing content and specification provisions.

[0036] The annotated dataset is then divided into a training set and a test set, both of which include complete drawings and blurred / incomplete drawings. The engineering drawings and task requirements in the training set are fed to the large model, instructing it to perform the corresponding tasks and output the corresponding review results in the prescribed format. The annotated ground truth information from these engineering drawings is then fed back to the large model, instructing it to compare the discrepancies between its output review results and the annotated content. This helps adjust the model's reasoning and review direction, ultimately enabling it to gradually acquire a certain level of precision and stability in its review capabilities.

[0037] Finally, we used the test set to test several large models with stable image review capabilities to compare their respective image review capabilities. The tests also included three tasks: construct recognition, information extraction, and common sense judgment.

[0038] S140. Based on the test results of each task, evaluate the drawing analysis and comprehension capabilities, information recognition capabilities, and professional common sense judgment capabilities of each major model.

[0039] Specifically, based on the test results of each large model, perform the following operations respectively:

[0040] S1-1. Based on the test results of the component identification task, according to the information output by the large model and the labeled information, the comprehensiveness of the model in identifying the types of components and / or spaces (such as water wells and electric wells) is counted to obtain the element coverage of the large model; the accuracy of the hierarchical division of the building and / or equipment system of any large model is evaluated to obtain the hierarchical understanding ability score of the large model; the accuracy of the large model in identifying the positional relationship of components is calculated to obtain the spatial topological relationship score of the large model. Finally, the element coverage, hierarchical understanding ability score and spatial topological relationship score are weighted and summed to obtain the drawing parsing and understanding ability score of the large model. Optionally, the score range of element coverage, hierarchical understanding ability and spatial topological relationship is [0,100], and the weighted weights are all in the range of [0,1].

[0041] S1-2. Based on the test results of the information extraction task, the size and quantity recognized by the large model are compared with the marked values ​​to calculate the numerical accuracy score of the large model; the consistency of the large model in recognizing graphic information in the information extraction task is verified to obtain the multimodal matching score of the large model; based on the robustness of the large model in recognizing fuzzy / defective drawings in the information extraction task, the anti-interference ability score of the large model is determined. Finally, the numerical accuracy score, multimodal matching score and anti-interference ability score are weighted and summed to obtain the information recognition ability score of any large model. Optionally, the score range of numerical accuracy, multimodal matching and anti-interference ability is [0,100], and the weighted weights are all in the range of [0,1].

[0042] S1-3. Based on the test results of the common sense judgment task, the accuracy of the large model's reference to the standard provisions is evaluated according to the output information and annotation information of the large model to obtain the specification association ability score of the large model; the number and defects of the design defects found by the large model are counted to obtain the logical reasoning ability score of the large model; the accuracy of the problem location of the large model is evaluated to obtain the error tracing ability score of the large model. Finally, the standard association ability score, logical reasoning ability score and error tracing ability score are weighted and summed to obtain the professional common sense judgment ability score of the large model. Optionally, the score range of the specification association ability, logical reasoning ability and error tracing ability is [0,100], and the weighted weight range is [0,1].

[0043] After performing the above operations on each large model, the drawing comprehension ability score, information recognition ability score, and professional common sense judgment ability score of each large model can be obtained respectively.

[0044] S150. Perform weighted summation of the assessment scores of each capability to determine the engineering drawing review capability of each major model.

[0045] By taking the weighted sum of the drawing comprehension ability score, information recognition ability score, and professional common sense judgment ability score of each large model, we can obtain the engineering drawing review ability score of each large model, thereby realizing the quantitative evaluation of the engineering drawing review ability of multimodal large models.

[0046] Furthermore, by summarizing the training, testing, and evaluation processes in the above method, we can derive a three-dimensional, nine-indicator evaluation system as shown in the following table:

[0047]

[0048] The highest score of each secondary indicator is the weight of each secondary indicator multiplied by 100. In actual applications, the scores of each secondary indicator in the test results can be given according to these highest scores. These scores can be added together to obtain the comprehensive score of the engineering drawing review ability of the large model. In order to better understand the method of this embodiment, the evaluation purpose of each indicator is explained in more detail below:

[0049] 1. Ability to analyze and understand drawings:

[0050] (1) Element coverage: This assessment uses the model to identify the types of specific components or spaces (such as water wells and electrical well rooms) in the drawings and evaluates the comprehensiveness of the identification scope.

[0051] (2) Hierarchical comprehension ability: This assessment requires the model to accurately divide the hierarchical structure of the building or equipment system to ensure the depth and breadth of understanding.

[0052] (3) Spatial topological relationship: This assessment requires the model to accurately identify the positional relationship between components and evaluate the accuracy of spatial understanding.

[0053] 2. Information recognition accuracy:

[0054] (1) Numerical accuracy: This evaluation compares the size and quantity recognized by the model with the actual value, calculates the error rate, and evaluates the recognition accuracy.

[0055] (2) Multimodal matching: This evaluation verifies whether the image and text information recognized by the model are consistent, ensuring the accuracy and consistency of the information.

[0056] (3) Anti-interference ability: This evaluation requires the model to identify blurred or incomplete drawings and evaluate the robustness and adaptability of the model.

[0057] 3. Professional common sense judgment:

[0058] (1) Standards association ability: This assessment requires the model to accurately quote relevant standard provisions and evaluate its ability to understand and apply the standards.

[0059] (2) Logical reasoning ability: This assessment evaluates the model's logical reasoning ability by identifying the number and types of design defects.

[0060] (3) Error tracing capability: This assessment requires the model to accurately locate problems and evaluate its problem analysis and problem solving capabilities.

[0061] Through the above-mentioned evaluations, the objectivity of large-scale model evaluation can be guaranteed, while the robustness of the evaluation can be improved to meet the recognition needs of fuzzy or incomplete drawings; the ability to associate standards and logical reasoning can be introduced to improve the comprehensiveness and depth of the evaluation.

[0062] After the evaluation is complete, a comprehensive evaluation report can be generated. This report includes: Score Details: listing the model's specific scores for each indicator; Strengths and Weaknesses: analyzing the model's strengths and weaknesses in drawing parsing, information recognition, and professional judgment; and Optimization Suggestions: proposing targeted improvements based on the evaluation results (such as enhancing the ability to recognize fuzzy drawings). The report can be formatted in PDF or HTML for easy viewing and analysis.

[0063] Based on the evaluation results, users can compare the performance of different multimodal large models in a targeted manner and select the most suitable model for use; they can also adjust the parameters, improve the algorithm and enhance the data of the large model according to the optimization suggestions, and re-evaluate the optimized model to form a closed-loop mechanism of dynamic optimization.

[0064] In a specific embodiment, the various evaluation indicators and weights in the above evaluation system can be determined by the following method:

[0065] Step 1. Through expert scoring, obtain the scores of multiple multimodal large models for multiple candidate indicators in multiple tasks, as well as the scores of engineering drawing review capabilities. Among them, the candidate indicators here correspond to the secondary evaluation indicators in the above-mentioned evaluation system, covering three categories: drawing parsing and understanding ability, information recognition accuracy, and professional common sense judgment. Specifically, a massive engineering drawing data set (covering fields such as architecture, structure, and electrical) can be collected, and initial candidate indicators can be extracted from the drawings (such as "component type recognition type", "dimensional error rate", "specification reference accuracy" and other 50 candidate indicators; each engineering drawing is used as a sample and input into multiple large models respectively. Each large model performs the component recognition task, information extraction task and common sense judgment task respectively. Experts score each candidate indicator corresponding to each engineering drawing based on the task execution results, and score the large model's review capability for the drawing.

[0066] Step 2: Perform principal component analysis on the scores of each candidate indicator in each task to obtain multiple eigenvectors and eigenvalues; and select multiple eigenvectors with the top eigenvalues ​​as multiple principal component vectors. Optionally, a scoring matrix can be constructed based on the above expert scores. Each row of the matrix corresponds to an engineering drawing, and each column corresponds to a candidate indicator. The matrix elements are the scores of the engineering drawings in the row on the candidate indicators in the column. Performing principal component analysis on the matrix can obtain multiple eigenvectors and eigenvalues. The dimensions of these eigenvectors are the same as the number of engineering drawings in the scoring matrix. The eigenvectors are mutually orthogonal and represent independent data dimensions in the scoring matrix. The larger the eigenvalue, the more important the eigenvector is in the entire data matrix. The characteristic vectors are sorted according to the eigenvalue. The eigenvectors that rank in the top are the most important data components in the scoring matrix.

[0067] Step three, select multiple indicators whose distances to the principal component vectors are less than a set threshold from the multiple candidate indicators as multiple evaluation indicators. Optionally, calculate the Euclidean distance between the column vector corresponding to each candidate indicator and each principal component vector in the score matrix, and use the candidate indicators whose distances are less than the set threshold as preliminary evaluation indicators. These candidate indicators are closest to the principal component vectors and are close to the main data components of the score matrix. Through principal component decomposition, redundant data components (such as collinear components, etc.) in the score matrix can be filtered out, and several candidate indicators with the most important information and the least mutual redundant components can be extracted as preliminary evaluation indicators.

[0068] Step 4: Quantify the correlation between each evaluation indicator score and the engineering drawing review ability score through mutual information calculation; retain multiple evaluation indicators with mutual information values ​​greater than the set threshold as the final multiple evaluation indicators. Through mutual information calculation, the correlation between each indicator and model performance (i.e., comprehensive score) can be quantified, and indicators with mutual information values ​​greater than 0.3 are retained as the final evaluation indicators.

[0069] In practical applications, the above indicator screening process can also be automatically implemented through the PCA module of the Scikit-learn library. The new code snippet is as follows:

[0070]

[0071] Step 5: Use the scores of each evaluation indicator corresponding to the same drawing as the independent variable and the score of the engineering drawing review ability corresponding to the same drawing as the dependent variable. Use the least squares method to fit the linear relationship between the dependent variable and the independent variables; determine the weight of each evaluation indicator based on the fitted linear relationship. Optionally, the linear relationship to be fitted can be expressed as follows:

[0072]

[0073] Among them, y represents the engineering drawing review ability score of the model, x i and w i are the scores and corresponding weights of each secondary evaluation indicator, and ∈ is the error term. By fitting the above relationship through the least squares method, a set of optimal weights w for each secondary evaluation indicator can be obtained. i By adding up the weights of each second-level evaluation indicator under the same type of first-level evaluation indicator, the weight of the first-level evaluation indicator can be obtained.

[0074] In summary, this embodiment provides a method for evaluating the engineering drawing review ability of multimodal large models, and provides a systematic evaluation system that covers the ability to parse and understand drawings, the accuracy of information recognition, and the ability to judge professional common sense. It also constructs corresponding model training tasks according to the evaluation purpose. Through training and feedback on large models, the large models are given a certain degree of accuracy and stability in engineering drawing review capabilities, and then these stable large models are evaluated. The ability level of the model is quantified through a scoring system to ensure the objectivity and scientific nature of the evaluation results. At the same time, the robustness of the evaluation is improved to meet the recognition needs of fuzzy or incomplete drawings. The evaluation of specification association ability and logical reasoning ability is introduced to improve the comprehensiveness and depth of the evaluation.

[0075] Figure 2 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention is shown in FIG. Figure 2As shown, the device includes a processor 60, a memory 61, an input device 62 and an output device 63; the number of processors 60 in the device can be one or more. Figure 2 In the embodiment, a processor 60 is used as an example; the processor 60, the memory 61, the input device 62 and the output device 63 in the device can be connected by a bus or other means. Figure 2 The bus connection is taken as an example.

[0076] Memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the engineering drawing review capability assessment method for multimodal large models in the embodiments of the present invention. Processor 60 executes the software programs, instructions, and modules stored in memory 61 to perform various functional applications and data processing of the device, thereby implementing the aforementioned engineering drawing review capability assessment method for multimodal large models.

[0077] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, and these remote memories may be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0078] The input device 62 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 63 may include a display device such as a display screen.

[0079] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for evaluating engineering drawing review capabilities for multimodal large models according to any embodiment.

[0080] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by an instruction execution system, device or device or used in combination with it.

[0081] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0082] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0083] Computer program code for performing the operations of the present invention can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A method for evaluating engineering drawing review capabilities for multimodal large models, characterized by: include: Acquiring a plurality of engineering drawings from different construction projects, the plurality of engineering drawings including architectural drawings, structural drawings, equipment drawings, and electrical drawings; Mark each engineering drawing, including component information, space information, specification information and defect information; Using the annotated engineering drawings, multiple large multimodal models to be evaluated are trained and tested for component recognition, information extraction, and common sense judgment tasks. Based on the test results of each task, the model's ability to interpret and understand drawings, recognize information, and make professional judgments was evaluated. The evaluation scores of each capability are weighted and summed to determine the engineering drawing review capabilities of each major model.

2. The method according to claim 1, characterized in that Based on the test results of each task, the model's drawing analysis and comprehension ability, information recognition ability, and professional common sense judgment ability are evaluated respectively, including: Counting the comprehensiveness of component and / or space type recognition of any large model in the component recognition task to obtain the element coverage score of any large model; Evaluate the accuracy of the hierarchical division of the building and / or equipment system by any one of the large models in the component identification task, and obtain a hierarchical understanding ability score of any one of the large models; Calculating the accuracy of any one of the large models in identifying the positional relationship of components in the component identification task, and obtaining a spatial topological relationship score of any one of the large models; The element coverage score, the hierarchical comprehension ability score and the spatial topological relationship score are weightedly summed to obtain the drawing parsing and comprehension ability score of any large model.

3. The method according to claim 1, wherein the evaluation of each model's drawing analysis and comprehension ability, information recognition ability, and professional common sense judgment ability based on the test results of each task includes: Comparing the sizes and quantities identified by any large model in the information extraction task with the labeled values, and calculating the numerical accuracy score of the large model; Verify the consistency of the image and text information recognition of any of the large models in the information extraction task, and obtain the multimodal matching score of any of the large models; Determine the anti-interference ability score of any of the large models according to the recognition robustness of any of the large models for blurry / defective drawings in the information extraction task; The numerical accuracy score, multimodal matching score and anti-interference ability score are weightedly summed to obtain the information recognition ability score of any large model.

4. The method according to claim 1, wherein the evaluation of each model's drawing parsing and understanding ability, information recognition ability, and professional common sense judgment ability based on the test results of each task includes: Evaluate the accuracy of the reference to standard provisions of any large model in the common sense judgment task, and obtain the normative relevance ability score of any large model; Counting the number and defects of design defects found by any of the large models in the common sense judgment task, and obtaining a logical reasoning ability score of any of the large models; Evaluate the accuracy of problem location of any of the large models in the common sense judgment task, and obtain an error tracing ability score of any of the large models; The standard association ability score, logical reasoning ability score and error tracing ability score are weighted and summed to obtain the professional common sense judgment ability score of any large model.

5. The method according to claim 1, wherein The drawing analysis and comprehension ability, information recognition ability and professional common sense judgment ability each include multiple evaluation indicators; Accordingly, before evaluating the drawing parsing and understanding ability, information recognition ability, and professional common sense judgment ability of each model based on the test results of each task, it also includes: Through expert scoring, the scores of multiple multimodal large models on multiple candidate indicators in multiple tasks were obtained. The multiple candidate indicators covered three categories: drawing analysis and comprehension ability, information recognition ability, and professional common sense judgment ability; Perform principal component analysis on the scores of each candidate indicator in each task to obtain multiple eigenvectors and eigenvalues; Select multiple eigenvectors with the highest eigenvalues ​​as multiple principal component vectors; A plurality of indicators whose distances to the principal component vectors are less than a set threshold are screened from the plurality of candidate indicators as the plurality of evaluation indicators.

6. The method according to claim 5, characterized in that The expert scoring is used to obtain scores of the multiple multimodal large models on multiple candidate indicators in multiple tasks, including: obtaining scores of the multiple multimodal large models on multiple candidate indicators in multiple tasks and scores of engineering drawing review capabilities; Correspondingly, after screening multiple indicators whose contributions to each principal component vector are greater than a set threshold from the multiple candidate indicators as multiple evaluation indicators, it also includes: quantifying the correlation between the scores of each evaluation indicator and the engineering drawing review ability score through mutual information calculation; retaining multiple evaluation indicators whose mutual information values ​​are greater than the set threshold as the final multiple evaluation indicators.

7. The method according to claim 1, characterized in that The drawing analysis and comprehension ability, information recognition ability, and professional common sense judgment ability each include multiple evaluation indicators, and the evaluation scores of the three abilities are obtained by weighted average of the scores of their respective evaluation indicators; Accordingly, before weighted summing of the assessment scores of each capability is performed to determine the engineering drawing review capability of each model, the following is also included: Through expert scoring, we obtained the scores of multiple multimodal large models on various evaluation indicators in multiple tasks, as well as the scores of engineering drawing review capabilities; The scores of various evaluation indicators corresponding to the same drawing are used as independent variables, and the scores of engineering drawing review ability corresponding to the same drawing are used as dependent variables. The least squares method is used to fit the linear relationship between the dependent variable and the independent variables. According to the fitted linear relationship, the weight of each evaluation indicator is determined.

8. The method according to claim 1, characterized in that The drawing parsing and comprehension ability includes three evaluation indicators: element coverage, hierarchical comprehension ability, and spatial topological relationship, with corresponding weights of 15%, 10%, and 15% respectively; The information recognition capability includes three evaluation indicators: numerical accuracy, multimodal matching and anti-interference ability, with corresponding weights of 20%, 10% and 5% respectively; The professional common sense judgment ability includes three evaluation indicators: norm association ability, logical reasoning ability and error tracing ability, with corresponding weights of 15%, 7% and 3% respectively.

9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the engineering drawing review capability assessment method for multimodal large models as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, which, when executed by a processor, implements the method for evaluating the engineering drawing review capability for multimodal large models as described in any one of claims 1-8.

Citation Information

Cited By

  • Electric power drawing understanding large model self-evolution training method

    CN121835820A