Method and system for evaluating application effect of large language model in building field
By constructing an architectural knowledge system and evaluating large language models using AO and COT methods, the problem of inaccurate evaluation in the existing technology is solved, and a comprehensive and reliable evaluation of LLM in the construction field is achieved, which improves its application effect in the construction process.
Patent Information
- Application Number
- CN202510265555.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-18
AI Technical Summary
The existing evaluation system fails to effectively measure the practical application performance of large language models (LLMs) in the construction field, especially in different project stages, which makes it difficult to objectively judge its overall optimization effect on the construction process, and lacks customized training and optimization, resulting in comprehension deviations and inaccurate application.
Build an architectural knowledge system, organize the test set, and ask questions on the large language model through AO and COT methods. Combined with paired sample t-test, determine the appropriate test set size, quantify the stability and accuracy of the model, and establish unified and industry-oriented evaluation standards.
A detailed evaluation of LLM's reliability, accuracy and demand adaptability in the construction field was achieved, and its application effect in complex building tasks was improved, providing a basis for further development and improvement.
Smart Images

Figure CN120336777A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method and system for evaluating the application effect of large language models in the construction field. Background Art
[0002] Currently, the application of large language models (LLMs) in the construction field mainly focuses on automated document generation and code parsing. Ding et al. reviewed the applications of NLP in the construction-related fields, introducing the applications of NLP in construction documents, accident analysis, safety management, automated building information modeling, risk management, etc. from three aspects: datasets and sources, technologies and tools, and applications and progress. Zheng et al. (2022) systematically explored transfer learning and fine-tuning techniques by developing large-scale domain corpora and pre-training domain-specific language models to improve the performance of natural language processing tasks in the architecture, engineering, and construction (AEC) field, providing a new perspective for automated document generation and code parsing. Alireza Shamshiri et al. also explored the application potential of text analysis from the perspective of construction management, while Saka A et al. preliminarily explored the application of GPT models in the construction project cycle, pointing out potential opportunities in material selection and optimization.
[0003] Another study, "Research on the BIM Forward Design Q&A System Based on Large Language Models" (Ding Zhikun et al., 2024), designed a set of BIM forward design Q&A systems based on large language models, aiming to explore the application potential of LLMs in BIM forward design. The system uses QLORA fine-tuning and combines with a local knowledge base to achieve the Q&A function, and verifies the performance of the system through subjective and objective performance evaluations and professional Q&A comparisons. However, only 100 test questions were extracted for evaluation during the test, with a small sample size, which is likely to lead to inaccurate parameter estimation, especially in statistical indicators such as mean and standard deviation, affecting the reliability of the evaluation and the applicability of the system.
[0004] The above studies have demonstrated the potential of LLMs in the construction field. However, the existing evaluation system fails to effectively measure the actual application performance of LLMs in construction projects, such as their specific contributions and adaptabilities in different project stages, making it difficult to objectively judge the overall optimization effect of LLMs on the construction process. In addition, the construction design and construction processes rely on rich professional knowledge and experience, involving a large number of codes and design principles, while the existing LLMs lack customized training and optimization, resulting in understanding deviations and inaccurate applications when dealing with professional content such as construction codes, industry standards, and design principles. It can be seen that the current research lacks a detailed evaluation of the reliability, accuracy, and demand adaptability of LLMs in the construction field, and also does not provide a comprehensive measurement standard for the evaluation of diverse and complex knowledge applications.
[0005] Therefore, there is an urgent need to develop a solution to evaluate the application effect of large language models in the construction field, in order to identify the deficiencies of LLMs, improve their application effect in the construction field, and provide a basis for the further development and improvement of large language models. Summary of the Invention
[0006] The object of the present invention is to provide a method and system for evaluating the application effect of large language models in the construction field. By establishing a unified and industry-specific evaluation standard, and quantitatively analyzing the accuracy and stability of large language models in complex construction tasks, the application effect in the construction field can be improved.
[0007] To achieve the above object, the present invention provides the following solution:
[0008] A method for evaluating the application effect of large language models in the construction field, comprising the following steps:
[0009] S1, construct an architectural knowledge system and sort out the question set of the architectural knowledge system;
[0010] S2, pre-experiment: extract a test set from the question set, conduct a preliminary test on the representative large language model for stability and efficiency, and determine the appropriate size of the test set;
[0011] S3, evaluation: use the size determined in S2 to re-extract the test set, use the AO and COT questioning methods to question each large language model to be tested respectively, calculate the mean difference of the correct rates of these two methods, and use the paired samples t-test to statistically verify the mean difference to obtain an objective conclusion on the application effect of the large language model in the construction field.
[0012] Further, in S2, the pre-experiment: extract a test set from the question set, conduct a preliminary test on the representative large language model for stability and efficiency, and determine the appropriate size of the test set, specifically including:
[0013] S2.1, extract some questions from the question set as the first test set, use the AO questioning method to repeatedly question the representative large language model, and judge the stability of the representative large language model. The stability refers to whether the answers to multiple questions are consistent;
[0014] S2.2, set up second test sets of different sizes in gradients, with multiple second test sets for each gradient, and then calculate the sample variance of the correct rates of these second test sets to determine the effective and stable size of the test set.
[0015] Further, the method further includes: using the AO questioning method to question the large language model to be tested, and quantifying the stability and accuracy of the large language model to be tested, specifically including:
[0016] Use the questioning method of AO to test the large language model to be tested. Use the sample variance of the correct rate of answers to the test sets with different contents as the standard for evaluating stability, and use the average value of the correct rate of answers as the standard for judging accuracy, so as to obtain the stability and accuracy of the large language model to be tested.
[0017] Furthermore, the architecture knowledge system at least includes architecture design, urban and rural planning, structural engineering, and environmental science.
[0018] The present invention also provides a system for evaluating the application effect of a large language model in the field of architecture, which is used to execute the method for evaluating the application effect of a large language model in the field of architecture, including:
[0019] A construction module, which is used to construct an architecture knowledge system and organize the question set of the architecture knowledge system;
[0020] A pre-experiment module, which is used to conduct a pre-experiment: extract a test set from the question set, conduct a preliminary test on the representative large language model for stability and efficiency, and determine the appropriate size of the test set;
[0021] An evaluation module, which is used to re-extract the test set with the size determined in the pre-experiment module, use the questioning methods of AO and COT to question each large language model to be tested respectively, calculate the mean difference of the correct rates of these two methods, and use the paired sample t-test to statistically verify the mean difference to obtain an objective conclusion on the application effect of the large language model in the field of architecture.
[0022] The present invention also provides a computer device, including: one or more processors;
[0023] The processor is used to store one or more programs; when the one or more programs are executed by the one or more processors, the method for evaluating the application effect of a large language model in the field of architecture is implemented.
[0024] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the method for evaluating the application effect of a large language model in the field of architecture is implemented.
[0025] According to the specific embodiments provided by the present invention, the following technical effects are disclosed: The method and system for evaluating the application effect of large language models in the construction field provided by the present invention first construct an architectural knowledge system and organize a question set of the architectural knowledge system; then, extract a test set from the question set to conduct a preliminary test on representative large language models for stability and efficiency, and determine the appropriate size of the test set; finally, re-extract the test set using the determined size, test the large language model to be tested in combination with the AO questioning method, and quantify the stability and accuracy of the large language model to be tested. The present invention establishes a unified and industry-specific evaluation standard, and truly and comprehensively reflects the applicability and reliability of the LLM by quantitatively evaluating the performance of the LLM in terms of knowledge understanding, reasoning ability, and practical application potential, and solves the problems of inaccurate performance evaluation and lack of professionalism faced by the current construction industry when applying the LLM; the present invention can identify the deficiencies of the LLM, improve its application effect in the construction field, and provide a basis for the further development and improvement of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts.
[0027] Figure 1 It is a flowchart of the method for evaluating the application effect of large language models in the construction field in the embodiments of the present invention;
[0028] Figure 2 It is a schematic diagram for evaluating the knowledge system in the embodiments of the present invention;
[0029] Figure 3 It is the proportion of questions in each subject in the question bank;
[0030] Figure 4 It is a graph of the output stability test results of Qwen-14B-Chat and gpt-3.5-turbo, where 4-1 is the graph of the output stability test results of Qwen-14B-Chat, and 4-2 is the graph of the output stability test results of gpt-3.5-turbo;
[0031] Figure 5 It is a graph of the sample variance of the correct rate of the output of Qwen-14B-Chat and gpt-3.5-turbo under different test set sizes, where 5-1 is the graph of the sample variance of the correct rate of the output of Qwen-14B-Chat under different test set sizes, and 5-2 is the graph of the sample variance of the correct rate of the output of gpt-3.5-turbo under different test set sizes;
[0032] Figure 6 For the stability test results of different LLMs outputs, among which, Figure 6-1 、 Figure 6-2 、 Figure 6-3 、 Figure 6-4 、 Figure 6-5 are respectively the stability test results corresponding to the 14 LLMs described in the embodiments;
[0033] Figure 7 For the accuracy test results of different LLMs outputs, among which, Figure 7-1 、 Figure 7-2 、 Figure 7-3 、 Figure 7-4 、 Figure 7-5 are respectively the accuracy test results corresponding to the 14 LLMs described in the embodiments;
[0034] Figure 8 For the difference results of the AO and COT output accuracy rates of some LLMs, among which, Figure 8-1 、 Figure 8-2 、 Figure 8-3 、 Figure 8-4 、 Figure 8-5 are respectively the difference results of the AO and COT output accuracy rates corresponding to the 5 LLMs in the embodiments;
[0035] Figure 9 Is the t-test result of the AO and COT sample data. Detailed implementation manners
[0036] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0037] The object of the present invention is to provide a method and system for evaluating the application effect of large language models in the construction field. By establishing a unified and industry-specific evaluation standard and quantitatively analyzing the accuracy and stability of large language models in complex construction tasks, the application effect in the construction field is improved.
[0038] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0039] Example 1
[0040] Such as Figure 1As shown in the figure, the method for evaluating the application effect of large language models in the field of architecture provided by the embodiments of the present invention includes the following steps:
[0041] S1. Construct an architecture knowledge system and organize a question set for the architecture knowledge system;
[0042] S2. Preliminary experiment: Extract a test set from the question set, conduct a preliminary test on representative large language models for stability and efficiency, and determine the appropriate size of the test set;
[0043] S3. Evaluation: Use the determined size in S2 to re-extract the test set, ask each large language model to be tested using the AO and COT questioning methods respectively, calculate the mean difference in the correct rates of these two methods, and use paired-sample t-test to statistically verify the mean difference to obtain an objective conclusion on the application effect of large language models in the field of architecture.
[0044] In this embodiment, for S1, constructing an architecture knowledge system, specifically, this system covers multiple fields such as architectural design, urban and rural planning, structural engineering, and environmental science, and the content extends from ancient architecture to modern building technology. As Figure 2 shown, it is mainly divided into the following four categories:
[0045] 1) Architectural design and theory: including architectural styles, planning principles, sustainable design, etc., to help architects balance aesthetics and practicality in design.
[0046] 2) Structural engineering and building physics: Focus on the stability of building structures and their interaction with the environment, covering seismic design, energy-saving systems, and thermal and acoustic performance control.
[0047] 3) Building codes, economy, and business management: including building laws and regulations, project management, etc., to provide compliance and economic decision-making support for building professionals.
[0048] 4) Planning and site design: Focus on urban planning and ecological protection, involving urban sociology, ecological environment, and infrastructure planning.
[0049] The construction of this architecture knowledge system provides a basis for classifying the collection and collation of the data set. The design of these questions often simulates real-world building and planning problems and is closely related to actual work requirements.
[0050] To ensure the scientificity and comprehensiveness of the test set, the sampling strategy of the evaluation system is strictly based on the distribution ratio of the questions in each subject in the overall question bank, aiming to ensure that the test set can comprehensively cover and represent the entire knowledge system, thereby providing an accurate and balanced data basis for evaluating the performance of language models.
[0051] In this embodiment, the purpose of the preliminary experiment stage is to verify the feasibility of the evaluation system, explore the basic performance of the LLM in the construction field, and determine the appropriate size of the test set to ensure the scientific nature of the system design and the smoothness of the evaluation process. Specifically, in S2, the preliminary experiment: extract a test set from the question set, conduct a preliminary test on the representative large language model for stability and efficiency, and determine the appropriate size of the test set, which specifically includes:
[0052] S2.1, evaluate the performance of the LLM under the same questions: extract part of the questions from the question set as the first test set, repeatedly ask the representative large language model multiple times using the AO questioning method, and judge the stability of the representative large language model. The stability refers to whether the answers to multiple questions are consistent;
[0053] This is the basis of the LLM model evaluation system. When answering the same question multiple times, the LLM should provide consistent answers, indicating that the model's understanding of specific inputs is stable and predictable
[0054] S2.2, determine the appropriate size of the test set: set different sizes of the second test set according to gradients, with multiple second test sets for each gradient, and then calculate the sample variance of the correct rates of these second test sets to determine the effective and stable test set size; this step aims to determine the appropriate scale of the data set to ensure comprehensive coverage of the knowledge system, reduce accidental errors in the evaluation, and improve the economy and efficiency of subsequent evaluations.
[0055] In this embodiment, the method further includes: using the AO questioning method to test the large language model to be tested, and quantifying the stability and accuracy results of the large language model to be tested, which specifically includes:
[0056] Use the AO questioning method to test the large language model to be tested. Use the sample variance of the correct rates of answers using test sets with different contents as the standard for evaluating stability, and use the mean of the answer correct rates as the standard for judging accuracy to obtain the stability and accuracy of the large language model to be tested, which can fairly reflect the answering level of the model. This step aims to obtain two goals for evaluating the LLM: stability and accuracy;
[0057] Regarding stability: When answering the same type of questions multiple times, the LLM should provide consistent answers, indicating that the model's understanding of specific inputs is stable and predictable. Stability helps reduce errors caused by randomness or model fluctuations, which is particularly important for the construction field that requires precise information.
[0058] Regarding accuracy: In the construction field, inaccurate information may lead to serious consequences, and the high accuracy of the LLM helps reduce the risks caused by incorrect information.
[0059] In this embodiment, the architectural knowledge system at least includes architecture design, urban and rural planning, structural engineering, and environmental science.
[0060] This evaluation method follows three principles: domain professionalism and comprehensiveness, dataset authority, and architectural practice applicability.
[0061] In addition, in the long-term research on exploring the prediction guidance method of large language models, researchers have particularly focused on the impact of prompt changes on obtaining the required answers. Therefore, to more comprehensively understand the role of prompts, this embodiment proposes the following method for evaluating the impact of AO and COT on LLM answers:
[0062] Use the questioning methods of AO and COT to question the LLMs respectively. Intuitively display the results by calculating the mean difference in the correct rates of the two methods, and use the paired-sample t-test to statistically verify the mean difference to obtain an objective conclusion.
[0063] Among them, AO refers to the Answer Only mode; COT refers to the Chain of Thought mode.
[0064] Embodiment 2
[0065] This embodiment of the present invention is a specific case of the method for evaluating the application effect of a large language model in the architectural field described in Embodiment 1.
[0066] 1. Step S1: Data collection and collation, constructing a knowledge system and a related question set
[0067] Collect the official questions of the registered architect and registered planner examinations published on public channels, and obtain paper test papers and practice questions from universities. The question bank includes actual examination questions over the years and selected simulation questions, aiming to cover the core knowledge points of professional qualification certification. The size of the entire test set is 10,440 questions, and these questions are all single-choice questions, with only one correct answer among the four options. Each sampling is strictly based on the distribution ratio of the questions in each subject in the overall question bank, as Figure 3 shown.
[0068] As shown in Table 1, in order to deeply explore the latest applications and progress of large language models (LLMs) in the field of Chinese architectural knowledge, we conducted a comprehensive evaluation of 14 high-performance LLMs that support Chinese input. The selection of these models is based on their widespread application and high usage frequency in the Chinese field, as well as the support of the professional development and operation and maintenance teams behind them, which not only ensures the authority and professionalism of the models but also guarantees their potential for future continuous development and optimization. Therefore, these models deserve our further attention and research.
[0069] Table 1 Large language models participating in the test
[0070]
[0071] 2. Step S2: Preliminary Experiment
[0072] 2.1. Preliminary Experiment ①: Evaluate the performance of LLMs under the same questions
[0073] Two models, Qwen-14B-Chat and GPT-3.5-turbo, were selected for testing. Among them, Qwen-14B is a locally deployed model, and GPT-3.5-turbo is an online model. They represent two different deployment methods of large language models (LLMs) respectively. A test set (Dataset 1) with a question volume of 500 was randomly selected from the total set, and the two models were repeatedly questioned 7 times. The questioning method was AO to streamline the question-and-answer time.
[0074] The experimental results are as Figure 4 shown. The answers of Qwen-14B-Chat in 7 answers are completely consistent, while the 7 answers of GPT-3.5-turbo contain inconsistent answers. The overall output of GPT-3.5-turbo is basically stable in 7 answers, and the difference in the correct rate is within 0.8%, and there is no obvious linear relationship between the rounds and the correct rate. That is, the correct rate of GPT-3.5-turbo does not increase significantly with time during the test period. The reason may be that the online model GPT-3.5-turbo continuously adjusts parameters through learning techniques during use, which results in different outputs at different time points. As a locally deployed model, Qwen-14B-Chat uses the same environment and parameter settings each time it runs, thus forming a relatively stable output result.
[0075] 2.2. Preliminary Experiment ②: Determine the appropriate test set size
[0076] As Figure 5 shown, the method of gradually increasing the test set size was adopted. Starting from 100 questions, it was gradually increased to 200 questions, 500 questions, 1000 questions, and finally 1500 questions to observe the change of sample variance. During the process, we found that when the test set size was between 500 and 1000 questions, the sample variance of the model's answering correct rate gradually tended to be stable and remained at a low level. So we used the bisection method to reduce the gradient and seek a better data set size between 500 and 1000. Finally, we decided to uniformly adopt 875 questions as the standard size of the test set in the subsequent experiments.
[0077] Under this test set size, the impact of the test set size on the model evaluation results is minimized, thus ensuring the stability and reliability of the evaluation. This optimization not only improves the efficiency of the evaluation process but also ensures the accuracy and consistency of the evaluation results.
[0078] 3. Step S3: Formal evaluation
[0079] 3.1. Conduct stability verification
[0080] When evaluating the stability of large language models, we use a test set of an appropriate scale determined in previous preliminary experiments to evaluate each model, and thus quantify the stability of the model by calculating the sample variance of the test result accuracy rate. The size of the sample variance is directly related to the performance fluctuations of the model on different test sets: a larger variance means greater fluctuations in the model's test results, reflecting lower stability; while a smaller variance means more consistent test results of the model, indicating higher stability.
[0081] Overall, the overall variances are all within an acceptable and relatively small range, indicating that the stabilities of the test models are all good.
[0082] As Figure 6 shown, according to our experimental data, although it is an online model, gpt-4-turbo performs best in overall stability, with a sample variance of 0.000060. ChatgIm2-6B and Qwen-7B-Chat follow closely, also showing good stability performance, with sample variances of 0.000068 and 0.000081 respectively. Other models such as ChatgIm2-6B, Qwen-7B-Chat, and Baichuan2-7B- also show good stability in specific fields. The chart further breaks down the model stabilities under different subjects: in the subjects of "Architectural Design and Theory" and "Planning and Site Design", the sample variances of the GPT-3.5-turbo, Qwen-14B-Chat, and Baichuan2-7B-Chat models are smaller, indicating their relatively stable performance in these subjects. In the subjects of "Building Codes, Economics, and Business Management" and "Structural Engineering and Building Physics", the Qwen-7B-Chat, Baichuan2-7B-Chat, and Qwen1.5-14B-Chat models show lower sample variances, meaning their stabilities in these fields are better.
[0083] 3.2. Conduct accuracy verification
[0084] As Figure 7The evaluation results shown indicate that the Qwen1.5-14B-Chat model performs particularly well in terms of accuracy, successfully surpassing the previously leading gpt-4-turbo model. In addition, other models in the Qianwen series also demonstrate remarkable performance.
[0085] Of particular note is that in the fields of building codes and economic and business management, all the models participating in the evaluation generally show a relatively high accuracy rate, indicating that in these fields, large language models already possess relatively mature application potential. However, in sharp contrast to this, even the top-ranked models fail to reach the average level in other fields in the subject of architectural design and theory, and this phenomenon is particularly evident among the top-ranked models. This not only reveals the possible limitations of large language models in this field but also provides an important reference for future model optimization and fine-tuning training.
[0086] 3.3. Comparison of the impact of using AO and COT prompting methods on the answer effects of the same model
[0087] Several models that performed outstandingly in different teams were selected and tested using the COT (Chain of Thought) answering method. Subsequently, the accuracy rates of the COT tests were compared and analyzed with those of the AO (direct answering) tests. By calculating the difference in the mean accuracy rates of the two answering methods, the result differences were visually presented, and the hypothesis testing method was used for objective statistical verification.
[0088] The difference in the output accuracy rates between COT and AO is as Figure 8 shown. Among the 5 tested LLMs, AO is overall better than COT in 4 LLMs. Looking at different fields, among the total 20 tested sub-subjects, COT is better than AO in only 5. It can be seen that using the COT answering method is not necessarily better than AO, which is exactly the opposite of the common perception that "the accuracy of COT may be better than AO".
[0089] Through analysis, it can be found that the different answering methods do have a certain impact on the model's prediction results, and the maximum difference in the mean accuracy rates between the two answering methods does not exceed 3%, but whether this impact is positive or negative is not clear. To further verify this phenomenon, we took the null hypothesis that "there is no significant difference in the accuracy rates of the results of large language models answering in the AO and COT methods in the evaluation" and performed a paired-sample t-test on the sample data, setting the p-value threshold to 0.05. As Figure 9 shown by the test results, in most cases, there is no significant statistical difference in the results obtained by the COT and AO answering methods.
[0090] In addition, during the experiment, we also noticed a phenomenon: the time required for the model to answer questions using the COT method is significantly longer than that using the AO method. For example, when asking questions in the Qwen1.5-14B-Chat model, the average answering time of AO is 2.38 s / question, while that of COT reaches 62.23 s / question. This finding suggests that when evaluating the efficiency of large language models, in addition to considering accuracy, the answering time should also be taken as an important consideration factor. Future research can further explore how to optimize the COT answering method to reduce the required time while maintaining or improving the accuracy of the answers.
[0091] Example 3
[0092] An embodiment of the present invention provides a system for evaluating the application effect of a large language model in the field of architecture, which is used to execute the method for evaluating the application effect of a large language model in the field of architecture described in Embodiment 1, including:
[0093] A construction module, configured to construct an architecture knowledge system and organize a question set of the architecture knowledge system;
[0094] A pre-experiment module, configured to conduct a pre-experiment: extract a test set from the question set, conduct a preliminary test on a representative large language model for stability and efficiency, and determine a suitable test set size;
[0095] An evaluation module, configured to re-extract a test set using the size determined in the pre-experiment module, and test the large language model to be tested in combination with the AO question method to quantify the stability and accuracy of the large language model to be tested.
[0096] Example 4
[0097] An embodiment of the present invention provides a computer device, including: one or more processors;
[0098] The processor is configured to store one or more programs; when the one or more programs are executed by the one or more processors, the method for evaluating the application effect of a large language model in the field of architecture described in Embodiment 1 is implemented.
[0099] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the method for evaluating the application effect of a large language model in the field of architecture described in Embodiment 1 is implemented.
[0100] In summary, the method and system for evaluating the application effect of large language models in the construction field provided by the present invention establish a unified and highly industry-targeted evaluation standard; ensure the reliability and adaptability of LLM in complex construction tasks through quantitative analysis, improve its accuracy in professional fields such as building codes and design principles, and comprehensively quantify and evaluate its application effect from all stages of the project.
[0101] For the remaining technical features in this embodiment, those skilled in the art can flexibly select them according to the actual situation to meet different specific actual needs. However, it is obvious to those of ordinary skill in the art that these specific details do not have to be adopted to implement the present invention. In other instances, well-known components, structures, or parts are not specifically described in order to avoid obscuring the present invention, and all are within the scope of the technical solutions claimed in the claims of the present invention.
[0102] Any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention. In the above description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it is obvious to those of ordinary skill in the art that these specific details do not have to be adopted to implement the present invention. In other instances, well-known technologies, such as specific construction details, working conditions, and other technical conditions, are not specifically described in order to avoid obscuring the present invention.
[0103] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, there will be changes in the specific implementation manners and application scopes according to the idea of the present invention. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for evaluating the application effect of large language models in the construction field, characterized in that, It includes the following steps: S1. Construct an architectural knowledge system and organize a question set for the architectural knowledge system; S2. Preliminary experiment: Extract a test set from the question set and conduct a preliminary test on representative large language models for stability and efficiency to determine an appropriate test set size; S3. Formal evaluation: Use the size determined in S2 to re-extract a test set, conduct stability evaluation and accuracy evaluation on each large language model to be tested, obtain an objective conclusion on the application effect of each large language model to be tested in the field of architecture, and finally select a large language model that meets the requirements.
2. The method for evaluating the application effect of a large language model in the construction field according to claim 1, wherein The S2, preliminary experiment: Extract a test set from the question set and conduct a preliminary test on representative large language models for stability and efficiency to determine an appropriate test set size, specifically including: S2.
1. Extract some questions from the question set as the first test set, repeatedly ask a representative large language model using the AO questioning method multiple times, and judge the stability of the representative large language model. The stability refers to whether the answers to multiple questions are consistent; S2.
2. Extract second test sets of different sizes in gradients, with multiple second test sets in each gradient, and then calculate the sample variance of the correct rates of these second test sets to determine an effective and stable test set size.
3. The method for evaluating the application effect of a large language model in the construction field according to claim 1, wherein The method further includes: S4. Ask the same large language model to be tested using different questioning methods to determine the appropriate questioning method for the large language model to be tested.
4. The method for evaluating the application effect of a large language model in the construction field according to claim 3, wherein, The questioning method is AO or COT.
5. The method for evaluating the application effect of a large language model in the construction field according to claim 1, wherein The method further includes: When conducting accuracy and stability evaluations, first extract test sets with different contents, use the sample variance of the correct rates of the answers to each test set as the standard for evaluating stability, and use the mean of the correct rates of the answers as the standard for evaluating accuracy to obtain the stability and accuracy results of the large language model to be tested.
6. The method for evaluating the application effect of a large language model in the construction field according to claim 1, wherein The architectural knowledge system at least includes those covering architectural design, urban and rural planning, structural engineering, and environmental science.
7. A system for evaluating the application effect of large language models in the construction field, which is used to execute the method for evaluating the application effect of large language models in the construction field according to any one of claims 1-6, characterized in that, It includes: A construction module for constructing an architectural knowledge system and organizing a question set for the architectural knowledge system; A preliminary experiment module for conducting a preliminary experiment: Extracting a test set from the question set and conducting a preliminary test on representative large language models for stability and efficiency to determine an appropriate test set size; An evaluation module for using the size determined in the preliminary experiment module to re-extract a test set, conduct stability evaluation and accuracy evaluation on each large language model to be tested, obtain an objective conclusion on the application effect of each large language model to be tested in the field of architecture, and finally select a large language model that meets the requirements.
8. A computer device, characterized in that, It includes: One or more processors; The processor is used to store one or more programs; When the one or more programs are executed by the one or more processors, it implements a method for evaluating the application effect of a large language model in the field of architecture as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that There is a computer program stored thereon, and when the computer program is executed, it implements a method for evaluating the application effect of a large language model in the field of architecture as described in any one of claims 1 to 6.