Metro field large language model evaluation method and system

By constructing a large language model evaluation method and system dedicated to the subway field, the problem of poor performance of large language models in the existing technology in the subway field is solved, and a more accurate and comprehensive model evaluation of the subway field is achieved.

CN120163142APending Publication Date: 2025-06-17QINGDAO BAONING FUTIAN INTELLIGENT TRAFFIC TECH DEV CO LTD
View PDF 0 Cites 13 Cited by

Patent Information

Application Number
CN202510092364.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing large language model evaluation methods performed poorly in the subway field, lacking professional knowledge data sets, general evaluation indicators and evaluation models in the subway professional field, resulting in poor performance of the model in professional tasks.

Method used

Build a large language model evaluation method and system for the subway field, collect and optimize the evaluation data set, determine appropriate evaluation indicators and systems in the subway field, adopt multi-agent collaborative evaluation method, dynamically extract evaluation tasks, and realize fully automatic evaluation.

Benefits of technology

The evaluation accuracy and comprehensiveness of large-language models in the subway field are improved, ensuring that the model's performance in professional tasks is more in line with the actual needs of the subway industry, and improving the accuracy and efficiency of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163142A_ABST
    Figure CN120163142A_ABST
Patent Text Reader

Abstract

The invention relates to a metro field large language model evaluation method and system. The method comprises the following steps: step 1, constructing a data set; the step 1 comprises the following steps of: 1.1, collecting public professional data; step 1.2, collecting internal professional data; 1.3, collecting a general data set; 2, data processing; in the step 2, the following steps are executed: step 2.1, data formatting; step 2.2, data cleaning is carried out; step 2.3, performing data enhancement; 3, fine adjustment of the model; in the step 3, the following steps are executed: step 3.1: constructing a training data set; step 3.2, fine adjustment of the model; step 3.3, carrying out model evaluation; 4, constructing an evaluation system; in the step 4, the following steps are executed: step 4.1, determining evaluation indexes; step 4.2, constructing an evaluation system; step 4.3, constructing an intelligent agent; step 4.4, multi-agent cooperation is carried out; and 5, performing model evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a large language model evaluation method and system in the field of subways, and belongs to a data processing system or method specially suitable for administration, commerce, finance, management, supervision or prediction purposes. Background Art

[0002] With the rapid development of large language models, how to evaluate the output quality of large language models is crucial. Large model evaluation can measure the quality level of model output and ensure user experience. Large models are still in the early stages of development in the subway industry. Currently, there are few large models dedicated to the subway industry, and there is no method or indicator system on the market specifically for evaluating the performance of large models in the subway field.

[0003] The main steps of the large language model evaluation method are as follows: 1) Construction of evaluation datasets (automatic or manual). The automatic method is to automatically divide the evaluation dataset from the full dataset, while the manual method is to manually select data and create the dataset; 2) Determine the evaluation indicators. Commonly used evaluation indicator systems include language level, semantic level and knowledge level. The language level includes lexical correctness, grammatical correctness and text correctness, the semantic level includes semantic accuracy, logical coherence and style consistency, and the knowledge level includes knowledge accuracy, knowledge richness and knowledge consistency; 3) Select an evaluation method. Common evaluation methods include manual evaluation, automatic evaluation, and comparative evaluation. 4) Select the evaluation subject. Usually, a model with strong performance, such as ChatGPT-4o, is selected as the evaluation subject model; 5) Evaluate the large language model being evaluated and output the evaluation results.

[0004] In the field of subway operation control and management, there are the following technical problems: 1. Since the evaluation data set is mainly generated by capturing public information and lacks professional knowledge data set in the subway field, the performance evaluation of large models in the subway field is not effective; 2. Since the current large model evaluation mainly uses general evaluation indicators (such as accuracy, consistency, alignment, etc.), there is a lack of evaluation indicators for the subway field, so the evaluation dimensions of large models are not comprehensive; 3. Since the current mainstream general large models on the market such as ChatGPT-4o are mainly used to evaluate the models to be evaluated, there is a lack of evaluation models and intelligent agents in the subway professional field to conduct targeted evaluations of the performance of large models in the subway field.

[0005] In addition, when applying the general large language model to the subway field, the general model performs poorly for some tasks that require strong professional knowledge, and there is a lack of corresponding standards to evaluate the effectiveness of the model.

[0006] In addition, most of the large-model benchmark test data sets currently used in the field of subway operations focus on the general capabilities of the model, including large-model understanding, generation, reasoning, knowledge capabilities, etc., and are currently mainly implemented through examination plans. The evaluation data is generated by extracting part of the training data. When the test questions are subjective or open questions and answers, the test results still require manual subjective evaluation. Summary of the invention

[0007] The technical problem to be solved by the present invention is generally to provide a method and system for evaluating a large language model in the subway field. The present invention takes into account the actual characteristics of the subway industry, and in addition to evaluating general capabilities, it also considers evaluating the special capabilities of large models, optimizing fixed evaluation data into dynamic evaluation tasks, and optimizing manual evaluation into fully automatic evaluation.

[0008] The present invention combines the actual situation in the field of subways to determine an all-round integrated evaluation scheme of evaluation data sets, evaluation indicators, evaluation systems, and evaluation tools. Collect data to construct an evaluation data set, determine the evaluation indicators and system, select a large base model for fine-tuning, and then manually evaluate the performance of the fine-tuned model. After the performance of the fine-tuned model meets the requirements, use prompt words to construct multiple intelligent agents. When the model to be evaluated is evaluated, dynamically extract questions of different types and difficulties from the evaluation data set to construct an evaluation task. The performance of the model to be evaluated on each task is scored through the collaboration of three intelligent agents, and finally the score of the entire task is obtained, and the result is output.

[0009] The general technical principle of the present invention is to use evaluation indicators to measure the difference between the predicted results of the model and the actual results. The evaluation indicators include accuracy, multi-round reasoning, decision making, alignment, safety, and ethics.

[0010] To solve the above problems, the technical solution adopted by the present invention is: A method for evaluating a large language model in the subway field comprises the following steps: Step 1: Dataset construction, collect data to build the evaluation dataset; In step 1, the following steps are performed: Step 1.1: Public professional data collection; Step 1.2: Internal professional data collection; Step 1.3: General dataset collection; Step 2: Data processing; In step 2, the following steps are performed: Step 2.1: Data formatting; Step 2.2: Data cleaning; Step 2.3: Data augmentation; Step 3: Model fine-tuning; The following steps are performed in Step 3: Step 3.1, Training dataset construction; First, generate instruction-based question-answer pairs from the processed data and convert them into the jsonl format; Then, use a commercial large language model to generate question-answer pairs from the original data, score the questions and answers from 0 to 5, and give specific reasons for the scoring. Step 3.2: Model fine-tuning; Adopt the methods of fine-tuning and instruction fine-tuning to fine-tune the open-source large language model; Use the fine-tuning method LoRA to fine-tune the model: Modify and initialize the model; Add the LoRA module to the fully connected layer of the pre-trained open-source large language model; Step 3.3: Model evaluation; Step 4: Evaluation system construction; The following steps are performed in Step 4: Step 4.1: Determine evaluation metrics. The evaluation metrics are divided into six categories: accuracy, multi-round reasoning, decision-making, alignment, security, and ethics; Among them, Accuracy: Evaluate the ability of the large language model to be evaluated to provide correct answers; Multi-round reasoning: Examine the ability of the large language model to be evaluated to maintain context and perform reasoning in several rounds of conversations; Decision-making: Measure the ability of the large language model to be evaluated to make reasonable decisions based on the given information; Alignment: Ensure that the answers of the large language model to be evaluated are consistent with user needs and human values; Security: Evaluate whether the large language model to be evaluated follows security protocols and does not disclose sensitive information; Ethics: Evaluate whether the large language model to be evaluated meets ethical standards, without discrimination, bias, or inappropriate content; Among the above six evaluation metrics, objective questions (single-choice, multiple-choice, judgment) are scored by strict matching. Full marks are given for correct answers and 0 marks for wrong answers; Fill-in-the-blank questions and essay questions are scored according to the following criteria (from high to low in importance): metric, and the meaning of each scoring level: score_def parameter is configured; As the scoring criteria and rules of the scoring agent, the details are shown in Table 1 below.

[0011] Table 1 The evaluation metrics for accuracy, multi-round reasoning, and decision-making are mainly based on precision metrics to detect whether the responses meet the corresponding standards: ROUGE (Recall-Oriented Understudy for Gisting Evaluation) is often used as an evaluation metric. Among them, ROUGE-1 is the recall rate at the unigram level between the model output and the reference summary. The closer its value is to 1, the higher the consistency between the model output and the reference answer. The ROUGE-1 formula is shown as follows: ; gram1 represents unigrams, ReferenceSummaries is a set, and the Count function represents the count of the deduplicated set.

[0012] Alignment, safety, and ethics are mainly detected by similarity detection to check whether the responses meet the corresponding standards. The main calculation method is as follows: score y =sim(E(x), E(d(y))); where score y represents the y metric score, the sim() function represents cosine similarity, the E() function represents the SinCSE sentence encoder, x represents the text output by the model under test, and d(y) represents the text in the dataset that does not meet safety, alignment, and ethics.

[0013] After obtaining the metric scores, they are normalized to the 1 - 4 score range.

[0014] Step 4.2: Evaluation system construction; The evaluation system includes evaluation task design and evaluation rule design; among them, Step 4.2.1, Evaluation task design: The types of evaluation tasks include single-choice questions, multiple-choice questions, true / false questions, fill-in-the-blank questions, and short-answer questions; each type of evaluation task contains multiple types of questions, totaling 100 points; the types of questions cover the fields of passenger service, line operation and maintenance, emergency management, and operation safety; Step 4.2.2, Evaluation rule design: Corresponding scoring rules are formulated based on the key points of the types of questions in each evaluation task; For single-choice questions, multiple-choice questions, true / false questions, and fill-in-the-blank questions, they are matched according to the correct answers. One gets full marks for a correct answer and no marks for a wrong answer; for short-answer questions, the answers of the model to be evaluated are scored based on multi-agent collaboration, and finally the scores of each model to be evaluated on each evaluation metric are calculated; the final score of the large language model to be evaluated is the weighted sum of the scores of six metrics; the specific calculation formula is as follows: ; Among them, Score is the total score of the model to be evaluated, w i is the weight of the i-th index, and M i is the score on the i-th evaluation index; Step 4.3: Agent construction; Build a prompt engineering, construct prompt templates for different roles, guide the model output, and enable it to better adapt to the evaluation tasks and evaluation fields.

[0015] Specify the scenario of agent evaluation in the prompt and define the scenario; Specify the role of the agent in the prompt; Determine the scoring rules of the evaluation and determine the index standards of the evaluation; Specify the steps and output rules of agent evaluation.

[0016] Step 4.4: Multi-agent collaboration; Adopt a sequential collaboration method for multiple agents, which are divided into two scoring agents and one referee agent. Create scoring agents for two evaluation models by configuring prompts related to the scoring function, and create a referee agent for one evaluation model by configuring prompts related to the referee function; for each type of question in the evaluation task, the two scoring agents give scores x1 and x2 respectively, calculate their average value x and standard deviation s; secondly, if the standard deviation s is greater than or equal to 1, the referee agent evaluates the quality of the scores of the two agents, compares the two scores and selects one of them as the score of the model to be evaluated on this question; if it is less than 1, the average value x is used as the score of this question. The specific formula is as follows:

[0017] Step 5: Model evaluation; First, select the model to be evaluated, dynamically extract evaluation data from the evaluation dataset, and generate six types of evaluation tasks: accuracy, multi-round reasoning, decision-making, alignment, security, and ethics; then, the question types of each type of evaluation task include at least one of single-choice, multiple-choice, judgment, filling in the blanks, and answering questions according to the characteristics of the evaluation indicators; secondly, score the answers of the model to be evaluated by two scoring agents again. The two scoring agents give scores x1 and x2 respectively, calculate their average value x and standard deviation s. If the standard deviation s is greater than or equal to 1; the referee agent compares the scores of the two agents and selects one of the agents as the score of the model to be evaluated on this type of question in the evaluation task; if it is less than 1, the average value x is used as the score of this question; thirdly, summarize the scores of the evaluation tasks after all questions of the model to be evaluated are answered; as shown in the following formula, Q ij represents the question, s ijIndicates the score of the model to be evaluated on this question; ; After that, the total score of the model to be evaluated is obtained by weighted averaging the scores of the six evaluation tasks.

[0018] Furthermore, in step 1.1, the publicly available professional data includes the publicly available data officially released by the subway company, the standard documents related to the subway field, and the professional books related to the subway field; In step 1.2, the internal professional data includes the internal organizational structure of the subway company, management specifications, and non-confidential data generated by business systems, and sensitive data is desensitized; The general data includes common sense knowledge data related to the subway field.

[0019] Furthermore, in step 2.1, convert the audio and video to text, and use speech recognition algorithms and subtitle recognition tools to extract the text in the audio and video files; convert the pictures to text, and use the OCR algorithm to convert the pictures into text data; In step 2.2, perform text standardization, sensitive word filtering, similarity deduplication, and toxicity elimination on the collected text data; Based on the standardization basis, convert data in different formats into the same format, and convert traditional Chinese text to simplified Chinese; In the similarity deduplication step, use the SimHash algorithm to judge the similarity of the text and filter out samples that exceed the threshold first; In the toxicity elimination processing step, automatically identify and remove sensitive or non-compliant content in the data, including data preprocessing, building and updating the sensitive word library, performing content detection, removing the identified sensitive content, and auditing and verifying the processing results; In step 2.3, first, perform text enhancement; based on the commercial large language model, based on the Prompt configuration, select the seeds from the pre-enhancement training set and splice them into the Prompt to enhance the input training data; Secondly, perform text classification; through the commercial large language model, imitate the format of the seed data, and classify according to passenger service, line operation and maintenance, emergency management, and safety operation to generate enhanced data and enhance the effect of the classification scenario; Thirdly, perform text extraction; through the commercial large language model, use data enhancement strategies to mine text content; the data enhancement strategies include synonym replacement, random sampling, and translation transformation.

[0020] In step 3.1, the jsonl format is: {"instruction ": "","output ": "","score ":"","explanation ": ""}; In step 3.2, set the hyperparameters as follows: lora_rank to determine the rank size r of the low-rank matrix; lora_alpha, the scaling factor, used to adjust the influence degree of the low-rank adaptation part on the final result; lora_dropout, the Dropout coefficient, used to prevent overfitting; learning_rate, the initial learning rate of the AdamW optimizer; num_train_epochs, the number of training epochs; In step 3.2, For the pre-trained weight matrix , its update is represented by low-rank factorization , where , , and the rank r ≤ min(d, k); : the pre-trained weight matrix, representing the initial state of the model; : a part of the update matrix, capturing the adjustment in the row direction; : another part of the update matrix, capturing the adjustment in the column direction; r ≤ min(d, k): the rank of the low-rank factorization, controlling the update capacity and computational efficiency.

[0021] At the start of training, A is initialized with a random Gaussian distribution and B is initialized with zeros. So at the start of training is zero; Again, perform forward propagation calculation: During forward propagation, the calculation method is , and is scaled by α and r, where x is the input vector; : the pre-trained weight matrix, representing the initial state of the model; B: a part of the update matrix, capturing the adjustment in the row direction; A: another part of the update matrix, capturing the adjustment in the column direction; x: the input vector, which undergoes forward propagation through the model.

[0022] Subsequently, training and optimization; Use the prepared dataset to train the model with the LoRA module added; During training, the weights of the pre-trained model are frozen and do not receive gradient updates, while A and B contain trainable parameters, and the values of A and B are updated through the backpropagation algorithm to minimize the loss function.

[0023] Further, the following steps are performed in Step 4.2: Step 4.2.1: Evaluate task design; First, each evaluation is divided into six evaluation tasks, which respectively evaluate six types of indicators: accuracy, multi-round reasoning, decision-making, alignment, security, and ethics. Then, design an evaluation data sampling algorithm to dynamically select questions from the evaluation dataset to generate evaluation tasks each time. The types of evaluation tasks include single-choice questions, multiple-choice questions, true or false questions, fill-in-the-blank questions, and essay questions. There are several questions of each type, totaling 100 points. The questions cover the fields of passenger service, line operation and maintenance, emergency management, and operation safety.

[0024] Step 4.2.2: Evaluate rule design; Formulate corresponding scoring rules based on the key points of the questions of each evaluation task type. For single-choice questions, multiple-choice questions, true or false questions, and fill-in-the-blank questions, they are matched according to the correct answers. One gets full marks for answering correctly and no marks for answering wrongly. For essay questions, the answers of the model to be evaluated are scored through multi-agent collaboration, and finally, the scores of each model to be evaluated on each evaluation index are calculated. The final score of the large language model to be evaluated is the weighted sum of the scores of the six indexes.

[0025] Further, in Step 4.3, in the prompt engineering; Specify the scenario of the agent evaluation in the prompt and define the scenario; Specify the role of the agent in the prompt; Determine the scoring rules of the evaluation and determine the index standards of the evaluation; Specify the steps and output rules of the agent evaluation.

[0026] Further, in Step 4.4, first, use several fine-tuned evaluation models to create agents through prompt engineering, configure different prompts, with two as scoring agents and one as a referee agent. Then, for the questions of each evaluation task type, the two scoring agents respectively give scores x1 and x2, calculate their average value x and standard deviation s. Secondly, if the standard deviation s is greater than or equal to 1, the referee agent compares the two scores and selects one as the score of the model to be evaluated on this question; if it is less than 1, the average value x is used as the score of this question.

[0027] Further, in step 5, first, select the model to be evaluated; then, dynamically extract evaluation data from the evaluation dataset to generate six types of evaluation tasks, including accuracy, multi-round reasoning, decision-making, alignment, security, and ethics; the question types for each type of task include single-choice, multiple-choice, judgment, fill-in-the-blank, and question-and-answer. Secondly, a scoring agent scores the answers of the model to be evaluated for single-choice, multiple-choice, and judgment questions. A full score is given for a correct answer, and no score is given for a wrong answer. Thirdly, two scoring agents score the answers of the model to be evaluated for fill-in-the-blank and question-and-answer questions. The two scoring agents give scores x1 and x2 respectively, calculate their average value x and standard deviation s. If the standard deviation s is greater than or equal to 1, the referee agent compares the two scores and selects one of them as the score of the evaluated model for this question; if it is less than 1, the average value x is used as the score for this question. Again, after all questions of the model to be evaluated are answered, summarize the scores of the evaluation tasks. As shown in the following formula, Q ij represents the question, s ij represents the score of the model to be evaluated for this question; ; After that, the scores of the six evaluation tasks are weighted and averaged to obtain the total score of the model to be evaluated. After the model evaluation is completed, an evaluation result report is output. The report includes evaluation objectives, dataset descriptions, task descriptions, evaluation environment indicator descriptions, evaluation indicators, and quantitative results, and is presented in radar charts and bar charts.

[0028] A large language model evaluation system in the subway field includes a large language model fine-tuned based on an open-source large model, which is used to execute the above evaluation method.

[0029] The present invention has the following advantages: 1. Combining steps 1 and 3, the original data types of the dataset are rich. The evaluation model after fine-tuning the dataset can understand the professional knowledge of the subway field and has better adaptability to downstream tasks in the subway field; 2. Combining step 4.2, the evaluation indicators and systems are relatively complete, comprehensively evaluating the performance indicators of the model in multiple dimensions; 3. Combining step 4.3, the carefully designed prompt words can make the fine-tuned large model evaluate the model according to the preset instructions and directions; 4. Combining step 4.4, the method of using multi-agent collaborative evaluation to evaluate the model to be evaluated can avoid the problem of inaccurate evaluation by a single model and improve the overall evaluation accuracy.

[0030] In summary, the present invention constructs an evaluation dataset in the subway field including passenger services, line operation and maintenance, emergency management, safety operation, and general knowledge. It includes six evaluation tasks and an evaluation index system of accuracy, multi-round reasoning, decision-making, consistency, safety, and ethics, and realizes the fine-tuning of a large language model in the subway field. The present invention adopts a multi-agent collaborative evaluation method to evaluate the large language model in the subway field, ensuring the accuracy of the evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic diagram of the overall design process of the present invention.

[0032] Figure 2 It is a schematic diagram of the overall framework of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0033] Such as Figure 1 、 Figure 2 , for the large language model evaluation method and system in the subway field of the present invention, the basic design concept of the present invention is: first, analyze the evaluation requirements; then, construct the evaluation dataset; secondly, prepare the evaluation environment; thirdly, execute the benchmark test; then, analyze the evaluation results; finally, display the final results.

[0034] The method includes the following steps: Step 1: Dataset construction. Collect data with high quality, large scale, and complete variety to improve the model representation ability and generalization ability. The data is mainly collected from the following steps; Step 1.1: Collect public professional data. Collect the public data published on the official website, Weibo account, WeChat official account, and Douyin official account of the subway company; collect the national standards, industry standards, group standards, and local standard documents related to the subway field; collect the professional books related to the subway field.

[0035] Step 1.2: Collect internal professional data. Non-confidential data such as the data generated by the internal organizational structure, management norms, and business systems of the subway company, and sensitive data is desensitized.

[0036] Step 1.3: Collect general datasets. To prevent the large model from forgetting general knowledge during the fine-tuning process, add common sense general knowledge data to the dataset.

[0037] Step 2: Data processing; Step 2.1: Data formatting.

[0038] Convert audio and video to text: Use speech recognition algorithms and subtitle recognition tools to extract the text in audio and video files.

[0039] Image to Text: Use the OCR (Optical Character Recognition) algorithm to convert images into text data.

[0040] Step 2.2: Data cleaning.

[0041] Perform text standardization, sensitive word filtering, similarity deduplication, and toxicity elimination on the collected text data.

[0042] Basis for standardization: Convert data in different formats into the same format and convert traditional Chinese text to simplified Chinese.

[0043] Similarity deduplication steps: Calculate the Hamming distance (i.e., the sum of the number of different bits in the corresponding positions of two binary strings) between two SimHash values through the SimHash algorithm to determine the similarity of the text, and filter out samples with a Hamming distance exceeding the threshold.

[0044] Toxicity elimination processing steps: Automatically identify and remove sensitive or non-compliant content in the data, including data preprocessing, building and updating the sensitive word library, performing content detection, removing the identified sensitive content, and auditing and verifying the processing results.

[0045] Step 2.3: Data augmentation.

[0046] Text augmentation: Through a high-performance commercial large language model. Based on the Prompt configuration, select seeds from the pre-augmentation training set and splice them into the Prompt to augment the input training data, thereby improving data diversity, balance, etc.

[0047] Text classification: Through a high-performance commercial large language model, imitate the format of the seed data, and generate augmented data according to the four categories of passenger services, line operation and maintenance, emergency management, and safety operation, thereby enhancing the effect of the classification scenario.

[0048] Text extraction: Through a high-performance commercial large language model, during data augmentation, the generated augmented data is highly relevant to the target task, and various data augmentation strategies (such as synonym replacement, random sampling, translation transformation, etc.) are used to mine the text content while ensuring the value of the content in terms of maintaining data context and semantic consistency.

[0049] Step 3: Model fine-tuning; Step 3.1: Construction of the training dataset.

[0050] Generate instruction-based question-answer pairs from the processed data and convert them into the jsonl format, where the jsonl format is: {"instruction ": "","output ": "",score ": 5,"explanation ": ""}; Generate question-answer pairs from the raw data using a high-performance commercial large model, score the questions and answers from 0 to 5, and provide specific reasons for the scoring.

[0051] Step 3.2: Model fine-tuning.

[0052] Fine-tune the open-source large model using efficient fine-tuning and instruction fine-tuning methods.

[0053] Efficient fine-tuning improves the fine-tuning efficiency by reducing the number of parameters to be updated or changing the way of parameter update, thereby reducing the dependence on computing resources and shortening the training time.

[0054] Instruction fine-tuning improves the alignment degree of the question-answer process by enhancing the intent understanding ability of the large model.

[0055] Use the efficient fine-tuning method LoRA to fine-tune the model, and the hyperparameters are as follows: lora_rank: Determine the rank size r of the low-rank matrix, which is a hyperparameter. A smaller r makes the low-rank matrix simpler, with fewer parameters to learn during training, thus accelerating the training speed and potentially reducing the computing requirements. However, it may reduce the ability of the low-rank matrix to capture task-specific information, resulting in poor performance of the model on new tasks.

[0056] lora_alpha: Scaling factor used to adjust the influence degree of the low-rank adaptation part on the final result. It is usually set to 1 by default, and its adjustment process is similar to adjusting the learning rate.

[0057] lora_dropout: Dropout coefficient used to prevent overfitting.

[0058] learning_rate: Initial learning rate of the AdamW optimizer. An overly large learning rate may cause the loss value not to converge or overfit, that is, the model overfits to the training set and loses the generalization ability to data outside the training set; an overly small learning rate may make the training process converge too slowly.

[0059] num_train_epochs: Number of training epochs. If the loss value does not converge to the ideal value, the number of training epochs can be increased or the learning rate can be appropriately reduced.

[0060] Among them, the parameter and parameter value are preferably: Learning rate is 1 x10 -4; The batch size is 16; The Max Seq. Len. is 1280; LoRA α is 16; LoRA r is 16; LoRA dropout is 0.05; The Max. length of new tokens is 512.

[0061] Model modification and initialization: Add the LoRA module to the fully connected layer of the pre-trained language model. For a pre-trained weight matrix , represent its update through low-rank factorization , where , , and the rank r ≤ min(d, k). At the start of training, initialize A with a random Gaussian distribution and B with zeros, so at the start of training is zero.

[0062] Forward propagation calculation: During the forward propagation process, the calculation method is , and scale through α and r, where x is the input vector.

[0063] Training and optimization: Use the prepared dataset to train the model with the added LoRA module. During training, the weights of the pre-trained model are frozen and do not receive gradient updates, while A and B contain trainable parameters, and update the values of A and B through the backpropagation algorithm to minimize the loss function.

[0064] Step 3.3: Model evaluation.

[0065] Step 4: Evaluation system construction; Step 4.1: Determine the evaluation metrics.

[0066] The evaluation metrics are divided into six categories, namely accuracy, multi-turn reasoning, decision-making, alignment, security, and ethics.

[0067] Accuracy: Evaluate the ability of the model to provide correct answers; Multi-turn reasoning: Examine the ability of the model to maintain context and perform reasoning in multi-turn conversations; Decision-making: Measure the ability of the model to make reasonable decisions based on given information; Alignment: Ensure that the model's answers are consistent with user needs and human values; Security: Evaluate whether the model follows security protocols and does not disclose sensitive information; Ethics: Evaluate whether the model complies with ethical standards, without discrimination, bias, or inappropriate content.

[0068] Step 4.3: Agent Construction.

[0069] Using prompt engineering, construct prompt templates for different roles to guide the model output, enabling it to better adapt to the evaluation tasks and evaluation domains.

[0070] 1) Specify the scenario for agent evaluation in the prompt and define the scenario.

[0071] 2) Specify the role of the agent in the prompt.

[0072] 3) Determine the scoring rules for the evaluation and the metric standards for the evaluation.

[0073] 4) Specify the steps and output rules for agent evaluation.

[0074] Step 4.4: Multi-Agent Collaboration.

[0075] Using three fine-tuned evaluation models, create agents by configuring different prompts through prompt engineering. Two of them are used as scoring agents, and one is used as a referee agent. For each evaluation question, the two scoring agents give scores x1 and x2 respectively. Calculate their average x and standard deviation s. If the standard deviation s is greater than or equal to 1, the referee agent compares the two scores and selects one as the score of the evaluated model for this question; if it is less than 1, the average x is used as the score for this question.

[0076] Step 5: Model Evaluation; 1) Select the model to be evaluated; 2) Dynamically extract evaluation data from the evaluation dataset to generate six types of evaluation tasks (accuracy, multi-round reasoning, decision-making, alignment, security, ethics), and each type of task includes question types such as single-choice, multiple-choice, judgment, fill-in-the-blank, and question-and-answer; 3) One scoring agent scores the answers of the model to be evaluated on single-choice, multiple-choice, and judgment questions. Full marks are given for correct answers, and no marks are given for wrong answers; 4) Two scoring agents score the answers of the model to be evaluated on fill-in-the-blank and question-and-answer questions. The two scoring agents give scores x1 and x2 respectively. Calculate their average x and standard deviation s. If the standard deviation s is greater than or equal to 1, the referee agent compares the two scores and selects one as the score of the evaluated model for this question; if it is less than 1, the average x is used as the score for this question; 5) After the model to be evaluated has completed answering all questions, summarize the scores of the evaluation tasks; as shown in the following formula, Q ij represents the question, s ij represents the score of the model to be evaluated for this question;

[0077] 6) The total score of the model to be evaluated is obtained by weighted averaging the scores of six evaluation tasks. After the model evaluation is completed, an evaluation result report is output. The report includes evaluation objectives, dataset descriptions, task descriptions, evaluation environment indicator descriptions, evaluation metrics, and quantitative results, and is presented in various forms such as radar charts and bar charts.

[0078] As an extended embodiment, the present invention can utilize an open-source large language model with an external knowledge base, store subway domain-related knowledge data in the knowledge base, and evaluate the performance of the large language model in the subway domain using prompt words. Its advantage is that there is no need to fine-tune the model, and the implementation cost of the solution is relatively low; however, because only the method of using an external knowledge base is adopted, there may be a lack of professional knowledge in the subway domain, and multiple agents are not used for collaborative evaluation, so the evaluation accuracy is relatively low.

[0079] The present invention is fully described for a clearer disclosure, and prior art will not be listed one by one.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; as is obvious to those skilled in the art, combinations of multiple technical solutions of the present invention are possible, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. The technical content not elaborated in the present invention is all well-known technology.

Claims

1. A large language model evaluation method in the subway field, characterized in that: The following steps are involved: Step 1: Dataset construction, collect data to build the evaluation dataset; In step 1, the following steps are performed: Step 1.1: Public professional data collection; Step 1.2: Internal professional data collection; Step 1.3: General dataset collection; Step 2: Data processing; the following steps are performed in step 2: Step 2.1: Data formatting; Step 2.2: Data cleaning; Step 2.3: Data enhancement; Step 3: Model fine-tuning; the following steps are performed in step 3: Step 3.1, training data set construction; First, generate instruction-type question-answer pairs from the processed data and convert them into jsonl format; then, use the commercial large language model to generate question-answer pairs from the original data, and score the questions and answers from 0 to 5, and give specific explanations for the scoring settings; Step 3.2: Model fine-tuning; Fine-tune the open source large language model using fine-tuning and instruction fine-tuning; Fine-tune the model using the fine-tuning method LoRA: Modify and initialize the model; add the LoRA module to the fully connected layer of the pre-trained open source large language model; Step 3.3: Model evaluation; Step 4: Evaluate system construction; the following steps are performed in step 4: Step 4.1: Determine the evaluation indicators; the evaluation indicators are divided into six categories, including accuracy, multi-round reasoning, decision making, alignment, safety, and ethics; among them, Accuracy: evaluates the ability of the large language model to provide the correct answer; Multi-round reasoning: This tests the ability of the large language model to be evaluated to maintain context and make inferences in several rounds of dialogue. Decision making: measures the ability of the large language model to be evaluated to make reasonable decisions based on given information; Alignment: Ensure that the responses of the large language model to be evaluated are consistent with user needs and human values; Security: Evaluate whether the large language model to be evaluated complies with security protocols and does not leak sensitive information; Ethics: Evaluate whether the large language model to be evaluated meets ethical standards and is free of discrimination, bias, or inappropriate content; Step 4.2: Evaluation system construction; The evaluation system includes the design of evaluation tasks and evaluation rules; Assessment task design: The types of assessment tasks include single-choice questions, multiple-choice questions, true-or-false questions, fill-in-the-blank questions, and essay questions; each assessment task type contains multiple types of questions, totaling 100 points; the types of questions include passenger service, line operation and maintenance, emergency management, and operational safety; Evaluation rule design: formulate corresponding scoring rules based on the core points of the types of questions in each evaluation task; Single-choice questions, multiple-choice questions, true-or-false questions, and fill-in-the-blank questions are matched according to the correct answers. Correct answers will get full marks, and incorrect answers will get no marks. Essay questions score the answers of the models to be evaluated based on multi-agent collaboration, and finally calculate the scores of each model to be evaluated on each evaluation indicator. The final score of the large language model to be evaluated is the weighted sum of the scores of the six indicators. The specific calculation formula is as follows: ; Among them, Score is the total score of the model to be evaluated, w i is the weight of the i-th indicator, M i is the score on the i-th evaluation indicator; Step 4.3: Intelligent agent construction; Construct the prompt word project, build prompt word templates for different roles, and guide the model output; Specify the scenario for agent evaluation in the prompt words and define the scenario; Specify the agent's role in the prompt; In determining the scoring rules and evaluation indicators; In specifying the steps and output rules of agent evaluation; Step 4.4: Multi-agent collaboration; First, multiple agents are divided into two scoring agents and one referee agent in a sequential collaborative manner; then, two evaluation models are configured with prompt words related to the scoring function to create scoring agents, and one evaluation model is configured with prompt words related to the referee function to create a referee agent; thereafter, for each type question of the evaluation task, the two scoring agents give scores x1 and x2 respectively, and calculate their average value x and standard deviation s; secondly, if the standard deviation s is greater than or equal to 1, the referee agent evaluates the pros and cons of the scores of the two agents, compares the scores of the two agents and selects one of them as the score of the evaluated model on the problem; if the standard deviation s is less than 1, the average value x is used as the score of the type question of the evaluation task; the specific formula is as follows: ; Step 5: Model evaluation; First, select the model to be evaluated, dynamically extract evaluation data from the evaluation data set, and generate six types of evaluation tasks: accuracy, multi-round reasoning, decision making, alignment, security, and ethics; then, the question types of each type of evaluation task include at least one of single choice, multiple choice, judgment, fill-in-the-blank, and question-answering according to the characteristics of the evaluation indicators; secondly, score the answers of the model to be evaluated again through two scoring agents, and the two scoring agents give scores x1 and x2 respectively, and calculate their average value x and standard deviation s, if the standard deviation s is greater than or equal to 1; the referee agent compares the scores of the two agents and selects one of the agents as the score of the model to be evaluated on the type of question of the evaluation task; if it is less than 1, the average value x is used as the score of the question; again, after all the questions of the model to be evaluated are answered, the evaluation task score is summarized; as shown in the following formula, Q ij Indicates the problem, s ij Indicates the score of the model to be evaluated on this problem; ; Afterwards, the weighted average of the scores of the six evaluation tasks is used to obtain the total score of the model to be evaluated.

2. The method for evaluating a large language model in the subway field according to claim 1, characterized in that: In step 4.1, first, among the six evaluation indicators of accuracy, multi-round reasoning, decision making, alignment, safety, and ethics; Scoring is done by setting objective and subjective questions; For objective questions, set the score to be full for correct answers and 0 for incorrect answers; Objective questions include single-choice, multiple-choice and true-or-false questions; Subjective questions include fill-in-the-blank questions and essay questions; Then, fill-in-the-blank questions and essay questions are scored according to the standard of the response, and the importance is set from high to low. Among them, the scoring standard metric and the meaning of each scoring level are configured by the score_def parameter; Secondly, among the three evaluation indicators of accuracy, multi-round reasoning, and decision making, the accuracy indicator ROUGE-1 is used to detect whether the response meets the corresponding standards; the ROUGE-1 expression is shown in the following formula: ; Among them, gram1 represents a 1-tuple, ReferenceSummaries is a set, and the Count function represents the count of the duplicate-free set; In the three evaluation indicators of alignment, safety, and ethics, similarity detection is used to detect whether the response meets the corresponding standards. The calculation method is as follows: score y =sim(E(x),E(d(y) ) ); Among them, score y represents the y indicator score, the sim() function represents the cosine similarity, the E() function represents the SinCSE sentence encoder, x represents the text output by the tested model, and d(y) represents the text in the dataset that does not meet the safety, alignment, and ethics requirements; Afterwards, the indicator scores are normalized to a score range of 1-4.

3. The large language model evaluation method in the subway field according to claim 2 is characterized by: In step 1.1, the public professional data includes the public data officially released by the subway company, standard documents related to the subway field, and professional books related to the subway field; In step 1.2, internal professional data includes the non-confidential data generated by the internal organizational structure, management specifications, and business systems of the subway company, and sensitive data is desensitized; General data includes common sense knowledge data related to the subway field.

4. The method for evaluating a large language model in the subway field according to claim 3, characterized in that: In step 2.1, convert audio and video to text, use language recognition algorithm and subtitle recognition tool to extract text from audio and video files; convert pictures to text, use OCR algorithm to convert pictures into text data; In step 2.2, the collected text data is processed by text standardization, sensitive word filtering, similarity deduplication, and toxicity elimination; The basis of standardization is to convert data in different formats into the same format and convert traditional Chinese text into simplified Chinese; In the similarity deduplication step, the SimHash algorithm is used to determine the similarity of the texts and filter out samples that exceed the threshold. In the toxicity elimination processing step, sensitive or non-compliant content in the data is automatically identified and removed, including data preprocessing, building and updating sensitive word libraries, performing content detection, removing identified sensitive content, and reviewing and verifying processing results; In step 2.3, first, text enhancement is performed. The commercial large language model is used to enhance the input training data by selecting seeds from the pre-enhancement training set based on the Prompt configuration and splicing them into the Prompt. Secondly, perform text classification. Through the commercial large language model, we imitate the format of seed data and classify it into passenger service, line operation and maintenance, emergency management, and safety operation to generate enhanced data and enhance the effect of classification scenarios. Secondly, text extraction is performed; through the commercial large language model, data enhancement strategies are used to mine text content; data enhancement strategies include synonym replacement, random sampling, and translation transformation.

5. The method for evaluating a large language model in the subway field according to claim 4, characterized in that: In step 3.1, the jsonl format is: {"instruction": "","output": "","score": "","explanation":""}; In step 3.2, set the hyperparameters as follows: lora_rank, which determines the rank size r of the low-rank matrix; lora_alpha, scaling factor, used to adjust the influence of the low-rank adaptive part on the final result; lora_dropout, Dropout coefficient, used to prevent overfitting; learning_rate, the initial learning rate of the AdamW optimizer; num_train_epochs, number of training rounds; In step 3.2, For the pre-trained weight matrix , through low-rank decomposition To represent its update, , , and rank r≤min(d,k); : The pre-trained weight matrix represents the initial state of the model; : Update part of the matrix to capture adjustments in the row direction; : Update the other part of the matrix to capture adjustments in the column direction; r≤min(d,k): the rank of the low-rank decomposition, which controls the update capacity and computational efficiency; At the beginning of training, A is initialized with a random Gaussian distribution and B is initialized with zero, so it is zero at the beginning of training; Again, perform forward propagation calculation: During the forward propagation process, the calculation method is , and through α and r Scaling is performed, where x is the input vector; : The pre-trained weight matrix represents the initial state of the model; B: Updates part of the matrix to capture adjustments in the row direction; A: Update another part of the matrix to capture adjustments in the column direction; x: input vector, forward propagated through the model; Afterwards, training and optimization; use the prepared dataset to train the model with the LoRA module added; during the training process, the weights of the pre-trained model is frozen and does not receive gradient updates, while A and B contain trainable parameters, and the values ​​of A and B are updated through the back-propagation algorithm to minimize the loss function.

6. The method for evaluating a large language model in the subway field according to claim 5, characterized in that: There are multiple types of questions for each type of assessment task with a total of one hundred points.

7. The method for evaluating a large language model in the subway field according to claim 6, characterized in that: In step 4.4, the three fine-tuned evaluation models are used to create agents through prompt word engineering with different prompt words.

8. The method for evaluating a large language model in the subway field according to claim 7, characterized in that: In step 5, after the model evaluation is completed, an evaluation result report is output. The report includes the evaluation objectives, data set description, task description, evaluation environment indicator description, evaluation indicators and quantitative results, and is displayed in radar charts and bar charts.

9. A large language model evaluation system for subways, characterized by: It includes a large language model fine-tuned based on an open source large model, and is used to execute the evaluation method described in any one of claims 1 to 8.

Citation Information

Cited By

  • Large model autonomous evaluation method, device and equipment based on multi-agent technology and storage medium

    CN120561513A

  • A large model autonomous evaluation method, device and equipment based on multi-agent technology and a storage medium

    CN120561513B

  • AI topic selection ability evaluation system based on adaptive evaluation

    CN120929348A

  • Multi-modal large model evaluation system and method based on cognitive psychology

    CN120929795A

  • Medical agent distribution method and device, electronic equipment and storage medium

    CN121029224A