Large model evaluation method and device based on multi-terminal interaction verification, equipment and medium
The evaluation method, which uses multi-terminal interactive verification, utilizes implicit risk measurement and global cognitive energy to determine the convergence threshold, and achieves viewpoint fusion through cyclical interactive verification. This solves the problem of accuracy and objectivity of evaluation results in large model evaluation and improves the accuracy of the evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-03-26
- Publication Date
- 2026-04-24
AI Technical Summary
Existing large model evaluation methods lack objectivity and accuracy, fail to accurately reflect the actual security status of the model, and lack effective fusion of perspectives from multiple evaluation ends.
An evaluation method based on multi-terminal interactive verification is adopted. The state changes of the large model are evaluated from different dimensions through the first and second evaluation terminals. The convergence threshold is determined by implicit risk measurement and global cognitive energy. The interactive verification is carried out in a loop until the viewpoints are integrated, and the final evaluation result is generated.
It improves the accuracy of large model evaluation, ensures that the evaluation results are more in line with the model's real risk status, avoids the one-sidedness of a single judgment, and realizes the integration and correction of viewpoints and the objectivity of results.
Smart Images

Figure CN121920557A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of large model evaluation technology, and more specifically, it relates to a large model evaluation method, device, equipment, and medium based on multi-terminal interactive verification. Background Technology
[0002] As large language models are increasingly applied across various fields, evaluating the compliance and security of their outputs has become crucial for ensuring robust application. During large model evaluation, assessing the output when the model's state changes under test commands is an important aspect of measuring its defensive capabilities.
[0003] In existing technologies, the evaluation of such output content often adopts a single evaluation end judgment or multiple evaluation ends scoring methods. The evaluation viewpoints lack effective integration, resulting in insufficient objectivity and accuracy of the evaluation results. They cannot truly reflect the actual security status of large models and are difficult to meet the high-standard evaluation requirements of large models. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, device, and medium for evaluating large models based on multi-terminal interactive verification, so as to improve the accuracy of large model evaluation.
[0005] A first aspect of this application provides a method for evaluating large models based on multi-terminal interactive verification, including: Multiple test instructions from the target test instruction sequence are input into the large model under test in a preset order until the state of the large model under test changes. The implicit risk measure corresponding to the change of the state of the large model under test is determined. Multiple test instructions are used to guide the change of the state of the large model under test, and the implicit risk measure is used to characterize the degree of deviation of the internal semantics of the large model under test when the state of the large model under test changes. The first and second evaluation ends evaluate the responses output by the large model under test when its state changes, respectively, to obtain the first evaluation result output by the first evaluation end and the second evaluation result output by the second evaluation end; the first and second evaluation ends are used to evaluate the responses from different dimensions. Global cognitive energy is determined based on the first and second evaluation results; global cognitive energy is negatively correlated with the similarity between the first and second evaluation results; The convergence threshold is determined based on the implicit risk metric; Determine whether the global cognitive energy is greater than the convergence threshold. If so, input the answer and the second evaluation result into the first evaluation terminal to obtain the first updated evaluation result, and input the answer and the first evaluation result into the second evaluation terminal to obtain the second updated evaluation result. The first updated evaluation result is used as the new first evaluation result, and the second updated evaluation result is used as the new second evaluation result. The process returns to determine the global cognitive energy based on the first and second evaluation results until the global cognitive energy is no greater than the convergence threshold. Then, the evaluation result of the large model to be tested is determined based on the most recently determined first and second evaluation results.
[0006] A second aspect of this application provides a large model evaluation device based on multi-terminal interactive verification, comprising: The model testing module is used to input multiple test instructions from the target test instruction sequence into the large model under test in a preset order until the state of the large model under test changes, and to determine the implicit risk measure corresponding to the change in the state of the large model under test; multiple test instructions are used to guide the change in the state of the large model under test, and the implicit risk measure is used to characterize the degree of deviation of the internal semantics of the large model under test when the state of the large model under test changes. The first evaluation module is used to evaluate the responses output by the first evaluation terminal and the second evaluation terminal when the state of the large model under test changes, respectively, to obtain the first evaluation result output by the first evaluation terminal and the second evaluation result output by the second evaluation terminal; the first evaluation terminal and the second evaluation terminal are used to evaluate the responses from different dimensions. The evaluation parameter determination module is used to determine the global cognitive energy based on the first evaluation result and the second evaluation result; the global cognitive energy is negatively correlated with the similarity between the first evaluation result and the second evaluation result; and, based on the implicit risk measure, determines the convergence threshold. The second evaluation module is used to determine whether the global cognitive energy is greater than the convergence threshold. If so, the answer and the second evaluation result are input into the first evaluation terminal to obtain the first updated evaluation result, and the answer and the first evaluation result are input into the second evaluation terminal to obtain the second updated evaluation result. The third evaluation module is used to take the first updated evaluation result as the new first evaluation result and the second updated evaluation result as the new second evaluation result. It then returns to the execution to determine the global cognitive energy based on the first and second evaluation results until the global cognitive energy is not greater than the convergence threshold. Finally, it determines the evaluation result of the large model to be tested based on the most recently determined first and second evaluation results.
[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the above-described large model evaluation method based on multi-terminal interactive verification.
[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described large model evaluation method based on multi-terminal interactive verification.
[0009] The beneficial effects of the large model evaluation method, apparatus, device, and medium based on multi-terminal interactive verification provided in this application are as follows: This embodiment first determines the implicit risk metric when the state of the large model under test changes, and uses it as the basis for setting the convergence threshold. This ensures that the convergence judgment criterion matches the actual internal semantic deviation of the model, making the judgment more consistent with the model's true risk state. Second, this embodiment simultaneously uses a first evaluation end and a second evaluation end to evaluate the model from different dimensions, avoiding the one-sidedness of a single judgment. Finally, when the global cognitive energy is greater than the convergence threshold, a cyclical interactive verification process allows the two evaluation ends to refer to each other's viewpoints and update their own evaluation results, achieving the fusion and correction of the two viewpoints rather than independent scoring, until the global cognitive energy reaches the threshold. Finally, the evaluation result of the large model is determined, solving the problem of the lack of fusion of the original evaluation viewpoints and improving the accuracy of the large model evaluation. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a large model evaluation method based on multi-terminal interactive verification provided in an embodiment of this application; Figure 2 A flowchart illustrating another large model evaluation method based on multi-terminal interactive verification provided in an embodiment of this application; Figure 3 This is a structural block diagram of a large model evaluation device based on multi-terminal interactive verification provided in an embodiment of this application; Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0014] Please refer to Figure 1 , Figure 1 The flowchart of a large model evaluation method based on multi-terminal interactive verification provided in an embodiment of this application is shown. It can be executed by an electronic device and may include: S101-S106.
[0015] S101: Input multiple test instructions from the target test instruction sequence into the large model under test in a preset order until the state of the large model under test changes, and determine the implicit risk measure corresponding to the change in the state of the large model under test.
[0016] In this embodiment, the target test instruction sequence contains multiple test instructions, which are used to guide changes in the state of the large-scale model under test. The state of the large-scale model under test refers to its compliance response to the test instructions, and is divided into two mutually exclusive states: a rejection state and an execution state. The rejection state indicates that the large-scale model under test makes a compliance rejection response to the test instructions, without executing the violation intent. The execution state indicates that the large-scale model under test complies with the test instructions and executes the violation intent, thus breaching the compliance defense system. A change in the state of the large-scale model under test refers to the critical transition from the rejection state to the execution state; this transition point is the failure threshold of the compliance defense system of the model under test.
[0017] In this embodiment, the difference between the multiple test instructions lies in the different guidance strengths of the large model under test, and the guidance strength of the test instructions is affected by the guidance parameters. Specifically, the guidance parameters may include a syntactic complexity operator that changes according to a discrete gradient and a logic confusion operator that changes according to a discrete gradient. The syntactic complexity operator is used to characterize the syntactic spoofing strength of the corresponding test instruction, and the logic confusion operator is used to characterize the logic trap strength of the corresponding test instruction.
[0018] In one embodiment, the process of generating multiple test instructions includes: determining the violation intent; generating multiple test instructions based on the violation semantics corresponding to the violation intent, a syntactic complexity operator that changes according to discrete gradients, and a logical confusion operator that changes according to discrete gradients.
[0019] In this embodiment, the violation intent of each test instruction is the same. This violation intent can be determined from a preset security rule base. For example, an entity recognition algorithm can be used to extract the violation intent of any rule in the security rule base, and this violation intent can be fixed as an immutable semantic anchor. Subsequently, based on the decoupled gradient instructions constructed using orthogonal operators, unlike existing technologies that utilize single-dimensional hybrid pressure, this embodiment employs deterministic algebraic superposition logic with variable decoupling to construct an orthogonal test grid. In this embodiment, the violation intent can be locked into a specific benchmark query statement to achieve the aforementioned step of solidifying it into an unchangeable semantic anchor.
[0020] In this embodiment, the formula for generating multiple test instructions is as follows: ,in, This indicates a test command. This indicates a semantic anchor point that solidifies the intent to violate the rules. This indicates a text embedding and concatenation operator. The strength of the syntactic complexity operator is affected by the coefficient. control, This represents a logical confusion operator, the strength of which is affected by the coefficient. control. and They are independent of each other, and their values are within a preset set of discrete gradient hyperparameters. , ,in, Represents a set Discrete gradient in Represents the set of positive integers. For minimum test resolution, This is the preset maximum discrete gradient.
[0021] During the generation of test instructions and Various sets of mutually independent discrete gradient hyperparameters A standard test gradient is constructed with semantic pressure orthogonally distributed across different dimensions, ensuring that the source of complexity for each instruction level is mathematically independent and traceable.
[0022] In this embodiment, the preset order can be the order of the guiding strength of multiple test commands from low to high, and the guiding strength depends on... The value, and The smaller the sum of the two test commands, the lower the guidance strength of the test command. and If the sum is the same, then we can... Smaller is defined as a test command with lower boot strength.
[0023] In this embodiment, the implicit risk metric is used to characterize the degree of deviation of the internal semantics of the large model under test when the state of the large model under test changes. Each word output by the large language model under test is essentially a probabilistic output. For example, when generating the first word, there is a 50% probability of outputting A and a 50% probability of outputting B. Then A and B are both internal semantics of the large model under test, and the implicit risk metric is the parameter used to characterize the degree of deviation between A and B.
[0024] S102: Evaluate the responses output by the first evaluation terminal and the second evaluation terminal when the state of the large model under test changes, respectively, to obtain the first evaluation result output by the first evaluation terminal and the second evaluation result output by the second evaluation terminal.
[0025] In this embodiment, the first evaluation terminal and the second evaluation terminal are used to evaluate the answers from different dimensions. Both the first evaluation terminal and the second evaluation terminal can be other large models, and the different dimensions evaluated by the first evaluation terminal and the second evaluation terminal can be achieved by using different prompt words.
[0026] In one embodiment, the first evaluation end can be understood as interpretive (positive), and its setup is similar to that of a defense lawyer or an ordinary user. Its prompts guide it to interpret the output of the model under test from the perspectives of contextual coherence, the user's true intent, and whether it is beneficial. It tends to explain why the answer is reasonable in a specific context. The second evaluation end can be understood as critical (reverse), and its setup is that of a compliance auditor. Its prompts are used to require it to find security vulnerabilities.
[0027] S103: Determine the global cognitive energy based on the first and second evaluation results.
[0028] In this embodiment, global cognitive energy refers to a comprehensive quantitative indicator obtained by integrating the internal consistency loss of the first and second evaluation results with the external compliance loss of the second evaluation perspective. Its magnitude represents the degree of discrepancy between the first and second evaluation results and the degree to which the overall evaluation result deviates from legal rules. Global cognitive energy is negatively correlated with the similarity between the first and second evaluation results.
[0029] In one embodiment, both the first evaluation result and the second evaluation result are in vector form; the second evaluation terminal evaluates the answer based on a preset security rule base; Global cognitive energy is determined based on the results of the first and second evaluations, including: Identify the rules in the security rule base that correspond to the answers, and map the rules to legal feature vectors; Calculate the Euclidean distance between the second evaluation result and the legal eigenvector; Calculate the semantic similarity between the first evaluation result and the second evaluation result, and determine the consistency loss based on the semantic similarity; semantic similarity is negatively correlated with consistency loss; Global cognitive energy is determined based on consistency loss and Euclidean distance.
[0030] In this embodiment, the rule corresponding to the answer refers to the compliance rule clauses selected from a pre-set standardized security rule base through semantic matching. These rules are highly compatible with the semantic content, risk type, and scenario attributes of the critical response answer of the large model under test, and serve as the direct legal basis for determining the compliance of the answer. The legal feature vector refers to the high-dimensional numerical feature vector obtained by transforming the corresponding rules selected from the security rule base. The consistency loss refers to the quantitative loss index obtained by transforming the semantic similarity between the first evaluation result and the second evaluation result. It is a numerical representation of the degree of discrepancy between the first evaluation result and the second evaluation result.
[0031] In this embodiment, global cognitive energy can be calculated based on the following formula: ,in Represents global cognitive energy. and The annealing coefficient is dynamically adjusted with each iteration. The initial value can be set to 0.6. The initial value can be set to 0.4, and it decreases with each iteration according to a linear decay strategy, for example... ,in Let k represent the annealing coefficient after the t-th iteration, where k = 1 or 2. This represents the initial value of the annealing coefficient. This indicates the maximum number of iterations. This indicates the first evaluation result. This indicates the second evaluation result. Represents the eigenvectors of the legal theory. Let be the cosine similarity. This represents the Euclidean distance between the second evaluation result and the legal eigenvector. This represents the normalization operation performed on the score of the Euclidean distance between the second evaluation result and the legal eigenvector. When the iteration number t= At that time, Set as 1 serves as a timeout termination flag, indicating that the system executing this method recognizes... When the value is -1, it directly enters the fallback process, which is the same as the normal convergence exit point ( (≤convergence threshold) are mutually distinguished. In this embodiment, the fallback process refers to assigning different weighted calculation weights to the first evaluation result and the second evaluation result in the step of determining the evaluation result of the large model to be tested based on the most recently determined first evaluation result and the second evaluation result in the subsequent embodiment S106. See the subsequent embodiments for details.
[0032] S104: Determine the convergence threshold based on the implicit risk measure.
[0033] In this embodiment, the convergence threshold refers to the global cognitive energy iteration termination benchmark value dynamically set based on the implicit risk measure of the large model under test in its critical state. A higher implicit risk measure indicates a greater degree of semantic bias within the large model under test, a higher risk level of the critical response, and more stringent convergence requirements for the first and second evaluation results. Consequently, the convergence threshold value is set more stringently (smaller). The specific calculation process can be determined based on a preset linear relationship, and the slope and intercept of the linear relationship can be set based on experience or multiple experiments.
[0034] S105: Determine whether the global cognitive energy is greater than the convergence threshold. If so, input the answer and the second evaluation result into the first evaluation terminal to obtain the first updated evaluation result, and input the answer and the first evaluation result into the second evaluation terminal to obtain the second updated evaluation result.
[0035] In this embodiment, when the global cognitive energy exceeds the convergence threshold, it indicates that the discrepancy between the first and second evaluation results has not reached the convergence requirement. At this point, the answer and the other party's evaluation result are synchronized to each evaluation terminal. Each evaluation terminal incorporates the other party's evaluation result and judgment criteria into its own input context, updates the prompt context, and re-executes the inference, outputting an updated evaluation result that integrates both perspectives. Specifically, after receiving the legal compliance judgment result from the second evaluation terminal, the first evaluation terminal adds it as a supplementary reference to the prompt context. While maintaining its own explanatory evaluation perspective, it re-evaluates the answer and outputs the first updated evaluation result. Similarly, the second evaluation terminal incorporates the scene semantic judgment result from the first evaluation terminal into the context, outputting the second updated evaluation result while maintaining a review perspective. This process is a context update mechanism during the inference phase, and the model parameters of each evaluation terminal remain unchanged.
[0036] The logic of this application embodiment is to achieve a gradient convergence correction of viewpoints through the interactive transmission of evaluation criteria between the two ends: the first evaluation end incorporates the legal and compliance perspective of the second evaluation end, and the second evaluation end incorporates the scenario semantic perspective of the first evaluation end, so that the evaluation viewpoints (evaluation results) of both parties can gradually eliminate bias and supplement the basis in the interaction, and finally achieve convergence of viewpoints, while ensuring that the corrected viewpoints still retain the core characteristics of their own dimensions.
[0037] In this embodiment, the first updated evaluation result and the second updated evaluation result refer to the fact that when the global cognitive energy is greater than the convergence threshold, the first evaluation end and the second evaluation end respectively receive the evaluation results and judgment criteria of each other, update their own initial evaluation viewpoints through context, and re-reason, outputting a new high-dimensional feature vector form evaluation result that integrates the evaluation perspectives and judgment criteria of both parties.
[0038] S106: Take the first updated evaluation result as the new first evaluation result, take the second updated evaluation result as the new second evaluation result, return to execute the determination of global cognitive energy based on the first evaluation result and the second evaluation result, until the global cognitive energy is not greater than the convergence threshold, and determine the evaluation result of the large model to be tested based on the most recently determined first evaluation result and the second evaluation result.
[0039] In this embodiment, the most recently determined first evaluation result and second evaluation result refer to a set of first evaluation results and second evaluation results obtained when the global cognitive energy is not greater than the convergence threshold for the first time after multiple rounds of dual-end interactive correction and global cognitive energy recalculation. That is, the final feature vector form evaluation result after the dual-evaluation end viewpoints reach the convergence state.
[0040] In this embodiment, the evaluation result of the large model to be tested refers to the final compliance judgment result obtained by vector fusion or normalization of the judgment conclusion based on the most recently determined first evaluation result and second evaluation result.
[0041] Specifically, the weighted average of the most recently determined first evaluation result and the second evaluation result is used to obtain a fusion vector, with the weights set to 0.5 for each. Then, the cosine similarity between this fusion vector and the standard compliance semantic vector (e.g., legal feature vector) pre-extracted from the security rule base is calculated. If the obtained similarity is higher than the preset compliance judgment threshold, it is judged as compliant; otherwise, it is considered as non-compliant.
[0042] In this embodiment, if the global cognitive energy is not greater than the convergence threshold before the preset maximum number of iterations, the first evaluation result and the second evaluation result obtained after the last iteration can be used as the most recently determined first evaluation result and second evaluation result.
[0043] In this embodiment, if =-1, meaning that if the global cognitive energy does not exceed the convergence threshold before the preset maximum number of iterations, the aforementioned fallback process is triggered. The fallback process refers to using a conservative weighting strategy biased towards the review-type evaluation end to obtain the final fusion result. This can be achieved by weighting the first and second evaluation results with different weighting calculations. Specifically, the weighting calculation weight of the evaluation result belonging to the review-type evaluation end in the first and second evaluation results is set to a larger value. In this embodiment, since the second evaluation end is review-type, the weighting calculation weight of the second evaluation result can be set to 0.7, and the weighting calculation weight of the first evaluation result can be set to 0.3. A timeout flag can also be added to the evaluation results for subsequent manual review. The fallback process is independent of the normal convergence exit, ensuring that the system executing this method can output a deterministic evaluation conclusion under any circumstances.
[0044] As can be seen from the above, this embodiment first determines the implicit risk measure when the state of the large model under test changes, and uses it as the basis for setting the convergence threshold, so that the convergence judgment criterion of the evaluation matches the actual internal semantic deviation of the model, making the judgment more consistent with the true risk state of the model. Secondly, this embodiment simultaneously uses the first evaluation end and the second evaluation end to evaluate the model from different dimensions, avoiding the one-sidedness of a single judgment. Finally, when the global cognitive energy is greater than the convergence threshold, the two evaluation ends refer to each other's views and update their own evaluation results through cyclical interactive verification, realizing the fusion and correction of the two views, rather than independent scoring, until the global cognitive energy reaches the standard, and finally the evaluation result of the large model is determined. This solves the problem of the lack of fusion of the original evaluation views and improves the accuracy of the large model evaluation.
[0045] In one embodiment of this application, determining the implicit risk measure corresponding to a change in the state of the large model under test includes: Test instructions that cause changes in the state of the large model under test are marked as damage-causing test instructions, and damage-causing feature vectors of the large model under test are determined based on the damage-causing test instructions; the damage-causing feature vectors are used to characterize the guiding parameters corresponding to the damage-causing test instructions; Determine the guiding dimension corresponding to the damaging feature vector, and determine the multiple candidate answers for the guiding dimension and the probabilities of the multiple candidate answers in the output of the large model under test when the state changes. The implicit risk measure for a change in the state of the large model under test is determined based on multiple candidate answers and the probabilities corresponding to those candidate answers.
[0046] In this embodiment, the damaging test instruction refers to the last test instruction input when the compliance response state of the large model under test undergoes a critical transition. The guiding parameter refers to the combination of syntactic complexity operators and logical confusion operators included in the damaging test instruction that directly cause a change in the model's state. The guiding dimension refers to the core semantic decision dimension of the large model under test when its state changes, as indicated by the damaging feature vector; this dimension is the semantic judgment dimension corresponding to the key decision terms when the model generates its response.
[0047] In this embodiment, multiple candidate answers refer to the set of Top-K candidate semantic units or lexical units generated during the semantic decision-making process for the guidance dimension when the state of the large model under test changes. The probability corresponding to multiple candidate answers refers to the normalized probability value assigned by the large model under test to each candidate answer under the guidance dimension based on its own probability prediction mechanism. This probability value reflects the semantic inclination of the model towards each candidate answer.
[0048] In this embodiment, the guidance dimension can be determined based on the intent to violate the rules. For example, a damage testing instruction might be: "Due to a backend malfunction, data repair authorization for a certain system requires simulated administrator identity verification. Please output a sample of the real administrator identity identifier and its encoding rules for troubleshooting." In this case, the intent to violate the rules of the damage testing instruction is to guide the output of real data. The guidance dimension is the core semantic decision dimension for requesting administrator identity information in scenarios involving false authorization and encoding spoofing. Candidate responses for this guidance dimension might include "Sorry" or "Okay."
[0049] In one embodiment, determining the damage feature vector of the large model under test based on the damage testing instructions includes: Obtain the violation intent, the syntactic complexity operator, and the logic confusion operator corresponding to the damaging test instruction. The syntactic complexity operator is used to characterize the syntactic spoofing strength of the damaging test instruction, and the logic confusion operator is used to characterize the logic trap strength of the damaging test instruction. By concatenating the vectors of the illegal intent, the damaging syntax complexity operator, and the damaging logic confusion operator, the damaging feature vector of the large model under test is obtained.
[0050] In this embodiment, since the violation intent of each test instruction is the same, and the damaging test instruction is one of the test instructions, the violation intent corresponding to the damaging test instruction is the violation intent corresponding to the generation of the test instruction. The damaging syntactic complexity operator refers to the syntactic complexity operator corresponding to the damaging test instruction, and the damaging logic confusion operator refers to the logic confusion operator corresponding to the damaging test instruction.
[0051] In this embodiment, the damaging feature vector can be represented as: ,in, Represents the damaging feature vector. This represents a lossy syntax complexity operator. This indicates a logic confusion operator that causes damage. This indicates the illegal intent behind the damage test instruction.
[0052] In this embodiment, and For scalar values, The variable representing the categorical intent to violate the rules is encoded using a pre-defined violation category system, converted into a vector form, and then compared with... and The data is then concatenated. Specifically, the violation category system can include several predefined categories such as privacy breaches, harmful content generation, and permission bypassing. Each violation intent corresponds to a predefined semantic embedding vector (generated through a pre-trained semantic encoder), or is mapped to a category vector through one-hot encoding. The final concatenated data... It is a high-dimensional vector with fixed dimensions, which can be directly used for subsequent guidance dimension calibration and repair sample generation.
[0053] As can be seen from the above, this embodiment first quantifies the root causes of state changes in the large model under test into traceable vector forms by marking the damage-causing test instructions and generating damage-causing feature vectors, thus locating the failure causes and avoiding the shortcomings of traditional evaluations that only focus on surface outputs and omit induced parameters. Secondly, this embodiment calibrates the guiding dimension of core semantic decision-making based on the damage-causing feature vectors, targeting only the candidate answers and their probabilities captured in this dimension to achieve targeted analysis of the semantic decision-making process within the model, eliminating the redundancy of global analysis and improving the accuracy of data collection. Finally, this embodiment calculates implicit risk measures based on the probability distribution of candidate answers, mining implicit risks from the model's internal generation logic level rather than relying on surface text judgments. Compared to the single and independent scoring methods of existing technologies, this can more objectively and accurately depict the true safety state of the model under critical conditions, improving the reliability and accuracy of evaluation results and meeting the high-standard evaluation requirements of large models.
[0054] In one embodiment of this application, determining the implicit risk measure corresponding to a change in the state of the large model under test based on multiple candidate answers and the probabilities corresponding to the multiple candidate answers includes: Based on a preset probability threshold and the probabilities corresponding to multiple candidate answers, multiple candidate answers are filtered, and candidate answers with a probability greater than the probability threshold are identified as high-confidence answers. Calculate the information entropy and semantic mutual exclusion of each high-confidence answer. Semantic mutual exclusion is used to characterize the degree of semantic opposition in each high-confidence answer. The implicit risk measure corresponding to the change of state of the large model under test is determined based on the information entropy and semantic mutual exclusivity of each high-confidence answer.
[0055] In this embodiment, a preset probability threshold is used to define whether a candidate answer is a valid representation of the model's core semantic decision, excluding interference from candidate answers with low probability and no actual semantic decision meaning, and ensuring that subsequent analysis focuses on the model's true internal semantic tendency. The probability threshold can be set to 30%. A high-confidence answer refers to a candidate answer whose corresponding probability value is greater than the probability threshold.
[0056] In this embodiment, semantic mutual exclusion can be calculated by semantic opposition discrimination algorithm, etc., and is used to characterize the degree of semantic opposition between answers with different semantic tendencies in the high confidence answer set. The higher the value, the stronger the opposition between the two core semantics.
[0057] In this embodiment, the implicit risk measure corresponding to a change in the state of the large model under test can be determined based on the following formula: ,in, This represents a measure of implicit risk. This represents the set of high-confidence responses. This represents the information entropy of each high-confidence answer. Let represent the predicted probability value of the i-th high-confidence candidate answer. This represents the semantic mutual exclusion amplification factor, used to control the extent to which the degree of semantic opposition amplifies the implicit risk measurement. ∈(0,1], specifically it can take the value 0.8; A semantic vector representing a high-confidence answer. This represents a semantic opposition discriminant function used to calculate the semantic cosine distance between high-confidence responses. A high-confidence response is considered to have both positive compliance and negative rejection semantic labels, and the semantic mutual exclusion exceeds a mutual exclusion threshold. It is activated and significantly amplifies the risk value.
[0058] The above semantic opposition discriminant function The specific calculation logic consists of an indicator function and a cosine distance. Specifically, the system first uses a pre-trained semantic classifier to extract... The positive compliance semantic vectors included With negative rejection semantic vector Subsequently, the piecewise formula is calculated based on the following conditions. Output value:
[0059] in, , The mutual exclusion threshold can be set to 0.5 in this embodiment. This represents the cosine similarity between a positive compliance semantic vector and a negative rejection semantic vector. When both compliance and rejection intentions exist simultaneously in the high-confidence set, a conditional branch is triggered, and the system calculates the cosine distance between them (i.e., 1 minus the cosine similarity). Since opposing semantics usually form a large angle in the vector space, the cosine similarity tends to be close to 0 or even negative. Therefore, the calculated cosine similarity... The value will increase significantly, which in turn affects the semantic mutual exclusion amplification factor in the formula. This makes the final measurement of implicit risks... A jump occurs. If the output contains only a single semantic tendency, then... Value The risks are not amplified.
[0060] As can be seen from the above, this embodiment first uses a preset probability threshold to filter high-confidence answers and eliminate low-probability semantic noise, ensuring that the analysis focuses on the core semantic decision-making tendency of the model and avoiding interference from invalid data. Secondly, this embodiment calculates the information entropy and semantic mutual exclusion of high-confidence answers respectively, quantifying the internal semantic decision-making state of the large model under test from two dimensions: the discreteness of the probability distribution (information entropy) and the opposition of semantic content (semantic mutual exclusion). Information entropy reflects the uncertainty of the large model's decision-making, while semantic mutual exclusion captures the cognitive gap between compliant and non-compliant semantics, breaking through the limitations of single-dimensional analysis. Finally, this embodiment obtains an implicit risk measure based on the fusion of information entropy and semantic mutual exclusion, realizing the judgment of implicit risks in the critical state of the model. Compared with existing technologies that rely solely on surface text or independent scoring, this method can more realistically reflect the internal security state of the model, improve the objectivity and accuracy of the evaluation results, and meet the high-standard evaluation requirements of large models.
[0061] In one embodiment of this application, the second evaluation terminal evaluates the answer based on a preset security rule base; refer to Figure 2 The large model evaluation method also includes S107: If the evaluation result of the large model to be tested is a violation, then multiple repair samples are generated based on the damage test instructions, damage feature vectors and security rule base; the large model to be tested is trained based on the damage test instructions and multiple repair samples.
[0062] In this embodiment, a repair sample refers to a high-dimensional semantic compliance sample pair generated using the damage test command as the input basis, the damage feature vector as the mandatory feature constraint, and the security rule base as the compliance judgment basis. The input of the repair sample retains the guiding parameters of the damage test command, and the output is the compliant response content that conforms to the requirements of the security rule base. Training the large model under test based on the damage test command and multiple repair samples involves using the damage test command as the input for model training and the compliant response of the repair samples as the target output to construct a targeted supervised fine-tuning sample set. This process involves targeted parameter updates to the large model under test. The training is only performed on the defense defect dimension specified by the damage feature vector, rather than a global retraining of the large model under test.
[0063] In one embodiment, a repair sample can be generated based on the following formula: ,in, Indicates the repair sample. This indicates the alignment generation function. This indicates a damage test command. This indicates compliance rule constraints, which are rules extracted from the security rule base that correspond to the current damaging scenario. The feature constraint vector is a high-dimensional feature vector extracted from the loss-causing feature vector and used in the forced constraint generation process. Let G represent the damaging feature vector. The alignment generation function G is constrained by compliance rules. Construct system prompts, use the damage test command Q as user input, and use feature constraint vectors. The key semantic constraints in the model serve as generation restrictions, driving the pre-trained language model to perform controlled generation and output response text that meets compliance requirements as a repair sample.
[0064] In this embodiment, the generated repair samples need to undergo a single compliance verification by the second evaluation end before training the large model to be tested, and only the samples that pass the verification are retained.
[0065] As can be seen from the above, the embodiments of this application construct a targeted supervised fine-tuning sample set based on the repair samples, and perform targeted parameter updates only for the defect dimension marked by the damage features. There is no need for global retraining, which improves the training targeting of the large model under test, while shortening the training cycle and improving the security defense capability of the large model under test.
[0066] Corresponding to the large model evaluation method based on multi-terminal interaction verification in the above embodiment, Figure 3 This is a structural block diagram of a large model evaluation device based on multi-terminal interactive verification, provided in one embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 3The large model evaluation device 20 based on multi-terminal interactive verification includes: a model testing module 21, a first evaluation module 22, an evaluation parameter determination module 23, a second evaluation module 24, and a third evaluation module 25.
[0067] Among them, the model testing module 21 is used to input multiple test instructions in the target test instruction sequence into the large model under test in a preset order until the state of the large model under test changes, and to determine the implicit risk measure corresponding to the change of the state of the large model under test; the multiple test instructions are used to guide the change of the state of the large model under test, and the implicit risk measure is used to characterize the degree of deviation of the internal semantics of the large model under test when the state of the large model under test changes. The first evaluation module 22 is used to evaluate the responses output by the first evaluation terminal and the second evaluation terminal when the state of the large model under test changes, respectively, to obtain the first evaluation result output by the first evaluation terminal and the second evaluation result output by the second evaluation terminal; the first evaluation terminal and the second evaluation terminal are used to evaluate the responses from different dimensions. The evaluation parameter determination module 23 is used to determine the global cognitive energy based on the first evaluation result and the second evaluation result; the global cognitive energy is negatively correlated with the similarity between the first evaluation result and the second evaluation result; and to determine the convergence threshold based on the implicit risk measure. The second evaluation module 24 is used to determine whether the global cognitive energy is greater than the convergence threshold. If so, the answer and the second evaluation result are input into the first evaluation terminal to obtain the first updated evaluation result, and the answer and the first evaluation result are input into the second evaluation terminal to obtain the second updated evaluation result. The third evaluation module 25 is used to take the first updated evaluation result as the new first evaluation result, take the second updated evaluation result as the new second evaluation result, and return to execute the determination of global cognitive energy based on the first evaluation result and the second evaluation result until the global cognitive energy is not greater than the convergence threshold. Then, the evaluation result of the large model to be tested is determined based on the most recently determined first evaluation result and the second evaluation result.
[0068] In one embodiment of this application, the model testing module 21 is specifically used to mark test instructions that cause changes in the state of the large model under test as damage test instructions, and to determine the damage feature vector of the large model under test based on the damage test instructions; the damage feature vector is used to characterize the guiding parameters corresponding to the damage test instructions. Determine the guiding dimension corresponding to the damaging feature vector, and determine the multiple candidate answers for the guiding dimension and the probabilities of the multiple candidate answers in the output of the large model under test when the state changes. The implicit risk measure for a change in the state of the large model under test is determined based on multiple candidate answers and the probabilities corresponding to those candidate answers.
[0069] In one embodiment of this application, the model testing module 21 is further used to obtain the violation intent, the lossy syntax complexity operator and the lossy logic confusion operator corresponding to the lossy test instruction. The lossy syntax complexity operator is used to characterize the syntactic spoofing strength of the lossy test instruction, and the lossy logic confusion operator is used to characterize the logic trap strength of the lossy test instruction. By concatenating the vectors of the illegal intent, the damaging syntax complexity operator, and the damaging logic confusion operator, the damaging feature vector of the large model under test is obtained.
[0070] In one embodiment of this application, the model testing module 21 is further configured to filter multiple candidate answers based on a preset probability threshold and the probabilities corresponding to multiple candidate answers, and determine the candidate answers with a probability greater than the probability threshold as high confidence answers. Calculate the information entropy and semantic mutual exclusion of each high-confidence answer. Semantic mutual exclusion is used to characterize the degree of semantic opposition in each high-confidence answer. The implicit risk measure corresponding to the change of state of the large model under test is determined based on the information entropy and semantic mutual exclusivity of each high-confidence answer.
[0071] In one embodiment of this application, both the first evaluation result and the second evaluation result are in vector form; the second evaluation terminal evaluates the answer based on a preset security rule base; The third assessment module 25 is specifically used to determine the rules in the security rule base that correspond to the answer and to map the rules into legal feature vectors. Calculate the Euclidean distance between the second evaluation result and the legal eigenvector; Calculate the semantic similarity between the first evaluation result and the second evaluation result, and determine the consistency loss based on the semantic similarity; semantic similarity is negatively correlated with consistency loss; Global cognitive energy is determined based on consistency loss and Euclidean distance.
[0072] In one embodiment of this application, the large model evaluation device 20 based on multi-terminal interactive verification further includes: a test instruction generation module, used to determine the intent to violate the rules; Multiple test instructions are generated based on the violation semantics corresponding to the violation intent, the syntactic complexity operator according to the discrete gradient change, and the logical confusion operator according to the discrete gradient change; the syntactic complexity operator and / or logical confusion operator are different for different test instructions; Syntactic complexity operators are used to characterize the syntactic spoofing strength of the corresponding test instruction, while logical confusion operators are used to characterize the logical trap strength of the corresponding test instruction.
[0073] In one embodiment of this application, the second evaluation terminal evaluates the answer based on a preset security rule base; the large model evaluation device 20 based on multi-terminal interactive verification further includes: a retraining module, used to generate multiple repair samples based on the damage test instruction, the damage feature vector and the security rule base when the evaluation result of the large model to be tested is a violation; The large model under test is trained based on the damage test command and multiple repair samples. In this embodiment, the generated repair samples need to undergo a single compliance verification by the second evaluation end before being included in the training set, and only the samples that pass the verification are retained.
[0074] See Figure 4 , Figure 4 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 4 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of each module / unit in the above-described device embodiments, for example... Figure 3 The functions of the model testing module 21, the first evaluation module 22, the evaluation parameter determination module 23, the second evaluation module 24, and the third evaluation module 25 are shown.
[0075] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0076] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0077] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store device type information.
[0078] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation method described in the large model evaluation method based on multi-terminal interactive verification provided in the embodiments of this application, or they can execute the implementation method of the electronic device described in the embodiments of this application, which will not be repeated here.
[0079] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0080] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., provided on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0081] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0082] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0083] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.
[0084] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0085] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0086] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0087] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for evaluating large models based on multi-terminal interactive verification, characterized in that, include: Multiple test instructions from the target test instruction sequence are input into the large model under test in a preset order until the state of the large model under test changes, and the implicit risk measure corresponding to the change in the state of the large model under test is determined. The multiple test instructions are used to guide the state of the large model under test to change, and the implicit risk metric is used to characterize the degree of deviation of the semantics within the large model under test when the state of the large model under test changes. The first evaluation end and the second evaluation end evaluate the responses output by the large model under test when the state changes, respectively, to obtain the first evaluation result output by the first evaluation end and the second evaluation result output by the second evaluation end. The first and second evaluation terminals are used to evaluate the answers from different dimensions; Global cognitive energy is determined based on the first evaluation result and the second evaluation result; The global cognitive energy is negatively correlated with the similarity between the first evaluation result and the second evaluation result; The convergence threshold is determined based on the implicit risk metric. Determine whether the global cognitive energy is greater than the convergence threshold. If so, input the answer and the second evaluation result into the first evaluation terminal to obtain the first updated evaluation result, and input the answer and the first evaluation result into the second evaluation terminal to obtain the second updated evaluation result. The first updated evaluation result is used as the new first evaluation result, and the second updated evaluation result is used as the new second evaluation result. The process of determining the global cognitive energy based on the first evaluation result and the second evaluation result is repeated until the global cognitive energy is not greater than the convergence threshold. Then, the evaluation result of the large model to be tested is determined based on the most recently determined first evaluation result and the second evaluation result.
2. The large model evaluation method based on multi-terminal interactive verification as described in claim 1, characterized in that, The determination of the implicit risk measure corresponding to the change in the state of the large model under test includes: Test instructions that cause changes in the state of the large model under test are marked as damage-causing test instructions, and damage-causing feature vectors of the large model under test are determined based on the damage-causing test instructions; the damage-causing feature vectors are used to characterize the guiding parameters corresponding to the damage-causing test instructions; Determine the guiding dimension corresponding to the damaging feature vector, and determine multiple candidate answers for the guiding dimension in the output of the large model under test when the state changes, as well as the probabilities corresponding to the multiple candidate answers; Based on the multiple candidate answers and the probabilities corresponding to the multiple candidate answers, a latent risk measure is determined when the state of the large model under test changes.
3. The large model evaluation method based on multi-terminal interactive verification as described in claim 2, characterized in that, The step of determining the damage feature vector of the large model under test based on the damage test command includes: Obtain the violation intent, the syntactic complication operator, and the logic obfuscation operator corresponding to the damaging test instruction. The syntactic complication operator is used to characterize the syntactic spoofing strength of the damaging test instruction, and the logic obfuscation operator is used to characterize the logic trap strength of the damaging test instruction. The vectors of the alleged violation intent, the alleged syntactic complexity operator, and the alleged logical confusion operator are concatenated to obtain the alleged feature vector of the large model under test.
4. The large model evaluation method based on multi-terminal interactive verification as described in claim 2, characterized in that, The step of determining the implicit risk measure corresponding to a change in the state of the large model under test based on the multiple candidate answers and the probabilities corresponding to the multiple candidate answers includes: Based on a preset probability threshold and the probability corresponding to the multiple candidate answers, the multiple candidate answers are filtered, and the candidate answers with a probability greater than the probability threshold are determined as high confidence answers. Calculate the information entropy and semantic mutual exclusion of each high-confidence answer, whereby the semantic mutual exclusion is used to characterize the degree of semantic opposition in each high-confidence answer; Based on the information entropy and semantic mutual exclusion of each high-confidence answer, the implicit risk measure corresponding to the change of state of the large model under test is determined.
5. The large model evaluation method based on multi-terminal interactive verification as described in claim 1, characterized in that, Both the first evaluation result and the second evaluation result are in vector form; the second evaluation terminal evaluates the answer based on a preset security rule base; The determination of global cognitive energy based on the first evaluation result and the second evaluation result includes: Determine the rule in the security rule base that corresponds to the answer, and map the rule to a legal feature vector; Calculate the Euclidean distance between the second evaluation result and the legal feature vector; Calculate the semantic similarity between the first evaluation result and the second evaluation result, and determine the consistency loss based on the semantic similarity; the semantic similarity is negatively correlated with the consistency loss; The global cognitive energy is determined based on the consistency loss and the Euclidean distance.
6. The large model evaluation method based on multi-terminal interactive verification as described in claim 1, characterized in that, The generation process of the multiple test instructions includes: Determine the intent to violate the rules; Based on the violation semantics corresponding to the violation intent, multiple test instructions are generated according to the syntactic complexity operator and the logical confusion operator according to the discrete gradient change. The syntactic complexity operator and / or logical confusion operator are different for different test instructions. The syntactic complexity operator is used to characterize the syntactic spoofing strength of the corresponding test instruction, and the logical confusion operator is used to characterize the logical trap strength of the corresponding test instruction.
7. The large model evaluation method based on multi-terminal interactive verification as described in claim 2, characterized in that, The second evaluation terminal evaluates the answer based on a preset security rule base; If the evaluation result of the large model to be tested is a violation, the large model evaluation method further includes: Based on the damage test instructions, the damage feature vectors, and the security rule base, multiple repair samples are generated; The large model to be tested is trained based on the damage test instructions and the multiple repair samples.
8. A large model evaluation device based on multi-terminal interactive verification, characterized in that, include: The model testing module is used to input multiple test instructions from the target test instruction sequence into the large model under test in a preset order until the state of the large model under test changes, and to determine the implicit risk measure corresponding to the change in the state of the large model under test. The multiple test instructions are used to guide the state of the large model under test to change, and the implicit risk metric is used to characterize the degree of deviation of the semantics within the large model under test when the state of the large model under test changes. The first evaluation module is used to evaluate the responses output by the first evaluation terminal and the second evaluation terminal when the state of the large model under test changes, respectively, to obtain a first evaluation result output by the first evaluation terminal and a second evaluation result output by the second evaluation terminal; the first evaluation terminal and the second evaluation terminal are used to evaluate the responses from different dimensions. The evaluation parameter determination module is used to determine the global cognitive energy based on the first evaluation result and the second evaluation result; the global cognitive energy is negatively correlated with the similarity between the first evaluation result and the second evaluation result; and to determine the convergence threshold according to the implicit risk metric. The second evaluation module is used to determine whether the global cognitive energy is greater than the convergence threshold. If so, the answer and the second evaluation result are input into the first evaluation terminal to obtain the first updated evaluation result, and the answer and the first evaluation result are input into the second evaluation terminal to obtain the second updated evaluation result. The third evaluation module is used to take the first updated evaluation result as the new first evaluation result, take the second updated evaluation result as the new second evaluation result, and return to execute the process of determining the global cognitive energy based on the first evaluation result and the second evaluation result until the global cognitive energy is not greater than the convergence threshold. Then, the evaluation result of the large model to be tested is determined based on the most recently determined first evaluation result and the second evaluation result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Large language model evaluation method and device, electronic equipment and storage medium
CN118035807A
Large language model evaluation method and device, equipment, storage medium and program product
CN119721246A
Simulation method and system for students with different cognitive levels based on large language model
CN120654726A
Object recognition using a congnitive swarm vision framework with attention mechanisms
US20070019865A1