A method for evaluating the capabilities of large command and decision-making models
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]目前,现有的大模型评估方式存在评估数据集难以获取、评估维度过于单一且无法适配指挥决策场景等问题
[0022]根据本发明的另一方面,提供了一种计算机程序产品,包括计算机程序,该计算机程序在被处理器执行时实现如本发明任一实施例所述的评估指挥决策大模型能力的方法。
Smart Images

Figure CN122571014A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method for evaluating the capabilities of large command and decision-making models. Background Technology
[0002] Generative artificial intelligence technologies, exemplified by large language models, rely on neural network architectures with large-scale parameters, multi-head attention mechanisms, and high-performance computing power. They can autonomously generalize and infer based on input information to generate output content that matches the input. To ensure that large language models can be used correctly, model evaluation is usually required.
[0003] Currently, existing large-scale model evaluation methods suffer from problems such as difficulty in obtaining evaluation datasets, overly simplistic evaluation dimensions, and inability to adapt to command and decision-making scenarios. In other words, existing evaluation methods cannot guarantee the effectiveness and accuracy of model evaluation, hindering the application of large language models in command and decision-making scenarios. Summary of the Invention
[0004] This invention provides a method for evaluating the capabilities of a large command and decision-making model, ensuring the effectiveness and accuracy of the evaluation, and guiding the adjustment of the large command and decision-making model based on the evaluation results, so that the large command and decision-making model can adapt to the application requirements of command and decision-making scenarios.
[0005] According to one aspect of the present invention, a method for evaluating the capabilities of a large command and decision-making model is provided, the method comprising:
[0006] Obtain a test sample set, wherein the test sample set includes multiple test samples under at least one type of command and decision simulation scenario, and the test samples include at least: scenario information and first text information, wherein the scenario information includes at least: environmental information, situation information, interference information and command information, and the first text information is used to characterize the task processing instructions under the command and decision simulation scenario;
[0007] Based on the large command and decision model to be evaluated, each test sample in the test sample set is processed to obtain the data to be used corresponding to the test sample; wherein, the data to be used includes: command and decision output data and model operation data;
[0008] Based on the data to be used corresponding to multiple test samples in the same type of command and decision simulation scenario and multiple pre-set evaluation dimensions, the command and decision large model to be evaluated is evaluated to determine the fusion evaluation attributes of the command and decision large model under each evaluation dimension; wherein, the multiple evaluation dimensions include: data adaptability dimension, anti-interference dimension, output data accuracy dimension, and data compliance dimension.
[0009] Based on the evaluation attributes to be integrated under at least some of the evaluation dimensions, determine the target evaluation attributes of the command and decision-making big model to be evaluated in the command and decision-making simulation scenario.
[0010] Based on the target evaluation attributes corresponding to at least one of the command and decision simulation scenarios, the target evaluation result of the command and decision large model to be evaluated is determined.
[0011] According to another aspect of the present invention, an apparatus for evaluating the capability of a large command and decision-making model is provided, the apparatus comprising:
[0012] The test sample acquisition module is used to acquire a test sample set, wherein the test sample set includes multiple test samples under at least one type of command and decision simulation scenario, and the test sample includes at least: scenario information and first text information, wherein the scenario information includes at least: environmental information, situation information, interference information and command information, and the first text information is used to characterize the task processing instructions under the command and decision simulation scenario;
[0013] The sample processing module is used to process each test sample in the test sample set based on the large command and decision model to be evaluated, to obtain the data to be used corresponding to the test sample; wherein, the data to be used includes: command and decision output data and model running data;
[0014] The large model evaluation module is used to evaluate the large command and decision model to be evaluated based on the data to be used corresponding to multiple test samples in the same type of command and decision simulation scenario and multiple pre-set evaluation dimensions, and to determine the fusion evaluation attributes of the large command and decision model to be evaluated under each evaluation dimension; wherein, the multiple evaluation dimensions include: data adaptability dimension, anti-interference dimension, output data accuracy dimension, and data compliance dimension.
[0015] The target evaluation attribute determination module is used to determine the target evaluation attributes of the command and decision-making big model to be evaluated in the command and decision-making simulation scenario based on the evaluation attributes to be fused under at least some of the evaluation dimensions.
[0016] The target evaluation result determination module is used to determine the target evaluation result of the command and decision simulation model to be evaluated based on the target evaluation attributes corresponding to at least one type of command and decision simulation scenario.
[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program that can be executed by the at least one processor, such that the at least one processor can perform the method for evaluating the capabilities of a large command and decision-making model as described in any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer-readable storage medium is provided that stores computer instructions for causing a processor to execute and implement the method for evaluating the capability of a large command and decision model as described in any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method for evaluating the capabilities of a large command and decision-making model as described in any embodiment of the present invention.
[0023] The technical solution of this invention obtains multiple test samples under at least one type of command and decision simulation scenario; processes each test sample in the test sample set based on the large command and decision model to be evaluated to obtain the data to be used corresponding to the test sample. This ensures that the data to be used fits the actual command and decision scenario, and solves the problem in the prior art that the evaluation results are not in line with the command and decision scenario and the model evaluation is distorted due to the difficulty in obtaining model processing data under the real command and decision scenario. The data to be used provides reliable data support for the subsequent evaluation of the large model. Based on the test samples and evaluation dimensions of multiple test samples belonging to the same type of command and decision-making simulation scenario, the large-scale command and decision-making model to be evaluated is evaluated to determine the evaluation attributes to be integrated under each evaluation dimension. Based on the evaluation attributes to be integrated under at least some evaluation dimensions, the target evaluation attributes of the large-scale command and decision-making model under the command and decision-making simulation scenario are determined. This achieves the evaluation of the interpretability, robustness, and compliance dimensions of the large-scale command and decision-making model under evaluation, quantifies the output stability and decision transparency of the large-scale command and decision-making model under evaluation in the corresponding command and decision-making scenario, and solves the problem that existing evaluation dimensions are too singular and unsuitable for corresponding command and decision-making scenarios. This enables accurate measurement of the model performance of the large-scale command and decision-making model under evaluation in command and decision-making scenarios. Based on the target evaluation attributes corresponding to at least one type of command and decision-making simulation scenario, the target evaluation result of the large-scale command and decision-making model under evaluation is determined, realizing the model performance evaluation of the large-scale command and decision-making model under evaluation in at least one type of command and decision-making scenario. This allows the large-scale command and decision-making model adjusted based on the target evaluation result to be adaptable to applications in various command and decision-making scenarios. This invention solves the problems of difficulty in obtaining evaluation datasets, overly simplistic evaluation dimensions, and inability to adapt to command and decision-making scenarios in existing technologies. It ensures the effectiveness and accuracy of large-scale model evaluation and facilitates adjustments to the large-scale command and decision-making model based on the evaluation results, so that the large-scale command and decision-making model can adapt to the application needs of various command and decision-making scenarios.
[0024] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1This is a flowchart of a method for evaluating the capabilities of a large command and decision-making model provided in an embodiment of the present invention;
[0027] Figure 2 This is a flowchart of another method for evaluating the capabilities of a large command and decision-making model provided in an embodiment of the present invention;
[0028] Figure 3 This is a schematic diagram of the structure of a device for evaluating the capabilities of a large command and decision-making model, provided in an embodiment of the present invention.
[0029] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the method for evaluating the capabilities of a large-scale command and decision-making model according to embodiments of the present invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0033] Example 1
[0034] Figure 1This is a flowchart of a method for evaluating the capabilities of a large-scale command and decision-making model according to Embodiment 1 of the present invention. This embodiment is applicable to situations where a large-scale command and decision-making model to be evaluated is being evaluated. This method can be executed by a device for evaluating the capabilities of a large-scale command and decision-making model. This device can be implemented in hardware and / or software, and can be configured in electronic devices such as mobile phones, computers, or servers. Figure 1 As shown, the method includes:
[0035] S110. Obtain a test sample set, wherein the test sample set includes multiple test samples under at least one type of command and decision simulation scenario.
[0036] It should be noted that before obtaining the test sample set, realistic command and decision-making simulation scenarios can be built based on modeling and simulation technology and a simulation platform. These scenarios can be categorized into land, sea, air, space, and cyberspace, resulting in multiple types of simulation scenarios. Different types of simulation scenarios correspond to different scenario information. Based on this, by clarifying the rules, parameters, processes, and test sample acquisition nodes of the simulation, and relying on the simulation platform, the construction and dynamic simulation of command and decision-making scenarios can be achieved.
[0037] Since the large model evaluation scenario can evaluate the model performance of a large model in a certain command and decision simulation scenario, or it can evaluate the model performance of a large model in multiple command and decision simulation scenarios, the test sample set includes multiple test samples from at least one type of command and decision simulation scenario.
[0038] The test samples include at least: scenario information and first text information. Scenario information includes at least: environmental information, situational information, interference information, and command information. Environmental information can be weather information and terrain information within the command and decision-making simulation scenario. Situational information can include the troop deployment, operational status, and operational intentions of both sides. For example, situational information can determine whether the two sides are in an offensive, defensive, or stalemate state. Interference information can be information generated artificially or unintentionally in the command and decision-making simulation scenario that affects command decisions. For example, interference information can include: environmental interference information, infiltration behavior data, and cognitive deception behavior data. Command information can include command hierarchy information within the command and decision-making simulation scenario. Command hierarchy information, from highest to lowest, is: strategic command hierarchy information, operational command hierarchy information, and tactical command hierarchy information. The first text information is used to characterize the task processing instructions within the command and decision-making simulation scenario. For example, taking the two sides as friendly and hostile, the first text information can be instructions to attack the hostile side.
[0039] Specifically, based on at least one type of pre-constructed command and decision simulation scenario, multiple test samples associated with each type of command and decision simulation scenario are identified so that the large command and decision model to be evaluated can process the test samples to obtain data to be used for evaluating the model's performance.
[0040] S120. Based on the large command and decision-making model to be evaluated, each test sample in the test sample set is processed to obtain the data to be used corresponding to the test sample.
[0041] The large-scale command and decision-making model to be evaluated can be understood as a large language model applied to command and decision-making scenarios. The data to be used includes: command and decision-making output data and model operation data. Command and decision-making output data may include situational assessment information, resource allocation information, and decision command issuance data for the command and decision-making simulation scenario. Model operation data may include: the response time and computational efficiency of the large-scale command and decision-making model to generate command and decision-making output data.
[0042] Specifically, the large-scale command and decision model to be evaluated is connected to the simulation platform. Based on the large-scale command and decision model, test samples corresponding to the command and decision simulation scenarios in the simulation platform are processed in real time online to obtain the corresponding command and decision output data and model operation data. Alternatively, the test sample set is input into the large-scale command and decision model to be evaluated to obtain the command and decision output data and model operation data corresponding to each test sample of each type of command and decision simulation scenario.
[0043] It should be noted that command and decision-making simulations can also be performed on each test sample in the test sample set to determine the expected command and decision-making data corresponding to each test sample. This expected data can be the standard command and decision-making data corresponding to the test sample in the command and decision-making simulation scenario. Optionally, to ensure the reliability, compliance, and validity of the data to be used, and to improve the evaluation accuracy of the subsequent large-scale command and decision-making model to be evaluated, all data to be used can undergo data preprocessing and data quality control. Specifically, this processing can include: word segmentation, part-of-speech tagging, syntactic analysis, entity extraction, semantic alignment, and format standardization to obtain preprocessed data. Based on the above, non-standardized command and decision-making output data in the data to be used can be converted into standardized data that meets the evaluation requirements, and the text information in the data to be used can be aligned and converted with the relevant regulations and terminology to ensure the professionalism and standardization of the data. It should be noted that this only involves the replacement of standard terminology and does not involve any change in the specific content.
[0044] The preprocessed data to be used undergoes outlier detection, data integrity verification, sensitivity review, and logical consistency verification to eliminate invalid, abnormal, and logically contradictory data. Simultaneously, the preprocessed data is categorized and labeled to distinguish between sensitive and normal data. This facilitates subsequent evaluation of the command and decision-making model based on data compliance, determining the number of unauthorized accesses based on the categorized sensitive and normal data. Based on the above data quality control processing of the preprocessed data, quality-controlled data is obtained, which is then used to evaluate the command and decision-making model.
[0045] The above approach constructs a command and decision-making simulation scenario and uses the large-scale command and decision-making model to be evaluated to simulate and extrapolate the test samples in the simulation scenario, thereby obtaining the corresponding data to be used. This solves the problem of insufficient data in real command and decision-making scenarios and provides reliable data support for subsequent large-scale model evaluation.
[0046] S130. Based on the data to be used corresponding to multiple test samples in the same type of command and decision simulation scenario and multiple pre-set evaluation dimensions, evaluate the large command and decision model to be evaluated and determine the fusion evaluation attributes of the large command and decision model to be evaluated under each evaluation dimension.
[0047] The evaluation includes several dimensions: data fit, robustness to interference, accuracy of output data, and data compliance. Data fit is used to quantitatively assess the interpretability of the command and decision-making model under evaluation. Interpretability characterizes the model's ability to interpret and understand test samples. Robustness is used to quantitatively assess the robustness of the model under evaluation. Robustness characterizes the model's ability to output reliable command and decision-making data even in command and decision-making simulation scenarios containing interference.
[0048] The output data accuracy dimension can be used to quantitatively evaluate the basic performance of the large-scale command and decision-making model under evaluation, such as assessing its output accuracy, precision, and recall. The data compliance dimension can be used to evaluate the compliance of the command and decision-making output data from the large-scale model under evaluation. The attributes to be fused and evaluated can be used to characterize the model performance of the large-scale command and decision-making model under the corresponding evaluation dimensions.
[0049] Specifically, for the data to be used corresponding to multiple test samples in the same type of command and decision simulation scenario, the data to be used for multiple test samples is processed according to each evaluation dimension in order to quantitatively determine the fusion evaluation attributes of the large command and decision model to be evaluated under each evaluation dimension.
[0050] S140. Based on the evaluation attributes to be integrated under at least some evaluation dimensions, determine the target evaluation attributes of the large command and decision-making model to be evaluated in the command and decision-making simulation scenario.
[0051] The evaluation dimensions can include at least two of the following: data fit, robustness to interference, output data accuracy, and data compliance. The target evaluation attributes characterize the performance of the large-scale command and decision-making model under evaluation in a class of command and decision-making simulation scenarios.
[0052] Specifically, the evaluation attributes to be integrated under at least some evaluation dimensions are weighted and summed to determine the target evaluation attributes of the command and decision-making big model under the command and decision-making simulation scenario.
[0053] S150. Based on the target evaluation attributes corresponding to at least one type of command and decision simulation scenario, determine the target evaluation result of the large command and decision model to be evaluated.
[0054] The target evaluation results are used to characterize the performance of the large command and decision model to be evaluated in at least one type of command and decision simulation scenario.
[0055] Specifically, the target evaluation attributes corresponding to at least one type of command and decision simulation scenario are weighted and summed to obtain the target evaluation result of the large command and decision model to be evaluated, and the target evaluation result is used to determine whether to perform iterative optimization processing on the large command and decision model to be evaluated.
[0056] The technical solution of this embodiment obtains multiple test samples under at least one type of command and decision simulation scenario; processes each test sample in the test sample set based on the large command and decision model to be evaluated to obtain the data to be used corresponding to the test sample. This ensures that the data to be used fits the actual command and decision scenario, and solves the problem in the prior art that the evaluation results are not in line with the command and decision scenario and the model evaluation is distorted due to the difficulty in obtaining model processing data under the real command and decision scenario. The data to be used provides reliable data support for the subsequent evaluation of the large model. Based on the test samples and evaluation dimensions of multiple test samples belonging to the same type of command and decision-making simulation scenario, the large-scale command and decision-making model to be evaluated is evaluated to determine the evaluation attributes to be integrated under each evaluation dimension. Based on the evaluation attributes to be integrated under at least some evaluation dimensions, the target evaluation attributes of the large-scale command and decision-making model under evaluation in the command and decision-making simulation scenario are determined. This achieves the evaluation of the interpretability, robustness, and compliance dimensions of the large-scale command and decision-making model under evaluation, quantifies the output stability and decision transparency of the large-scale command and decision-making model under evaluation in the corresponding command and decision-making simulation scenario, and solves the problem that existing evaluation dimensions are too singular and unsuitable for corresponding command and decision-making scenarios. This enables accurate measurement of the model performance of the large-scale command and decision-making model under evaluation in the command and decision-making simulation scenario. Based on the target evaluation attributes corresponding to at least one type of command and decision-making simulation scenario, the target evaluation result of the large-scale command and decision-making model under evaluation is determined, realizing the model performance evaluation of the large-scale command and decision-making model under evaluation in at least one type of command and decision-making simulation scenario. This allows the large-scale command and decision-making model adjusted based on the target evaluation result to be adaptable to applications in various command and decision-making simulation scenarios. This invention solves the problems of difficulty in obtaining evaluation datasets, overly simplistic evaluation dimensions, and inability to adapt to command and decision-making scenarios in existing technologies. It ensures the effectiveness and accuracy of large-scale model evaluation and guides adjustments to the large-scale command and decision-making model based on the evaluation results, so that the large-scale command and decision-making model can adapt to the application needs of various command and decision-making scenarios.
[0057] Example 2
[0058] Figure 2 This is a flowchart illustrating a method for evaluating the capabilities of a large command and decision-making model according to Embodiment 2 of the present invention. This embodiment is a preferred embodiment of the above embodiments. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here. Figure 2 As shown, the method includes:
[0059] S210. Obtain a test sample set, wherein the test sample set includes multiple test samples under at least one type of command and decision simulation scenario.
[0060] The test samples include at least: scenario information and first text information. The scenario information includes at least: environmental information, situational information, interference information and command information. The first text information is used to characterize the task processing instructions in the command and decision simulation scenario.
[0061] S220. Based on the large command and decision-making model to be evaluated, each test sample in the test sample set is processed to obtain the data to be used corresponding to the test sample.
[0062] The data to be used includes: command and decision-making output data and model operation data.
[0063] S230. Based on the data to be used corresponding to multiple test samples in the same type of command and decision simulation scenario and multiple pre-set evaluation dimensions, evaluate the large command and decision model to be evaluated and determine the fusion evaluation attributes of the large command and decision model to be evaluated under each evaluation dimension.
[0064] The evaluation dimensions include: data adaptability, anti-interference, output data accuracy, and data compliance.
[0065] Optionally, the attributes to be evaluated under the data fit dimension can be determined in the following way: Based on the command and decision output data corresponding to multiple test samples in the same type of command and decision simulation scenario, determine the target heterogeneity ratio, target judgment coefficient, target feature contribution attribute, target dispersion coefficient, and target overlap information corresponding to the large command and decision model to be evaluated; Based on the target heterogeneity ratio, target judgment coefficient, target feature contribution attribute, target dispersion coefficient, target overlap information, and the first evaluation function, determine the attributes to be evaluated under the data fit dimension of the large command and decision model to be evaluated.
[0066] The target heterogeneity ratio is used to characterize the stability of command and decision output data under the same type of command and decision simulation scenarios. This can be understood as determining the stability of the command and decision output data of the large-scale command and decision model under evaluation under the same type of command and decision simulation scenarios through the target heterogeneity ratio, thus avoiding sudden changes in command and decision output data due to a single data point in the test sample.
[0067] The target decision coefficient is used to characterize the degree of fit between the command and decision output data and the test samples. It can be understood as determining the degree of fit between the command and decision output data of the large command and decision model to be evaluated and the situation corresponding to the test samples.
[0068] The target feature contribution attribute is used to characterize the causal logical relationship between command and decision output data and scenario information. It can be understood as determining the impact of scenario features in the test sample on the command and decision output data through the target feature contribution attribute, and measuring the causal logical rationality between the command and decision output data and the test sample.
[0069] The target dispersion coefficient is used to characterize the rationality of the distribution of tactical types in the command and decision output data under the same type of command and decision simulation scenario. It can be understood as measuring the explanatory sufficiency of the large command and decision model to be evaluated, in order to determine whether the large command and decision model to be evaluated can adapt to different command and decision needs.
[0070] Target overlap information is used to characterize the semantic translation accuracy of the large command and decision-making model to be evaluated. The first evaluation function can be an evaluation function corresponding to the data fit dimension.
[0071] Specifically, based on the command and decision output data corresponding to multiple test samples in the same type of command and decision simulation scenario, the target heterogeneity ratio, target judgment coefficient, target feature contribution attribute, target dispersion coefficient, and target overlap information corresponding to the large command and decision model to be evaluated are determined. These target heterogeneity ratios, target judgment coefficients, target feature contribution attributes, target dispersion coefficients, and target overlap information are then substituted into the first evaluation function corresponding to the data fit dimension to obtain the fusion evaluation attributes of the large command and decision model to be evaluated in the data fit dimension.
[0072] The target determination coefficient is determined as follows: the target determination coefficient is determined based on the command and decision output data and the corresponding command and decision expectation data corresponding to multiple test samples under the same type of command and decision simulation scenario; wherein the command and decision expectation data is the expected output data determined based on the test samples corresponding to the command and decision output data.
[0073] The target determination coefficient can be determined using the following determination coefficient determination function:
[0074] ;
[0075] in, This represents the target determination coefficient, with a value range of [0,1]. The closer the target determination coefficient is to 1, the stronger the explanatory power of the command and decision-making model being evaluated in terms of situational evolution. Generally, when... At that time, it is determined that the model fit of the large command and decision-making model to be evaluated meets the command and decision-making requirements. i represents the sequence number of the test sample, and n represents the number of test samples. This represents the actual value quantified from the expected command and decision data for the i-th test sample. This represents the quantified average value corresponding to the command and decision-making expectation data of n test samples. This represents the predicted value obtained by quantifying the command decision output data under the i-th test sample. For example, the command decision output data can be quantified and scored to determine the corresponding predicted value.
[0076] Specifically, for command and decision output data corresponding to multiple test samples in the same type of command and decision simulation scenario, the multiple command and decision output data are quantified to obtain the predicted value corresponding to each command and decision output data. Similarly, the multiple command and decision expectation data are quantified to obtain the actual value corresponding to each command and decision expectation data. The multiple predicted values and the actual values corresponding to the predicted values are then substituted into the determination function to obtain the target determination coefficient.
[0077] Accordingly, the target feature contribution attribute is determined as follows: For multiple test samples in the same type of command and decision simulation scenario, a scenario feature correlation analysis is performed on the command and decision output data and scenario information corresponding to the test samples and multiple unprocessed scenario features to determine the first feature attribute of each unprocessed scenario feature; wherein, the unprocessed scenario feature is the feature determined by feature extraction of scenario information; based on the first feature attribute of each unprocessed scenario feature, a preset number of unused scenario features and second feature attributes of the unused scenario features are determined; based on the first feature attribute and the second feature attribute, the unused feature contribution attribute corresponding to the test sample is determined; based on the unused feature contribution attributes of multiple test samples, the target feature contribution attribute is determined.
[0078] The scene features to be processed can be features determined by feature extraction from scene information. For example, taking environmental information in scene information as an example, features can be extracted from the environmental information in scene information to obtain multiple scene features corresponding to the environmental information.
[0079] The first feature attribute can be used to characterize the impact of the scene features to be processed on the command and decision output data. Optionally, the first feature attribute can be determined as follows: For multiple scene features to be processed corresponding to the scene information of each test sample, one scene feature to be processed is selected as the current scene feature by controlling variables, while other scene features are fixed. By changing the current scene feature, multiple associated samples corresponding to the test sample are obtained. Based on the large command and decision model to be evaluated, the multiple associated samples are processed, and the causal correlation between the current scene feature and the command and decision output data is determined according to the difference information of the command and decision output data corresponding to all associated samples. Based on the causal correlation, the first feature attribute of the current scene feature is determined. The above process is repeated to obtain the first feature attribute corresponding to each scene feature to be processed.
[0080] The preset number of scenario features to be used can be at least a subset of scenario features selected from multiple scenario features to be processed. The selection method can be based on actual needs, choosing scenario features that are more important in the command and decision-making simulation scenario, or it can be based on the magnitude of the first feature attribute. For example, the scenario features corresponding to the top K largest first feature attributes can be selected as scenario features to be used, arranged in descending order of the first feature attribute. Correspondingly, the first feature attribute of the scenario features to be used is used as the second feature attribute.
[0081] The feature contribution attribute to be used can be obtained through the following contribution attribute determination function:
[0082] ;
[0083] in, This represents the feature contribution attribute to be used, with a value range of [0,1]. This indicates that the greater the impact of the scenario characteristics to be used on the command and decision-making output data, the stronger the causal logical correlation of the model's decisions; k represents the sequence number of the scenario characteristics to be used, and K represents the preset quantity. Let represent the second feature attribute of the k-th scenario feature to be used, where i represents the index of the scenario feature to be processed, and N represents the number of scenario features to be processed. This represents the i-th scene feature to be processed.
[0084] The feature contribution attribute to be used characterizes the degree of influence of the scenario features to be used on the command and decision output data under the current test sample. The target feature contribution attribute characterizes the degree of influence of the scenario features to be used on the command and decision output data under the same type of command and decision simulation scenario. Optionally, the target feature contribution attribute can be determined based on multiple feature contribution attributes to be used. For example, the target feature contribution attribute can be obtained by averaging multiple feature contribution attributes to be used, or the median or extreme value among multiple feature contribution attributes to be used can be used as the target feature contribution attribute.
[0085] Specifically, for multiple test samples in the same type of command and decision simulation scenario, feature extraction is performed on the scenario information of each test sample to determine multiple scenario features to be processed corresponding to each test sample. Scenario feature correlation analysis is then performed on the command and decision output data corresponding to each test sample and the multiple scenario features to be processed to obtain the first feature attribute of each scenario feature to be processed. Based on the first feature attribute of each scenario feature to be processed, a preset number of scenario features to be used and their second feature attributes are determined.
[0086] Substituting the second feature attributes of a preset number of scenario features to be used and the first feature attributes of each scenario feature to be processed into the above contribution attribute determination function, the contribution attributes of the features to be used corresponding to the test samples are obtained. Based on the contribution attributes of the features to be used from multiple test samples, the target feature contribution attribute is determined.
[0087] Accordingly, the target overlap information corresponding to the large command and decision model to be evaluated is determined in the following way: For multiple test samples under the same type of command and decision simulation scenario, based on the command text information in the command and decision output data corresponding to the test sample, the first jump word double-word set and the first single-word set corresponding to the command text information are determined; based on the expected text information in the expected command and decision data, the second jump word double-word set and the second single-word set corresponding to the expected text information are determined; based on the first jump word double-word set, the first single-word set, the second jump word double-word set, and the second single-word set, the first overlap information corresponding to the test sample is determined; based on the first overlap information of multiple test samples, the target overlap information is determined.
[0088] The command text information can be the text content contained in the command decision output data. The first skip-bigram set can be the skip-bigram set corresponding to the command text information. The first unigram set can be the unigram set corresponding to the command text information. Optionally, the first skip-bigram set can be the set of all non-adjacent ordered bigram pairs obtained after extracting the command text information according to a preset maximum skip interval. The first unigram set can be the set of all independent words after segmenting the command text information.
[0089] Correspondingly, the expected text information can be the text information contained in the command and decision-making expected data. The second jump word double-word set can be the jump word double-word set corresponding to the expected text information. The second single-word set can be the single-word set corresponding to the expected text information. It should be noted that the second jump word double-word set and the second single-word set can also be determined based on a standard command terminology database.
[0090] The first overlap information is used to characterize the degree of overlap and matching between the first set of two-word pairs and the first set of single-word pairs, and the second set of two-word pairs and the second set of single-word pairs, thus characterizing the semantic understanding accuracy of the command and decision-making model to be evaluated. Optionally, the first overlap information can be obtained through the following overlap determination function:
[0091] ;
[0092] in, This represents the first degree of overlap, with a value range of [0,2]. The larger the value, the stronger the explanatory and normative nature of the command and decision-making model to be evaluated. This represents the set of two-word phrases that are the first jump words. Represents the first set of words. Indicates a set of two-word phrases for the second jump. This represents the second set of single-word words; adding one to the denominator avoids the case where the divisor is 0.
[0093] The target overlap information can be determined based on multiple first overlap information and is used to characterize the semantic understanding accuracy of the large command and decision model to be evaluated in a certain type of command and decision simulation scenario.
[0094] Specifically, for multiple test samples in the same type of command and decision-making simulation scenario, the command text information corresponding to the command and decision-making output data in each test sample is determined, as well as the expected text information corresponding to the command and decision-making expected data corresponding to the command and decision-making output data is determined. The command text information is processed to obtain a first set of two-word pairs and a first set of single-word pairs, and the expected text information is processed to obtain a second set of two-word pairs and a second set of single-word pairs. These sets are then substituted into the aforementioned overlap determination function to obtain the first overlap information corresponding to the test sample. Based on the first overlap information corresponding to multiple test samples, the target overlap information is determined.
[0095] Accordingly, the target heterogeneity ratio is determined as follows: the command and decision output data corresponding to multiple test samples under the same type of command and decision simulation scenario are processed by decision classification to obtain command and decision output data under at least one decision type; the number of targets is determined based on the number of command and decision output data under at least one decision type; and the target heterogeneity ratio is determined based on the total number of command and decision output data under all decision types and the number of targets.
[0096] The decision type can be pre-defined, corresponding to the command decision output data. Optionally, clustering can be performed on multiple command decision output data to determine the command decision output data under each of the m decision types.
[0097] The target quantity is used to characterize the number of data points corresponding to the decision type that generates the most command and decision output data. For example, if there are 'a' command and decision output data points under the first decision type and 'b' command and decision output data points under the second decision type, and 'a' is greater than 'b', then the target quantity is 'a'. The total number of command and decision output data points can be the total number of all command and decision output data points across all decision types.
[0098] Optionally, the target dissident ratio can be determined using the following dissident ratio determination function:
[0099] ;
[0100] in, This represents the target audience dissimilarity ratio, with a value range of [0,1]. The smaller the value, the stronger the consistency of the command and decision-making output data of the model to be evaluated. This represents the number of objectives, and m represents the m decision types. This indicates the total amount of data output from command and decision-making processes.
[0101] Specifically, for command and decision output data corresponding to multiple test samples in the same type of command and decision simulation scenario, clustering is performed on the multiple command and decision output data to determine at least one command and decision output data under each decision type. Based on the number of command and decision output data included in each decision type, the target quantity is determined. Substituting the total number of command and decision output data and the target quantity into the aforementioned heterogeneity ratio determination function yields the target heterogeneity ratio.
[0102] Accordingly, the target dispersion coefficient is determined as follows: tactical classification processing is performed on the command decision output data corresponding to multiple test samples under the same type of command decision simulation scenario to determine the command decision output data under at least one tactical type; based on the number of command decision output data under at least one tactical type, the standard deviation and mean value corresponding to the tactical type are determined; based on the standard deviation and mean value corresponding to the tactical type, the target dispersion coefficient is determined.
[0103] The tactical type can be determined based on command and decision output data, and the current command and decision simulation scenario is based on the tactical approach adopted by the current test sample. For example, tactical types can include: offensive tactics, defensive tactics, maneuver tactics, and harassment and containment tactics, etc. Optionally, multiple command and decision output data can be clustered to determine the command and decision output data for each tactical type under at least one tactical type.
[0104] The standard deviation and mean corresponding to each tactical type can be determined based on the amount of command decision output data for each tactical type. The target coefficient of variation is used to reflect the reasonableness of the distribution of tactical types in the same type of command decision simulation scenario for the large-scale command decision model to be evaluated. The more reasonable the distribution of tactical types, the stronger the explanatory power of the large-scale command decision model to be evaluated. Optionally, the target coefficient of variation can be obtained through the following coefficient of variation determination function:
[0105] ;
[0106] in, Represents the target dispersion coefficient. The smaller the value, the more reasonable the tactical type distribution of the large command and decision model to be evaluated under the same type of command and decision simulation scenario, and the stronger the sufficiency of the model's explanation. This represents the standard deviation corresponding to the tactical type. This represents the average value corresponding to the tactical type.
[0107] Specifically, for command decision output data from multiple test samples within the same command decision simulation scenario, tactical classification is performed on these multiple command decision output data to determine the command decision output data belonging to each tactical type. Based on the number of command decision output data under each tactical type, the standard deviation and mean corresponding to that tactical type are determined. Substituting the standard deviation and mean into the aforementioned coefficient of variation determination function yields the target coefficient of variation.
[0108] Optionally, after obtaining the aforementioned target heterogeneity ratio, target determination coefficient, target feature contribution attribute, target dispersion coefficient, and target overlap information, the data fit dimension evaluation attribute can be determined based on the following first evaluation function. The first evaluation function can be expressed as:
[0109] ;
[0110] in, This represents the attributes to be evaluated in the data fit dimension. The value range is [0,1]. The larger the value, the better the interpretability of the command and decision-making model to be evaluated, and the higher the degree of trust in the model. Indicates the target audience's dissimilarity ratio. This represents the weighting coefficient corresponding to the target audience's dissimilarity ratio. Represents the target determination coefficient. This represents the weighting coefficient corresponding to the target determination coefficient. Indicates the contribution attribute of the target feature. This represents the weight coefficient corresponding to the contribution attribute of the target feature. Represents the target dispersion coefficient. This represents the weighting coefficient corresponding to the target dispersion coefficient. Indicates target overlap information. This represents the weighting coefficient corresponding to the overlap information with the target. Wherein, Each weighting coefficient can be adjusted according to actual needs.
[0111] Optionally, the anti-interference dimension includes a disturbance fluctuation sub-dimension and a disturbance stability sub-dimension. The evaluation attributes to be fused under the anti-interference dimension can be determined as follows: based on the data to be used and reference data corresponding to multiple test samples in the same type of command and decision simulation scenario, determine the first evaluation attribute under the disturbance fluctuation sub-dimension and the second evaluation attribute under the disturbance stability sub-dimension; wherein, the reference data is obtained by processing the test samples with disturbance information removed by the large command and decision model to be evaluated; based on the first evaluation attribute, the second evaluation attribute, and the second evaluation function, determine the evaluation attributes to be fused under the anti-interference dimension.
[0112] For each test sample, the scenario information can be adjusted to obtain a test sample with interference information removed. Then, based on the large command and decision model to be evaluated, each test sample with interference information removed is processed to obtain reference data. The reference data is used to characterize the command and decision output data corresponding to the test sample with interference information removed.
[0113] The anti-interference dimension is used to evaluate the robustness of the large command and decision-making model to be evaluated. That is, to evaluate the defense performance and stability of the large command and decision-making model to be evaluated in complex command and decision-making simulation scenarios, including at least one type of interference information such as adversarial sample attacks, environmental interference, illegal penetration, and cognitive deception, so as to determine whether the large command and decision-making model to be evaluated can output reliable command and decision-making output data in the above-mentioned complex command and decision-making simulation scenarios.
[0114] The perturbation fluctuation sub-dimension is used to assess the degree of difference between the reference data corresponding to the test sample with perturbation information removed and the command and decision output data corresponding to the test sample containing perturbation information, in order to quantitatively evaluate the robustness of the large command and decision model under evaluation. Accordingly, the first evaluation attribute can be the evaluation attribute corresponding to the perturbation fluctuation dimension.
[0115] The perturbation stability sub-dimensional is used to assess the degree of difference between the test sample containing the command decision output data with anomalies and the test sample containing the corresponding reference data. It measures the tolerance of the large command decision model to be evaluated to perturbations in the command decision simulation scenario. Correspondingly, the second evaluation attribute is the evaluation attribute corresponding to the perturbation stability sub-dimensional. The second evaluation function can be used to process the first and second evaluation attributes to determine the evaluation attribute to be fused corresponding to the anti-interference dimension.
[0116] Specifically, for multiple test samples and corresponding reference data in the same type of command and decision simulation scenario, based on the multiple test samples and corresponding reference data, a first evaluation attribute under the disturbance fluctuation sub-dimensional and a second evaluation attribute under the disturbance stability sub-dimensional are determined. Substituting the first and second evaluation attributes into the second evaluation function yields the fusion evaluation attribute under the interference resistance dimension of the large command and decision model to be evaluated.
[0117] The first evaluation attribute under the perturbation fluctuation sub-dimension is determined as follows: based on the data to be used and the reference data corresponding to multiple test samples, the first processing accuracy and the first processing time corresponding to the data to be used, and the second processing accuracy and the second processing time corresponding to the reference data are determined; the first evaluation attribute under the perturbation fluctuation sub-dimension is determined based on the first processing accuracy, the first processing time, the second processing accuracy, and the second processing time.
[0118] The first processing accuracy can be the output accuracy of the large-scale command and decision-making model to be evaluated, determined based on multiple datasets to be used. The first processing time can be used to characterize the average time taken for the large-scale command and decision-making model to be evaluated to process test samples corresponding to the datasets to be used. Correspondingly, the second processing accuracy can be the output accuracy of the large-scale command and decision-making model to be evaluated, determined based on multiple reference datasets. The second processing time can be used to characterize the average time taken for the large-scale command and decision-making model to be evaluated to process test samples corresponding to the reference datasets. The first evaluation attribute can be determined by the following function:
[0119] ;
[0120] in, This represents the disturbance volatility, i.e., the primary assessment attribute. The range of values is [0, +∞). The smaller the value, the stronger the stability of the command and decision-making model under the influence of interference information, and the more stable the command and decision-making effectiveness. It can be a first value quantified based on the first processing accuracy and the first processing time, used to characterize the performance of the command and decision-making large model to be evaluated under interference information disturbance. It can be a second value determined based on the second processing accuracy and the second processing time, used to characterize the performance of the large command and decision-making model to be evaluated without being disturbed by interference information.
[0121] Specifically, for multiple test samples belonging to the same type of command and decision simulation scenario, corresponding to the data to be used and multiple test samples with removed interference information, the data to be used is evaluated to determine the first processing accuracy, and the original processing time corresponding to each data to be used is determined based on the model running data in the data to be used. The average of all original processing times is applied to obtain the first processing time. Correspondingly, multiple reference data are evaluated to determine the second processing accuracy. The second processing time is determined based on the output time of each reference data from the command and decision model to be evaluated and the number of reference data. The first processing accuracy and the first processing time are processed to determine the first value. The second processing accuracy and the second processing time are processed to determine the second value. Optionally, the processing accuracy and processing time can be weighted and summed to obtain the corresponding value. Substituting the first value and the second value into the above function, the first evaluation attribute under the perturbation fluctuation sub-dimension is obtained.
[0122] Accordingly, the second evaluation attribute under the perturbation stability sub-dimension is determined as follows: anomaly detection is performed on the data to be used corresponding to multiple test samples to identify at least one data to be processed; based on the test sample to which the at least one data to be processed belongs and the test sample to which the corresponding reference data belongs, test sample difference information associated with the perturbation type is determined; based on at least one test sample difference information associated with the perturbation type, the second evaluation attribute under the perturbation stability sub-dimension is determined.
[0123] The data to be processed can be anomalous data identified from multiple datasets to be used. For example, if the data to be used contains problems such as decision-making errors, logical contradictions, or decision-making delays, then that data will be considered as data to be processed.
[0124] The test sample difference information associated with the perturbation type can be determined as follows: Based on the perturbation information of at least one test sample to which the data to be processed belongs, determine the perturbation type corresponding to each test sample to which the data to be processed belongs. For example, the perturbation type may include: environmental perturbation type, data noise perturbation type, illegal penetration perturbation type, etc. For the test sample corresponding to at least one data to be processed under each perturbation type, evaluate the difference between the test sample corresponding to the data to be processed and the test sample to which the corresponding reference data belongs, and determine the test sample difference information corresponding to each test sample to which the data to be processed belongs under each perturbation type. Optionally, the test sample difference information can be determined using the following function:
[0125] when hour, ;
[0126] in, Indicates the data to be used. This indicates the preset label information corresponding to the reference data, used to assess whether the data to be used is abnormal. This indicates that the data to be used contains problems such as decision-making errors, logical contradictions, or decision-making delays. This refers to the test samples corresponding to the reference data, which can be understood as test samples with interference information removed, i.e., the original samples. This represents the test sample corresponding to the data to be processed. express and The distance between them corresponds to the test sample difference information mentioned above. The index indicates the type of disturbance. ,in It represents the set of all disturbance types, including environmental disturbance types, data noise disturbance types, illegal infiltration disturbance types, and other disturbance types specific to command and decision-making scenarios.
[0127] In addition, when hour, ;
[0128] in, This indicates that the data to be used does not contain issues such as decision-making errors, logical contradictions, or decision-making delays. In other words, under these circumstances, the large command and decision-making model to be evaluated has not experienced performance degradation. This indicates the degree of difference between the test sample containing the data to be used and the test sample containing the reference data.
[0129] The second evaluation attribute can be used to reflect the perturbation stability of the large-scale command and decision-making model being evaluated. The second evaluation attribute can be determined using the following function:
[0130] ;
[0131] in, This represents the stability of the disturbance, i.e., the second evaluation attribute, with a value range of [0, +∞). The larger the value, the stronger the disturbance resistance and the better the robustness of the command and decision-making model to be evaluated. Indicates the difference information of the test samples. This indicates the number of test samples to which at least one piece of data to be processed belongs, or the number of test samples to which the corresponding reference data belongs.
[0132] Specifically, based on reference data corresponding to multiple test samples with removed interference information, anomaly assessment is performed on the data to be used corresponding to the test samples to identify at least one data point to be processed. Based on the interference information of the test samples to which the at least one data point to be processed belongs, the perturbation type corresponding to each test sample to which the data point to be processed belongs is determined. For each perturbation type, the test samples corresponding to the at least one data point to be processed are compared with the test samples to which the corresponding reference data belongs, and the difference information between the test samples to be processed and the test samples to which the corresponding reference data belongs is determined. Based on the difference information of at least one test sample associated with the perturbation type, the minimum difference information of the test samples is determined as the second evaluation attribute under the perturbation stability sub-dimension.
[0133] Optionally, after obtaining the first evaluation attribute and the second evaluation attribute mentioned above, the evaluation attribute to be fused under the anti-interference dimension can be determined based on the following second evaluation function:
[0134] ;
[0135] in, This represents the attribute to be fused and evaluated under the anti-interference dimension, used to reflect the robustness of the large command and decision-making model to be evaluated. The value range is [0,1]. The larger the value, the stronger the robustness of the command and decision-making model to be evaluated, and the higher the reliability of decision-making in complex command environments. This represents the disturbance volatility, which is the first assessment attribute. This represents the weight coefficient corresponding to the first evaluation attribute. This represents the stability of the disturbance, i.e., the second evaluation attribute. This represents the weighting coefficient corresponding to the second evaluation attribute. and It can be adjusted according to actual needs, for example, in a command and decision-making simulation scenario with high interference. It could be 0.3. It can be 0.7. In a typical command and decision-making simulation scenario with interference, It could be 0.4. It can be 0.6. Used to normalize the first evaluation attribute to [0,1]. Used to normalize the second evaluation attribute to [0,1].
[0136] It should be noted that before determining the attributes to be evaluated in terms of the accuracy of the output data of the large command and decision model to be evaluated, positive and negative samples can be determined in the following way: based on multiple test samples in the same type of command and decision simulation scenario and the command and decision output data corresponding to the test samples, the test samples are divided to obtain the number of true positives, the number of true negatives, the number of false positives, and the number of false negatives.
[0137] Among them, true positives and true negatives correspond to test samples where the command decision output data is correct, while false positives and false negatives correspond to test samples where the command decision output data is incorrect.
[0138] In this context, a true example can be understood as a correct positive sample. For instance, if the command decision output data corresponding to a test sample correctly determines the opponent's true intentions or correctly identifies the opponent's target, then that test sample can be considered a true example. Optionally, a true example can be represented as TP (True Positive).
[0139] A true negative example can be understood as a correct negative sample. For example, if the command decision output data corresponding to the test sample correctly identifies the enemy's camouflage or feigned posture, and its output is correct, then the test sample is considered a true negative example. Optionally, a true negative example can be represented as TN (True Negative).
[0140] A false positive can be understood as an incorrect positive sample. For example, if the command decision output data of a test sample misclassifies the enemy as friendly or invalid situational information as valid, then the test sample is considered a false positive. Optionally, a false positive can be represented as FP (False Positive).
[0141] A false negative example can be understood as an incorrect negative sample. For example, if the command decision output data of a test sample incorrectly identifies a friendly force as an adversary or determines valid situational information as invalid, then that test sample is considered a false negative example. Optionally, a false negative example can be represented as FN (False Negative).
[0142] Specifically, for command and decision output data from multiple test samples within the same command and decision simulation scenario, each command and decision output data is quantized to obtain a quantized value. Based on a pre-set classification threshold, the quantized values of the multiple command and decision output data are classified to distinguish between positive and negative samples in the test samples. Furthermore, based on whether the command and decision output data corresponding to the test sample is correct, correct positive samples, correct negative samples, incorrect positive samples, and incorrect negative samples are distinguished to obtain the number of true positives, true negatives, false positives, and false negatives. It should be noted that the number of true positives, true negatives, false positives, and false negatives differs under different classification thresholds.
[0143] Optionally, the output data accuracy dimension includes: a decision-making tendency sub-dimension, a text similarity sub-dimension, and at least one precision sub-dimension. The attributes to be fused and evaluated under the output data accuracy dimension can be determined as follows: Based on the number of true positives, true negatives, false positives, and false negatives, determine at least one third evaluation attribute under the precision sub-dimension; perform command-making tendency analysis on the command-making output data and command-making expectation data corresponding to multiple test samples to determine the first decision probability distribution data corresponding to the command-making output data and the second decision probability distribution data corresponding to the command-making expectation data; determine the fourth evaluation attribute under the decision-making tendency sub-dimension based on multiple first decision probability distribution data and multiple second decision probability distribution data; perform text similarity evaluation processing on the command text information in the command-making output data and the expected text information in the command-making expectation data corresponding to multiple test samples to determine the fifth evaluation attribute under the text similarity sub-dimension; and perform a weighted summation of the fourth evaluation attribute, the fifth evaluation attribute, and at least one third evaluation attribute to determine the attribute to be fused and evaluated corresponding to the output data accuracy dimension.
[0144] Among them, the output data accuracy dimension measures the basic performance of the command and decision-making model under evaluation by assessing the output accuracy of the model.
[0145] At least one precision sub-dimension includes one or more of the following: model accuracy, model precision, model recall, model error rate, model harmonic mean, model classification, model average precision, and model gain. The third evaluation attribute under the model accuracy sub-dimension can be understood as the accuracy of the large command and decision-making model to be evaluated, which can be obtained through the following accuracy determination function:
[0146] ;
[0147] Here, ACC represents accuracy, which is the proportion of correct output samples from the large command and decision-making model being evaluated out of the total number of samples. It reflects the decision-making accuracy of the large command and decision-making model being evaluated. The value of ACC ranges from [0,1]. The larger the ACC, the higher the accuracy of the large command and decision-making model being evaluated. It corresponds to the third evaluation attribute under the accuracy sub-dimension mentioned above. TP represents the number of true positives, TN represents the number of true negatives, FP represents the number of false positives, and FN represents the number of false negatives.
[0148] Correspondingly, the third evaluation attribute under the model accuracy sub-dimension can be understood as the accuracy of the large command and decision-making model to be evaluated, which can be determined by the following function:
[0149] ;
[0150] Where P represents the third evaluation attribute under the model accuracy sub-dimension, namely accuracy, with a value range of [0,1]. The larger P is, the higher the recognition accuracy of the large command decision model to be evaluated. It is suitable for command decision scenarios such as enemy command intent recognition and key situation information extraction. TP represents the number of true positives and FP represents the number of false positives.
[0151] Correspondingly, the third evaluation attribute under the model recall sub-dimension can be understood as the recall of the large command and decision-making model to be evaluated, which can be obtained through the following recall determination function:
[0152] ;
[0153] Where R represents the third evaluation attribute under the model recall sub-dimension, namely recall, with a value range of [0,1]. The larger R is, the stronger the recall capability of the command and decision-making model to be evaluated, avoiding the omission of important information and reflecting the comprehensiveness of the command and decision-making model to be evaluated. TP represents the number of true positives and FN represents the number of false negatives.
[0154] Correspondingly, the third evaluation attribute under the model error rate sub-dimension can be understood as the error rate of the large command and decision-making model to be evaluated, which can be obtained through the following error rate determination function:
[0155] ;
[0156] Here, ACC represents accuracy, and Err represents error rate, which is the third evaluation attribute under the model error rate sub-dimension. Err takes the value range of [0,1]. The smaller Err is, the lower the decision error rate of the large command and decision model to be evaluated, and the higher the decision reliability.
[0157] Correspondingly, the third evaluation attribute under the harmonic mean sub-dimension can be understood as the harmonic mean of precision and recall, taking into account both the precision and recall of the large-scale command and decision-making model under evaluation, and comprehensively reflecting the recognition performance of the large-scale command and decision-making model under evaluation. The third evaluation attribute under the harmonic mean sub-dimension can be determined by the following function:
[0158] ;
[0159] Here, F1 represents the third evaluation attribute under the harmonic mean sub-dimension, i.e., the harmonic mean. The value of F1 ranges from [0,1]. The larger the F1 value, the better the overall recognition performance of the command and decision-making model being evaluated, which is suitable for scenarios where both precision and recall need to be considered. R represents the third evaluation attribute under the recall sub-dimension, i.e., recall; P represents the third evaluation attribute under the precision sub-dimension, i.e., precision.
[0160] Correspondingly, the third evaluation attribute under the model classification sub-dimension can be the area under the ROC curve. The ROC curve is used to reflect the relationship between the true positive rate (TPR) and false positive rate (FPR) of the large command and decision-making model under evaluation at different classification thresholds, and is used to comprehensively measure the classification ability of the large command and decision-making model under evaluation. The ROC curve is a model classification performance curve with the false positive rate (FPR) as the horizontal axis and the true positive rate (TPR) as the vertical axis.
[0161] The false positive rate (FPR) can be determined by the following function:
[0162] ;
[0163] Where FPR represents the false positive rate, FP represents the number of false positives, and TN represents the number of true negatives.
[0164] The True Rate of Return (TPR) can be determined using the following function:
[0165] ;
[0166] Where TPR represents the true positive rate, TP represents the number of true positives, and FN represents the number of false negatives.
[0167] Correspondingly, the third evaluation attribute under the model classification sub-dimension can be determined by the following function:
[0168] ;
[0169] Here, AUC represents the third evaluation attribute under the model classification sub-dimension, i.e., the area under the ROC curve. The AUC value ranges from [0.5, 1]. The larger the AUC, the stronger the decision classification performance of the large command and decision model to be evaluated. When AUC ≥ 0.8, it meets the requirements of the command and decision scenario. TPR represents the true positive rate, and FPR represents the false positive rate.
[0170] Correspondingly, the third evaluation attribute under the model's average precision sub-dimension can be understood as the area under the PRC curve, i.e., average precision (AP). The PRC curve (precision-recall curve) is determined by precision (P) and recall (R) at different classification thresholds. The PRC curve has recall (R) on the horizontal axis and precision (P) on the vertical axis. Average precision is the area under the PRC curve, used to reflect the relationship between precision and recall of the large-scale command and decision-making model under evaluation at different classification thresholds, and is used to measure the performance of the large-scale command and decision-making model under evaluation on multiple test samples. Optionally, the third evaluation attribute under the model's average precision sub-dimension can be determined by the following function:
[0171] ;
[0172] Here, AP represents the third evaluation attribute under the sub-dimension of model average precision, i.e., average precision. AP ranges from [0,1]. The larger the AP, the better the performance of the large command and decision model to be evaluated in the imbalanced command and decision dataset, indicating that the large command and decision model has better recognition ability and stronger robustness in command and decision scenarios with few enemy targets and many interference samples. k represents the index of the classification threshold, and the value of k ranges from [1,n], where n represents the total number of classification thresholds. This represents the recall rate at the k-th classification threshold. This represents the recall rate at the (k-1)th classification threshold. This represents the accuracy at the k-th classification threshold.
[0173] Correspondingly, the third evaluation attribute under the model gain value sub-dimension can be the difference data determined by the CRC curve and the random curve, i.e., the gain value G. The CRC curve (cumulative response curve) is used to display the true positive rate and positive prediction percentage in command and decision output data across multiple classification thresholds. The cumulative response curve reflects the filtering efficiency of the large command and decision model under evaluation in massive amounts of data. The gain value G is the difference between the CRC curve and the random curve, normalized to [0,1] and used for comprehensive evaluation. The larger the gain value, the stronger the performance of the large command and decision model under evaluation in quickly filtering key information and prioritizing command and decision decisions, and the more accurately it can focus on core objectives in unbalanced and highly redundant command and decision scenarios.
[0174] The decision propensity sub-dimension is used to assess whether the overall decision propensity of the command and decision-making output data of the large-scale command and decision-making model under evaluation conforms to the laws of command and decision-making. Command and decision expectation data can be understood as theoretical command and decision-making data under the same test sample. The first decision probability distribution data is used to characterize the decision probability distribution corresponding to the command and decision output data. Correspondingly, the second decision probability distribution data is used to characterize the decision probability distribution corresponding to the command and decision expectation data.
[0175] The fourth evaluation attribute under the decision propensity sub-dimension can be used to characterize the degree of difference between the probability distributions of the first and second command decisions. Optionally, the fourth evaluation attribute under the decision propensity sub-dimension can be characterized by KL divergence, i.e., implemented through the following function:
[0176] ;
[0177] in, Represents the i-th test sample The corresponding second decision probability distribution, Represents the i-th test sample The corresponding first decision probability distribution, This refers to the KL divergence, which is the fourth evaluation attribute under the aforementioned decision-making tendency sub-dimension. The range of values is [0, +∞). The smaller the value, the closer the decision distribution output by the large command decision model to be evaluated is to the actual command scenario distribution, and the better the fitting effect. The larger the value, the more it indicates that the decision deviates from practical logic and the lower its credibility.
[0178] The text similarity sub-dimension is used to evaluate the similarity between command text information and expected text information, in order to determine the accuracy of the text output of the command decision-making model under evaluation. Command text information can be understood as the text information in the command decision output data of the command decision-making model under evaluation. Expected text information can be the standard text information in the expected command decision data.
[0179] The fifth evaluation attribute can be used to characterize the similarity between the command text information and the expected text information. Optionally, the fifth evaluation attribute can be determined by the following function:
[0180] ;
[0181] in, This represents the fifth evaluation attribute, used to assess the performance of the large command and decision-making model being evaluated in generating natural language. The range of values is [0,1]. The larger the value, the higher the similarity between the command text information generated by the command decision-making model to be evaluated and the expected text information, and the better the standardization of the output data of the command decision-making model to be evaluated. BP is the brevity penalty factor, used to penalize excessively short text information, BP=exp(1-standard text length of expected text information / generated text length of command text information); To determine the number of times an n-gram appears in the text information; This represents the expected number of occurrences of an n-gram in the text. A sentence is divided into sliding window segments of n consecutive words (or characters), and each segment is called an n-gram.
[0182] Specifically, based on the number of true positives, true negatives, false positives, and false negatives, and the functions under the corresponding precision sub-dimensions mentioned above, a third evaluation attribute under at least one precision sub-dimension is determined.
[0183] Command decision propensity analysis was performed on the command decision output data and command decision expectation data corresponding to multiple test samples to determine the first decision probability distribution data corresponding to the command decision output data and the second decision probability distribution data corresponding to the command decision expectation data. Based on multiple first decision probability distribution data and multiple second decision probability distribution data, as well as the function used to determine the fourth evaluation attribute, the fourth evaluation attribute under the decision propensity sub-dimension was determined.
[0184] Based on multiple command text information, corresponding expected text information, and the function used to determine the fifth evaluation attribute, the fifth evaluation attribute under the text similarity sub-dimensional is obtained. The fourth evaluation attribute, the fifth evaluation attribute, and at least one third evaluation attribute are weighted and summed to determine the evaluation attribute to be fused corresponding to the output data accuracy dimension. Optionally, the evaluation attribute to be fused corresponding to the output data accuracy dimension can be obtained using the following function:
[0185] ;
[0186] Where A represents the attribute to be evaluated in the accuracy dimension of the output data, and the value of A ranges from [0,1]. The larger the value of A, the stronger the performance of the large command and decision-making model algorithm to be evaluated, indicating that it can provide reliable basic support for the effectiveness of command and decision-making. ACC represents the third evaluation attribute under the model accuracy sub-dimension, namely accuracy. This represents the weighting coefficient corresponding to accuracy; P represents the third evaluation attribute under the model accuracy sub-dimension, namely accuracy. Represents the weight coefficient corresponding to precision; R represents the third evaluation attribute under the model recall sub-dimension, namely recall. Represents the weight coefficient corresponding to recall; Err represents the third evaluation attribute under the model error rate sub-dimension, i.e., error rate; The weight coefficients corresponding to the error rate are represented; F1 represents the third evaluation attribute under the harmonic mean sub-dimension, i.e., the harmonic mean. This represents the weighting coefficient corresponding to the harmonic mean; This represents the fourth evaluation attribute under the decision-making tendency sub-dimension, namely the KL divergence. The weight coefficients corresponding to the KL divergence are represented; AUC represents the third evaluation attribute under the model classification sub-dimension, i.e., the area under the ROC curve. The area under the ROC curve represents the weighting coefficient; AP represents the third evaluation attribute under the average precision sub-dimension, i.e., the average precision. This represents the weighting coefficient corresponding to the average precision; G represents the third evaluation attribute under the model gain value sub-dimension, i.e., the gain value, normalized to [0,1]. This represents the weighting coefficient corresponding to the gain value; Indicates the fifth evaluation attribute; This represents the weighting coefficient corresponding to the fifth evaluation attribute; , It can be dynamically adjusted according to specific command and decision-making needs.
[0187] Optionally, the data compliance dimension includes a data quality control sub-dimension and a data compliance sub-dimension. The attributes to be evaluated under the data compliance dimension can be determined as follows: Perform data verification processing on the command and decision output data corresponding to multiple test samples to determine the data integrity attribute, data accuracy attribute, and data standardization attribute; perform weighted summation on the data integrity attribute, data accuracy attribute, and data standardization attribute to obtain the sixth evaluation attribute under the data quality control sub-dimension; perform evaluation processing on multiple test samples and the corresponding data to be used to determine the number of unauthorized data accesses and the total number of data accesses; determine the seventh evaluation attribute under the data compliance sub-dimension based on the number of unauthorized data accesses and the total number of data accesses; and perform a weighted summation on the sixth and seventh evaluation attributes to determine the attribute to be evaluated under the data compliance dimension.
[0188] The data compliance dimension is used to assess the data standardization and security of the data involved in the process of generating command and decision output data based on test samples for the large-scale command and decision model under evaluation. The data quality control sub-dimension is used to quantitatively evaluate the accuracy, completeness, and standardization of the command and decision output data. Correspondingly, the data completeness attribute characterizes the completeness of the command and decision output data; the data accuracy attribute characterizes the accuracy of the command and decision output data; and the data standardization attribute characterizes the standardization of the command and decision output data. The sixth evaluation attribute is the weighted sum of the data completeness, data accuracy, and data standardization attributes.
[0189] Unauthorized data access counts can be understood as the number of times unauthorized data access occurs during the process of the large-scale command and decision-making model under evaluation generating command and decision-making output data based on test samples. Total data access counts can be understood as the total number of data accesses during the process of the large-scale command and decision-making model under evaluation generating command and decision-making output data based on test samples. The seventh evaluation attribute can be used to characterize the compliance level during the process of the large-scale command and decision-making model under evaluation generating command and decision-making output data based on test samples.
[0190] Specifically, for command and decision output data corresponding to multiple test samples within the same type of command and decision simulation scenario, each command and decision output data is validated to determine the number of missing data characters corresponding to each output data. The missing data ratio is determined based on the proportion of missing data characters to the total number of command and decision output data characters. The complement of the missing data ratio is used as the unused complete attribute corresponding to the command and decision output data. All unused complete attributes are averaged to obtain the data completeness attribute.
[0191] Accordingly, based on the data accuracy verification rules, the accuracy of each command decision output data is verified to obtain the accuracy attribute to be used corresponding to each command decision output data. The accuracy attribute to be used can be represented by an accuracy rate. The mean of the accuracy attributes to be used for all command decision output data is then processed to determine the data accuracy attribute.
[0192] Accordingly, based on the data standardization verification rules, standardization verification is performed on each command and decision output data to obtain the standardization attributes to be used for each command and decision output data. The standardization attributes to be used can be represented by the ratio of the number of standard characters in the command and decision output data to the total number of characters in the command and decision output data. The average of the standardization attributes to be used for all command and decision output data is then processed to determine the data standardization attributes.
[0193] The data completeness attribute, data accuracy attribute, and data standardization attribute are weighted and summed to obtain the sixth evaluation attribute under the data quality control sub-dimension. Optionally, the sixth evaluation attribute can be determined by the following function:
[0194] ;
[0195] in, Indicates data integrity attributes, This represents the weight coefficient corresponding to the data integrity attribute. Indicates the accuracy attribute of the data. This represents the weighting coefficient corresponding to the accuracy attribute of the data. Indicates data specification attributes, This represents the weight coefficient corresponding to the data specification attribute. Q represents the sixth evaluation attribute, with a value range of [0,1]. The larger the Q value, the better the quality of the command and decision output data, and the more realistic the model evaluation results. It should be noted that... , , , The default value is 1 / 3, which can be adjusted according to actual needs.
[0196] Based on multiple test samples and the corresponding data to be used, determine the number of unauthorized data accesses and the total number of data accesses during the process of the command and decision-making large-scale model to be evaluated generating command and decision-making output data based on the test samples. Substitute the number of unauthorized data accesses and the total number of data accesses into the seventh evaluation attribute determination function below to obtain the seventh evaluation attribute:
[0197] C = 1 - Number of unauthorized data accesses / Total number of data accesses;
[0198] Wherein, C represents the seventh evaluation attribute, with a value range of [0,1]. The larger C is, the stronger the compliance. When a serious security problem occurs, C=0, to ensure that the process of generating command and decision output data based on the test sample by the large command and decision model to be evaluated complies with data security management regulations.
[0199] The sixth and seventh evaluation attributes are weighted and summed to determine the attributes to be evaluated under the data compliance dimension. Optionally, the attributes to be evaluated under the data compliance dimension can be determined using the following function:
[0200] ;
[0201] Where S represents the attribute to be evaluated under the data compliance dimension, with a value range of [0,1]. The larger S is, the more standardized and secure the model processing of the command and decision-making big model to be evaluated is. Q represents the sixth evaluation attribute. This is the weighting coefficient corresponding to the sixth evaluation attribute, and C represents the seventh evaluation attribute. These are the weighting coefficients corresponding to the seventh evaluation attribute. It should be noted that... In command and decision-making scenarios requiring a high level of compliance, the default is... 0.4 The value is 0.6, while the default value for typical scenarios is 0.5. By evaluating the data compliance dimension, it can be determined that the model processing procedure complies with data security management regulations, avoids unauthorized data access, and meets the security requirements of command and decision-making.
[0202] S240. Based on the evaluation attributes to be integrated under at least some evaluation dimensions, determine the target evaluation attributes of the large command and decision-making model to be evaluated in the command and decision-making simulation scenario.
[0203] Optionally, based on the above, the following weighted summation is performed on the attributes to be evaluated under the data adaptability dimension, the interference resistance dimension, the output data accuracy dimension, and the data compliance dimension to obtain the target evaluation attributes:
[0204] ;
[0205] in, This represents the target evaluation attribute, with a value range of [0,1]. The larger the value, the better the command and decision-making effectiveness of the large-scale command and decision-making model to be evaluated, and the better it meets the requirements for adapting to command and decision-making scenarios. E represents the attribute to be fused and evaluated under the data adaptability dimension. R represents the weight coefficient under the data fit dimension, and R represents the attribute to be fused and evaluated under the interference resistance dimension. The values represent the weighting coefficients under the anti-interference dimension, and A represents the attribute to be fused and evaluated corresponding to the accuracy dimension of the output data. This represents the weighting coefficient corresponding to the accuracy dimension of the output data, and S represents the attribute to be integrated and evaluated under the data compliance dimension. This represents the weighting coefficient under the data compliance dimension. The above weighting coefficients can be adjusted according to actual needs.
[0206] S250. Based on the target evaluation attributes corresponding to at least one type of command and decision simulation scenario, determine the target evaluation result of the large command and decision model to be evaluated.
[0207] Optionally, if the model performance evaluation is performed only for one type of command and decision-making simulation scenario, then the target evaluation attribute is the target evaluation result. If the model performance evaluation is performed for multiple types of command and decision-making simulation scenarios, then the target evaluation result is obtained by weighted summation of multiple target evaluation attributes.
[0208] S260. When the target evaluation results meet the preset conditions, the large-scale command and decision model to be evaluated is iteratively optimized to obtain the target command and decision model.
[0209] The preset conditions can be pre-set conditions that require adjustment and optimization of the command and decision-making model to be evaluated. For example, if the target evaluation result is less than the preset evaluation threshold, or if the target evaluation result is not suitable for the latest command and decision-making scenario, then the command and decision-making model to be evaluated will be iteratively optimized. The target command and decision-making model can be the optimized command and decision-making model.
[0210] Specifically, based on the standardized requirements for the selection, research, iteration, and application of generative artificial intelligence command and decision-making models, the target evaluation results are graded to determine the model effectiveness level to which the target evaluation results belong. An evaluation report is generated based on the model effectiveness level of the target evaluation results and the evaluation attributes corresponding to each evaluation dimension. This evaluation report explains the strengths and weaknesses of the command and decision-making model under evaluation across various evaluation dimensions, identifying problems such as decision-making bias and insufficient anti-interference capabilities in command and decision-making simulation scenarios. Based on the evaluation report, it is determined whether the current command and decision-making model under evaluation needs iterative optimization. This iterative optimization process yields the target command and decision-making model. Based on the above, a closed-loop mechanism for iterative research is formed, which is conducive to continuously improving the effectiveness and adaptability of the command and decision-making model.
[0211] The technical solution of this embodiment obtains multiple test samples under at least one type of command and decision simulation scenario; processes each test sample in the test sample set based on the large command and decision model to be evaluated to obtain the data to be used corresponding to the test sample. This ensures that the data to be used fits the actual command and decision scenario, and solves the problem in the prior art that the evaluation results are not in line with the command and decision scenario and the model evaluation is distorted due to the difficulty in obtaining model processing data under the real command and decision scenario. The data to be used provides reliable data support for the subsequent evaluation of the large model. Based on the test samples and evaluation dimensions of multiple test samples belonging to the same type of command and decision-making simulation scenario, the large-scale command and decision-making model to be evaluated is evaluated to determine the evaluation attributes to be integrated under each evaluation dimension. Based on the evaluation attributes to be integrated under at least some evaluation dimensions, the target evaluation attributes of the large-scale command and decision-making model under evaluation in the command and decision-making simulation scenario are determined. This realizes the evaluation of the interpretability, robustness, and compliance of the large-scale command and decision-making model under evaluation, quantifies the output stability and decision transparency of the large-scale command and decision-making model under evaluation in the corresponding command and decision-making scenario, solves the problem that the existing evaluation dimensions are too single and unsuitable for the corresponding command and decision-making scenario, and realizes the accurate measurement of the model performance of the large-scale command and decision-making model under evaluation in the command and decision-making scenario. Based on the target evaluation attributes corresponding to at least one type of command and decision-making simulation scenario, the target evaluation result of the large-scale command and decision-making model under evaluation is determined, realizing the quantitative evaluation of the model performance of the large-scale command and decision-making model under evaluation in at least one type of command and decision-making scenario, so that the large-scale command and decision-making model adjusted based on the target evaluation result can be adapted to applications in various command and decision-making scenarios. When the target evaluation results meet preset conditions, the large-scale command and decision-making model to be evaluated undergoes iterative optimization to obtain the target large-scale command and decision-making model. A closed-loop iterative mechanism is constructed to support the continuous optimization of the large-scale command and decision-making model, which is conducive to continuously improving its effectiveness and adaptability. This invention solves the problems of difficulty in obtaining evaluation datasets, overly singular evaluation dimensions, and inability to adapt to command and decision-making scenarios in existing technologies. It ensures the effectiveness and accuracy of the large-scale model evaluation, and the evaluation results can guide the adjustment direction of the large-scale command and decision-making model, enabling it to adapt to the application needs of various command and decision-making scenarios.
[0212] Example 3
[0213] Figure 3 This is a schematic diagram of the structure of a device for evaluating the capabilities of a large command and decision-making model, provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: a test sample acquisition module 310, a sample processing module 320, a large model evaluation module 330, a target evaluation attribute determination module 340, and a target evaluation result determination module 350.
[0214] The test sample acquisition module 310 is used to acquire a test sample set, wherein the test sample set includes multiple test samples under at least one type of command and decision simulation scenario. Each test sample includes at least: scenario information and first text information. The scenario information includes at least: environmental information, situational information, interference information, and command information. The first text information is used to characterize the task processing instructions under the command and decision simulation scenario. The sample processing module 320 is used to process each test sample in the test sample set based on the large command and decision simulation model to be evaluated, obtaining the data to be used corresponding to the test sample. The data to be used includes: command and decision output data and model running data. The large model evaluation module 330 is used to evaluate data based on the same type of command and decision simulation scenario. The system uses multiple test samples corresponding to the data to be used and multiple pre-defined evaluation dimensions to evaluate the command and decision-making model to be evaluated, and determines the fusion evaluation attributes of the command and decision-making model under each evaluation dimension. The multiple evaluation dimensions include: data adaptability dimension, anti-interference dimension, output data accuracy dimension, and data compliance dimension. A target evaluation attribute determination module 340 is used to determine the target evaluation attributes of the command and decision-making model under the command and decision-making simulation scenario based on the fusion evaluation attributes under at least some of the evaluation dimensions. A target evaluation result determination module 350 is used to determine the target evaluation result of the command and decision-making model to be evaluated based on the target evaluation attributes corresponding to at least one type of command and decision-making simulation scenario.
[0215] The technical solution of this embodiment obtains multiple test samples under at least one type of command and decision simulation scenario; processes each test sample in the test sample set based on the large command and decision model to be evaluated to obtain the data to be used corresponding to the test sample. This ensures that the data to be used fits the actual command and decision scenario, and solves the problem in the prior art that the evaluation results are not in line with the command and decision scenario and the model evaluation is distorted due to the difficulty in obtaining model processing data under the real command and decision scenario. The data to be used provides reliable data support for the subsequent evaluation of the large model. Based on the test samples and evaluation dimensions of multiple test samples belonging to the same type of command and decision-making simulation scenario, the large-scale command and decision-making model to be evaluated is evaluated to determine the evaluation attributes to be integrated under each evaluation dimension. Based on the evaluation attributes to be integrated under at least some evaluation dimensions, the target evaluation attributes of the large-scale command and decision-making model under the command and decision-making simulation scenario are determined. This achieves the evaluation of the interpretability, robustness, and compliance dimensions of the large-scale command and decision-making model under evaluation, quantifies the output stability and decision transparency of the large-scale command and decision-making model under evaluation in the corresponding command and decision-making scenario, and solves the problem that existing evaluation dimensions are too singular and unsuitable for corresponding command and decision-making scenarios. This enables accurate measurement of the model performance of the large-scale command and decision-making model under evaluation in command and decision-making scenarios. Based on the target evaluation attributes corresponding to at least one type of command and decision-making simulation scenario, the target evaluation result of the large-scale command and decision-making model under evaluation is determined, realizing the model performance evaluation of the large-scale command and decision-making model under evaluation in at least one type of command and decision-making scenario. This allows the large-scale command and decision-making model adjusted based on the target evaluation result to be adaptable to applications in various command and decision-making scenarios. This invention solves the problems of difficulty in obtaining evaluation datasets, overly simplistic evaluation dimensions, and inability to adapt to command and decision-making scenarios in existing technologies. It ensures the effectiveness and accuracy of large-scale model evaluation and can guide the adjustment direction of the large-scale command and decision-making model based on the evaluation results, so that the large-scale command and decision-making model can adapt to the application needs of various command and decision-making scenarios.
[0216] Based on the above embodiments, optionally, the large model evaluation module includes: a data fit dimension evaluation submodule, which includes: a data information determination unit, used to determine, based on the command decision output data corresponding to multiple test samples in the same type of command decision simulation scenario, the target heterogeneity ratio, target judgment coefficient, target feature contribution attribute, target dispersion coefficient, and target overlap information corresponding to the large command decision model to be evaluated; wherein, the target heterogeneity ratio is used to characterize the stability of the command decision output data in the same type of command decision simulation scenario, and the target judgment coefficient is used to characterize the fit between the command decision output data and the test samples. The target feature contribution attribute is used to characterize the causal logical correlation between the command decision output data and the scenario information; the target dispersion coefficient is used to characterize the rationality of the distribution of tactical types in the command decision output data under the same type of command decision simulation scenario; the target overlap information is used to characterize the semantic conversion accuracy of the command decision large model to be evaluated; the data fit dimension evaluation attribute determination unit is used to determine the fusion evaluation attributes of the command decision large model to be evaluated in the data fit dimension based on the target heterogeneity ratio, the target judgment coefficient, the target feature contribution attribute, the target dispersion coefficient, the target overlap information, and the first evaluation function.
[0217] Optionally, the data information determination unit includes: a target determination coefficient determination subunit, used to determine the target determination coefficient based on the command decision output data and corresponding command decision expectation data corresponding to multiple test samples under the same type of command decision simulation scenario; wherein, the command decision expectation data is the expected output data determined based on the test sample corresponding to the command decision output data; and a target feature contribution attribute determination subunit, used to perform scene feature correlation analysis on multiple test samples under the same type of command decision simulation scenario, on the command decision output data corresponding to the test sample and multiple scene features to be processed corresponding to the scene information, and determine the first feature attribute of each scene feature to be processed; wherein, the scene feature to be processed is a feature determined by feature extraction of the scene information; and based on the first feature attribute of each scene feature to be processed, determine a preset number of scene features to be used and the second feature attribute of the scene features to be used. Based on the first feature attribute and the second feature attribute, the feature contribution attribute to be used corresponding to the test sample is determined; based on the feature contribution attributes to be used of multiple test samples, the target feature contribution attribute is determined; the target overlap information determination subunit is used to determine, for multiple test samples under the same type of command and decision simulation scenario, a first skip word double-word group set and a first single-word set corresponding to the command text information in the command and decision output data corresponding to the test sample; based on the expected text information in the command and decision expected data, a second skip word double-word group set and a second single-word set corresponding to the expected text information are determined; based on the first skip word double-word group set, the first single-word set, the second skip word double-word group set, and the second single-word set, the first overlap information corresponding to the test sample is determined; based on the first overlap information of multiple test samples, the target overlap information is determined.
[0218] Optionally, the data information determination unit includes: a target heterogeneity ratio determination subunit, used to perform decision classification processing on command decision output data corresponding to multiple test samples under the same type of command decision simulation scenario to obtain command decision output data under at least one decision type; determine the target quantity based on the number of command decision output data under at least one decision type; wherein, the target quantity is used to characterize the number of data corresponding to the decision type with the most command decision output data; determine the target heterogeneity ratio based on the total number of command decision output data under all decision types and the target quantity; and a target dispersion coefficient determination subunit, used to perform tactical classification processing on command decision output data corresponding to multiple test samples under the same type of command decision simulation scenario to determine command decision output data under at least one tactical type; determine the standard deviation and mean value corresponding to the tactical type based on the number of command decision output data under at least one tactical type; and determine the target dispersion coefficient based on the standard deviation and mean value corresponding to the tactical type.
[0219] Optionally, the anti-interference dimension includes a disturbance fluctuation sub-dimension and a disturbance stability sub-dimension. The large model evaluation module includes an anti-interference dimension evaluation sub-module, which includes a sub-dimension evaluation unit, used to determine the first evaluation attribute under the disturbance fluctuation sub-dimension and the second evaluation attribute under the disturbance stability sub-dimension based on the data to be used and reference data corresponding to multiple test samples in the same type of command and decision simulation scenario; wherein, the reference data is obtained by the command and decision large model to be evaluated processing test samples with interference information removed; and an anti-interference dimension evaluation attribute determination unit, used to determine the evaluation attribute to be fused under the anti-interference dimension based on the first evaluation attribute, the second evaluation attribute, and the second evaluation function.
[0220] Optionally, the sub-dimension evaluation unit is configured to: determine a first processing accuracy and a first processing time corresponding to the data to be used, and a second processing accuracy and a second processing time corresponding to the reference data, based on the data to be used and reference data corresponding to multiple test samples; determine a first evaluation attribute under the disturbance fluctuation sub-dimension based on the first processing accuracy, the first processing time, the second processing accuracy, and the second processing time; perform anomaly detection on the data to be used corresponding to multiple test samples to determine at least one data to be processed; determine test sample difference information associated with the disturbance type based on the test sample to which the at least one data to be processed belongs and the test sample to which the corresponding reference data belongs; and determine a second evaluation attribute under the disturbance stability sub-dimension based on the test sample difference information associated with the disturbance type.
[0221] Optionally, the large model evaluation module further includes: a positive and negative sample partitioning unit, used to partition test samples based on multiple test samples in the same type of command and decision simulation scenario and the command and decision output data corresponding to the test samples, to obtain the number of true positives, the number of true negatives, the number of false positives, and the number of false negatives; wherein the true positives and the true negatives correspond to test samples with correct command and decision output data, and the false positives and false negatives correspond to test samples with incorrect command and decision output data.
[0222] Optionally, the output data accuracy dimension includes: a decision tendency sub-dimension, a text similarity sub-dimension, and at least one precision sub-dimension. The large model evaluation module includes: an output data accuracy dimension evaluation sub-module, used to determine a third evaluation attribute under at least one precision sub-dimension based on the number of true positives, the number of true negatives, the number of false positives, and the number of false negatives; wherein, the at least one precision sub-dimension includes one or more of the following: model accuracy sub-dimension, model precision sub-dimension, model recall sub-dimension, model error rate sub-dimension, model harmonic mean sub-dimension, model classification sub-dimension, model average precision sub-dimension, and model gain value sub-dimension; for multiple test samples corresponding to command decision output data and command decision expectation numbers... Based on separate command decision tendency analysis, a first decision probability distribution data corresponding to the command decision output data and a second decision probability distribution data corresponding to the command decision expectation data are determined. Based on multiple first decision probability distribution data and multiple second decision probability distribution data, a fourth evaluation attribute under the decision tendency sub-dimension is determined. Text similarity evaluation processing is performed on the command text information in the command decision output data and the expected text information in the command decision expectation data corresponding to multiple test samples to determine a fifth evaluation attribute under the text similarity sub-dimension. A weighted sum is performed on the fourth evaluation attribute, the fifth evaluation attribute, and at least one of the third evaluation attributes to determine the evaluation attribute to be fused corresponding to the output data accuracy dimension.
[0223] Optionally, the data compliance dimension includes: a data quality control sub-dimension and a data compliance sub-dimension. The large model evaluation module includes: a data compliance dimension evaluation sub-module, used to perform data verification processing on command and decision output data corresponding to multiple test samples to determine data integrity attributes, data accuracy attributes, and data standardization attributes; to perform weighted summation processing on the data integrity attributes, the data accuracy attributes, and the data standardization attributes to obtain the sixth evaluation attribute under the data quality control sub-dimension; to perform evaluation processing on multiple test samples and the data to be used corresponding to the test samples to determine the number of unauthorized data accesses and the total number of data accesses; to determine the seventh evaluation attribute under the data compliance sub-dimension based on the number of unauthorized data accesses and the total number of data accesses; and to perform weighted summation of the sixth evaluation attribute and the seventh evaluation attribute to determine the evaluation attribute to be integrated under the data compliance dimension.
[0224] Optionally, the device further includes a model optimization module, used to iteratively optimize the large-scale command and decision model to be evaluated when the target evaluation result meets preset conditions, so as to obtain the target command and decision model.
[0225] The apparatus for evaluating the capabilities of a large command and decision-making model provided in this embodiment of the invention can execute the method for evaluating the capabilities of a large command and decision-making model provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0226] Example 4
[0227] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0228] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0229] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0230] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as methods for evaluating the capabilities of large models of command and decision-making.
[0231] In some embodiments, the method for evaluating the capabilities of a large command and decision model can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for evaluating the capabilities of a large command and decision model described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the method for evaluating the capabilities of a large command and decision model by any other suitable means (e.g., by means of firmware).
[0232] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0233] Computer programs used to implement the method for evaluating the capabilities of large-scale command and decision-making models of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0234] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.
[0235] Example 5
[0236] Embodiment 5 of the present invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute a method for evaluating the capabilities of a large command and decision-making model, the method comprising:
[0237] Obtain a test sample set, wherein the test sample set includes multiple test samples under at least one type of command and decision simulation scenario, and the test samples include at least: scenario information and first text information, wherein the scenario information includes at least: environmental information, situation information, interference information and command information, and the first text information is used to characterize the task processing instructions under the command and decision simulation scenario;
[0238] Based on the large command and decision model to be evaluated, each test sample in the test sample set is processed to obtain the data to be used corresponding to the test sample; wherein, the data to be used includes: command and decision output data and model operation data;
[0239] Based on the data to be used corresponding to multiple test samples in the same type of command and decision simulation scenario and multiple pre-set evaluation dimensions, the command and decision large model to be evaluated is evaluated to determine the fusion evaluation attributes of the command and decision large model under each evaluation dimension; wherein, the multiple evaluation dimensions include: data adaptability dimension, anti-interference dimension, output data accuracy dimension, and data compliance dimension.
[0240] Based on the evaluation attributes to be integrated under at least some of the evaluation dimensions, determine the target evaluation attributes of the command and decision-making big model to be evaluated in the command and decision-making simulation scenario.
[0241] Based on the target evaluation attributes corresponding to at least one of the command and decision simulation scenarios, the target evaluation result of the command and decision large model to be evaluated is determined.
[0242] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0243] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0244] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0245] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0246] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the method for evaluating the capabilities of large command and decision-making models as provided in any embodiment of this application.
[0247] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider). This program product belongs to the same inventive concept as the methods for evaluating the capabilities of large command and decision-making models disclosed in the embodiments of this application, and therefore will not be described further here.
[0248] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0249] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for evaluating the capabilities of a large command and decision-making model, characterized in that, include: Obtain a test sample set, wherein the test sample set includes multiple test samples under at least one type of command and decision simulation scenario, and the test samples include at least: scenario information and first text information, wherein the scenario information includes at least: environmental information, situation information, interference information and command information, and the first text information is used to characterize the task processing instructions under the command and decision simulation scenario; Based on the large command and decision model to be evaluated, each test sample in the test sample set is processed to obtain the data to be used corresponding to the test sample; wherein, the data to be used includes: command and decision output data and model operation data; Based on the data to be used corresponding to multiple test samples in the same type of command and decision simulation scenario and multiple pre-set evaluation dimensions, the command and decision large model to be evaluated is evaluated to determine the fusion evaluation attributes of the command and decision large model under each evaluation dimension; wherein, the multiple evaluation dimensions include: data adaptability dimension, anti-interference dimension, output data accuracy dimension, and data compliance dimension. Based on the evaluation attributes to be integrated under at least some of the evaluation dimensions, determine the target evaluation attributes of the command and decision-making big model to be evaluated in the command and decision-making simulation scenario. Based on the target evaluation attributes corresponding to at least one of the command and decision simulation scenarios, the target evaluation result of the command and decision large model to be evaluated is determined.
2. The method according to claim 1, characterized in that, The attributes to be fused and evaluated in the data fit dimension of the command and decision-making big model to be evaluated include: Based on the command and decision output data corresponding to multiple test samples in the same type of command and decision simulation scenario, the target heterogeneity ratio, target judgment coefficient, target feature contribution attribute, target dispersion coefficient, and target overlap information corresponding to the command and decision large model to be evaluated are determined. Specifically, the target heterogeneity ratio characterizes the stability of the command and decision output data in the same type of command and decision simulation scenario; the target judgment coefficient characterizes the fit between the command and decision output data and the test samples; the target feature contribution attribute characterizes the causal logical correlation between the command and decision output data and the scenario information; the target dispersion coefficient characterizes the rationality of the distribution of tactical types in the command and decision output data in the same type of command and decision simulation scenario; and the target overlap information characterizes the semantic conversion accuracy of the command and decision large model to be evaluated. Based on the target heterogeneity ratio, the target determination coefficient, the target feature contribution attribute, the target dispersion coefficient, the target overlap information, and the first evaluation function, the attributes to be integrated and evaluated in the data fit dimension of the command and decision-making big model to be evaluated are determined.
3. The method according to claim 2, characterized in that, The target determination coefficient is determined in the following manner: The target determination coefficient is determined based on the command decision output data and corresponding command decision expectation data corresponding to multiple test samples in the same type of command decision simulation scenario; wherein, the command decision expectation data is the expected output data determined based on the test samples corresponding to the command decision output data. Accordingly, the target feature contribution attribute is determined in the following manner: For multiple test samples in the same type of command and decision simulation scenario, a scenario feature correlation analysis is performed on multiple unprocessed scenario features corresponding to the command and decision output data and scenario information of the test samples to determine the first feature attribute of each unprocessed scenario feature; wherein, the unprocessed scenario feature is the feature determined by feature extraction of the scenario information; Based on the first feature attribute of each scene feature to be processed, a preset number of scene features to be used and the second feature attribute of the scene features to be used are determined; Based on the first feature attribute and the second feature attribute, determine the feature contribution attribute to be used corresponding to the test sample; Determine the target feature contribution attribute based on the feature contribution attributes to be used from multiple test samples; Accordingly, the target overlap information corresponding to the large command and decision model to be evaluated is determined in the following manner: For multiple test samples in the same type of command and decision simulation scenario, the first set of two word groups and the first set of single words corresponding to the command text information in the command and decision output data corresponding to the test sample are determined. Based on the expected text information in the command decision expectation data, determine the second jump word double word set and the second single word set corresponding to the expected text information; Based on the first set of skip word pairs, the first set of single-character words, the second set of skip word pairs, and the second set of single-character words, determine the first overlap information corresponding to the test sample; The target overlap information is determined based on the first overlap information of multiple test samples.
4. The method according to claim 2, characterized in that, The target dissimilar ratio is determined in the following manner: The command and decision output data corresponding to multiple test samples under the same type of command and decision simulation scenario are classified and processed to obtain command and decision output data under at least one decision type. A target quantity is determined based on the number of command decision output data under at least one decision type; wherein, the target quantity is used to characterize the number of data corresponding to the decision type with the most command decision output data; The target heterogeneity ratio is determined based on the total number of command decision output data under all decision types and the number of targets. Accordingly, the target discrete coefficients are determined in the following manner: Tactical classification processing is performed on the command decision output data corresponding to multiple test samples under the same type of command decision simulation scenario to determine the command decision output data under at least one tactical type. Based on the quantity of command decision output data under at least one of the tactical types, determine the standard deviation and mean corresponding to the tactical type; The target dispersion coefficient is determined based on the standard deviation and mean value corresponding to the tactical type.
5. The method according to claim 1, characterized in that, The anti-interference dimension includes a disturbance fluctuation sub-dimension and a disturbance stability sub-dimension. The attributes to be fused and evaluated in the anti-interference dimension of the large command and decision-making model to be evaluated are determined, including: Based on the data to be used and the reference data corresponding to multiple test samples in the same type of command and decision simulation scenario, the first evaluation attribute under the disturbance fluctuation sub-dimension and the second evaluation attribute under the disturbance stability sub-dimension are determined; wherein, the reference data is obtained by the command and decision large model to be evaluated by processing the test samples after removing interference information; Based on the first evaluation attribute, the second evaluation attribute, and the second evaluation function, the evaluation attribute to be fused under the anti-interference dimension is determined.
6. The method according to claim 5, characterized in that, The process of determining the first evaluation attribute under the disturbance fluctuation sub-dimension and the second evaluation attribute under the disturbance stability sub-dimension based on the data to be used and reference data corresponding to multiple test samples in the same type of command and decision simulation scenario includes: Based on the data to be used and reference data corresponding to multiple test samples, determine the first processing accuracy and the first processing time corresponding to the data to be used, and the second processing accuracy and the second processing time corresponding to the reference data; Based on the first processing accuracy, the first processing time, the second processing accuracy, and the second processing time, determine the first evaluation attribute under the perturbation fluctuation sub-dimension; Anomaly detection is performed on the data to be used corresponding to multiple test samples to identify at least one data to be processed. Based on at least one test sample to which the data to be processed belongs and the test sample to which the corresponding reference data belongs, determine the test sample difference information associated with the disturbance type; A second evaluation attribute under the perturbation stability sub-dimension is determined based on at least one test sample difference information associated with the perturbation type.
7. The method according to claim 1, characterized in that, Before determining the attributes to be fused and evaluated in the dimension of output data accuracy of the large command and decision model to be evaluated, the method further includes: Based on multiple test samples in the same type of command and decision simulation scenario and the command and decision output data corresponding to the test samples, the test samples are divided to obtain the number of true positives, the number of true negatives, the number of false positives, and the number of false negatives. Wherein, the true positive and the true negative examples correspond to test samples where the command decision output data is correct, and the false positive and the false negative examples correspond to test samples where the command decision output data is incorrect.
8. The method according to claim 7, characterized in that, The output data accuracy dimension includes: a decision-making tendency sub-dimension, a text similarity sub-dimension, and at least one precision sub-dimension. The attributes to be fused and evaluated in the output data accuracy dimension of the large-scale command decision-making model to be evaluated are determined, including: Based on the number of true positives, the number of true negatives, the number of false positives, and the number of false negatives, a third evaluation attribute under at least one precision sub-dimension is determined; wherein, the at least one precision sub-dimension includes one or more of the following: model accuracy sub-dimension, model precision sub-dimension, model recall sub-dimension, model error rate sub-dimension, model harmonic mean sub-dimension, model classification sub-dimension, model average precision sub-dimension, and model gain value sub-dimension. Command decision tendency analysis is performed on the command decision output data and command decision expectation data corresponding to multiple test samples to determine the first decision probability distribution data corresponding to the command decision output data and the second decision probability distribution data corresponding to the command decision expectation data. Based on multiple first decision probability distribution data and multiple second decision probability distribution data, determine the fourth evaluation attribute under the decision tendency sub-dimension; Text similarity evaluation is performed on the command text information and the expected text information of the command decision output data corresponding to multiple test samples, and the fifth evaluation attribute under the text similarity sub-dimension is determined. The fourth evaluation attribute, the fifth evaluation attribute, and at least one of the third evaluation attributes are weighted and summed to determine the evaluation attribute to be fused corresponding to the accuracy dimension of the output data.
9. The method according to claim 1, characterized in that, The data compliance dimension includes: a data quality control sub-dimension and a data compliance sub-dimension. The attributes to be integrated and evaluated under the data compliance dimension of the command and decision-making big model to be evaluated are determined, including: Data verification processing was performed on the command and decision output data corresponding to multiple test samples to determine the data integrity attribute, data accuracy attribute, and data standardization attribute; The data integrity attribute, the data accuracy attribute, and the data standardization attribute are weighted and summed to obtain the sixth evaluation attribute under the data quality control sub-dimension. Multiple test samples and the corresponding data to be used are evaluated to determine the number of unauthorized data accesses and the total number of data accesses. The seventh evaluation attribute under the data compliance sub-dimension is determined based on the number of unauthorized accesses and the total number of data accesses. The sixth and seventh evaluation attributes are weighted and summed to determine the evaluation attributes to be integrated under the data compliance dimension.
10. The method according to claim 1, characterized in that, After determining the target evaluation result of the command and decision-making big data model to be evaluated, the method further includes: When the target evaluation result meets the preset conditions, the large-scale command and decision model to be evaluated is iteratively optimized to obtain the target command and decision model.