Military large model maturity evaluation method based on enhanced ANP-FCE

By constructing a performance evaluation index system for military large models based on enhanced ANP-FCE, and utilizing the ANP model and fuzzy comprehensive evaluation method, the interpretability and reliability issues of military large models are solved, achieving scientific maturity evaluation and improved deployment effectiveness.

CN121502256APending Publication Date: 2026-02-10SYST OVERALL RES INST INST OF SYST ENG ACAD OF MILITARY SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511668364.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing military large-scale models lack interpretability and reliability, and their application and deployment costs are high, with a lack of effective maturity evaluation methods.

Method used

A performance evaluation index system for a military big data model is constructed using a reinforced ANP-FCE approach, including indicators such as usability, credibility, reliability, interpretability, and controllability. The model is evaluated through the ANP model framework, and the maturity level is determined using a fuzzy comprehensive evaluation method.

Benefits of technology

It enabled a comprehensive evaluation of the large military model, improved the model's interpretability and reliability, reduced deployment costs, and enhanced the scientific rigor and accuracy of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502256A_ABST
    Figure CN121502256A_ABST
Patent Text Reader

Abstract

The invention provides a military large model evaluation method based on an enhanced ANP-FCE model. In a control layer, military large model efficiency is taken as a control target, military demand traction and intelligent technology promotion are taken as two major criteria, and logic consistency of top-down decomposition reduction and bottom-up aggregation verification is realized. In a network layer, five index groups including availability, credibility, reliability, interpretability and controllability and eight index elements are set, deep coupling and complex linking of ecology in a large model technology are considered, all the indexes are subjected to full-interaction analysis, the influence relation of all the index elements is traversed through intensified calculation, and logic conflicts of span advantages among the indexes are eliminated. Through a weight relation of logic science, five ability levels including a starting level, a development level, a robust level, an excellent level and an excellent level are mapped, and effective evaluation of the maturity of the military large model is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a maturity evaluation method for large military models based on reinforced ANP-FCE. Background Technology

[0002] Generative artificial intelligence (AIGC), represented by multimodal large language models (MLLMs), learns data patterns and rules based on deep neural networks and appropriately generalizes to generate low-cost, high-efficiency, and high-quality original content. It has rapidly become a new growth point for generating new combat capabilities. With the widespread application of large language models, the evaluation of large language models has become crucial.

[0003] The importance of evaluation in the development and application of large-scale military models lies in the following: First, the interpretability of large-scale military models is not strong, and evaluation is needed to improve interpretability; second, the reliability of large-scale military models cannot be effectively guaranteed, and evaluation is needed to improve trustworthiness; third, the application and deployment of large-scale military models is very costly, and evaluation is needed to improve efficiency.

[0004] Therefore, how to evaluate the maturity of large-scale military models has become a technical problem that urgently needs to be solved by those in the field. Summary of the Invention

[0005] Therefore, it is necessary to provide a maturity evaluation method for military large models based on enhanced ANP-FCE to address the aforementioned technical problems.

[0006] A maturity assessment method for large military models based on enhanced ANP-FCE, the method comprising: S10, Constructing an effectiveness evaluation index system for a large military model; wherein, the effectiveness evaluation index system includes five index groups: usability, credibility, reliability, interpretability, and controllability, and each primary index group contains multiple index elements; S20. Based on the objectives and criteria of the military big model, and the effectiveness evaluation index system, establish the ANP model structure framework; wherein, the ANP model structure framework includes a control layer and a network layer, the control layer is constructed based on the objectives and criteria, and the network layer is constructed based on the effectiveness evaluation index system, the objective is the comprehensive effectiveness of the military big model, and the criteria are military demand-driven and intelligent technology-driven. S30, using the criteria in the control layer as the primary criteria, for each primary criterion, using an indicator element in a certain indicator group in the network layer as the secondary criterion, comparing the influence degree of the indicator elements in another indicator group pairwise, constructing an evaluation judgment matrix, and calculating the first ranking vector based on the evaluation judgment matrix. S40, construct an unweighted hypermatrix based on the first sorting vector; S50, under each main criterion, compare the importance of each indicator group to other indicator groups to obtain the corresponding second ranking vector, and construct a weighting matrix based on the second ranking vector; S60, the unweighted supermatrix is ​​weighted according to the weighted matrix to obtain a weighted supermatrix; S70, perform self-multiplication on the weighted hypermatrix until the product converges to obtain the limiting hypermatrix, and extract column vectors from the limiting hypermatrix as the final weights of each indicator element. S80, construct the factor set and evaluation set for fuzzy comprehensive evaluation, generate a weight set according to the final weights, perform single-factor evaluation on each factor, obtain the membership degree matrix of each factor to the evaluation set, calculate the membership degree of the overall target of fuzzy comprehensive evaluation according to the membership degree matrix and the weight set, and determine the evaluation result of the maturity of the military big model according to the principle of maximum membership degree.

[0007] In one embodiment, S30 includes: Using the criteria in the control layer as the primary criteria, for each primary criterion, an indicator element from a certain indicator group in the network layer is used as a secondary criterion to compare the influence of the indicator elements in another indicator group pairwise. The Delphi method is used to construct the evaluation judgment matrix by expert scoring and a 1-9 scale method.

[0008] In one embodiment, the pairwise comparison of the influence of the index elements in another index group specifically includes: Indirect dominance comparison is performed on the degree of influence of the various indicator elements in another indicator group.

[0009] In one embodiment, S30 further includes: Based on the evaluation judgment matrix, the corresponding first sorting vector is calculated using the eigenvalue method.

[0010] In one embodiment, S40 includes: In the main principle The unweighted hypermatrix is ​​as follows: in, This represents an unweighted supermatrix.

[0011] In one embodiment, the weighting matrix is ​​as follows: in, Represents a weighted matrix. This represents the second sorting vector.

[0012] In one embodiment, S60 includes: For unweighted hypermatrix Perform weighted normalization, then let We can obtain: ; in, A weighted hypermatrix is ​​a column random matrix where the sum of the elements in each column is 1.

[0013] In one embodiment, calculating the overall objective membership degree of the fuzzy comprehensive evaluation based on the membership matrix and the weight set includes: Calculate the membership degree of the overall objective in the fuzzy comprehensive evaluation: ; in, This represents the membership degree of the overall objective in the fuzzy comprehensive evaluation. Represents the set of weights. This represents the membership matrix.

[0014] The aforementioned military large-scale model evaluation method based on the enhanced ANP-FCE model, at the control layer, uses the effectiveness of the military large-scale model as the control objective and military demand-driven and intelligent technology-driven as the two main criteria, achieving logical consistency between top-down decomposition and bottom-up aggregation and verification. At the network layer, it sets up 18 indicator elements in 5 indicator groups (availability, credibility, reliability, interpretability, and controllability), focusing on the deep coupling and complex links within the large-scale model's internal technological ecosystem. It performs full interactive analysis on all indicators, eliminating logical conflicts arising from the advantages of different indicators by traversing the influence relationships of all indicator elements through enhanced computation. Through logically sound weight relationships, it maps five capability levels—starting level, development level, robust level, excellent level, and outstanding level—to achieve an effective assessment of the maturity of the military large-scale model. Attached Figure Description

[0015] Figure 1 A schematic diagram of the military large model ANP evaluation framework provided in one embodiment of the present invention; Figure 2 This is a weighted ranking diagram of various indicators provided in one embodiment of the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0017] In one embodiment, a maturity assessment method for a large military model based on enhanced ANP-FCE is provided, which specifically includes the following steps: S10, a performance evaluation index system for building a large-scale military model.

[0018] The performance evaluation index system includes five primary index groups (availability, credibility, reliability, interpretability, and controllability) and 18 secondary index elements.

[0019] Specifically, establishing a scientific and reasonable indicator system is the prerequisite and foundation for completing performance evaluation. Through systematic analysis of the needs of different fields of military big data models in empowering combat effectiveness, equipment systems, and national defense and military construction, this paper establishes five indicator groups: usability, credibility, reliability, interpretability, and controllability. These indicators include the main tasks of the military big data model (classification, regression, clustering, detection, etc.) and sub-tasks such as text classification, named entity recognition, information extraction, mathematical reasoning, causal reasoning, and common sense reasoning. The performance of the military big data model's output is considered, along with comprehensive factors such as basic support services, algorithm indicators, and security protection. Table 1 shows the performance evaluation indicator system for the military big data model.

[0020] Table 1. Military Large Model Effectiveness Evaluation Index System Availability. The availability of large-scale military models refers to their ability to be widely accessed and used across military verticals. This includes the ability to support visual drag-and-drop layout programming services for manipulating AI services, and the ability to combine various data sources, components, algorithms, models, and evaluation modules. Key aspects include: distributed storage capabilities (supported storage capacity, number of data entries, etc., and support for vector databases); support for multiple data types and programming languages; support for multiple storage protocols such as NFS, SMB, FTP, HDFS, S3, and POSIX, with each protocol maintaining the same semantic compatibility as its native counterpart; support for distributed architecture cluster systems with load awareness and reduced operational complexity; implementation of data replication or erasure coding redundancy strategies; and tenant-level QoS, quota, user authentication domain, and business network segment isolation. High-performance computing capabilities are also crucial, supporting complex computing tasks such as ultra-large-scale distributed computing, batch computing, streaming computing, graph computing, machine learning computing, and edge computing. The theoretical computing power of computing devices used for training tasks should be at least 50 PFLOPS (FP16 precision) for models with tens of billions of tokens, and at least 1000 PFLOPS (FP16 precision) for models with hundreds of billions or trillions of tokens. For computing devices used for inference tasks with high real-time requirements, the average token output latency should not exceed 100ms. For computing devices used for inference tasks with lower real-time requirements, the average token output latency should not exceed 300ms. Under inference tasks, the total throughput of a single GPU for models with tens of billions of tokens should not be less than 5 tokens / s. Regarding the efficiency of basic support services, the development and operation of large-scale military models require a robust hardware and software infrastructure environment for storage, computation, data structuring, and decision-making. Therefore, large-scale models, in addition to requiring certain computational performance from training chips, also have high requirements for hardware specifications, such as memory size, memory access bandwidth, and communication bandwidth. To achieve efficient training and inference of large-scale models, it is necessary to achieve hardware adaptation and deep collaborative optimization through deep learning frameworks, using low-cost, high-efficiency hardware adaptation solutions. Efficiency evaluation should include, but is not limited to, the following metrics: Average processing time: Used to define and evaluate the time consumption of a deep learning algorithm model in processing the same task under the same test environment. During the testing phase, it includes test metrics such as the execution time of a single training epoch, the execution time of multiple training epochs, and the execution time to reach a specific level of accuracy. Average resource cost: Used to define and evaluate the amount of resources consumed by a deep learning algorithm model in processing the same task under the same test environment. During the testing phase, it includes test metrics such as computational power consumption, storage consumption, and bandwidth consumption during algorithm execution.

[0021] Credibility. The credibility of a large-scale military model refers to the degree to which users and decision-makers trust and accept the model's performance in military mission scenarios. Key metrics for accuracy include: Precision (the proportion of correctly predicted samples in the training set); Accuracy (the proportion of samples predicted as positive that actually belong to the positive class); Recall (the ratio of correctly predicted positive samples in the training set); and Error Rate (the ratio of incorrectly predicted samples to the total number of samples in a given dataset). The F1 score is the harmonic mean of precision and recall, a metric for measuring the precision of a binary classification model, balancing both accuracy and recall. KL divergence is an asymmetric measure of the difference between two probability distributions, comparing the difference between the true distribution and the theoretical (fitted) distribution. Characteristic curves mainly include: ROC curves (Receiving Operator Characteristic curves), which are response curves plotted from the true positive rate and false positive rate under different settings, and are comprehensive indicators reflecting continuous variables of sensitivity and specificity; PRC curves (Precision and Recall curves), which are graphical methods to simultaneously display the precision and recall of deep learning algorithms at different thresholds. Generally, the x-axis represents recall and the y-axis represents precision; and CRC curves (Cumulative Response curves), also known as gain curves or gain plots, are graphical methods to display the true positive rate and the percentage of positive predictions in total data across multiple thresholds.

[0022] Reliability, also known as robustness, in the context of large-scale military models refers to the ability of a deep learning algorithm to maintain performance comparable to its experimental performance when facing various uncertainties and anomalies, particularly when dealing with non-adversarial augmented samples. It is a crucial indicator for evaluating the quality and practicality of large-scale models, especially in real-world applications. The aim is to... Robustness evaluation metrics should include, but are not limited to, performance volatility and perturbation stability. Performance volatility describes the performance difference between the model on the original test dataset and the new test dataset after non-adversarial perturbation. This metric quantifies the model's stability in the face of perturbations; a smaller value indicates higher stability. Perturbation stability describes the minimum distance between a sample where the model's performance degrades after non-adversarial perturbation and its corresponding original sample. This metric quantifies the maximum variation the model can tolerate in the face of perturbations; a larger value indicates stronger resilience to perturbations, thus providing a measure of the model's robustness.

[0023] Interpretability. The interpretability of a large model refers to the ability of people to understand the reasons and basis for the model's specific predictions or decisions. Interpretability is used to evaluate an algorithm's ability to explain and understand its results, and is a key element in ensuring that the model's decision-making process and results are transparent and understandable to human users. The evaluation of interpretability should include, but is not limited to, the following: Interpretive consistency: For interpretability testing methods of local substitute models, the decision results of the deep learning algorithm to be explained must be consistent with the output results obtained through interpretability methods; that is, output consistency. This is the basis for the interpretability consistency of deep learning algorithms. Interpretive validity: The explanation must accurately reflect the decision-making logic of the deep learning algorithm. A valid explanation should include the information on which the deep learning algorithm made its predictions. Interpretive validity can be evaluated using the coefficient of determination, which reflects the proportion of all disturbances in the dependent variable that can be explained by the independent variables through the regression relationship. Interpretive causality: The generated explanation must have a causal relationship with the predictions of the deep learning algorithm to be explained. A causal relationship means that the features included in the explanation are the cause of the prediction result. The more explanations that are causally related to the prediction result, the better the interpretability. Explanatory sufficiency: Explanatory sufficiency can be evaluated using the coefficient of variation. The coefficient of variation is the ratio of the standard deviation to the mean of data, and is used to compare the degree of dispersion of data in different categories.

[0024] Controllability. The controllability of large military models refers to the ability of military operators to effectively manage and control the behavior of these models, ensuring they operate according to predetermined goals and constraints. As the size and complexity of models increase, ensuring model controllability becomes increasingly important. The assessment of controllability should include the following: Attack Success Rate, describing the ratio between the number of samples where the model fails to predict the target data and the total number of samples in a new test dataset constructed using attack methods. This metric quantifies the model's security under external attacks; a smaller value indicates higher resistance to attacks, thus providing a measure of the model's ability to withstand attacks. Model Theft Degree, describing the performance difference between the original model and a surrogate model constructed through methods such as model distillation. This metric quantifies how well the surrogate model approximates the original model; a larger value indicates that the model can better approximate or replicate the functionality of the original model, thus providing a measure of the model's security or difficulty in being replicated. Model Decision Separation Degree, measuring the difference in the probability of the model making incorrect predictions across different sensitive attribute groups when the true class is a specific value. This metric focuses on the model's incorrect decisions under a specific true class and compares whether the distribution of these errors is balanced across different sensitive attribute groups. Model decision sufficiency measures the difference in the probability of the model correctly predicting a label when the model predicts a specific value across different sensitive attribute groups. This metric focuses on correct decisions when the model predicts a specific class and compares whether the distribution of these correct predictions is balanced across different sensitive attribute groups. Sensitive attribute independence measures the maximum difference between the proportions of specific predictions made by the algorithm for different sensitive attribute groups. This metric aims to measure the potential influence of protected attributes in the algorithm's predictions. Ideally, a fair algorithm should make the prediction distribution as similar as possible for all protected attribute groups, i.e., the sensitive attribute has a small influence on the algorithm's predictions. Lower values ​​indicate more consistent model predictions across different groups, reflecting higher fairness.

[0025] S20. Based on the objectives and principles of the military big model, as well as the aforementioned effectiveness evaluation index system, an ANP model structural framework is established.

[0026] The ANP model framework includes a control layer and a network layer. The control layer is built based on objectives and criteria, while the network layer is built based on an effectiveness evaluation index system. The objective is the comprehensive effectiveness of the military big model, and the criteria are driven by military needs and promoted by intelligent technologies.

[0027] Specifically, the control layer establishes the objectives and criteria of the ANP model. The model objective is to achieve the overall effectiveness of the large-scale military model. This is the top-level criterion for evaluation. As the intersection and convergence point of the military and technological revolutions, the construction and development of the military large-scale model must adhere to the dual pull of military demand-driven and intelligent technology-driven development. By analyzing national security and development strategies and military plans, it is necessary to determine and propose scientifically accurate military requirements, and positively design the development coordinates and path of the military large-scale model. By accurately grasping the development pulse of intelligent technologies and deep learning and other software and hardware technologies, advanced technologies should be applied to the military large-scale model first, radiating and driving the development of science and technology industries. As a primary criterion for performance evaluation, namely military demand-driven ( ) and intelligent technology drive ( Each of them contributes to the overall goal. The contribution (i.e., weight) are respectively , ,in , This can be determined through pairwise comparisons of direct advantages. A Delphi method survey, conducted by convening relevant military requirements experts and intelligent technology experts, showed that the two criteria contribute roughly the same to the overall objective; here, we take... Both can be adjusted according to the development focus and the actual situation at different times and stages of development.

[0028] Network Layer. Based on the aforementioned indicator analysis, in the performance evaluation of the large-scale military model, its comprehensive effectiveness can be assessed from five primary indicators and eighteen secondary indicators: availability, credibility, reliability, interpretability, and controllability. Regarding the interrelationships within the network layer, due to the ubiquitous, emergent, and evolutionary characteristics of intelligent technologies, all indicator elements have the potential to interact and influence each other. Therefore, by strengthening the ANP model, all relationships between elements are traversed, and the model relationship diagram is drawn using the dedicated ANP model software yaanp. The strengthened model is as follows: Figure 2 As shown.

[0029] S30, using the criteria in the control layer as the primary criteria, for each primary criterion, using an indicator element in a certain indicator group in the network layer as the secondary criterion, the influence degree of the indicator elements in another indicator group is compared pairwise to construct an evaluation judgment matrix, and the first ranking vector is calculated based on the evaluation judgment matrix.

[0030] Specifically, a judgment matrix is ​​constructed to control a certain criterion in the layer. As a primary criterion, a specific element from a group of elements in a network layer is used as a secondary criterion to compare the influence of elements in another group pairwise. The mutual influences between all groups are then used to construct an unweighted hypermatrix based on the order vector, typically obtained through indirect dominance comparison. For example, this can be driven by task requirements. The primary criterion is credibility. Accuracy in For secondary criteria, regarding interpretability The indicators in the data are determined according to their respective influences. The magnitude of influence was indirectly compared using the Delphi method with expert scoring, and a judgment matrix was constructed using a 1-9 scale. The results were then listed. The evaluation judgment matrix under the two-layer criterion is shown in Table 2. The corresponding ranking vector is calculated by the eigenvalue method. .

[0031] Table 2 Pairs of elements under the criteria Comparison of importance Similarly, it can be concluded that in The evaluation judgment matrix under the two-layer criterion has the following sorting vector: Therefore, the principal criterion is obtained. The following sections discuss the indicator groups respectively. The evaluation matrix is ​​used to indirectly compare the degree of influence in order to obtain the ranking vector.

[0032] Military demand-driven principles Below, indicator group The internal evaluation and judgment matrix is ​​shown in Table 3. Table 3 Under the guidelines right Evaluation and judgment matrix table S40, construct an unweighted supermatrix based on the first sorting vector.

[0033] Specifically, from the above, we can conclude that: ; In the formula, Indicates the main criterion The main criterion is to... Each indicator is a sub-criteria, for Indirect dominance comparison is performed on each element, where each column vector is... medium indicators , , , right A vector arranging the degree of influence of each element in the vector; According to the table Similarly, we can obtain ; Considering the ubiquity and emergent nature of the indicators in the large-scale military model, an improved and strengthened ANP model was used to traverse the correlations of all elements, constructing 88 comparison matrices under each criterion, for a total of 176. According to the ANP definition of a hypermatrix, the unweighted hypermatrix under the first-level criterion can be obtained as follows: .

[0034] According to the ANP method, the 18×18 unweighted hypermatrix can be obtained from equation (3). As shown in Table 4: Table 4 Unweighted hypermatrix under the criterion 0.000000 0.250000 0.500000 0.595379 0.571429 0.333333 0.333333 0.333333 0.416061 0.600000 0.707117 0.569541 0.690835 0.333333 0.333333 0.333333 0.333333 0.648329 0.200000 0.000000 0.500000 0.128271 0.142857 0.333333 0.333333 0.333333 0.126005 0.100000 0.070155 0.097390 0.160329 0.333333 0.333333 0.333333 0.333333 0.122020 0.800000 0.750000 0.000000 0.276350 0.285714 0.333333 0.333333 0.333333 0.457934 0.300000 0.222728 0.333069 0.148836 0.333333 0.333333 0.333333 0.333333 0.229651 0.277181 0.250000 0.250000 0.000000 0.092419 0.648329 0.549946 0.250000 0.596844 0.076818 0.097878 0.097878 0.470453 0.250000 0.250000 0.250000 0.250000 0.250000 0.160088 0.250000 0.250000 0.654807 0.000000 0.122020 0.209844 0.250000 0.072778 0.268450 0.360454 0.360454 0.125235 0.250000 0.250000 0.250000 0.250000 0.250000 0.467296 0.250000 0.250000 0.095338 0.484410 0.000000 0.240211 0.250000 0.164352 0.423936 0.376970 0.376970 0.123352 0.250000 0.250000 0.250000 0.250000 0.250000 0.095435 0.250000 0.250000 0.249856 0.423171 0.229651 0.000000 0.250000 0.166025 0.230796 0.164698 0.164698 0.280960 0.250000 0.250000 0.250000 0.250000 0.250000 0.500000 0.500000 0.500000 0.500000 0.666667 0.500000 0.333333 0.000000 1.000000 0.750000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.333333 0.500000 0.666667 1.000000 0.000000 0.250000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.500000 0.139950 0.250000 0.250000 0.550428 0.470453 0.428062 0.140880 0.250000 0.456746 0.000000 0.633708 0.633708 0.584170 0.097878 0.450050 0.097878 0.488133 0.097878 0.519520 0.250000 0.250000 0.072032 0.125235 0.175099 0.455408 0.250000 0.081048 0.250000 0.000000 0.174371 0.184002 0.360454 0.091415 0.360454 0.116315 0.360454 0.080770 0.250000 0.250000 0.089766 0.123352 0.232778 0.262833 0.250000 0.103006 0.500000 0.174371 0.000000 0.231828 0.376970 0.131875 0.376970 0.121078 0.376970 0.259760 0.250000 0.250000 0.287773 0.280960 0.164062 0.140880 0.250000 0.359199 0.250000 0.191921 0.191921 0.000000 0.164698 0.326661 0.164698 0.274474 0.164698 0.200000 0.200000 0.249021 0.099789 0.510435 0.197682 0.361528 0.578777 0.459098 0.387367 0.048372 0.460803 0.048062 0.000000 0.097878 0.470453 0.076818 0.562369 0.200000 0.200000 0.259218 0.216864 0.044453 0.197682 0.095494 0.094958 0.082119 0.073047 0.213894 0.065699 0.210955 0.380539 0.000000 0.125235 0.268450 0.159432 0.200000 0.200000 0.200588 0.131134 0.052410 0.197682 0.110513 0.041947 0.052612 0.061963 0.408352 0.055760 0.437042 0.094871 0.360454 0.000000 0.423936 0.100391 0.200000 0.200000 0.217607 0.469070 0.189940 0.174729 0.246598 0.070693 0.105837 0.132258 0.235810 0.117202 0.208295 0.134374 0.376970 0.123352 0.000000 0.177809 0.200000 0.200000 0.073566 0.083143 0.202762 0.232223 0.185867 0.213625 0.300334 0.345364 0.093571 0.300536 0.095647 0.390216 0.164698 0.280960 0.230796 0.000000 S50, under each main criterion, compare the importance of each indicator group to other indicator groups to obtain the corresponding second ranking vector, and construct a weighting matrix based on the second ranking vector. Specifically, for Its column vectors are normalized, but The column vectors are not normalized, so they need to be weighted and further processed. (In the master criterion) Next, the indicators of each group will be compared with the criteria. The importance of each factor is compared to obtain the corresponding ranking vector, which in turn yields the evaluation judgment matrix table shown in Table 5. Table 5 Evaluation judgment matrix table of indicator family under the criteria Based on the evaluation judgment matrix table, determine the weighting matrix: ; S60, the unweighted supermatrix is ​​weighted according to the weighted matrix to obtain a weighted supermatrix.

[0035] Specifically, for unweighted matrices Perform weighted normalization, then let We can obtain: ; in, A weighted hypermatrix is ​​a column random matrix where the sum of the elements in each column is 1.

[0036] According to the ANP method, the 18×18 weighted hypermatrix can be obtained through calculation, as shown in Table 6: Table 6 Unweighted hypermatrix under the criterion 0 0.007174 0.014347 0.016921 0.016241 0.009474 0.009474 0.010572 0.013196 0.019973 0.023539 0.01896 0.022997 0.010985 0.010985 0.010985 0.010985 0.021366 0.005739 0 0.014347 0.003646 0.00406 0.009474 0.009474 0.010572 0.003997 0.003329 0.002335 0.003242 0.005337 0.010985 0.010985 0.010985 0.010985 0.004021 0.022955 0.021521 0 0.007854 0.00812 0.009474 0.009474 0.010572 0.014524 0.009987 0.007414 0.011088 0.004955 0.010985 0.010985 0.010985 0.010985 0.007568 0.030371 0.027393 0.027393 0 0.009353 0.065609 0.055653 0.053654 0.128092 0.011089 0.014129 0.014129 0.067911 0.01984 0.01984 0.01984 0.01984 0.01984 0.017541 0.027393 0.027393 0.066265 0 0.012348 0.021236 0.053654 0.015619 0.038751 0.052032 0.052032 0.018078 0.01984 0.01984 0.01984 0.01984 0.01984 0.051203 0.027393 0.027393 0.009648 0.049021 0 0.024309 0.053654 0.035273 0.061196 0.054416 0.054416 0.017806 0.01984 0.01984 0.01984 0.01984 0.01984 0.010457 0.027393 0.027393 0.025285 0.042824 0.02324 0 0.053654 0.035632 0.033316 0.023774 0.023774 0.040557 0.01984 0.01984 0.01984 0.01984 0.01984 0.101061 0.101061 0.101061 0.176552 0.235402 0.176552 0.117701 0 0.109421 0.192505 0.128337 0.128337 0.128337 0.084751 0.084751 0.084751 0.084751 0.084751 0.101061 0.101061 0.101061 0.176552 0.117701 0.176552 0.235402 0.109421 0 0.064168 0.128337 0.128337 0.128337 0.084751 0.084751 0.084751 0.084751 0.084751 0.030468 0.054426 0.054426 0.085072 0.072711 0.066159 0.021774 0.048448 0.088514 0 0.06108 0.06108 0.056305 0.023734 0.109132 0.023734 0.118367 0.023734 0.113101 0.054426 0.054426 0.011133 0.019356 0.027063 0.070386 0.048448 0.015707 0.024096 0 0.016807 0.017735 0.087406 0.022167 0.087406 0.028205 0.087406 0.017584 0.054426 0.054426 0.013874 0.019065 0.035977 0.040622 0.048448 0.019962 0.048192 0.016807 0 0.022345 0.091411 0.031978 0.091411 0.02936 0.091411 0.05655 0.054426 0.054426 0.044477 0.043424 0.025357 0.021774 0.048448 0.06961 0.024096 0.018498 0.018498 0 0.039938 0.079212 0.039938 0.066557 0.039938 0.088382 0.088382 0.110044 0.036196 0.185146 0.071704 0.131134 0.260713 0.206803 0.181792 0.022701 0.216255 0.022555 0 0.04656 0.223791 0.036542 0.267514 0.088382 0.088382 0.114551 0.078661 0.016124 0.071704 0.034638 0.042774 0.036991 0.034281 0.100381 0.030833 0.099001 0.181019 0 0.059574 0.1277 0.07584 0.088382 0.088382 0.088641 0.047565 0.01901 0.071704 0.040086 0.018895 0.023699 0.029079 0.19164 0.026168 0.205104 0.045129 0.171465 0 0.201663 0.047755 0.088382 0.088382 0.096163 0.170142 0.068896 0.063378 0.089446 0.031844 0.047675 0.062069 0.110666 0.055003 0.097753 0.063921 0.179322 0.058677 0 0.084582 0.088382 0.088382 0.03251 0.030158 0.073546 0.084233 0.067418 0.096228 0.135287 0.16208 0.043913 0.141042 0.044887 0.185623 0.078345 0.13365 0.109788 0 As can be seen from equation (6), the weighted hypermatrix is ​​a column random matrix with a column sum of 1. To ensure that each column is stable, it is multiplied by itself until the product converges. The value of each column is then the final sorted vector. In ANP, the results calculated using the yaanp software are shown in Table 7.

[0037] Table 7 Unweighted Limit Hypermatrix under Criterion 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.014307 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.007462 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.010314 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.036396 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.027937 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.03076 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.027231 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.10219 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.095756 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.054977 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.046122 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.048256 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.043733 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.133416 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.075252 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.071807 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.073432 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 0.100651 Based on comprehensive weight The normalized weighted sort can be obtained, as can be seen from the above. The sorting results are as follows Figure 2 As shown.

[0038] S80, construct the factor set and evaluation set for fuzzy comprehensive evaluation, generate a weight set according to the final weights, perform single-factor evaluation on each factor, obtain the membership degree matrix of each factor to the evaluation set, calculate the membership degree of the overall target of fuzzy comprehensive evaluation according to the membership degree matrix and the weight set, and determine the evaluation result of the maturity of the military big model according to the principle of maximum membership degree.

[0039] Specifically, by introducing fuzzy comprehensive evaluation method based on network analysis, the subjective influence of the evaluation process can be effectively reduced when constructing the judgment matrix, thereby improving the accuracy of the evaluation results. The factor set for fuzzy comprehensive evaluation is constructed using the indicator element set of ANP. , ,in The number of evaluation indicators, i.e. Establish the weight set. , , The weights of each factor are obtained using the ANP algorithm. An evaluation set is a collection constructed based on the different possible outcomes of an object. , , The number of evaluation results, i.e. Here, "Superior" refers to a military large-scale model that provides superior AI services, meeting the highest requirements in both the AI ​​basic service capability domain and the AI ​​business service capability domain. "Excellent" refers to a military large-scale model that provides sufficient AI services, meeting the highest requirements in both the AI ​​basic service capability domain and the AI ​​business service capability domain. "Robust" refers to a military large-scale model that provides stable AI services, meeting the medium requirements in both the AI ​​basic service capability domain and the AI ​​business service capability domain. "Developing" refers to a military large-scale model that provides scalable AI services, meeting the medium requirements in both the AI ​​basic service capability domain and the basic requirements in the AI ​​business service capability domain. "Starting" refers to a military large-scale model that provides basic AI services, meeting the minimum requirements in both the AI ​​basic service capability domain and the AI ​​business service capability domain.

[0040] The evaluation focus is on the first Evaluation factors As the evaluation object, The first evaluation set element The membership degree is Similarly, we can conclude that For the evaluation set The membership degree of each factor is a fuzzy set matrix. By conducting single-factor evaluations of each factor, the membership matrix of each factor to the comment set can be obtained. Single-factor fuzzy evaluation only reflects the influence of a single factor on the evaluation object. The calculation of the overall objective membership degree of the fuzzy comprehensive evaluation is as follows: but , According to the principle of maximum membership It can be determined that the comprehensive evaluation result of the maturity of the military large model in this round of expert scoring is at the initial level.

[0041] This application proposes a military large-scale model evaluation method based on a reinforced ANP-FCE model. At the control layer, the method uses the effectiveness of the military large-scale model as the control objective and military demand-driven and intelligent technology-driven approaches as two main criteria, achieving logical consistency between top-down decomposition and bottom-up aggregation and verification. At the network layer, five indicator groups and eight indicator elements are set up: availability, credibility, reliability, interpretability, and controllability. Focusing on the deep coupling and complex links within the large-scale model's technological ecosystem, a full interactive analysis is performed on all indicators. By reinforcing computation to traverse the influence relationships of all indicator elements, logical conflicts arising from the advantages of different indicators are eliminated. Through logically sound weight relationships, five capability levels—starting level, development level, robust level, excellent level, and superior level—are mapped, enabling an effective assessment of the maturity of the military large-scale model.

[0042] Implementation principle of enhanced ANP-FCE model The Analytic Hierarchy Process (AHP), proposed in 1996 by Professor T.S. Thaaty, a member of the National Academy of Engineering, fully considers the relationships between the components (elements) within the evaluation object and the degree of their mutual influence. It then uses a hypermatrix to conduct a comprehensive analysis of these relationships and interactions, ultimately deriving the mixed weights of each element. Fuzzy Comprehensive Evaluation (FCE) is a method that transforms qualitative evaluation into quantitative evaluation. Based on the principle of fuzzy transformation and using fuzzy mathematical membership theory, it conducts a comprehensive evaluation of factors for which significant quantitative boundaries cannot be clearly defined, thereby arriving at scientific qualitative conclusions.

[0043] The general steps to strengthen the ANP-FCE model are as follows: 0.1 Establishing a typical ANP structure ANP first divides the system elements into two main parts. The first part is called the control factor layer, which includes the problem objective and decision criteria. All decision criteria are considered to be independent of each other and only governed by the objective element. There may be no decision criteria in the control factors, but there must be at least one objective. The weight of each criterion in the control layer can be obtained using the traditional AHP method. The second part is the network layer, which is composed of all elements governed by the control layer. Its internal structure is a network structure that influences each other.

[0044] 0.2 Dominance Comparison A key step in implementing the AHP model is to perform pairwise comparisons of its subordinate elements under a specific criterion, and to represent these comparison results by constructing a judgment matrix. If the elements being compared are independent of each other, a direct dominance comparison is performed, that is, the importance of the two elements to the given criterion is directly assessed. Conversely, if there is interdependence or influence between the elements being compared, an indirect dominance comparison is performed, that is, the influence of the two elements on a third element is compared under the given criterion.

[0045] 0.3 Hypermatrix and Weighted Hypermatrix of ANP Structure Suppose that there are elements in the control layer of ANP. Below the control layer, the network layer contains element groups. ,in There are elements Control layer elements The main criterion is to... medium elements For this criterion, the element group Elements in the middle according to their pairs The magnitude of influence is used to indirectly compare the degree of dominance, i.e., to construct a judgment matrix. Down: The sorting vector is obtained by the eigenvalue method. remember for: (1) here The column vector is medium elements right medium elements The influence ranking vector, if The middle element is not affected The influence of elements, then This will ultimately yield the following results. Lower Hypermatrix (2) There are a total of such hypermatrices There are 100 sub-blocks of the supermatrix, all of which are nonnegative matrices. It is column normalization, but However, it is not a normalized list, therefore... As a standard, for The following groups of elements are criteria Comparing their importance Down and The sorting vector components corresponding to irrelevant element groups are zero, thus yielding the weighted matrix. (3) For hypermatrix Weighted sum of elements, we get ,in A weighted hypermatrix whose column sum is 1 is called a column random matrix (13). For simplicity, the following hypermatrixes are all weighted hypermatrixes, and the notation will still be used. express.

[0046] 0.4 Limit Relative Sort Vector Let the weighted hypermatrix be The elements are ,but The size reflects the element For elements The advantage of one step right The advantage can also be used The result is called the two-step advantage, which is... elements, It is still normalized, when When it exists, The The list is Each element in the lower network layer is related to the element The limiting relative sorting vector.

[0047] 0.5 Determine the evaluation object factor set and comment set Introducing fuzzy comprehensive evaluation method based on network analysis can effectively reduce the subjective influence of the evaluation process and improve the accuracy of evaluation results when constructing the judgment matrix. The factor set for fuzzy comprehensive evaluation is constructed using the indicator element set of ANP. , ,in To determine the number of evaluation indicators, a weight set is established. , , The weights of each factor are obtained using the ANP algorithm. The evaluation set is a collection constructed based on the different possible outcomes of the evaluated object. , , The number of evaluation results.

[0048] 0.6 Single-factor fuzzy rating and membership degree of the overall objective The evaluation focus is on the first Evaluation factors As the evaluation object, The first evaluation set element The membership degree is Similarly, we can conclude that Collection of comments The membership degree of each factor is a fuzzy set matrix. By conducting single-factor evaluations of each factor, the membership matrix of each factor to the comment set can be obtained. (4) Single-factor fuzzy evaluation only reflects the impact of a single factor on the evaluation object. The calculation of the membership degree of the overall objective in fuzzy comprehensive evaluation is as follows: (5) According to the principle of maximum membership Corresponding evaluation This is a comprehensive evaluation conclusion for the corresponding indicators.

[0049] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0050] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0051] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A maturity evaluation method for large military models based on enhanced ANP-FCE, characterized in that, The method includes: S10, Constructing an effectiveness evaluation index system for a large military model; wherein, the effectiveness evaluation index system includes five index groups: usability, credibility, reliability, interpretability, and controllability, and each primary index group contains multiple index elements; S20. Based on the objectives and criteria of the military big model, and the effectiveness evaluation index system, establish the ANP model structure framework; wherein, the ANP model structure framework includes a control layer and a network layer, the control layer is constructed based on the objectives and criteria, and the network layer is constructed based on the effectiveness evaluation index system, the objective is the comprehensive effectiveness of the military big model, and the criteria are military demand-driven and intelligent technology-driven. S30, using the criteria in the control layer as the primary criteria, for each primary criterion, using an indicator element in a certain indicator group in the network layer as the secondary criterion, comparing the influence degree of the indicator elements in another indicator group pairwise, constructing an evaluation judgment matrix, and calculating the first ranking vector based on the evaluation judgment matrix. S40, construct an unweighted hypermatrix based on the first sorting vector; S50, under each main criterion, compare the importance of each indicator group to other indicator groups to obtain the corresponding second ranking vector, and construct a weighting matrix based on the second ranking vector; S60, the unweighted supermatrix is ​​weighted according to the weighted matrix to obtain a weighted supermatrix; S70, perform self-multiplication on the weighted hypermatrix until the product converges to obtain the limiting hypermatrix, and extract column vectors from the limiting hypermatrix as the final weights of each indicator element. S80, construct the factor set and evaluation set for fuzzy comprehensive evaluation, generate a weight set according to the final weights, perform single-factor evaluation on each factor, obtain the membership degree matrix of each factor to the evaluation set, calculate the membership degree of the overall target of fuzzy comprehensive evaluation according to the membership degree matrix and the weight set, and determine the evaluation result of the maturity of the military big model according to the principle of maximum membership degree.

2. The method according to claim 1, characterized in that, S30 includes: Using the criteria in the control layer as the primary criteria, for each primary criterion, an indicator element from a certain indicator group in the network layer is used as a secondary criterion to compare the influence of the indicator elements in another indicator group pairwise. The Delphi method is used to construct the evaluation judgment matrix by expert scoring and a 1-9 scale method.

3. The method according to claim 1, characterized in that, The pairwise comparison of the influence of the index elements in another index group specifically includes: Indirect dominance comparison is performed on the degree of influence of the indicator elements in another indicator group.

4. The method according to claim 2, characterized in that, The S30 also includes: Based on the evaluation judgment matrix, the corresponding first sorting vector is calculated using the eigenvalue method.

5. The method according to claim 1, characterized in that, S40 includes: In the main principle The unweighted hypermatrix is ​​as follows: in, This represents an unweighted supermatrix.

6. The method according to claim 4, characterized in that, The weighting matrix is ​​as follows: in, Represents a weighted matrix. This represents the second sorting vector.

7. The method according to claim 1, characterized in that, The S60 includes: For unweighted hypermatrix Perform weighted normalization, then let We can obtain: ; in, A weighted hypermatrix is ​​a column random matrix where the sum of the elements in each column is 1.

8. The method according to claim 1, characterized in that, The step of calculating the overall objective membership degree of the fuzzy comprehensive evaluation based on the membership degree matrix and the weight set includes: Calculate the membership degree of the overall objective in the fuzzy comprehensive evaluation: ; in, This represents the membership degree of the overall objective in the fuzzy comprehensive evaluation. Represents the set of weights. This represents the membership matrix.