Multi-mode-oriented intelligent computing network management and control operation and maintenance intelligent agent system capability maturity evaluation method

By adopting a capability maturity assessment method for intelligent computing network management and operation and maintenance intelligent agent system oriented towards multimodality, the problem of insufficient functional assessment in existing technologies is solved, and the system achieves comprehensive and dynamic assessment and optimization, thereby improving the performance of intelligent computing networks.

CN120979920APending Publication Date: 2025-11-18CHINA ACADEMY OF INFORMATION & COMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511312052.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies lack scientific and comprehensive methods for evaluating the functions of intelligent agents for the management and operation of multimodal intelligent computing networks, making it difficult to optimize and improve the system in a targeted manner, thus affecting the overall performance improvement and development of intelligent computing networks.

Method used

This paper presents a capability maturity assessment method for intelligent computing network management and operation and maintenance intelligent agent systems oriented towards multimodality. By determining evaluation dimensions, setting weights, building test environments, executing test steps, scoring rules, and comprehensive scores, combined with dynamic adaptation and iterative optimization mechanisms, a comprehensive evaluation of the system can be achieved.

Benefits of technology

It enables multi-angle and multi-dimensional evaluation of the intelligent computing network management and operation system, which can truly approximate actual needs, has timeliness and practicality, and supports system optimization and performance improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979920A_ABST
    Figure CN120979920A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal-oriented intelligent computing network management and control operation and maintenance intelligent agent system capability maturity evaluation method, and relates to the technical field of intelligent computing networks, and the method comprises the following steps: S1, determining an evaluation dimension; s2, setting subordinate test items for each evaluation dimension, and respectively endowing the subordinate test items with weights; s3, building a test environment; s4, executing a preset test step for each test item to obtain a test result; s5, scoring is carried out according to a preset scoring rule; s6, obtaining a comprehensive score of the system; and S7, determining a system capability maturity level based on the comprehensive score. According to the multi-mode-oriented capability maturity evaluation method for the intelligent computing network management and control operation and maintenance agent system, maturity testing is carried out from all dimensions, a testing method for all indexes is provided, maturity evaluation of the intelligent computing network management and control operation and maintenance agent system can be standardized in an omnibearing and multi-angle mode, and the evaluation accuracy of the intelligent computing network management and control operation and maintenance agent system is improved. And the maturity evaluation of the intelligent calculation network management and control operation and maintenance intelligent agent system can really approach the actual demand.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent computing network, in particular to a multi-modal intelligent computing network management and operation agent system capability maturity evaluation method. BACKGROUND

[0002] With the surging tide of digitalization, as the key infrastructure for data transmission and processing, intelligent computing network plays a crucial role in modern information society. The multi-modal intelligent computing network management and operation agent system, with its ability to integrate multiple data modalities (including images, speech, text, etc.) and achieve intelligent management, has become the core force driving the efficient operation of intelligent computing network, providing an ideal solution for high concurrency and complex task processing of intelligent computing network. To ensure the stable and reliable operation of the multi-modal intelligent computing network management and operation agent system and meet the growing demand for diversified services, it is particularly important to evaluate its functions. The stability of system functions needs to be tested and verified comprehensively in real complex business scenarios. Standardizing the evaluation of system functions is of great significance to the stable operation and service quality of intelligent computing network.

[0003] However, there is currently a lack of a scientific and complete evaluation method for the multi-modal intelligent computing network management and operation agent system in the field of intelligent computing network. Most existing network operation-related patents focus on network operation methods, devices, storage media and electronic equipment, and do not involve the evaluation of network operation agent function maturity level. This situation hinders the in-depth evaluation of the multi-modal intelligent computing network management and operation agent system, making it difficult to optimize and improve the system, and thus affecting the overall performance improvement and development of intelligent computing network. In the increasingly complex environment of multi-modal data processing, traditional evaluation methods cannot fully consider the system's ability to handle dynamic changes in multi-modal data and cross-modal collaborative processing, nor can they accurately evaluate the system's actual performance in handling high concurrency and low latency requirements, nor can they effectively identify potential risks in the process of integrating different modal data. SUMMARY

[0004] To address the deficiencies in the prior art, the present application provides a multi-modal intelligent computing network management and operation agent system capability maturity evaluation method, which solves the problems raised in the background art.

[0005] To achieve the above purpose, the present application is implemented by the following technical solution: a multi-modal intelligent computing network management and operation agent system capability maturity evaluation method, comprising the following steps: S1: Determine the evaluation dimensions, which include large model capabilities, memory capabilities, tool capabilities, operation and maintenance capabilities, and openness capabilities; S2: Set up subordinate test items for each evaluation dimension, and assign weights based on the importance of each dimension and test item; S3: Set up the test environment, which includes both hardware and software environments; S4: Execute preset test steps for each test item and obtain test results; S5: Score the test results of each test item according to the preset scoring rules; S6: The scoring results are weighted and summed according to the weight of each test item to obtain the overall system score; S7: Determine the system capability maturity level based on the comprehensive score, which includes 1-star, 2-star and 3-star.

[0006] Furthermore, in step S1, the sub-test items for the large model capability include text parsing accuracy, semantic understanding depth, contextual coherence, model inference performance, instruction compliance, and task completion capability; the sub-test items for the memory capability include storage capacity, knowledge storage and retrieval, information updating, and forgetting mechanism; the sub-test items for the tool capability include the ability to guide multi-tenant isolated deployment of intelligent agents and the pre-training GPU card aggregation communication performance of intelligent agents; the sub-test items for the operation and maintenance capability include log analysis intelligent agents, fault location intelligent agents, and root cause analysis intelligent agents; and the sub-test items for the open capability include API function testing, user interface testing, and long-term operation testing.

[0007] Furthermore, in step S2, the weight of the large model capability is 40%, of which the weights of text parsing accuracy, semantic understanding depth, and contextual coherence are 20%, the weight of model inference performance is 10%, and the weight of instruction compliance and task completion capability is 10%; the weight of the memory capability is 20%, of which the weights of storage capacity, knowledge storage and retrieval, information update, and forgetting mechanism are each 5%; the weight of the tool capability is 15%, of which the weight of agent multi-tenant isolation deployment guidance capability is 8%, and the weight of agent pre-training GPU card aggregation communication performance is 7%; the weight of the operation and maintenance capability is 15%, of which the weights of log analysis agent, fault location agent, and root cause analysis agent are each 5%; and the weight of the open capability is 10%, of which the weight of API function testing is 3%, the weight of user interface testing is 3%, and the weight of long-term operation testing is 4%.

[0008] Furthermore, in step S3, the hardware environment includes: a server equipped with a high-performance CPU, such as an Intel Xeon Gold series CPU; large-capacity memory, such as 256GB or more; and high-speed storage devices, such as NVMe SSDs; a GPU, using NVIDIA brand high-performance GPUs, such as A100 and H100; network devices, including switches and routers that support VLAN segmentation and network isolation to ensure low latency and high bandwidth data transmission; the software environment includes: an operating system, using a mainstream Linux operating system, such as Ubuntu 20.04LTS; programming languages ​​and frameworks, installing compatible versions of languages ​​and frameworks according to the model development requirements; and testing tools, including Postman for API function testing, JMeter for performance testing, and ELKS tack for log analysis.

[0009] Furthermore, in step S4, the testing steps for text parsing accuracy, semantic understanding depth, and contextual coherence include: constructing diverse input texts, including question-and-answer, dialogue, long documents, complex sentence structures, and domain terms, to verify the keyword extraction accuracy and semantic classification correctness; inputting sentences containing polysemous words or ambiguous references to verify the reference parsing success rate; and providing multilingual mixed input to verify translation consistency and cross-language semantic retention rate.

[0010] Furthermore, in step S4, the testing steps for model inference performance include: performing inference using standard text datasets, such as SQuAD and GLUE, and recording the number of sentences processed per second (QPS), memory usage, and GPU utilization; gradually increasing the length of the input text to verify whether the system's performance degradation is ≤10% when the load increases; and adjusting model parameters, such as batch size and sequence length, to verify the balance between memory usage and computational efficiency.

[0011] Furthermore, in step S5, the scoring rules include: text parsing accuracy: keyword extraction accuracy ≥95% earns 7 points, 90%-94% earns 5 points, <90% earns 0 points; semantic understanding depth: semantic classification accuracy ≥90% earns 7 points, 80%-89% earns 4 points, <80% earns 0 points; contextual coherence: reference parsing success rate ≥90% earns 6 points, 80%-89% earns 3 points, <80% earns 0 points; in model inference performance, inference speed meets the standard and GPU usage is reasonable, each earns 2 points, load stability earns 3 points, and resource optimization meets the standard, each earns 3 points; passing the storage capacity test earns 5 points, failing earns 0 points; knowledge storage and retrieval answer accuracy earns 5 points, relatively accurate earns 3 points, and inaccurate earns 0 points.

[0012] Furthermore, in step S7, the comprehensive score range corresponding to 1 star is 0-39, the comprehensive score range corresponding to 2 stars is 40-79, and the comprehensive score range corresponding to 3 stars is 80-100. The comprehensive score needs to be determined after multiple measurements, such as 3-10 times.

[0013] Furthermore, the capability maturity assessment method for the multimodal intelligent computing network management and operation intelligent agent system also includes the following steps: S8: Establish a dynamic adaptation and iterative optimization mechanism, including: real-time collection of multimodal dynamic data of the intelligent computing network operation, and adjustment of the weight coefficients of each test item based on changes in data characteristics; periodic execution of the re-evaluation process, and generation of maturity trend curves based on historical evaluation results; and optimization of test items and scoring rules based on user feedback and actual application scenario requirements.

[0014] Furthermore, in step S8, the real-time collected multimodal dynamic data includes network traffic fluctuation data, multimodal interaction request frequency, and fault type change trend; the weight coefficient adjustment is based on the entropy method or the hierarchical analysis method, and the weight update is triggered when the change rate of a certain type of data feature exceeds a preset threshold, such as 20%, for example, the change rate of voice command data features exceeds 20%; the periodic reassessment period can be set to 1 month, 3 months, or 6 months, and is dynamically configured based on the iteration speed of intelligent computing network services.

[0015] This invention provides a capability maturity assessment method for intelligent computing network management and operation intelligent agent systems oriented towards multimodality, which has the following beneficial effects: 1. This maturity assessment method for a multimodal intelligent computing network management and operation intelligent agent system evaluates the system's capabilities from multiple dimensions, including model inference, multi-tenant isolation, fault location agent, root cause analysis, and API functionality. It assesses the system's large-scale model capabilities, memory capabilities, tool capabilities, operation and maintenance capabilities, and openness capabilities. The method provides testing methods for each indicator, offering diverse indicators to standardize the maturity assessment. It assigns scores to each indicator based on their importance, providing a scoring reference standard and assigning different weights to each indicator to ensure the maturity evaluation accurately reflects actual needs.

[0016] 2. This capability maturity assessment method for intelligent computing network management and operation and maintenance intelligent agent system for multimodal applications utilizes dynamic adaptation and iterative optimization mechanisms. By adjusting weights through real-time data feedback, periodically reassessing and analyzing trends, and optimizing rules based on user feedback, the assessment method can adapt to the dynamic changes in intelligent computing network business scenarios and the iterative upgrades of multimodal data characteristics. This ensures the timeliness and practicality of the assessment results and solves the problem of insufficient adaptability of traditional static assessment methods. Attached Figure Description

[0017] Fig. 1 This is a schematic diagram of the test items for the capability maturity assessment method of a multimodal intelligent computing network management and operation and maintenance intelligent agent system according to the present invention; Fig. 2 This is a schematic diagram of the scoring criteria for the capability maturity assessment method of a multimodal intelligent computing network management and maintenance intelligent agent system according to the present invention. Detailed Implementation

[0018] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.

[0019] like Figs. 1-2 As shown, the present invention provides a technical solution: a method for assessing the capability maturity of a multimodal intelligent computing network management and control operation and maintenance intelligent agent system, comprising the following steps: S1: Determine the evaluation dimensions, which include large model capabilities, memory capabilities, tool capabilities, operation and maintenance capabilities, and openness capabilities; The sub-tests for the large model capability include text parsing accuracy, semantic understanding depth, contextual coherence, model inference performance, instruction compliance, and task completion capability; the sub-tests for the memory capability include storage capacity, knowledge storage and retrieval, information updating, and forgetting mechanism; the sub-tests for the tool capability include the ability to guide multi-tenant isolated deployment of agents and the pre-training GPU card aggregation communication performance of agents; the sub-tests for the operation and maintenance capability include log analysis agents, fault location agents, and root cause analysis agents; and the sub-tests for the open capability include API function testing, user interface testing, and long-term operation testing. S2: Set up subordinate test items for each evaluation dimension, and assign weights based on the importance of each dimension and test item; The weight of large model capabilities is 40%, with text parsing accuracy, semantic understanding depth, and contextual coherence each accounting for 20%, model inference performance accounting for 10%, and instruction compliance and task completion capabilities accounting for 10%. The weight of memory capabilities is 20%, with storage capacity, knowledge storage and retrieval, information updating, and forgetting mechanisms each accounting for 5%. The weight of tool capabilities is 15%, with agent multi-tenant isolation deployment guidance capabilities accounting for 8%, and agent pre-training GPU card aggregation communication performance accounting for 7%. The weight of operation and maintenance capabilities is 15%, with log analysis agents, fault location agents, and root cause analysis agents each accounting for 5%. The weight of open capabilities is 10%, with API function testing accounting for 3%, user interface testing accounting for 3%, and long-term runtime testing accounting for 4%. S3: Set up the test environment, which includes both hardware and software environments; The hardware environment includes: a server equipped with a high-performance CPU, such as an Intel Xeon Gold series CPU; large-capacity memory, such as 256GB or more; and high-speed storage devices, such as NVMe SSDs; a GPU using a high-performance NVIDIA GPU, such as the A100 or H100; and network equipment, including switches and routers supporting VLAN segmentation and network isolation to ensure low latency and high bandwidth data transmission. If the hardware specifications of the test environment are lower than the above standards, the performance benchmark deviation must be noted in the test report. For example, if the GPU memory is insufficient, the inference speed correction factor should be equal to the actual memory usage. Standard video memory × 0.8; The software environment includes: operating system, using a mainstream Linux operating system, such as Ubuntu 20.04LTS; programming language and framework, installing compatible versions of the language and framework according to the model development requirements, supporting mainstream frameworks compatible with the model development language such as TensorFlow, MindSpore, etc.; before testing, the compatibility between the framework version and the model must be verified, such as confirming compatibility through a basic functional test case pass rate ≥ 95%; testing tools, including Postman for API functional testing, JMeter for performance testing, and ELKS tack for log analysis; S4: Execute preset test steps for each test item and obtain test results; The testing steps for text parsing accuracy, semantic understanding depth, and contextual coherence include: constructing diverse input texts, including questions and answers, dialogues, long documents, complex sentences, and domain terms, to verify keyword extraction accuracy and semantic classification correctness; inputting sentences containing polysemous words or ambiguous references to verify the reference parsing success rate; and providing multilingual mixed input to verify translation consistency and cross-language semantic preservation rate. The testing steps for model inference performance include: performing inference using standard text datasets such as SQuAD and GLUE, and recording the number of sentences processed per second (QPS), memory usage, and GPU utilization; gradually increasing the length of the input text to verify whether the system's performance degradation is ≤10% when the load increases; and adjusting model parameters such as batch size and sequence length to verify the balance between memory usage and computational efficiency. S5: Score the test results of each test item according to the preset scoring rules; The scoring rules include: Text parsing accuracy: Keyword extraction accuracy ≥95% earns 7 points, 90%-94% earns 5 points, <90% earns 0 points; Semantic understanding depth: Semantic classification accuracy ≥90% earns 7 points, 80%-89% earns 4 points, <80% earns 0 points; Contextual coherence: Referential resolution success rate ≥90% earns 6 points, 80%-89% earns 3 points, <80% earns 0 points; In model inference performance, inference speed meets the standard and GPU usage is reasonable, each earns 2 points, load stability earns 3 points, and resource optimization meets the standard, each earns 3 points; Passing the storage capacity test earns 5 points, failing earns 0 points; Knowledge storage and retrieval answer accuracy earns 5 points, relatively accurate earns 3 points, and inaccurate earns 0 points; S6: The scoring results are weighted and summed according to the weight of each test item to obtain the overall system score; S7: Determine the system capability maturity level based on the comprehensive score, which includes 1-star, 2-star and 3-star levels; The comprehensive score range for 1 star is 0-39, for 2 stars it is 40-79, and for 3 stars it is 80-100. The comprehensive score needs to be determined after multiple measurements, such as 3-10 times. The specific technical solution is as follows: Determine the evaluation dimensions: including five dimensions: large model capability, memory capability, tool capability, operation and maintenance capability, and openness capability, which comprehensively cover the core functional characteristics of the system; Test Items and Weights: Test items are set for each evaluation dimension, and corresponding weights are assigned based on the importance of each dimension and test item. Specifically: Large model capability has a weight of 40%, with the following sub-items weighted at 20% each: text parsing accuracy, etc.; model inference performance, etc., 10%; instruction compliance and task completion capability, 10%; memory capability has a weight of 20%, with the following four sub-items weighted at 5% each; tool capability has a weight of 15%, with the following sub-items weighted at 8% for multi-tenant isolation deployment guidance and 7% for pre-training GPU card aggregation communication performance; operation and maintenance capability has a weight of 15%, with the following three intelligent agents weighted at 5% each; and openness capability has a weight of 10%, with the following three sub-items weighted at 3%, 3%, and 4% respectively. Setting up the testing environment: The hardware environment includes servers equipped with high-performance CPUs such as Intel Xeon Gold series, large-capacity memory (256GB or more), high-speed storage devices such as NVMe SSDs, high-performance NVIDIA GPUs such as A100 and H100, and network devices that support VLAN segmentation and network isolation; the software environment includes mainstream Linux operating systems such as Ubuntu 20.04LTS, compatible programming languages ​​and frameworks, and testing tools such as Postman, JMeter, and ELKS tack. Execution of test steps: Design specific test steps for each test item. For example, text parsing accuracy test verifies keyword extraction accuracy by constructing diverse input texts; model inference performance test verifies performance by inference on standard datasets, increasing load, and adjusting parameters; storage capacity test verifies storage integrity by writing enterprise private operation and maintenance data. Scoring rules: Scoring is based on test results according to preset rules. For example, the three sub-items such as text parsing accuracy are scored with 7 points, 7 points, and 6 points respectively; model inference performance is scored according to inference speed, GPU usage, etc.; storage capacity test is scored as follows: 5 points for passing and 0 points for failing. Comprehensive scoring and rating: The scores of each test item are weighted and summed to obtain a comprehensive score. Based on the score, maturity levels are divided into 1-3 stars, where 1 star corresponds to 0-39 points, 2 stars correspond to 40-79 points, and 3 stars correspond to 80-100 points. The comprehensive score needs to be determined through multiple measurements.

[0020] Five evaluation dimensions are set, and the test items for each dimension are as follows: Large model capabilities, weighted at 40%: including text parsing accuracy, semantic understanding depth, and contextual coherence, weighted at 20%; model inference performance, weighted at 10%; instruction compliance and task completion capabilities, weighted at 10%. Memory ability, weighted at 20%: including storage capacity (5%), knowledge storage and retrieval (5%), information updating (5%), and forgetting mechanism (5%). Tool capabilities, weighted at 15%: including agent multi-tenant isolation deployment guidance capability (8%), agent pre-training GPU card aggregation communication performance (7%); Operational capabilities, weighted at 15%, include log analysis agent (5%), fault location agent (5%), and root cause analysis agent (5%). Open capabilities, weighted at 10%: including API function testing (3%), user interface testing (3%), and long-term runtime testing (4%).

[0021] Specific test steps for each test item: Large model capability test: Text parsing accuracy, semantic understanding depth, and contextual coherence testing: Construct diverse texts including network fault reports, maintenance instructions, and multimodal descriptions, such as "Server temperature is too high, log shows 'ERROR: CPU overheat', please analyze the cause," for a total of 100 test samples. Verify keyword extraction accuracy (≥95%), i.e., the percentage of samples that correctly extract key information such as "Server temperature is too high, CPU overheat"; semantic classification accuracy, such as sentiment analysis and intent recognition consistency (≥90%); input sentences with ambiguous references, such as "Device A connection is abnormal, its indicator light is red," verify the reference parsing success rate (≥90%); provide mixed Chinese and English text, such as "Please check if server disk usage exceeds 80%," verify translation consistency and cross-language semantic retention rate (≥90%). Model inference performance testing: Standard inference was performed using the SQuAD dataset containing 100,000 question-answer samples, recording QPS ≥ 50, GPU utilization ≤ 80%, and memory usage ≤ 128GB; the input text length was gradually increased from 100 characters to 1000 characters, with a total of 5 gradients, each gradient tested 10 times, verifying that the performance degradation (QPS decrease) ≤ 10%; the batch size was adjusted to 8, 16, and 32, and the sequence length was adjusted to 512 and 1024, testing the balance between memory usage and computational efficiency under different settings, ensuring that QPS ≥ 30 and no memory overflow occurred when batch size = 16 and sequence length = 1024; Command compliance and task completion capability testing: Explicit command testing (e.g., summarizing a 1000-word maintenance report into three network faults ranked by impact within 100 words, verifying command compliance accuracy ≥ 95%; complex command testing, such as prioritizing database-related faults while ignoring temporary alarms), assessing the accuracy of interpreting explicit and implicit requirements; domain-specific task testing, such as generating network topology diagram description code and analyzing database deadlock causes, verifying the professionalism and accuracy of the results ≥ 90%. Memory ability test: Storage capacity test: Prepare enterprise private operation and maintenance data containing 5,000 network topology structures, 2,000 device configuration information, and 3,000 fault records, with a total data volume of 5GB. After writing the data to the intelligent agent, verify it through index check and random sampling of 100 data. If all data is stored completely without loss or error, it passes. Knowledge storage and retrieval test: Write 1000 pieces of operation and maintenance knowledge, including fault handling procedures and equipment parameter specifications. Use keyword queries such as switch port fault handling steps and semantic queries such as how to solve latency problems caused by insufficient bandwidth 50 times each. If ≥90% of the query answers are accurate, you will get 5 points, 70%-89% will get 3 points, and <70% will get 0 points. Information update test: Modify 200 pieces of knowledge base data, such as updating the device firmware version, and observe whether the agent completes the update within 1 hour; ask 50 questions based on the updated data, and if ≥95% of the answers are timely, it passes; Forgetting mechanism test: Delete 100 outdated pieces of information, such as maintenance guides for old models of equipment, and observe whether the agent deletes the corresponding storage within 2 hours; ask 50 questions based on the updated knowledge. If ≥95% of the answers do not contain outdated information, you get 5 points; 80%-94% get 3 points; and <80% get 0 points. Tool proficiency test: Multi-tenant isolation deployment guidance capability test: Ask the agent how to configure a network isolation deployment scheme for 3 tenants. If the generated guidance includes key steps such as VLAN division, resource quota, and security policy and is accurate, it will score 8 points; if it includes some key steps, it will score 5 points; if it does not include key steps, it will score 0 points. Pre-training GPU card communication performance test: The instruction agent executes the GPU card communication test. If it can correctly call the NCCL test tool and output test results with bandwidth ≥200GB / s and latency ≤5us and the analysis is accurate, it will get 7 points. If the result is relatively accurate, it will get 3 points. If it is inaccurate, it will get 0 points. Operation and maintenance capability test: Log analysis agent testing: Inject 100,000 mixed logs containing 100 preset anomalies, such as sudden increases in error codes and cross-service errors, including text, Syslog, and JSON. Verify that the accuracy of key field extraction is ≥95% and the recall rate of abnormal events is ≥90%. Insert 20% irrelevant information, such as garbled characters and non-standard timestamps. Test that the effective data retention rate after cleaning is ≥90%. Load test 100,000 logs / second, and monitor the processing latency ≤1 second. Fault location agent test: Simulate a fault chain, with service A timeout → service B circuit breaker 5 times, verify that the location result includes service A and the priority is correct ≥4 times; inject contradictory indicators (CPU is normal but error rate is increased) 10 times, test the multi-dimensional verification of logical correctness ≥9 times; dynamically update the service dependency graph (add / delete nodes) 5 times, verify the real-time adaptability ≥4 times. Root Cause Analysis Agent Test: Input logs and metrics data for 10 known faults. If ≥9 root cause conclusions are consistent with historical records, 2 points are awarded. Insert 5 sets of irrelevant change data. If ≥4 sets can filter out noise, 1 point is awarded. Execute 5 remediation suggestions in a sandbox. If ≥4 actually solve the problem, 2 points are awarded. That is, if the fault recurrence rate decreases by ≥90% and the service recovery time is ≤5 minutes after executing ≥4 suggestions, 2 points are awarded. Open Capability Test: API Functionality Test: Use Postman to send 100 API requests containing normal parameters, error parameters, and null values. If all responses meet expectations (i.e., parameters are passed correctly, return values ​​are standardized, and exception handling is reasonable), you will get 3 points; otherwise, you will get 0 points. User interface test: 10 users complete operations such as asking questions, querying, and exporting reports through the interface. If ≥9 users find the interface intuitive and easy to use, 3 points are awarded; 5-8 users receive 2 points; and <5 users receive 0 points. Long-term operation test: Let the agent run continuously for 24 hours, send 100 requests per hour, and record the response stability; intentionally interrupt the network 5 times, each time for 10 minutes, and simulate GPU failure 2 times. If there is no crash and the failure recovery time is ≤5 minutes, you will get 4 points; if you recover part of the failure, you will get 2 points; if you are unstable or cannot recover, you will get 0 points. Scoring and rating: Scoring Summary: After obtaining the scores for each test item according to the above testing steps, the scores are weighted and calculated. For example, the total score for the large model capability is calculated as follows: score for sub-items such as text parsing × 20% + score for model inference performance × 10% + score for instruction compliance capability × 10%. The other dimensions are similar. Finally, the comprehensive score is obtained by summing the scores. Rating: The average score is calculated after three repeated tests. 0-39 points is 1 star (lowest maturity), 40-79 points is 2 stars, and 80-100 points is 3 stars (highest maturity).

[0022] Based on the above description, this invention conducts maturity testing on the intelligent computing network management and maintenance intelligent agent system's functions from multiple dimensions, including model inference, multi-tenant isolation, fault location agent, root cause analysis, and API functionality, focusing on its large model capabilities, memory capabilities, tool capabilities, operation and maintenance capabilities, and openness capabilities. It also provides testing methods for various indicators, offering diverse test indicators to standardize the maturity assessment of the intelligent computing network management and maintenance intelligent agent system from all angles. Scoring is provided for each indicator, and scoring reference standards are given, assigning different weights based on the importance of each indicator. This ensures that the maturity evaluation of the intelligent computing network management and maintenance intelligent agent system truly reflects actual needs.

[0023] The capability maturity assessment method for the intelligent computing network management and operation and maintenance agent system oriented towards multimodality also includes the following steps: S8: Establish a dynamic adaptation and iterative optimization mechanism, including: real-time collection of multimodal dynamic data of the intelligent computing network operation, and adjustment of the weight coefficients of each test item based on changes in data characteristics; periodic execution of the re-evaluation process, and generation of maturity trend curves based on historical evaluation results; optimization of test items and scoring rules based on user feedback and actual application scenario requirements. The real-time collected multimodal dynamic data includes network traffic fluctuation data, including peak traffic, bandwidth utilization, and latency jitter characteristics; multimodal interaction request frequency; and fault type change trends. The weight coefficient adjustment is based on the entropy method or the analytic hierarchy process (AHP). The entropy method calculates the weight coefficients of each data item by... The information entropy of a feature, such as the entropy value H = -Σpi×lnpi of the proportion of voice commands, where pi is the proportion of the i-th type of command, quantifies the importance of the feature. The analytic hierarchy process (AHP) constructs a judgment matrix using a 1-9 scale. For example, if the importance of voice commands is higher than that of text commands, the scale is 3 to determine the weight allocation. When the change rate of a certain type of data feature exceeds a preset threshold, such as 20%, a weight update is triggered. For example, if the change rate of voice command data features exceeds 20%, the periodic re-evaluation period can be set to 1 month, 3 months, or 6 months, dynamically configured based on the iteration speed of intelligent computing network services. For example, if ≥3 new functional modules are added each month, the period is 1 month; if 1-2 new modules are added, the period is 3 months. The specific technical solution is as follows: Real-time data acquisition and weight adjustment: Deploy a data acquisition module to collect multimodal dynamic data in real time during the operation of the intelligent computing network, including network traffic fluctuation data such as peak traffic period distribution, multimodal interaction request frequency such as changes in the proportion of text and voice commands, and fault type change trends such as the proportion of new faults. Calculate the data feature change rate based on the entropy method. When the change rate of a certain type of data feature exceeds 20%, trigger the weight adjustment mechanism. Reallocate the weight of the corresponding test item through the analytic hierarchy process. For example, when the proportion of voice commands increases from 10% to 30%, increase the weight of the semantic understanding depth sub-item. Periodic reassessment and trend analysis: Based on the iteration speed of intelligent computing network services, the reassessment cycle is configured, such as a 1-month cycle for e-commerce scenarios and a 3-month cycle for financial scenarios. After each reassessment, a maturity trend curve is generated by combining historical assessment results to identify bottlenecks in capability improvement, such as root cause analysis of the agent's capabilities being consistently below average, and key areas for optimization are marked. User feedback closed-loop optimization: Build a user feedback interface to collect feedback from enterprise users on problems encountered in actual applications, such as the lack of cross-vendor device adaptation instructions in multi-tenant isolation deployment guidance. Perform cluster analysis on the feedback content every quarter, and add or optimize test items, such as supplementing cross-vendor device isolation compatibility test sub-items and scoring rules, such as increasing the scoring weight of compatibility indicators. Implementation steps of dynamic adaptation and iterative optimization mechanism: Real-time data acquisition and weight adjustment: The data acquisition module captures network traffic data every 5 minutes, recording peak traffic and protocol type percentages; the multimodal interaction logs statistically analyze the number and percentage of text, voice, and image commands; and the fault alarm data categorizes and statistically analyzes the number of hardware, software, and network faults.

[0024] Calculate the rate of change of data characteristics each week. For example, if the proportion of voice commands increases from 15% last week to 25% this week, the rate of change = (25% - 15%). 15%≈66.7%>20% triggers the adjustment of the weight of the semantic understanding depth sub-item. The weight is increased from the original 7 points to 8 points through the hierarchical analysis method, which corresponds to the redistribution of the sub-item weight under the large model capability dimension.

[0025] Periodic reassessment and trend analysis: For e-commerce scenarios, a 1-month re-evaluation cycle is configured because at least 3 new functional modules are added each month. For financial scenarios, a 3-month cycle is configured because 1-2 new functional modules are added each month. The complete evaluation process is executed on the start date of the cycle. A trend curve is generated by comparing the results of the last 3 evaluations. If the root cause analysis agent score is below 60 points for 2 consecutive months (out of 100), it will be marked as a key optimization item in the trend report, and the developer will be prompted to supplement historical failure case training data. User feedback closed-loop optimization: Users submitted feedback through the feedback interface regarding issues where the multi-tenant isolation deployment guidelines did not cover hybrid networking scenarios involving Huawei and Cisco equipment. The collected feedback was clustered quarterly, revealing that cross-vendor compatibility-related feedback accounted for 30%. A new sub-project, cross-vendor device isolation compatibility test, has been added and incorporated into the tool capability dimension, with a weight of 3%. This sub-project is separated from the original weight of the intelligent agent multi-tenant isolation deployment guidance capability. The scoring rules now include an indicator of the accuracy of guidance in hybrid networking scenarios, which accounts for 40% of the score for this sub-project.

[0026] Based on the above description, this invention utilizes a dynamic adaptation and iterative optimization mechanism to adjust weights through real-time data feedback, periodically re-evaluate and analyze trends, and optimize rules based on user feedback. This enables the evaluation method to adapt to the dynamic changes in intelligent computing network business scenarios and the iterative upgrades of multimodal data characteristics, ensuring the timeliness and practicality of the evaluation results and solving the problem of insufficient adaptability of traditional static evaluation methods.

[0027] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.

Claims

1. A method for assessing the capability maturity of a multimodal intelligent computing network management and control operation and maintenance intelligent agent system, characterized in that: Includes the following steps: S1: Determine the evaluation dimensions, which include large model capabilities, memory capabilities, tool capabilities, operation and maintenance capabilities, and openness capabilities; S2: Set up subordinate test items for each evaluation dimension, and assign weights based on the importance of each dimension and test item; S3: Set up the test environment, which includes both hardware and software environments; S4: Execute preset test steps for each test item and obtain test results; S5: Score the test results of each test item according to the preset scoring rules; S6: The scoring results are weighted and summed according to the weight of each test item to obtain the overall system score; S7: Determine the system capability maturity level based on the comprehensive score, which includes 1-star, 2-star and 3-star.

2. The capability maturity assessment method for a multimodal intelligent computing network management and maintenance intelligent agent system according to claim 1, characterized in that: In step S1, the sub-test items for the large model capability include text parsing accuracy, semantic understanding depth, contextual coherence, model inference performance, instruction compliance, and task completion capability; the sub-test items for the memory capability include storage capacity, knowledge storage and retrieval, information updating, and forgetting mechanism; the sub-test items for the tool capability include the ability to guide multi-tenant isolated deployment of intelligent agents and the pre-training GPU card aggregation communication performance of intelligent agents; the sub-test items for the operation and maintenance capability include log analysis intelligent agents, fault location intelligent agents, and root cause analysis intelligent agents; and the sub-test items for the open capability include API function testing, user interface testing, and long-term operation testing.

3. The capability maturity assessment method for a multimodal intelligent computing network management and maintenance intelligent agent system according to claim 1, characterized in that: In step S2, the weight of the large model capability is 40%, of which the weights of text parsing accuracy, semantic understanding depth, and contextual coherence are 20%, model inference performance is 10%, and instruction compliance and task completion capability is 10%; the weight of the memory capability is 20%, of which the weights of storage capacity, knowledge storage and retrieval, information update, and forgetting mechanism are each 5%; the weight of the tool capability is 15%, of which the weight of agent multi-tenant isolation deployment guidance capability is 8%, and the weight of agent pre-training GPU card aggregation communication performance is 7%; the weight of the operation and maintenance capability is 15%, of which the weights of log analysis agent, fault location agent, and root cause analysis agent are each 5%; and the weight of the open capability is 10%, of which the weight of API function testing is 3%, the weight of user interface testing is 3%, and the weight of long-term operation testing is 4%.

4. The capability maturity assessment method for a multimodal intelligent computing network management and maintenance intelligent agent system according to claim 1, characterized in that: In step S3, the hardware environment includes: a server equipped with a high-performance CPU, such as an Intel Xeon Gold series CPU; large-capacity memory, such as 256GB or more; and high-speed storage devices, such as NVMe SSDs; a GPU, using NVIDIA brand high-performance GPUs, such as A100 and H100; network devices, including switches and routers that support VLAN segmentation and network isolation to ensure low latency and high bandwidth data transmission; and a software environment including: an operating system, using a mainstream Linux operating system, such as Ubuntu 20.04LTS; programming languages ​​and frameworks, with compatible versions of languages ​​and frameworks installed according to the model development requirements; and testing tools, including Postman for API function testing, JMeter for performance testing, and ELKS tack for log analysis.

5. The capability maturity assessment method for a multimodal intelligent computing network management and maintenance intelligent agent system according to claim 2, characterized in that: In step S4, the testing steps for text parsing accuracy, semantic understanding depth, and contextual coherence include: constructing diverse input texts, including question-and-answer, dialogue, long documents, complex sentence structures, and domain terms, to verify the accuracy of keyword extraction and semantic classification; inputting sentences containing polysemous words or ambiguous references to verify the success rate of reference parsing; and providing multilingual mixed input to verify translation consistency and cross-language semantic retention rate.

6. The capability maturity assessment method for a multimodal intelligent computing network management and maintenance intelligent agent system according to claim 2, characterized in that: In step S4, the testing steps for model inference performance include: performing inference using standard text datasets such as SQuAD and GLUE, and recording the number of sentences processed per second (QPS), memory usage, and GPU utilization; gradually increasing the length of the input text to verify whether the system's performance degradation is ≤10% when the load increases; and adjusting model parameters such as batch size and sequence length to verify the balance between memory usage and computational efficiency.

7. The capability maturity assessment method for a multimodal intelligent computing network management and maintenance intelligent agent system according to claim 2, characterized in that: In step S5, the scoring rules include: Text parsing accuracy: Keyword extraction accuracy ≥95% earns 7 points, 90%-94% earns 5 points, <90% earns 0 points; Semantic understanding depth: Semantic classification accuracy ≥90% earns 7 points, 80%-89% earns 4 points, <80% earns 0 points; Contextual coherence: Referential parsing success rate ≥90% earns 6 points, 80%-89% earns 3 points, <80% earns 0 points; In model inference performance, inference speed meets the standard and GPU usage is reasonable, each earns 2 points, load stability earns 3 points, and resource optimization meets the standard, each earns 3 points; Passing the storage capacity test earns 5 points, failing earns 0 points; Knowledge storage and retrieval answer accuracy earns 5 points, relatively accurate earns 3 points, and inaccurate earns 0 points.

8. The capability maturity assessment method for a multimodal intelligent computing network management and maintenance intelligent agent system according to claim 1, characterized in that: In step S7, the comprehensive score range for 1 star is 0-39, the comprehensive score range for 2 stars is 40-79, and the comprehensive score range for 3 stars is 80-100. The comprehensive score needs to be determined after multiple measurements, such as 3-10 times.

9. The capability maturity assessment method for a multimodal intelligent computing network management and maintenance intelligent agent system according to claim 1, characterized in that: The capability maturity assessment method for the intelligent computing network management and operation and maintenance intelligent agent system oriented towards multimodality also includes the following steps: S8: Establish a dynamic adaptation and iterative optimization mechanism, including: real-time collection of multimodal dynamic data of the intelligent computing network operation, and adjustment of the weight coefficients of each test item based on changes in data characteristics; periodic execution of the re-evaluation process, and generation of maturity trend curves based on historical evaluation results; and optimization of test items and scoring rules based on user feedback and actual application scenario requirements.

10. The capability maturity assessment method for a multimodal intelligent computing network management and maintenance intelligent agent system according to claim 9, characterized in that: In step S8, the real-time collected multimodal dynamic data includes network traffic fluctuation data, multimodal interaction request frequency, and fault type change trend; the weight coefficient adjustment is based on the entropy method or the hierarchical analysis method, and the weight update is triggered when the change rate of a certain type of data feature exceeds a preset threshold, such as 20%, for example, the change rate of voice command data features exceeds 20%; the periodic re-evaluation period can be set to 1 month, 3 months or 6 months, and is dynamically configured based on the iteration speed of intelligent computing network services.

Citation Information

Cited By

  • End-to-end AI agent evaluation method, system, equipment and medium

    CN121900800A