An intelligent evaluation method and system for network security and informationization double evaluation scenarios
Patent Information
- Application Number
- CN202611316289.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-28
- Publication Date
- 2026-09-29
AI Technical Summary
一、AI评价结论可信度可量化。通过四维置信度融合算法,从知识库、自评估、上下文和知识图谱四个维度对AI生成的评价结论进行置信度评估,按路由决策树将结论分流至直接输出、人工审核或强制确认,有效解决AI幻觉导致的合规风险,使评价结论具备可解释性和可追溯性。
Smart Images

Figure CN122845294A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information network security technology, specifically to an intelligent assessment method and system for scenarios involving both network security and information technology assessment. Background Technology
[0002] In large enterprise groups (especially in critical infrastructure industries such as energy, finance, and telecommunications), cybersecurity assessment and information technology assessment (referred to as "dual assessment") are routine core tasks to ensure the company's compliant operation. Taking a large energy group as an example, its assessment covers 132 directly affiliated units, involving 126 clauses of cybersecurity standards and 41 information technology assessment indicators. There are 5,166 potential mapping relationships between the two sets of standards, and each unit needs to prepare more than 2,800 documents per round of assessment.
[0003] In existing technologies, the dual-evaluation process mainly relies on manual methods, which has the following significant drawbacks: 1. The credibility of AI conclusions cannot be quantified. Pure large-scale model solutions lack a confidence assessment mechanism, which poses a risk of AI illusion. They cannot conduct multi-dimensional confidence assessments of AI-generated evaluation conclusions, resulting in high-risk conclusions not being effectively intercepted.
[0004] 2. Multi-source standard ontologies cannot be automatically aligned. Existing technologies only construct ontologies for a single standard, which cannot accommodate the complete entity-relationship space of two standard systems. They also lack an automatic mapping mechanism for control items. Evidence for the same control item needs to be collected separately in two evaluations, resulting in duplicate evaluations.
[0005] 3. Risk assessment cannot be identified early. Existing solutions are all post-assessment methods, lacking an early warning mechanism based on a multi-factor weighted model. High-risk items are often only discovered in the later stages of the assessment, and the accuracy rate of manual identification is only 60-70%.
[0006] 4. Evaluation materials cannot be generated intelligently. The material retrieval list relies entirely on manual compilation. Each unit needs to manually sort through thousands of material items per round, which takes more than 3 days. Furthermore, historical drafts cannot be reused, resulting in serious duplication of work.
[0007] In summary, existing technologies cannot simultaneously solve the four core problems of credible quantification of AI conclusions, automatic alignment of multi-source standard ontology, early identification of evaluation risks, and intelligent generation of evaluation materials. Summary of the Invention
[0008] To address the shortcomings of existing technologies, this invention provides an intelligent assessment method and system for dual assessment scenarios of network security and information technology. It adopts an architecture that coordinates four core algorithms and five intelligent agents, and can simultaneously achieve credible quantification, dual standard alignment, risk warning, and intelligent material generation.
[0009] This invention is achieved through the following technical solution: An intelligent assessment method for dual assessment scenarios of cybersecurity and information technology is provided, including the following steps: Step A: Obtain network security evaluation indicators and information technology evaluation indicators and corresponding evidence materials. Input the evaluation indicators and evidence materials into the big language model, and the big language model will generate evaluation conclusions. Step B involves calculating confidence scores for the evaluation conclusions generated in Step A from four dimensions: knowledge base, self-assessment, context, and knowledge graph. These scores are then weighted and fused according to preset weights to obtain a comprehensive confidence score. Based on the comprehensive confidence score, the evaluation conclusions are routed to one of the following processing paths according to preset routing rules: direct output, manual review, or mandatory confirmation. Step C involves pairing control measures between network security standards and information evaluation standards, calculating semantic similarity and overlap of control objectives, and determining four types of semantic relationships—equivalence mapping, reinforcement mapping, complementary mapping, and mutually exclusive mapping—according to preset thresholds. Equivalence mapping triggers automatic mutual recognition and reuse of evidence, while mutually exclusive mapping triggers conflict detection and manual arbitration. Step D: For each evaluation indicator, calculate the weighted score of multiple risk factors in parallel. Risk factors include the impact of standard changes, historical non-compliance, lack of evidence, veto power, and low confidence. The total risk score is obtained by weighted summation of the contribution values of each factor, and then mapped to the risk level based on the total risk score. Step E: Load the set of evaluation indicators applicable to the task, obtain the evidence mapping tree corresponding to each indicator through graph database query. The evidence mapping tree includes a mandatory evidence layer, an enhanced evidence layer and a template layer. Merge historical draft information, call the large language model to suggest responsible persons and deadlines, and generate a material retrieval list.
[0010] Furthermore, in step B, The knowledge base dimension is calculated based on the recall and precision of vector retrieval; the self-evaluation dimension is obtained through the internal self-evaluation model of the large language model; the context dimension is obtained through historical case similarity matching; and the knowledge graph dimension is calculated as the ratio of the number of matched rules to the number of candidate rules. When the difference between any dimension and the overall confidence score exceeds a preset threshold, it is determined to be a multi-dimensional inconsistency. The overall confidence score is then downgraded by a preset downgrade coefficient and routed to manual review.
[0011] Furthermore, the knowledge base dimension is calculated as 0.6 × recall + 0.4 × precision, and the top-k=10 for vector retrieval; The context dimension is obtained by matching the top-k=5 similarity values of historical cases. The knowledge graph dimension employs AMIE3 Horn rule reasoning; If the difference between any dimension and the overall confidence score is greater than 0.25, the overall confidence score will be multiplied by 0.85 and downgraded.
[0012] Furthermore, the routing rules in step B are as follows: When the overall confidence level is ≥0.75, there are no multi-dimensional conflicts, and the number of pieces of evidence is greater than 0, the route is to the direct output; When 0.50 ≤ overall confidence level < 0.75 and there are no multi-dimensional conflicts, the result is routed to manual review; When the overall confidence level is less than 0.50 or the number of pieces of evidence is equal to 0, the route is to mandatory confirmation.
[0013] Furthermore, in step C, The criteria for determining equivalence class mapping are semantic similarity ≥ 0.85 and control target overlap ≥ 0.80; The criteria for determining reinforced class mapping are that the stringency of network security measures is greater than that of information technology measures and the semantic similarity is ≥0.60; The criterion for complementary mapping is 0.50 ≤ overlap of control targets ≤ 0.85; Mutual exclusion mappings are determined using the following conflict detection rules: Rule 1: When the access permission attribute of information control measures in the knowledge graph is "prohibited" and the access permission attribute of network security control measures is "allowed", and an administrator or privileged role is involved, it is considered a conflict; Rule 2: When the retention period attribute value of information control measures in the knowledge graph is less than the retention period attribute value of network security control measures, and audit logs or operation records are involved, it is determined to be a conflict; Rule 3: Use the large language model with a temperature parameter of 0.0 to determine substantial conflicts between the paired control measure texts; When any rule is determined to be in conflict, a mutual exclusion class mapping is triggered, and the conflict detection result is written into the cross-standard benchmarking layer of the knowledge graph, triggering a manual arbitration process.
[0014] Furthermore, in step D, the risk level mapping rule is as follows: When the total risk score is ≥0.60, the risk level is severe; When 0.40 ≤ Total Risk Score < 0.60, the risk level is high; When 0.20 ≤ Total Risk Score < 0.40, the risk level is medium. When the total risk score is less than 0.20, the risk level is low. When the risk level is high or severe, an alarm is triggered and pushed to the compliance checker's intelligent agent.
[0015] Furthermore, in step E, The mandatory evidence layer consists of the evidence materials that must be provided to complete the evaluation of this indicator; the enhanced evidence layer consists of supplementary evidence materials that improve the completeness of the evaluation; and the template layer consists of the format templates for the evaluation materials. Historical draft information consists of compliant materials collected in the previous year, which are automatically matched and reused with indicator identifiers stored in the knowledge graph; The responsible person and the deadline are suggested by the big language model based on the organizational context, which includes the unit and department structure, personnel positions and historical responsibility assignment records obtained from the relational database. The output of the big language model is written into the material retrieval list and persisted to the relational database.
[0016] This invention also provides a system using an intelligent assessment method for dual assessment scenarios of network security and information technology, employing a layered architecture, including: The presentation layer provides integration interfaces for web workbench, mobile devices, and collaborative office platforms; The intelligent agent layer includes five intelligent agents: issue planner, compliance checker, real-time minutes officer, action item manager, and case extractor, which are used to collaboratively execute evaluation tasks. The capability engine layer includes four core algorithm engines: a four-dimensional confidence fusion module, a bi-standard ontology alignment module, a high-risk item identification module, and a material retrieval list generation module. Among them, the four-dimensional confidence fusion module is used to execute step A; the bi-standard ontology alignment module is used to execute step B; the high-risk item identification module is used to execute step C; and the material retrieval list generation module is used to execute step D. The IT innovation foundation layer includes large language models, vector libraries, knowledge graphs, and relational databases, providing computing power and data support for the upper layers; The output of the bi-standard ontology alignment module serves as the grading basis for the material retrieval list generation module. The completion rate of the material retrieval list serves as an influencing factor for the recall rate of the knowledge base dimension in the four-dimensional confidence fusion module. The output of the four-dimensional confidence fusion module serves as the input for the low confidence factor in the high-risk item identification module.
[0017] Furthermore, in the foundational layer of information technology innovation, the large language model adopts a domestically developed large model, the vector library adopts a domestically developed vector database, the embedding model adopts a domestically developed pre-trained language model, the knowledge graph adopts a domestically developed graph database, the relational database adopts a domestically developed relational database, and the national cryptographic module adopts at least one of the domestic cryptographic algorithms SM2 / SM3 / SM4.
[0018] Among them: the domestic large model is Qwen2-7B model fine-tuned by QLoRA INT4; the domestic vector database is Milvus 2.4.6 domestic version; the domestic pre-trained language model is BGE-large-zh-v1.5 1024-dimensional vector; the domestic graph database is Neo4j 5.16 domestic version; the domestic relational database is DM8; the national cryptographic module adopts one of the domestic cryptographic algorithms SM2, SM3 and SM4.
[0019] Furthermore, the collaborative relationship among the five agents is as follows: The issue planner agent is responsible for evaluating task planning and indicator decomposition. The compliance verification agent is responsible for verifying the compliance of the evaluation conclusions based on the four-dimensional confidence fusion results. The real-time minutes officer AI is responsible for evaluating the real-time recording and minutes generation of the meeting; The Action Item Manager intelligent agent is responsible for tracking and closed-loop management of rectification tasks; The case refiner agent is responsible for extracting reusable experiences and templates from historical evaluation cases; When the compliance checker receives a high-risk alert, it triggers a manual review process.
[0020] The beneficial effects of this invention are: I. The credibility of AI evaluation conclusions can be quantified. Through a four-dimensional confidence fusion algorithm, the confidence of AI-generated evaluation conclusions is assessed from four dimensions: knowledge base, self-assessment, context, and knowledge graph. The conclusions are then routed to direct output, manual review, or mandatory confirmation according to the routing decision tree, effectively solving the compliance risks caused by AI illusions and making the evaluation conclusions interpretable and traceable.
[0021] II. Automatic Alignment of Dual Standards and Cross-System Evidence Recognition. Through a dual-standard ontology alignment algorithm, the semantic similarity and overlap of control objectives are calculated for the pairing of control measures between cybersecurity standards and information technology evaluation standards, automatically determining four types of mapping relationships: equivalence, reinforcement, complementarity, and mutual exclusion. Equivalence mappings automatically recognize and reuse evidence, avoiding duplicate collection; mutual exclusion mappings trigger conflict detection and manual arbitration to ensure consistency in evaluation conclusions.
[0022] Third, early warning of risk assessment. Through a high-risk item identification algorithm, based on five risk factors in parallel calculations—impact of standard changes, historical non-compliance, lack of evidence, veto power, and low confidence—the risk level is mapped to a risk level through weighted summation, enabling early identification of high-risk assessment items and transforming passive rectification into proactive prevention.
[0023] IV. One-click generation of evaluation materials. Through an automatic material retrieval list generation algorithm, the system retrieves the three-layer evidence mapping tree corresponding to the indicators based on graph database queries, automatically merges historical drafts, and calls a large language model to suggest responsible persons and deadlines, thereby realizing the intelligent generation of the material retrieval list.
[0024] V. Full-Process Intelligentization and Full-Stack Information Technology Innovation Deployment. This invention adopts a PCE-A4 layered architecture. The capability engine layer deploys four core algorithms to form a data closed loop—the dual-standard alignment results provide a grading basis for the material list, the completion status of the material list is fed back to the knowledge base dimension of confidence fusion, and the confidence assessment results serve as input for risk identification. The intelligent agent layer deploys five intelligent agents (issue planner, compliance checker, real-time minutes officer, action item manager, and case extractor) for layered collaboration, realizing full-process intelligentization of dual evaluation. The information technology innovation foundation layer supports full-stack domestic components and domestic cryptographic algorithms to meet data security compliance requirements. Attached Figure Description
[0025] Figure 1 This is a diagram of the overall system architecture of the present invention.
[0026] Figure 2 This is a flowchart of the four-dimensional confidence fusion algorithm of the present invention.
[0027] Figure 3 This is a flowchart of the bi-label ontology alignment algorithm of the present invention.
[0028] Figure 4 This is a flowchart of the high-risk item identification algorithm of the present invention.
[0029] Figure 5 A timing diagram is automatically generated for the material retrieval list of this invention.
[0030] Figure 6 This is a diagram showing the collaborative closed-loop relationship of the four algorithms in this invention. Detailed Implementation
[0031] To clearly illustrate the technical features of this solution, the following detailed implementation method will be used to explain the solution.
[0032] Example 1: System Overall Architecture The system of this invention adopts a PCE-A4 four-layer architecture, which consists of the presentation layer, agent layer, capability engine layer, and foundation layer from top to bottom, with each layer being decoupled from the others.
[0033] The presentation layer provides integration interfaces for web workbench, mobile devices, and collaborative office platforms (such as Lark and DingTalk), providing users with an interactive interface.
[0034] The agent layer consists of five agents that work together to perform the evaluation task: Issue planner: Responsible for evaluating task planning and indicator breakdown; Compliance Verifier: Responsible for verifying the compliance of evaluation conclusions based on the four-dimensional confidence fusion results; Real-time minutes officer: Responsible for evaluating the real-time recording and minutes generation of meetings; Action Project Manager: Responsible for tracking and closed-loop management of rectification tasks; Case study specialist: Responsible for extracting reusable experiences and templates from historical evaluation cases.
[0035] The capability engine layer deploys four core algorithm engines: Four-dimensional confidence fusion module: used to execute the four-dimensional confidence fusion algorithm; Bi-label ontology alignment module: used to execute the bi-label ontology alignment algorithm; High-risk item identification module: used to execute the high-risk item identification algorithm; Material retrieval list generation module: Used to execute the automatic generation algorithm for the material retrieval list.
[0036] The four core algorithm engines form a collaborative closed loop: the output of the bi-standard ontology alignment module serves as the grading basis for the material retrieval list generation module, the completion status of the material retrieval list is fed back to the knowledge base dimension of the four-dimensional confidence fusion module, and the output of the four-dimensional confidence fusion module serves as the input for the high-risk item identification module.
[0037] The IT innovation foundation layer includes a large language model (LLM), a vector library (VDB), a knowledge graph (KG), and a relational database (DB), providing computing power and data support for the upper layers.
[0038] The overall system architecture is as follows Figure 1 As shown ( Figure 1 This is a diagram of the overall system architecture, showing the positional relationships of the PCE-A4 four-layer architecture, four core algorithms, and five intelligent agents.
[0039] Example 2: Implementation Environment (Full-Stack Domestic IT Innovation Deployment) The implementation environment of this invention adopts a full-stack information technology stack, and the specific technology selection is shown in Table 1.
[0040] Table 1. Selection Table for Domestic Deployment Technologies AI computing power Ascend 910B 128GB HBM / Cambricon MLU370 BF16 has a computing power of ≥256 TFLOPS and supports INT8 / INT4 quantization deployment. CPU Hygon C86-3G / Kunpeng 920 — operating system Galaxy Kirin V10 SP3 / Tongxin UOS V20 Kernel version ≥ 4.19, containerization (Docker) + Kubernetes orchestration relational database DM8 v8.4.3.62 Oracle / PostgreSQL / MySQL protocol compatible, transparent switching of ORM middleware Vector library Milvus 2.4.6 (China's domestic IT innovation version) Storage Vector Embedding knowledge graph Neo4j 5.16 (China's domestic IT innovation version) Stored bi-rated knowledge graph entities and relations Large Language Model Qwen2-7B fine-tuned with QLoRA INT4 Unified API call, supporting flexible switching of model vendors Embedded Model BGE-large-zh-v1.5 (1024 dimensions) Used for semantic similarity calculation National cryptographic module SM2 / SM3 / SM4 SM2 digital signature + key exchange; SM3 integrity verification; SM4 encrypted storage The system is deployed and runs in a fully domestically produced environment. It uses SM4 encryption to store core evaluation data, and the cryptographic algorithm implementation complies with national cryptographic standards such as GB / T 32918 and GB / T 32907, fully meeting data security compliance requirements.
[0041] Example 3: Four-dimensional confidence fusion algorithm First, obtain network security evaluation indicators and information technology evaluation indicators and corresponding evidence materials. Then, input the evaluation indicators and evidence materials into the big language model, and the big language model will generate evaluation conclusions. like Figure 2 As shown, a comprehensive confidence score (0~1) is calculated for the evaluation conclusion using a four-dimensional confidence fusion algorithm, which is used to determine the routing strategy (direct output, manual review, or mandatory confirmation) to solve the trust problem of AI black box conclusions.
[0042] The four confidence dimensions and their weights are shown in Table 2.
[0043] Table 2. Four-dimensional confidence levels and their weights KB (Knowledge Base) 0.30 0.6 × Recall + 0.4 × Precision, top-k=10 for vector retrieval Authoritative Standard Clause Library Self-Eval 0.20 LLM Internal Self-Assessment Model Internal assessment of LLM Context 0.20 Historical case similarity matching top-k=5 Historical Evaluation Case Library KG (Knowledge Graph) 0.30 Number of matched rules / number of candidate rules, AMIE3 Horn rules knowledge graph Weighted fusion formula: Overall confidence level = 0.30 × KB + 0.20 × Self + 0.20 × Context + 0.30 × KG; Conflict resolution rules: When the difference between any dimension and the overall score is greater than 0.25, it is judged as a multi-dimensional inconsistency, automatically downgraded (overall confidence level × 0.85), and routed to manual review.
[0044] Routing decision tree rules: HIGH (Direct Output): Overall confidence level ≥ 0.75, no multi-dimensional conflicts, number of pieces of evidence > 0; MEDIUM (manual review recommended): 0.50 ≤ overall confidence level < 0.75, no multi-dimensional conflicts; LOW (Forced Confirmation): Overall confidence level < 0.50, or number of pieces of evidence = 0.
[0045] Boundary condition handling (7 items): If the evaluation conclusion text is empty, throw an "Evaluation conclusion text is empty" exception and terminate the current confidence calculation. If the evaluation indicator identifier does not exist, throw an "Evaluation indicator identifier does not exist" exception and terminate the current confidence calculation. When the knowledge base retrieval has no recall, the knowledge base dimension is scored as 0.5 (neutral score). When the historical case library is empty, the context dimension is counted as 0.7 (default value). When the knowledge graph is unavailable, the knowledge graph dimension is 0.6 (based on rule base degradation). When the number of pieces of evidence is 0, forced routing is initiated to forced confirmation. When the self-assessment model is unavailable, the base model is used as a fallback.
[0046] Performance metrics: Calculation time for a single conclusion <100ms.
[0047] Input data: Evaluation conclusion text: "In the 2025 cybersecurity evaluation of this unit, the access control policy documents were complete, the permission approval records were complete, and the access control measures were effective." Task ID: 10001; Evaluation indicator label: 8001: List of evidence cited: ["ev_001", "ev_002", "ev_003"].
[0048] Calculation process: KB dimension (knowledge base dimension): top-10 vector retrieval, recall rate 0.95, precision rate 0.88, KB score = 0.6 × 0.95 + 0.4 × 0.88 = 0.922.
[0049] Self dimension (self-evaluation dimension): LLM (Large Language Model) self-score: 0.88.
[0050] Context dimension: Top-5 historical case similarity matches, with a similarity of 0.85.
[0051] KG dimension (Knowledge Graph dimension): Number of matched rules 18 / Number of candidate rules 20 = 0.90.
[0052] Weighted fusion: The overall confidence score (final_score) = 0.30 × 0.922 + 0.20 × 0.88 + 0.20 × 0.85 + 0.30 × 0.90 = 0.8926.
[0053] Conflict detection: max(0.922, 0.88, 0.85, 0.90) - min = 0.922 - 0.85 = 0.072 < 0.25, no conflict.
[0054] Routing decision: If the overall confidence score (final_score) is 0.8926 ≥ 0.75, route to the direct output.
[0055] Example 4: Bi-label Ontology Alignment Algorithm The dual-standard ontology alignment algorithm is used to identify the semantic relationship between two sets of standards: the cybersecurity standard (126 clauses) and the information evaluation standard (41 indicators), and outputs four mappings (equivalence, reinforcement, complementarity, and mutual exclusion) to reduce redundant material preparation.
[0056] The four mapping types and their determination rules are shown in Table 3.
[0057] Table 3. Types and Judgment Rules for Bi-standard Ontology Alignment Mapping EQUIVALENT (equivalent) Semantic similarity ≥ 0.85 AND overlap (controlling target overlap) ≥ 0.80 ≥0.85 Automatic mutual recognition and reuse of evidence REINFORCE Cybersecurity strictness > Information technology AND similarity (semantic similarity) ≥ 0.60 ≥0.60 Prioritize meeting cybersecurity requirements COMPLEMENTARY 0.50 ≤ overlap (controlling target overlap) ≤ 0.85 0.50~0.85 Collect double standard evidence packages separately EXCLUSIVE (mutually exclusive) The `detect_conflict` function returns `True`. — Triggering manual arbitration Conflict detection employs three rules plus an LLM fallback: Rule 1 (Access Permission Direction Conflict): When the access permission attribute of information control measures in the knowledge graph is prohibited and the access permission attribute of network security control measures is allowed, and an administrator or privileged role is involved, it is determined to be a conflict; Rule 2 (Retention Period Conflict): When the retention period attribute value of information control measures in the knowledge graph is less than the retention period attribute value of network security control measures, and audit logs or operation records are involved, it is determined to be a conflict; Rule 3 (LLM fallback): Invoke the large language model to determine substantial conflicts between the paired control measure texts with a temperature parameter of 0.0; When any rule is determined to be in conflict, a mutual exclusion class mapping is triggered, and the conflict detection result is written into the cross-standard benchmarking layer of the knowledge graph, triggering a manual arbitration process.
[0058] Output merging strategy: EQUIVALENT: Automatic mutual recognition and reuse of evidence; UI prompts that this indicator is equivalent to XXX; REINFORCE: Prioritize cybersecurity requirements; IT requirements are automatically met. COMPLEMENTARY: Collect unique evidence separately; output double-standard evidence package; EXCLUSIVE: Triggers manual arbitration; UI pop-up indicates conflict.
[0059] Performance metrics: 41 × 126 = 5166 mappings pre-computed at once, pre-computed time < 5 minutes (based on BGE-large-zh + LLM check); runtime query < 50ms (cache hit).
[0060] The determination process of the bi-standard ontology alignment algorithm is as follows: Figure 3 As shown.
[0061] Take the alignment of the information technology evaluation indicator "audit log retention" with the cybersecurity standard clause "log retention" as an example.
[0062] The information technology evaluation indicator is named "Audit Log Retention," and its requirement is that "the audit log retention period shall not be less than 6 months." The network security standard clause is named "Log Retention," and its requirement is that "the audit log retention period shall not be less than 12 months, in accordance with Level 3 Information Security Protection Requirements."
[0063] First, the semantic similarity between the two texts was calculated. Using a Chinese text vector model, both texts were converted into vectors, and then the cosine similarity was calculated, resulting in a semantic similarity of 0.89. Second, the overlap of the control targets was calculated, yielding an overlap of 0.92. Furthermore, the stringency of the two standards was compared: the cybersecurity standard requires retention for 12 months, while the information technology standard requires retention for 6 months, indicating that the cybersecurity requirement is more stringent.
[0064] According to the judgment rules, a semantic similarity of 0.89 is greater than or equal to the threshold of 0.85, thus the condition is met; and a control target overlap of 0.92 is greater than or equal to the threshold of 0.80, thus the condition is met. Therefore, this information technology evaluation index and the cybersecurity standard clauses are determined to be an equivalence class mapping.
[0065] In terms of business output, the audit log materials for this IT evaluation indicator can directly reuse materials from cybersecurity standard clauses, avoiding duplicate data collection. The system interface displays a message stating, "This indicator is equivalent to the log retention of cybersecurity clauses." The audit log materials for this IT indicator can directly reuse materials from cybersecurity clauses, avoiding duplicate data collection.
[0066] Example 5: High-Risk Item Identification Algorithm The high-risk item identification algorithm is based on a multi-factor weighted model to achieve early identification of high-risk evaluation items, with an accuracy rate of over 92%.
[0067] The five-factor weighted scoring model is shown in Table 4.
[0068] Table 4. Five-Factor Weighted Model for High-Risk Item Identification Impact of standard changes 0.25 The standard for this indicator may change within 180 days (the new standard is stricter). New standards will be implemented within 180 days. Historical non-compliance 0.20 This indicator scored less than 60 last year. Last year's score <60 Degree of lack of evidence 0.20 Number of collected evidence < Number of required evidence Insufficient amount of evidence Veto 0.30 This indicator is a veto item under GA / T 2380. This is a veto item. low confidence 0.15 The confidence level of this indicator in four dimensions is <0.6. Confidence level < 0.6 The total risk score is obtained by summing the contribution values of the five risk factors (i.e., the product of each factor's weight and whether that factor was triggered). The risk level is then determined based on the total risk score. When the total risk score is greater than or equal to 0.60, the risk level is severe. When the total risk score is greater than or equal to 0.40 and less than 0.60, the risk level is high. When the total risk score is greater than or equal to 0.20 and less than 0.40, the risk level is medium. When the total risk score is less than 0.20, the risk level is low.
[0069] Algorithm Flow: After loading the indicators, the five factors are calculated in parallel, and the following steps are performed in sequence: standard change check, historical score query, evidence quantity comparison, veto judgment, and four-dimensional confidence query. Then, the weighted sum is calculated. When the score is ≥0.40, it is marked as high risk and an alarm is pushed to the compliance checker agent.
[0070] Accuracy metrics: Based on practical data from 132 organizations, the accuracy rate for identifying high-risk items is ≥92%.
[0071] like Figure 4 As shown, taking the data encryption transmission evaluation indicator as an example, this indicator is a veto item, and a new national standard was released within the last 180 days, resulting in a standard change. The historical score of this indicator in the previous evaluation cycle was 55 points, which is lower than 60 points; the number of pieces of evidence collected in this round of evaluation is 2, while the required number of pieces of evidence is 5, which is insufficient; in addition, the 4-dimensional confidence assessment result of this indicator is 0.55, which is lower than the threshold of 0.6.
[0072] Based on the above input, the five risk factors are determined as follows: First, the standard change impact factor was triggered because a new standard was released for this indicator within 180 days, contributing 0.25. Second, the historical non-compliance factor was triggered because last year's score of 55 was less than 60, contributing 0.20. Third, the evidence deficiency factor was triggered because the number of collected evidences (2) was less than the required 5, contributing 0.20. Fourth, the veto factor was triggered because this indicator falls under the veto category stipulated in GA / T 2380, contributing 0.30. Fifth, the low confidence factor was triggered because the four-dimensional confidence level of 0.55 was less than 0.6, contributing 0.15.
[0073] The contribution values of the five factors are summed to obtain a total risk score of 1.10, which is greater than 0.60. Therefore, the risk level is determined to be severe. The system marks this indicator as high risk, triggers an alert from the compliance checker's AI agent, and automatically pushes it to the evaluation working group leader, initiating a manual review process.
[0074] Example 6: Algorithm for Automatic Generation of Material Requisition List The automatic material retrieval list generation algorithm is based on the "indicator-evidence inverse mapping tree" and automatically generates a material retrieval list with responsible persons and deadlines.
[0075] The indicator-evidence mapping tree adopts a 3-layer structure: Indicator layer (e.g., access control under the information security compliance system) Required evidence layers: Access control policy file (PDF), screenshot of system permission configuration, permission approval flow record; Enhanced evidence layer (optional): Permission audit report, role-permission matrix; Template layer: Access control policy v2.3 template; Key steps of the algorithm: The loading task applies to a set of indicators (132 units, 167 indicators, and approximately 5,000 material items). Obtain the evidence mapping tree from the knowledge graph using Cypher queries; Merge historical documentation (automatically reuse compliance materials collected in the previous year); The large language model was used to suggest responsible departments and individuals; Recommend a deadline based on the organizational context; Persist to the material retrieval list.
[0076] Performance metrics: Generation time for 5000 material items <30 seconds; AI-suggested accuracy for responsible persons and deadlines ≥85%.
[0077] like Figure 5 As mentioned above, taking a complete task of retrieving dual evaluation materials as an example, the task identifier is 10001, which needs to cover all 167 evaluation indicators.
[0078] The algorithm execution process is as follows: The first step is to load all 167 evaluation indicators. The second step is to perform a query on the graph database for each indicator to obtain the corresponding evidence type information, including the evidence type name, whether it is mandatory, and its priority. The third step is to query the historical documentation database, automatically matching and reusing compliance materials collected in the previous year based on the indicator identifier to avoid duplicate collection. The fourth step is to use a large language model to suggest the responsible department and person for each new evidence item. The fifth step is to persistently store all results in the material retrieval list table.
[0079] After the algorithm completed execution, it generated more than 5,000 material retrieval list records. Each record includes the associated evaluation indicator identifier, evidence level marker (mandatory evidence, enhanced evidence, or template), responsible department (suggested by the large language model), responsible person identifier (suggested by the large language model), deadline (calculated based on organizational context), priority, and whether it was generated by AI.
[0080] In terms of performance, the generation time for a single batch of 5,000 material items is 28.3 seconds, which meets the target requirement of less than 30 seconds.
[0081] Example 7: Collaborative Closed Loop of Four Major Algorithms like Figure 6 As shown, the four algorithms form a complete collaborative closed loop in the business flow: Bi-standard alignment algorithm - Bill of Materials algorithm: The bi-standard alignment result provides a hierarchical basis for "mandatory / enhanced / template" in the generation of the bill of materials; Bill of Materials Algorithm - Confidence Algorithm: Bill of Materials Completeness is used as an influencing factor on the recall rate of the KB dimension in the confidence algorithm; Confidence Algorithm - Risk Identification Algorithm: The confidence assessment result serves as the input to the "low confidence" factor in the risk identification algorithm; Risk Identification Algorithm - Business Flow: High-risk item identification results trigger manual review routing or directly output decision.
[0082] Complete collaborative workflow: Evaluation task initiation - Bill of materials algorithm generates bill of materials - Material collection - Bi-standard alignment algorithm achieves bi-standard alignment - Self-assessment report generation - Confidence algorithm completes four-dimensional confidence assessment - Risk identification algorithm completes high-risk item identification: When it is high-risk, it is pushed to manual review; if it is determined not to be high-risk, the report is directly output.
[0083] Furthermore, the four-dimensional confidence fusion algorithm of this invention can also be applied to the confidence assessment of meeting minutes generation. When the confidence level is ≥0.75, the minutes are sent directly; when it is 0.50~0.75, they are sent to the secretary for review; and when it is <0.50, they are sent to the supervisor for confirmation. The dual-standard ontology alignment algorithm can also be applied to the alignment of "internal control compliance standards" and "anti-money laundering standards" in the financial industry, forming 2184 mapping pairs and reducing duplicate evidence collection by 25~35%. The high-risk item identification algorithm can also be applied to the risk assessment before the IT system goes live, with high-risk items automatically added to the review list of the "going live approval expert committee". The automatic generation algorithm for the material retrieval list can also be applied to the preparation of annual audit materials, generating more than 5000 material items in less than 30 seconds, and automatically suggesting the responsible department.
[0084] Of course, the above description is not limited to the examples above. Technical features not described in this invention can be implemented by or using existing technology, and will not be repeated here. The above embodiments and drawings are only used to illustrate the technical solutions of this invention and are not intended to limit this invention. This invention has been described in detail with reference to preferred embodiments. Those skilled in the art should understand that any changes, modifications, additions or substitutions made by those skilled in the art within the scope of this invention do not depart from the spirit of this invention and should also fall within the scope of protection of the claims of this invention.
Claims
1. An intelligent assessment method for dual assessment scenarios of network security and information technology, characterized in that: Includes the following steps: Step A: Obtain network security evaluation indicators and information technology evaluation indicators and corresponding evidence materials. Input the evaluation indicators and evidence materials into the big language model, and the big language model will generate evaluation conclusions. Step B involves calculating confidence scores for the evaluation conclusions generated in Step A from four dimensions: knowledge base, self-assessment, context, and knowledge graph. These scores are then weighted and fused according to preset weights to obtain a comprehensive confidence score. Based on the comprehensive confidence score, the evaluation conclusions are routed to one of the following processing paths according to preset routing rules: direct output, manual review, or mandatory confirmation. Step C involves pairing control measures between network security standards and information evaluation standards, calculating semantic similarity and overlap of control objectives, and determining four types of semantic relationships—equivalence mapping, reinforcement mapping, complementary mapping, and mutually exclusive mapping—according to preset thresholds. Equivalence mapping triggers automatic mutual recognition and reuse of evidence, while mutually exclusive mapping triggers conflict detection and manual arbitration. Step D: For each evaluation indicator, calculate the weighted score of multiple risk factors in parallel. Risk factors include the impact of standard changes, historical non-compliance, lack of evidence, veto power, and low confidence. The total risk score is obtained by weighted summation of the contribution values of each factor, and then mapped to the risk level based on the total risk score. Step E: Load the set of evaluation indicators applicable to the task, obtain the evidence mapping tree corresponding to each indicator through graph database query. The evidence mapping tree includes a mandatory evidence layer, an enhanced evidence layer and a template layer. Merge historical draft information, call the large language model to suggest responsible persons and deadlines, and generate a material retrieval list.
2. The intelligent assessment method for dual assessment scenarios of network security and informatization as described in claim 1, characterized in that: In step B, The knowledge base dimension is calculated based on the recall and precision of vector retrieval; the self-evaluation dimension is obtained through the internal self-evaluation model of the large language model; the context dimension is obtained through historical case similarity matching; and the knowledge graph dimension is calculated as the ratio of the number of matched rules to the number of candidate rules. When the difference between any dimension and the overall confidence score exceeds a preset threshold, it is determined to be a multi-dimensional inconsistency. The overall confidence score is then downgraded by a preset downgrade coefficient and routed to manual review.
3. The intelligent assessment method for dual assessment scenarios of network security and informatization as described in claim 2, characterized in that: The knowledge base dimension is calculated as 0.6 × recall + 0.4 × precision, and the top-k=10 for vector retrieval; The context dimension is obtained by matching the top-k=5 similarity values of historical cases. The knowledge graph dimension employs AMIE3 Horn rule reasoning; If the difference between any dimension and the overall confidence score is greater than 0.25, the overall confidence score will be multiplied by 0.85 and downgraded.
4. The intelligent assessment method for dual assessment scenarios of network security and informatization as described in claim 1, characterized in that: The routing rules in step B are as follows: When the overall confidence level is ≥0.75, there are no multi-dimensional conflicts, and the number of pieces of evidence is greater than 0, the route is to the direct output; When 0.50 ≤ overall confidence level < 0.75 and there are no multi-dimensional conflicts, the result is routed to manual review; When the overall confidence level is less than 0.50 or the number of pieces of evidence is equal to 0, the route is to mandatory confirmation.
5. The intelligent assessment method for dual assessment scenarios of network security and informatization as described in claim 1, characterized in that: In step C, The criteria for determining equivalence class mapping are semantic similarity ≥ 0.85 and control target overlap ≥ 0.80; The criteria for determining reinforced class mapping are that the stringency of network security measures is greater than that of information technology measures and the semantic similarity is ≥0.60; The criterion for complementary mapping is 0.50 ≤ overlap of control targets ≤ 0.85; Mutual exclusion mappings are determined using the following conflict detection rules: Rule 1: When the access permission attribute of information control measures in the knowledge graph is "prohibited" and the access permission attribute of network security control measures is "allowed", and an administrator or privileged role is involved, it is considered a conflict; Rule 2: When the retention period attribute value of information control measures in the knowledge graph is less than the retention period attribute value of network security control measures, and audit logs or operation records are involved, it is determined to be a conflict; Rule 3: Use the large language model with a temperature parameter of 0.0 to determine substantial conflicts between the paired control measure texts; When any rule is determined to be in conflict, a mutual exclusion class mapping is triggered, and the conflict detection result is written into the cross-standard benchmarking layer of the knowledge graph, triggering a manual arbitration process.
6. The intelligent assessment method for dual assessment scenarios of network security and informatization as described in claim 1, characterized in that: In step D, the risk level mapping rule is as follows: When the total risk score is ≥0.60, the risk level is severe; When 0.40 ≤ Total Risk Score < 0.60, the risk level is high; When 0.20 ≤ Total Risk Score < 0.40, the risk level is medium. When the total risk score is less than 0.20, the risk level is low. When the risk level is high or severe, an alarm is triggered and pushed to the compliance checker's intelligent agent.
7. The intelligent assessment method for dual assessment scenarios of network security and informatization as described in claim 1, characterized in that: In step E, The mandatory evidence layer consists of the evidence materials that must be provided to complete the evaluation of this indicator; the enhanced evidence layer consists of supplementary evidence materials that improve the completeness of the evaluation; and the template layer consists of the format templates for the evaluation materials. Historical draft information consists of compliant materials collected in the previous year, which are automatically matched and reused with indicator identifiers stored in the knowledge graph; The responsible person and the deadline are suggested by the big language model based on the organizational context, which includes the unit and department structure, personnel positions and historical responsibility assignment records obtained from the relational database. The output of the big language model is written into the material retrieval list and persisted to the relational database.
8. A system using the intelligent assessment method for dual assessment scenarios of network security and informatization as described in any one of claims 1 to 7, characterized in that: A layered architecture is adopted, including: The presentation layer provides integration interfaces for web workbench, mobile devices, and collaborative office platforms; The intelligent agent layer includes five intelligent agents: issue planner, compliance checker, real-time minutes officer, action item manager, and case extractor, which are used to collaboratively execute evaluation tasks. The capability engine layer includes four core algorithm engines: a four-dimensional confidence fusion module, a bi-standard ontology alignment module, a high-risk item identification module, and a material retrieval list generation module. Among them, the four-dimensional confidence fusion module is used to execute step A; the bi-standard ontology alignment module is used to execute step B; the high-risk item identification module is used to execute step C; and the material retrieval list generation module is used to execute step D. The IT innovation foundation layer includes large language models, vector libraries, knowledge graphs, and relational databases, providing computing power and data support for the upper layers; The output of the bi-standard ontology alignment module serves as the grading basis for the material retrieval list generation module. The completion rate of the material retrieval list serves as an influencing factor for the recall rate of the knowledge base dimension in the four-dimensional confidence fusion module. The output of the four-dimensional confidence fusion module serves as the input for the low confidence factor in the high-risk item identification module.
9. The intelligent assessment system for dual assessment scenarios of network security and informatization as described in claim 8, characterized in that: In the foundational layer of the information technology innovation platform, the large language model adopts a domestically developed large model, the vector library adopts a domestically developed vector database, the embedding model adopts a domestically developed pre-trained language model, the knowledge graph adopts a domestically developed graph database, the relational database adopts a domestically developed relational database, and the national cryptographic module adopts at least one of the domestic cryptographic algorithms SM2 / SM3 / SM4.
10. The intelligent evaluation system for dual evaluation scenarios of network security and informatization as described in claim 8, characterized in that: The collaborative relationship among the five agents is as follows: The issue planner agent is responsible for evaluating task planning and indicator decomposition. The compliance verification agent is responsible for verifying the compliance of the evaluation conclusions based on the four-dimensional confidence fusion results. The real-time minutes officer AI is responsible for evaluating the real-time recording and minutes generation of the meeting; The Action Item Manager intelligent agent is responsible for tracking and closed-loop management of rectification tasks; The case refiner agent is responsible for extracting reusable experiences and templates from historical evaluation cases; When the compliance checker receives a high-risk alert, it triggers a manual review process.