An artificial intelligence-based data security risk automatic evaluation method and system

By employing an AI-based automated data security risk assessment method, utilizing NLP and machine learning algorithms and a private large-scale model, the problem of low efficiency and poor adaptability in traditional assessments is solved. This enables efficient and accurate assessment of massive data assets and complex business scenarios, providing personalized risk management recommendations and supporting dynamic risk monitoring.

CN122113126APending Publication Date: 2026-05-29GOLDEN SHIELD TESTING TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GOLDEN SHIELD TESTING TECH CO LTD
Filing Date
2026-04-28
Publication Date
2026-05-29

Smart Images

  • Figure CN122113126A_ABST
    Figure CN122113126A_ABST
Patent Text Reader

Abstract

The application discloses a kind of data security risk automatic evaluation method and system based on artificial intelligence, it is related to AI model calling technical field, load evaluation index and preset task parameter, the initialization matching of to-be-tested environment is carried out, and evaluation benchmark model is obtained;Standardized data asset list is obtained by processing multi-source original data using NLP and machine learning algorithm, and compliance comparison analysis is carried out with the index in evaluation benchmark model, and preliminary risk list is obtained;The preliminary risk list is automatically desensitized, and the implicit risk mining and causal inference are carried out in the private large model, and the deep risk research result is obtained, then weighted calculation is carried out, and the quantized risk grade numerical value is obtained;According to the quantized risk grade numerical value, reasoning processing is carried out, and risk disposal suggestion priority and standardization evaluation report are obtained;Through evaluation report rectification retest, dynamic security risk situation data is obtained, realizes the timely early warning of data, solves the short board of traditional static evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of AI model invocation technology, and in particular to an automatic data security risk assessment method and system based on artificial intelligence. Background Technology

[0002] With the rapid development of the digital economy, data has become a core production factor. The scale of data in various organizations is growing explosively, the types of data are becoming increasingly complex, and the data flow process is becoming more diversified. As a result, the risks and challenges to data security are also intensifying. Problems such as data leakage, unauthorized access, data tampering, and lack of compliance occur frequently, seriously threatening the operation and development of organizations and the information security of users.

[0003] Data security risk assessment is a core means of identifying data security risks, quantifying risk levels, and formulating protection strategies. However, there are still many technical pain points in the field of data security risk assessment, which make it difficult to meet the large-scale, routine, and intelligent assessment needs of organizations in the digital age.

[0004] Traditional data security risk assessments rely heavily on manual asset inventory, risk identification, and risk level determination. This requires a large number of professionals, and the assessment process is cumbersome and time-consuming. For massive data assets and complex business scenarios, manual assessments are extremely inefficient and prone to omissions and errors. Secondly, manual assessments are limited by the professional experience and knowledge of the assessors, making it difficult to accurately identify hidden and related risks across different stages and scenarios. Furthermore, the assessment of the likelihood and impact of risks is mostly qualitative, lacking scientific quantitative evidence, thus limiting the objectivity and reference value of the assessment results. While existing assessment tools can partially comply with national and industry standards, they are mostly fixed templates and cannot adapt to different situations. Traditional risk indicator configurations are poorly tailored to an organization's business characteristics, data types, and security needs, and struggle to quickly respond to changes in standards and industry regulatory requirements. Furthermore, traditional assessments merely list general remedial suggestions based on risk points, failing to develop customized and implementable solutions tailored to the organization's actual technical architecture, protection capabilities, and operational levels. This results in assessment results being disconnected from actual security remediation efforts, failing to effectively guide data security protection work. Data assets and business scenarios are constantly changing, and traditional assessments, being phased static evaluations, cannot perceive new risks arising from data flow, security policy adjustments, and changes in external threats in real time, making dynamic risk monitoring and continuous assessment difficult. Summary of the Invention

[0005] The technical problem solved by this invention is that existing technologies are difficult to conduct efficient non-manual assessments of massive data assets and complex business scenarios, lack effective guidance for data security protection, and cannot achieve personalized risk indicator configuration according to the content needs of different organizations.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] Firstly, an automatic data security risk assessment method based on artificial intelligence includes the following steps:

[0008] Step S1: Load the data security assessment standard indicator library and preset task parameters, initialize and match the environment to be tested, and obtain the assessment benchmark model.

[0009] Step S2: Obtain multi-source raw data, and use NLP and machine learning algorithms to intelligently identify and classify the multi-source raw data to obtain a standardized data asset list;

[0010] Step S3: Compare and analyze the standardized data asset list with the indicators in the evaluation benchmark model to obtain a preliminary risk list;

[0011] Step S4: Automated desensitization processing is performed on the preliminary risk list, and the desensitized preliminary risk list is input into the private large model for implicit risk mining and causal inference to obtain in-depth risk assessment results. The in-depth assessment results are jointly verified by the preset risk rule base and the industry-specific small model, and the prompt words are dynamically adjusted for secondary inference. Finally, the quantitative risk level value is output.

[0012] Step S5: The private big model is used to reason based on the quantitative risk level values ​​to obtain risk disposal recommendations, priority, and standardized assessment reports.

[0013] Step S6 involves prioritizing risk management recommendations, conducting rectification and retesting, and carrying out regular patrol monitoring to obtain dynamic safety risk situation data.

[0014] Preferably, step S1 includes the following sub-steps:

[0015] Step S11: Obtain task parameters, which include the evaluation scope, industry-specific standards, and custom indicator weights.

[0016] Step S12: Match the data security evaluation standard indicator library with industry-specific standards one by one to construct a multi-dimensional indicator system. The multi-dimensional indicator system includes data asset identification, data classification and grading, security management system, technical protection measures, and personnel management and emergency response.

[0017] Step S13: Obtain the environment to be tested, adapt and map the multi-dimensional indicator system and custom indicator weights according to the business attributes of the environment to be tested, define the input and output rules for each stage of the evaluation process, and generate the evaluation benchmark model.

[0018] Preferably, step S2 includes the following sub-steps:

[0019] Step S21: Collect network data based on the principle of least privilege. The network data includes basic information, flow information, and attribute information. The collection methods include API interface, proxy collection, and log capture.

[0020] Step S22: Use NLP and feature extraction algorithms to parse network data, extract sensitive attribute features, and perform feature matching and recognition on encrypted network data to obtain feature labels;

[0021] Step S23: Input the network data and feature labels one-to-one into the pre-trained machine learning classification model to generate grade category labels, and save the grade category labels, network data and feature labels as a standardized data asset list.

[0022] Preferably, step S3 includes the following sub-steps:

[0023] Step S31: The standardized data asset list is automatically scanned using the indicator system of the evaluation benchmark model to obtain the life cycle stages. The life cycle stages include several nodes, specifically including collection, storage, use, processing, transmission, provision, disclosure and deletion.

[0024] Step S32: Detect the risk status of each node in the lifecycle process. The risk status includes a risk-free status and a risky status. The risky status includes unauthorized access, unencrypted data, and missing regulations.

[0025] Step S33: Extract the standardized data asset list corresponding to the nodes in a risky state to obtain a preliminary risk list.

[0026] Preferably, step S4 includes the following sub-steps:

[0027] Step S41: Perform automated desensitization processing on the preliminary risk list to obtain the desensitized preliminary risk list;

[0028] Step S42 involves inputting the anonymized preliminary risk list into the privatization model, specifically including:

[0029] By using a dedicated private interface adaptation layer and a lightweight calling protocol, the anonymized preliminary risk list is input into the locally deployed private large model in the intranet environment.

[0030] Preferably, step S4 further includes:

[0031] The privatization model performs impact risk mining and causal deduction on the initial risk list after anonymization to obtain in-depth risk assessment results, specifically including:

[0032] By combining the industry attributes of the environment under test, we can analyze the multi-dimensional impact of economic losses, reputational damage, and compliance penalties after the risk occurs, and summarize the results of the large model inference to generate in-depth analysis results.

[0033] The in-depth analysis results are double-verified to generate in-depth risk analysis results;

[0034] The processing logic for double-verification of in-depth analysis results is as follows:

[0035] The dual verification includes a first verification and a second verification;

[0036] The first verification includes:

[0037] The results of in-depth risk assessment are input into a preset risk assessment rule base to obtain rule verification labels. The rule verification labels include rule compliance and rule non-compliance. The risk assessment rule base is used to verify whether the results of in-depth risk assessment meet industry standard requirements.

[0038] The second verification includes:

[0039] A lightweight industry-specific small model is invoked to cross-validate the reasoning logic corresponding to the deep risk assessment results, and logical verification labels are obtained, including logical compliance and non-compliance.

[0040] If both the first and second checks match the tags, the result is considered passed, and the in-depth risk assessment result is output. Otherwise, the result is considered failed, the prompt words are automatically adjusted, and the private large model is triggered to perform secondary inference until both the first and second checks match, at which point the adjustment stops.

[0041] The results of the in-depth risk assessment are weighted and calculated to obtain a quantitative risk level value.

[0042] Preferably, the processing logic for weighted calculation of the in-depth risk assessment results to obtain a quantitative risk level value is as follows:

[0043] Based on the inducing factors and the degree of protective deficiencies in the in-depth risk assessment results, a first quantitative value is assigned to the probability of the risk occurring.

[0044] Based on the multi-dimensional impact of the in-depth risk assessment results, a second quantitative value is assigned to the degree of risk impact.

[0045] Multiply the first quantified value by the second quantified value to calculate the initial risk value of a single risk point;

[0046] For combined risk, the combined risk value is calculated using the superposition method;

[0047] The quantitative risk level is determined based on the preset grading range in which the initial risk value or combined risk value falls.

[0048] Preferably, step S5 specifically includes:

[0049] The quantitative risk level values, in-depth risk assessment results, and the actual technical architecture and protection capabilities of the environment under test are input into the private large model;

[0050] The large model outputs adapted technical and management rectification plans based on the actual technical architecture, and divides the rectification plans into three priorities: emergency rectification, time-limited rectification, and continuous optimization according to the quantitative risk level value, thus obtaining the priority of risk disposal suggestions.

[0051] The system automatically aggregates data from the entire process, generates a standardized assessment report that conforms to GB / T45577, and outputs risk management recommendations as rectification work orders.

[0052] Preferably, step S6 specifically includes:

[0053] Record the processing status of rectification work orders, including rectified and unrectified status. Automatically trigger a retesting mechanism for nodes marked as rectified to verify whether the risk has been eliminated.

[0054] Based on the preset timed assessment cycle, incremental multi-source raw data and external threat intelligence are collected regularly to identify changes in data assets and adjustments to security policies in the environment under test.

[0055] When the risk quantification indicator is detected to exceed the preset threshold, an alarm is triggered and the overall quantified risk level value is updated, generating dynamic security risk situation data.

[0056] Secondly, an AI-based automatic data security risk assessment system includes a standardized assessment module, a data asset intelligent sorting module, an AI risk analysis and judgment module, a risk level quantitative assessment module, an assessment result visualization module, and an automated scheduling module.

[0057] The standardized evaluation module is used to load the data security evaluation standard indicator library and preset task parameters, initialize and match the environment to be tested, and obtain the evaluation benchmark model.

[0058] The data asset intelligent sorting module is used to acquire multi-source raw data and use NLP and machine learning algorithms to intelligently identify and classify the multi-source raw data to obtain a standardized data asset list.

[0059] The AI ​​risk analysis and assessment module is used to compare and analyze the standardized data asset list with the indicators in the evaluation benchmark model to obtain a preliminary risk list.

[0060] The risk level quantification assessment module is used to automatically de-identify the preliminary risk list, input the de-identified preliminary risk list into the private big model for implicit risk mining and causal inference, obtain the in-depth risk assessment results, and perform weighted calculation on the in-depth risk assessment results to obtain the quantitative risk level value.

[0061] The assessment result visualization module is used to perform reasoning based on the quantitative risk level value through a private large model to obtain risk disposal recommendation priorities and standardized assessment reports.

[0062] The automated scheduling module is used to perform rectification and retesting of risk handling recommendations and to conduct periodic round-robin monitoring to obtain dynamic security risk situation data.

[0063] The beneficial effects of this invention are as follows: This invention strictly follows the evaluation process, indicator system, and quantitative methods of industry standards such as GB / T45577 to ensure the compliance of the evaluation results and meet regulatory and audit requirements; at the same time, it integrates the large-scale model deployed privately within the company into the core link of risk assessment, and uses the knowledge reasoning and correlation analysis capabilities of the large-scale model to uncover hidden and related risks that are difficult to discover by manual assessment, so as to achieve in-depth risk assessment and significantly improve the accuracy and depth of the evaluation results;

[0064] This invention achieves fully automated execution of the entire process of data collection, asset sorting, risk identification, AI analysis, risk quantification, and result output through an automated scheduling module, without the need for manual intervention, compared to traditional manual assessment. At the same time, it avoids the problems of omissions and errors in manual assessment, reduces the input cost of professional personnel, and can meet the large-scale assessment needs of organizations with massive data assets and complex business scenarios.

[0065] This invention designs a dedicated interface and data anonymization mechanism for large models deployed privately within the company. All evaluation data flows within the organization's intranet, with no data being sent out, completely solving the data privacy leakage problem in the AI ​​application process. At the same time, the lightweight calling protocol reduces model calling latency and improves the system's operating efficiency.

[0066] This invention leverages the customized reasoning capabilities of large-scale models, combined with an organization's technical architecture, protection capabilities, and industry attributes, to generate priority-based, customized, and actionable risk management recommendations, avoiding the generalization and vagueness of traditional assessments. It also supports integration with security work order systems, automatically generating rectification work orders to seamlessly connect assessment results with actual security rectification, effectively guiding the organization's data security protection efforts. Furthermore, this invention incorporates a timed assessment configuration unit and a dynamic risk update unit, enabling real-time awareness of changes in data assets, business scenarios, and security policies. This allows for dynamic and routine assessment of data security risks, timely detection of new risks, and triggering alerts, addressing the shortcomings of traditional static assessments and providing continuous risk monitoring support for the organization's data security. Attached Figure Description

[0067] Figure 1 A flowchart illustrating the steps of an automatic data security risk assessment method based on artificial intelligence, as provided in one embodiment of the present invention;

[0068] Figure 2 This is a basic flowchart of an automatic data security risk assessment system based on artificial intelligence, provided as an embodiment of the present invention. Detailed Implementation

[0069] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0070] Example 1, referring to Figure 1 This paper provides an automatic data security risk assessment method based on artificial intelligence, which includes the following steps:

[0071] Step S1: Load the data security assessment standard indicator library and preset task parameters, initialize and match the environment to be tested, and obtain the assessment benchmark model.

[0072] Step S2: Obtain multi-source raw data, and use NLP and machine learning algorithms to intelligently identify and classify the multi-source raw data to obtain a standardized data asset list;

[0073] Step S3: Compare and analyze the standardized data asset list with the indicators in the evaluation benchmark model to obtain a preliminary risk list;

[0074] Step S4: Automated desensitization processing is performed on the preliminary risk list, and the desensitized preliminary risk list is input into the private large model for implicit risk mining and causal inference to obtain in-depth risk assessment results. The in-depth assessment results are jointly verified by the preset risk rule base and the industry-specific small model, and the prompt words are dynamically adjusted for secondary inference. Finally, the quantitative risk level value is output.

[0075] Step S5: The private big model is used to reason based on the quantitative risk level values ​​to obtain risk disposal recommendations, priority, and standardized assessment reports.

[0076] Step S6 involves prioritizing risk management recommendations, conducting rectification and retesting, and carrying out regular patrol monitoring to obtain dynamic safety risk situation data.

[0077] Step S1 includes the following sub-steps:

[0078] Step S11: Obtain task parameters, which include the evaluation scope, industry-specific standards, and custom indicator weights.

[0079] Step S12: Match the data security assessment standard indicator library with industry-specific standards one by one to construct a multi-dimensional indicator system. The multi-dimensional indicator system includes data asset identification, data classification and grading, security management system, technical protection measures, and personnel management and emergency response.

[0080] Step S13: Obtain the environment to be tested, adapt and map the multi-dimensional indicator system and custom indicator weights according to the business attributes of the environment to be tested, define the input and output rules for each stage of the evaluation process, and generate the evaluation benchmark model.

[0081] In this embodiment, based on user-input task parameters, such as evaluating a specific financial transaction system, and by introducing the Cybersecurity Classified Protection 2.0 or financial industry data security standards, a multi-dimensional indicator system mapping matrix is ​​constructed. This matrix intersects and merges the six dimensions of asset identification, classification, system, technology, personnel, and emergency response in the national standard with industry-specific standards to eliminate duplicates. Simultaneously, it allows users to set custom indicator weights for specific core data. Finally, the adapted indicator system is bound to the actual network topology and business attributes of the organization under test to generate an evaluation benchmark model. This benchmark model strictly specifies the input and output constraints that must be met in each subsequent evaluation stage to prevent skipping steps or omissions in the evaluation.

[0082] Step S2 includes the following sub-steps:

[0083] Step S21: Collect network data based on the principle of least privilege. The network data includes basic information, flow information and attribute information. The collection methods include API interface, proxy collection and log capture.

[0084] Step S22: Use NLP and feature extraction algorithms to parse network data, extract sensitive attribute features, and perform feature matching and recognition on encrypted network data to obtain feature labels;

[0085] Step S23: Input the network data and feature labels one-to-one into the pre-trained machine learning classification model to generate grade category labels, and save the grade category labels, network data and feature labels as a standardized data asset list.

[0086] In this embodiment, the specific technical implementation process of the principle of least privilege (PoLP) is as follows:

[0087] When connecting to databases or business systems, we strictly allocate and use dedicated service accounts with only metadata reading permissions. Through API interfaces, proxy collection, or bypass traffic mirroring technology, we refuse to execute any entity data query commands with SELECT * FROM business tables. Instead, we selectively capture system tables (such as MySQL's information_schema), database table structures, field names, field comments, as well as header information of network traffic packets and a small number of anonymized sample packets. In this way, we obtain basic information about network data (such as table names and field types), flow information (such as source application IP and destination DB port), and attribute information, thus eliminating privacy leaks that may be caused by the data collection process itself from the physical and logical source.

[0088] After acquiring network data, a dual-track feature extraction strategy is adopted for both plaintext and encrypted data:

[0089] For NLP feature extraction of plaintext and structured metadata, this embodiment does not use the outdated BERT model, but instead employs the more advanced BGE pre-trained vector model combined with the DeBERTa-v3 architecture for semantic parsing. The specific processing procedure is as follows:

[0090] The collected database table names and field names are de-splittered, camelCase converted, and stop word removed to transform them into standard natural language sequences. The cleaned standard natural language sequences are then input into the BGE model, which maps them into 768-dimensional high-dimensional semantically dense vectors. Since the BGE model can more accurately understand the semantics of abbreviations in professional databases, cosine similarity calculation is then performed. This invention incorporates a high-dimensional sensitive word vector library, which includes nationally defined personal information and financial data as standard feature vectors. Feature extraction is performed by calculating the cosine similarity S between vectors:

[0091] ;

[0092] When the cosine similarity S is greater than the preset semantic threshold, the semantic threshold in this embodiment is set to 0.85. It is determined that the current field has the corresponding sensitive semantics, the standard natural language sequence corresponding to the high-dimensional semantic dense vector is extracted, and a personal identification label is assigned.

[0093] Because encrypted data loses semantic information, the traditional Shannon entropy algorithm has an extremely low recognition rate against modern encryption algorithms; this embodiment uses a byte-level sequence feature analysis algorithm based on a one-dimensional convolutional neural network:

[0094] The sampled obfuscated network data (ciphertext fragments) is converted into a sequence of ASCII byte arrays from 0 to 255, and then truncated or padded to a fixed length (the fixed length is set to 256 bytes).

[0095] Multiple one-dimensional convolutional kernels of different sizes (3, 5, 7) are slid across the byte sequence to extract hidden patterns and non-linear structural features between adjacent bytes, such as specific padding patterns in Base64 encoding and the character set distribution contours of FPE encryption. The extracted high-dimensional feature maps are then processed by global max pooling and input into a pre-trained Softmax feature classifier. This Softmax feature classifier does not decrypt the data but rather outputs the encryption probability distribution of the data stream by comparing the byte feature distribution. For example, it may determine that there is an 85% probability that the data is encrypted with AES and a 10% probability that it is encrypted with the Chinese national standard SM4.

[0096] The extracted semantic attribute features and encryption morphological features are concatenated into vectors to generate multimodal comprehensive feature labels;

[0097] The context features (such as data volume and access frequency) associated with the network data obtained in step S21 are matched one-to-one with the comprehensive feature labels output in step S22 to construct a standardized feature matrix.

[0098] The standardized feature matrix is ​​input into a pre-trained machine learning-based classification model (LightGBM). Compared with traditional random forests, LightGBM uses a histogram-based decision tree algorithm. The LightGBM model performs non-linear mapping on the input feature matrix according to the built-in GB / T35273 data classification rules, and finally outputs the exact level category label of the network data.

[0099] Finally, the acquired network data metadata, multimodal feature labels, and grade category labels are structurally encapsulated, saved, and output to the graph database to form a standardized data asset list with entity relationships and grading attributes.

[0100] Step S3 includes the following sub-steps:

[0101] Step S31: The standardized data asset list is automatically scanned using the indicator system of the evaluation benchmark model to obtain the life cycle stages. The life cycle stages include several nodes, specifically including collection, storage, use, processing, transmission, provision, disclosure and deletion.

[0102] Step S32: Detect the risk status of each node in the lifecycle process. The risk status includes a risk-free status and a risky status. The risky status includes unauthorized access, unencrypted data, and missing systems.

[0103] Step S33: Extract the standardized data asset list corresponding to the nodes in a risky state to obtain a preliminary risk list.

[0104] The evaluation benchmark model generated in step S1 is used to automatically scan the data asset list, covering eight key nodes in the entire data lifecycle: collection, storage, use, processing, transmission, provision, publication, and deletion. A rule engine (such as Drools) is used to execute the detection logic. For example, when detecting storage nodes, if a core data table is found to be without encrypted storage parameters, it is determined to be in a risky state (unencrypted data); if a usage node is not bound to an access control policy (RBAC), it is determined to be in a risky state (unauthorized access).

[0105] Extract all node data marked as being at risk and aggregate them to generate a preliminary risk list.

[0106] Step S4 includes the following sub-steps:

[0107] Step S41: Perform automated desensitization processing on the preliminary risk list to obtain the desensitized preliminary risk list;

[0108] Step S42 involves inputting the anonymized preliminary risk list into the privatization model, specifically including:

[0109] By using a dedicated private interface adaptation layer and a lightweight calling protocol, the anonymized preliminary risk list is input into the locally deployed private large model in the intranet environment.

[0110] In this embodiment, the preliminary risk list is standardized and denoised to generate JSON format data, and then the de-identification algorithm library is called:

[0111] By using regular expressions combined with Named Entity Recognition (NER) technology, the organization name is generalized to a certain institution, and personal information is masked, for example, replacing IP 192.168.1.100 with IP_Node_A. The risk logic relationship is preserved, and the anonymized preliminary risk list is streamed into the locally deployed private large model through the internal network's dedicated private interface layer and RESTful lightweight protocol.

[0112] Step S4 also includes:

[0113] The privatization model performs impact risk mining and causal deduction on the initial risk list after anonymization to obtain in-depth risk assessment results, specifically including:

[0114] By combining the industry attributes of the environment under test, we can analyze the multi-dimensional impact of economic losses, reputational damage, and compliance penalties after the risk occurs, and summarize the results of the large model inference to generate in-depth analysis results.

[0115] The in-depth analysis results are double-verified to generate in-depth risk analysis results;

[0116] The processing logic for double-verification of in-depth analysis results is as follows:

[0117] Dual verification includes a first verification and a second verification;

[0118] The first verification includes:

[0119] The results of in-depth risk assessment are input into a preset risk assessment rule base to obtain rule verification labels. The rule verification labels include rule compliance and rule non-compliance. The risk assessment rule base is used to verify whether the results of in-depth risk assessment meet industry standard requirements.

[0120] The second verification includes:

[0121] The lightweight industry-specific small model is used to cross-validate the reasoning logic corresponding to the deep risk assessment results, and logical verification labels are obtained. The logical verification labels include logical compliance and logical non-compliance.

[0122] If the labels of the first and second checks are both met, the result is considered passed and the deep risk assessment result is output. Otherwise, the result is considered failed, the prompt words are automatically adjusted and the private big model is triggered to perform secondary reasoning until the first and second checks are both met and the adjustment stops.

[0123] The results of in-depth risk assessment are weighted and calculated to obtain a quantitative risk level value.

[0124] Large-scale model latent risk mining and dual verification include a first verification and a second verification:

[0125] Privatization is based on prompt engineering, which not only analyzes single risks but also uses a long contextual attention mechanism to uncover cross-stage related risks. The output of in-depth analysis results must undergo a double-check interception mechanism.

[0126] First verification (rule baseline verification): The output logic is parsed into structured labels and input into the locally built "GB / T45577 Hard Compliance Rule Library". If the handling method suggested by the large model violates the national standard bottom line, then a rule non-compliance label is added; otherwise, the rule is compliant.

[0127] The second verification (small model logic cross-validation): calls a natural language reasoning model with a small number of parameters that has been fine-tuned for the vertical field of network security (in this embodiment, it is an NLI classification model based on the RoBERTa architecture), and decomposes the causal reasoning chain output by the large model into premise and conclusion pairs. The implication relationship is judged for each pair. If any pair is judged to be contradictory, it is marked as logically inconsistent; otherwise, it is marked as logically consistent, thus completing the second verification.

[0128] The final in-depth risk assessment result is output only if both the first and second checks are met; if either check is not met, the prompt word is automatically modified for a second inference, and the specific processing is as follows:

[0129] Extract the violated hard compliance rules from the first verification output and the logical contradictions from the second verification output as penalty feedback information. Embed the penalty feedback information into a preset error correction prompt word template to generate new prompt words, such as: The previous output has an error in [logical contradiction point]. Please strictly follow [compliance rules and clauses] and re-perform causal deduction.

[0130] The new prompt words are input into the private large model for secondary reasoning until the first and second checks are both met. Then, the final deep risk assessment result is output. To prevent infinite loops, a maximum retry threshold is set (e.g., 3 times). If the maximum retry threshold is exceeded, it will be switched to manual review.

[0131] The logic for weighted calculation of the in-depth risk assessment results to obtain the quantitative risk level value is as follows:

[0132] Based on the inducing factors and the degree of protective deficiencies in the in-depth risk assessment results, a first quantitative value is assigned to the probability of the risk occurring.

[0133] Based on the multi-dimensional impact of the in-depth risk assessment results, a second quantitative value is assigned to the degree of risk impact.

[0134] Multiply the first quantified value by the second quantified value to calculate the initial risk value of a single risk point;

[0135] For combined risk, the combined risk value is calculated using the superposition method;

[0136] The quantitative risk level is determined based on the preset grading range in which the initial risk value or combined risk value falls.

[0137] After the double verification passes, perform quantization calculation:

[0138] Assign a first quantization value P (value 1-5), and a second quantization value I (value 1-5);

[0139] Single initial risk value = P × I (value range 1-25);

[0140] If the large model determines risk point A With risk point B ( If a causal transmission chain exists, the summation formula is used to calculate the portfolio risk value. :

[0141] ;

[0142] in, The risk coupling coefficient output by the large model is set to a value between 0 and 1. This ensures that the combined risk value is non-linearly amplified and does not exceed the national standard upper limit of 25. The final quantitative risk level value is determined based on a preset range, which is:

[0143] 1-4: General risks;

[0144] 5-9: Significant risk;

[0145] 10-16: Significant risks;

[0146] 17-25: Particularly significant risks.

[0147] Step S5 specifically includes:

[0148] The quantitative risk level values, in-depth risk assessment results, and the actual technical architecture and protection capabilities of the environment under test are input into the private large model;

[0149] The large model outputs adapted technical and management rectification plans based on the actual technical architecture, and divides the rectification plans into three priorities: emergency rectification, time-limited rectification, and continuous optimization according to the quantitative risk level value, thus obtaining the priority of risk disposal suggestions.

[0150] The system automatically aggregates data from the entire process, generates a standardized assessment report that conforms to GB / T45577, and outputs risk management recommendations as rectification work orders.

[0151] In this embodiment, the above-mentioned quantitative values, judgment results, and the actual technical architecture of the enterprise are fed back into the large model as a context. The large model combines the actual architecture and refuses to generate vague and unrealistic suggestions. Instead, it outputs highly feasible technical rectification solutions such as "configure the TLS1.3 protocol at the K8s cluster Ingress gateway".

[0152] Based on the quantitative risk level, the recommendations are automatically classified into three levels: urgent (major / extremely major), time-limited (relatively significant), and continuous optimization (general).

[0153] Finally, the report engine is called to generate a PDF report that conforms to the national standard format, and the priority of the handling suggestions is directly pushed to the enterprise's Jira or self-developed rectification work order system via Webhook.

[0154] Step S6 specifically includes:

[0155] Record the processing status of rectification work orders, including rectified and unrectified status. Automatically trigger the retesting mechanism for nodes marked as rectified to verify whether the risks have been eliminated.

[0156] Based on the preset timed assessment cycle, incremental multi-source raw data and external threat intelligence are collected regularly to identify changes in data assets and adjustments to security policies in the environment under test.

[0157] When the risk quantification indicator is detected to exceed the preset threshold, an alarm is triggered and the overall quantified risk level value is updated, generating dynamic security risk situation data.

[0158] In this embodiment, when the callback interface of the work order monitoring system receives a status change to "rectified," the automated scheduling module automatically extracts the asset information of that node and triggers steps S2 and S3 for automated retesting to verify whether the vulnerability has been truly closed. Based on a Crontab scheduled task (preferred to be set every 24 hours), the system incrementally collects raw data and connects with external threat intelligence. If a new business interface or a sudden 0-day vulnerability is detected that causes the recalculated quantitative indicator ΔR of a certain type of asset to exceed a preset threshold, a system-level audible / visual / email alarm is immediately triggered, refreshing the dynamic security risk situation data on the command screen, thus achieving a leap from static assessment to normalized adaptive protection.

[0159] Example 2, refer to Figure 2 This provides an AI-based automatic data security risk assessment system, including a standardized assessment module, a data asset intelligent sorting module, an AI risk analysis and judgment module, a risk level quantitative assessment module, an assessment result visualization module, and an automated scheduling module.

[0160] The standardized evaluation module is used to load the data security evaluation standard indicator library and preset task parameters, perform initial matching of the environment to be tested, and obtain the evaluation benchmark model.

[0161] The data asset intelligent sorting module is used to acquire multi-source raw data and use NLP and machine learning algorithms to intelligently identify and classify multi-source raw data for an automatic assessment of data security risks based on artificial intelligence, so as to obtain a standardized data asset list.

[0162] The AI ​​risk analysis and assessment module is used to compare and analyze the compliance of a standardized data asset list based on artificial intelligence for automatic data security risk assessment with the indicators in the evaluation benchmark model to obtain a preliminary risk list.

[0163] The risk level quantification assessment module is used to automatically de-identify the preliminary risk list, and input the de-identified preliminary risk list into the private big model for hidden risk mining and causal inference to obtain in-depth risk assessment results. It also performs weighted calculation on the in-depth risk assessment results of an AI-based data security risk automatic assessment to obtain a quantitative risk level value.

[0164] The assessment results visualization module is used to perform reasoning based on the quantified risk level values ​​using a private large model to obtain risk disposal recommendations, priority, and standardized assessment reports.

[0165] The automated scheduling module is used to prioritize risk handling recommendations, conduct rectification and retesting, and perform periodic round-robin monitoring to obtain dynamic security risk situation data.

[0166] In this embodiment, the standardized evaluation module serves as the system configuration hub, including a task scheduler and a rule parser. It loads indicators such as GB / T45577-2025 and integrates them to generate an evaluation benchmark model containing mandatory enforcement rules. The intelligent data asset sorting module incorporates a multi-protocol data probe library, an NLP parsing engine, and a classification and grading processor integrating machine learning algorithms. It is responsible for the lossless collection, purification, and asset tagging of multi-source heterogeneous data. The AI ​​risk analysis and judgment module integrates a compliance baseline scanning engine to perform differential comparisons between the asset list and the benchmark model, delineating preliminary risk boundaries. The risk level quantitative assessment module serves as the core AI of the system. The brain layer, containing a data anonymization sandbox, a large model communication proxy, and a dual-verification service cluster, is responsible for performing deep context mining, completing the first and second verification feedback loops, and using matrix operators to perform non-linear superposition calculations of risk weights. The evaluation result visualization module, with an embedded document generation engine and BI data dashboard system, connects to the output stream of the private large model, renders and distributes differentiated rectification work orders and compliance reports. The automated scheduling module, as the nerve center of the entire chain, adopts a microservice architecture, is responsible for listening to work order callbacks, triggering automated retesting operations, and driving regular polling scans to maintain real-time security risk situation data awareness.

[0167] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0168] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the protection scope of the present invention.

Claims

1. An automatic data security risk assessment method based on artificial intelligence, characterized in that, Includes the following steps: Step S1: Load the data security assessment standard indicator library and preset task parameters, initialize and match the environment to be tested, and obtain the assessment benchmark model. Step S2: Obtain multi-source raw data, and use NLP and machine learning algorithms to intelligently identify and classify the multi-source raw data to obtain a standardized data asset list; Step S3: Compare and analyze the standardized data asset list with the indicators in the evaluation benchmark model to obtain a preliminary risk list; Step S4: Automated desensitization processing is performed on the preliminary risk list, and the desensitized preliminary risk list is input into the private large model for implicit risk mining and causal inference to obtain in-depth risk assessment results. The in-depth assessment results are jointly verified by the preset risk rule base and the industry-specific small model, and the prompt words are dynamically adjusted for secondary inference. Finally, the quantitative risk level value is output. Step S5: The private big model is used to reason based on the quantitative risk level values ​​to obtain risk disposal recommendations, priority, and standardized assessment reports. Step S6 involves prioritizing risk management recommendations, conducting rectification and retesting, and carrying out regular patrol monitoring to obtain dynamic safety risk situation data.

2. The method for automatically assessing data security risks based on artificial intelligence as described in claim 1, characterized in that, Step S1 includes the following sub-steps: Step S11: Obtain task parameters, which include the evaluation scope, industry-specific standards, and custom indicator weights. Step S12: Match the data security evaluation standard indicator library with industry-specific standards one by one to construct a multi-dimensional indicator system. The multi-dimensional indicator system includes data asset identification, data classification and grading, security management system, technical protection measures, and personnel management and emergency response. Step S13: Obtain the environment to be tested, adapt and map the multi-dimensional indicator system and custom indicator weights according to the business attributes of the environment to be tested, define the input and output rules for each stage of the evaluation process, and generate the evaluation benchmark model.

3. The method for automatically assessing data security risks based on artificial intelligence as described in claim 2, characterized in that, Step S2 includes the following sub-steps: Step S21: Collect network data based on the principle of least privilege. The network data includes basic information, flow information, and attribute information. The collection methods include API interface, proxy collection, and log capture. Step S22: Use NLP and feature extraction algorithms to parse network data, extract sensitive attribute features, and perform feature matching and recognition on encrypted network data to obtain feature labels; Step S23: Input the network data and feature labels one-to-one into the pre-trained machine learning classification model to generate grade category labels, and save the grade category labels, network data and feature labels as a standardized data asset list.

4. The method for automatically assessing data security risks based on artificial intelligence as described in claim 3, characterized in that, Step S3 includes the following sub-steps: Step S31: The standardized data asset list is automatically scanned using the indicator system of the evaluation benchmark model to obtain the life cycle stages. The life cycle stages include several nodes, specifically including collection, storage, use, processing, transmission, provision, disclosure and deletion. Step S32: Detect the risk status of each node in the lifecycle process. The risk status includes a risk-free status and a risky status. The risky status includes unauthorized access, unencrypted data, and missing regulations. Step S33: Extract the standardized data asset list corresponding to the nodes in a risky state to obtain a preliminary risk list.

5. The method for automatically assessing data security risks based on artificial intelligence as described in claim 4, characterized in that, Step S4 includes the following sub-steps: Step S41: Perform automated desensitization processing on the preliminary risk list to obtain the desensitized preliminary risk list; Step S42 involves inputting the anonymized preliminary risk list into the privatization model, specifically including: By using a dedicated private interface adaptation layer and a lightweight calling protocol, the anonymized preliminary risk list is input into the locally deployed private large model in the intranet environment.

6. The method for automatically assessing data security risks based on artificial intelligence as described in claim 5, characterized in that, Step S4 further includes: The privatization model performs impact risk mining and causal deduction on the initial risk list after anonymization to obtain in-depth risk assessment results, specifically including: By combining the industry attributes of the environment under test, we can analyze the multi-dimensional impact of economic losses, reputational damage, and compliance penalties after the risk occurs, and summarize the results of the large model inference to generate in-depth analysis results. The in-depth analysis results are double-verified to generate in-depth risk analysis results; The processing logic for double-verification of in-depth analysis results is as follows: The dual verification includes a first verification and a second verification; The first verification includes: The results of in-depth risk assessment are input into a preset risk assessment rule base to obtain rule verification labels. The rule verification labels include rule compliance and rule non-compliance. The risk assessment rule base is used to verify whether the results of in-depth risk assessment meet industry standard requirements. The second verification includes: A lightweight industry-specific small model is invoked to cross-validate the reasoning logic corresponding to the deep risk assessment results, and logical verification labels are obtained, including logical compliance and non-compliance. If both the first and second checks match the tags, the result is considered passed, and the in-depth risk assessment result is output. Otherwise, the result is considered failed, the prompt words are automatically adjusted, and the private large model is triggered to perform secondary inference until both the first and second checks match, at which point the adjustment stops. The results of the in-depth risk assessment are weighted and calculated to obtain a quantitative risk level value.

7. The method for automatically assessing data security risks based on artificial intelligence as described in claim 6, characterized in that, The logic for weighted calculation of the in-depth risk assessment results to obtain the quantitative risk level value is as follows: Based on the inducing factors and the degree of protective deficiencies in the in-depth risk assessment results, a first quantitative value is assigned to the probability of the risk occurring. Based on the multi-dimensional impact of the in-depth risk assessment results, a second quantitative value is assigned to the degree of risk impact. Multiply the first quantified value by the second quantified value to calculate the initial risk value of a single risk point; For combined risk, the combined risk value is calculated using the superposition method; The quantitative risk level is determined based on the preset grading range in which the initial risk value or combined risk value falls.

8. The method for automatically assessing data security risks based on artificial intelligence as described in claim 7, characterized in that, Step S5 specifically includes: The quantitative risk level values, in-depth risk assessment results, and the actual technical architecture and protection capabilities of the environment under test are input into the private large model; The large model outputs adapted technical and management rectification plans based on the actual technical architecture, and divides the rectification plans into three priorities: emergency rectification, time-limited rectification, and continuous optimization according to the quantitative risk level value, thus obtaining the priority of risk disposal suggestions. The system automatically aggregates data from the entire process, generates a standardized assessment report that conforms to GB / T45577, and outputs risk management recommendations as rectification work orders.

9. The method for automatically assessing data security risks based on artificial intelligence as described in claim 8, characterized in that, Step S6 specifically includes: Record the processing status of rectification work orders, including rectified and unrectified status. Automatically trigger a retesting mechanism for nodes marked as rectified to verify whether the risk has been eliminated. Based on the preset timed assessment cycle, incremental multi-source raw data and external threat intelligence are collected regularly to identify changes in data assets and adjustments to security policies in the environment under test. When the risk quantification indicator is detected to exceed the preset threshold, an alarm is triggered and the overall quantified risk level value is updated, generating dynamic security risk situation data.

10. An automatic data security risk assessment system based on artificial intelligence, applied in an automatic data security risk assessment method based on artificial intelligence as described in any one of claims 1-9, characterized in that, It includes a standardized assessment module, a data asset intelligent sorting module, an AI risk analysis and judgment module, a risk level quantitative assessment module, an assessment result visualization module, and an automated scheduling module; The standardized evaluation module is used to load the data security evaluation standard indicator library and preset task parameters, initialize and match the environment to be tested, and obtain the evaluation benchmark model. The data asset intelligent sorting module is used to acquire multi-source raw data and use NLP and machine learning algorithms to intelligently identify and classify the multi-source raw data to obtain a standardized data asset list. The AI ​​risk analysis and assessment module is used to compare and analyze the standardized data asset list with the indicators in the evaluation benchmark model to obtain a preliminary risk list. The risk level quantification assessment module is used to automatically de-identify the preliminary risk list, input the de-identified preliminary risk list into the private big model for implicit risk mining and causal inference, obtain the in-depth risk assessment results, and perform weighted calculation on the in-depth risk assessment results to obtain the quantitative risk level value. The assessment result visualization module is used to perform reasoning based on the quantitative risk level value through a private large model to obtain risk disposal recommendation priorities and standardized assessment reports. The automated scheduling module is used to perform rectification and retesting of risk handling recommendations and to conduct periodic round-robin monitoring to obtain dynamic security risk situation data.