Network security insurance questionnaire and risk assessment model adaptive linkage method
Through the adaptive linkage method of the network security questionnaire and risk assessment model based on the GIPDRR model, multi-source heterogeneous data and real-time threat intelligence are integrated, and the shortcomings of the accuracy and transparency of the evaluation results in the existing technology are solved, and a more comprehensive and dynamic network security risk assessment is achieved.
Patent Information
- Application Number
- CN202510155839.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-13
AI Technical Summary
The existing network security risk assessment methods are difficult to fully cover complex multi-source heterogeneous data, and lack comprehensive assessments of various dimensions of network security, resulting in insufficient accuracy and transparency of the evaluation results.
Adaptive linkage method of network security questionnaire and risk assessment model based on GIPDRR model is adopted, and through the combination of questionnaire management, data preprocessing, feature extraction and deep learning models, structured and unstructured data are integrated, threat intelligence is updated in real time, and network security status assessment is conducted.
It improves the comprehensiveness and accuracy of network security risk assessment, enhances the transparency of the system and user trust, and achieves comprehensive coverage and dynamic adjustment of complex risk scenarios.
Smart Images

Figure CN119991313A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cybersecurity technology, and more particularly to a method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model. In particular, the present invention relates to a method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model based on a GIPDRR model. Background Art
[0002] In the current market, solutions for cybersecurity risk assessment and insurance service recommendations are gradually gaining attention.
[0003] In terms of insurance service recommendations and risk assessment, many insurance companies and fintech companies have launched insurance service recommendation systems based on customer data analysis. These systems collect basic customer information and, combined with big data analysis and machine learning techniques, tailor insurance plans for each client. These products typically include questionnaires, data mining, risk assessment, and personalized recommendations to improve the quality of insurance services and transaction rates.
[0004] Patent document CN116756424A discloses an insurance service recommendation method, device, server and storage medium, wherein the method obtains the user's basic information, extracts the user's corresponding feature data based on the basic information and classifies the user to obtain the user's category label, and determines the target service strategy adapted to the user based on the category label; sends an insurance service questionnaire to the terminal device according to the target service strategy and obtains questionnaire feedback; determines insurance service keywords and insurance service themes based on the questionnaire feedback, and uses an image generation model to generate multiple theme cover images based on the insurance service keywords; generates target insurance service content based on the theme cover image and the insurance service theme, and sends the target insurance service content to the terminal device.
[0005] However, while existing insurance service recommendation systems can categorize and customize services based on user data, they often rely on pre-defined classification rules and static strategies. Simple classification models cannot fully capture and reflect users' diverse needs and complex risk profiles, resulting in limited personalization of recommendations that may not fully match users' actual needs.
[0006] Regarding large language models and data security risk assessment, products based on large language models (LLMs) are gradually emerging. These systems collect and analyze multiple data features, then use large language models to combine these features and calculate risk, thereby formulating data security management strategies. This approach can handle complex data features and provide highly accurate risk assessments, but it primarily focuses on the data security field.
[0007] Patent document CN118194358B discloses a data security risk assessment and management system based on a large language model, which relates to the field of data security technology. The method includes obtaining data feature types and forming a data feature set; randomly combining the data feature types in the data feature set to determine a combined feature set; defining a combined valid set and determining the number of combined features; determining the number of combined features with the largest value and determining a baseline risk coefficient based on the number of combined features, and determining different feature types based on the combined valid set and the data feature set; determining a risk variation coefficient based on all different feature types, and performing mean calculation to determine the mean variation coefficient, and determining a comprehensive risk coefficient based on the mean variation coefficient and the baseline risk coefficient; determining a risk management plan based on the comprehensive risk coefficient, and processing the current data according to the risk management plan.
[0008] However, the application of large language models (LLMs) in data security primarily focuses on processing and evaluating structured data features. However, the types of data in the cybersecurity field are more complex, including unstructured data (such as logs and communication content) and multi-source heterogeneous data. Existing large language models have limitations when processing this complex data and cannot fully cover all security risk scenarios. Due to the complexity and black-box nature of large language models, existing systems often lack transparency and explainability in risk assessment results. This makes it difficult for users and security experts to understand the basis for the model's decisions, increasing management and decision-making uncertainty.
[0009] Regarding cybersecurity risk assessment, some specialized cybersecurity products already offer risk assessment capabilities based on publicly available internet events and user system data. These products typically combine threat intelligence, vulnerability scanning, and risk assessment models to conduct comprehensive security analyses of user information systems, helping users identify potential security risks and recommend appropriate response strategies.
[0010] Patent document CN113542279A discloses a network security risk assessment method, system and device, including: determining possible network security risk scenarios based on network security incidents disclosed on the Internet and combined with the risks faced by the user's information system, where there are at least two network security risk scenarios; for each network security risk scenario, quantifying the threats, vulnerabilities and losses therein to obtain a risk value for the network security risk scenario; and calculating the network security risk value of the user's information system based on the risk values of each network security risk scenario.
[0011] However, existing cybersecurity risk assessment methods typically rely on publicly available data from cybersecurity incidents and user information systems. While this approach can provide a certain degree of visibility into the security status of a system, the single data source can overlook new or hidden risks, resulting in inaccurate and incomplete assessment results. Existing methods often assess risk using a single or limited set of dimensions, focusing on individual analyses of threats, vulnerabilities, or losses, while lacking a comprehensive assessment of all dimensions of cybersecurity, such as security management, risk identification, and security protection. This single-dimensional analysis may fail to fully reflect an enterprise's overall security posture, limiting its decision support capabilities.
[0012] In summary, existing similar products and processing methods have significant shortcomings in personalized recommendations, complex data processing, comprehensive risk assessment, and dynamic adjustment capabilities. These shortcomings directly affect the efficiency and security of users' decision-making in insurance services and network security management.
[0013] Therefore, the market needs a method that can more accurately capture user needs and security risks and achieve highly personalized and intelligent adaptive linkage between cybersecurity insurance questionnaires and risk assessment models. Summary of the Invention
[0014] In view of the defects in the prior art, the purpose of the present invention is to provide a method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model.
[0015] A method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model provided by the present invention includes:
[0016] Questionnaire management steps: Create a cybersecurity questionnaire based on the GIPDRR model, send it to the enterprise for completion, and then collect it back to obtain the corresponding questionnaire data;
[0017] Data collection and preprocessing step: collecting threat intelligence data from the threat intelligence database and preprocessing the questionnaire data in combination with the threat intelligence data;
[0018] Feature extraction step: extract key features from the preprocessed data, integrate all extracted and selected features into a unified feature vector, and use the feature vector as input to the risk rating step;
[0019] Risk rating step: Build a Transformer-based deep learning model, integrate the feature vectors and real-time updated threat intelligence to evaluate the network security status, obtain the corresponding risk score, and generate a risk assessment report.
[0020] Preferably, the GIPDRR model is based on the NIST Cybersecurity Framework, which includes six key dimensions of cybersecurity management: Governance, Risk Identification, Protect, Detect, Respond, and Recover.
[0021] We create a cybersecurity questionnaire with multiple specific questions for each dimension, and regularly adjust and optimize the questionnaire content based on the latest security intelligence, industry trends, and technological developments.
[0022] The cybersecurity questionnaire includes structured multiple-choice questions, scoring questions, and unstructured open-ended questions.
[0023] Preferably, the data collection and preprocessing steps include:
[0024] Step S2.1: Conduct a preliminary check on the received questionnaire data to confirm data integrity;
[0025] Step S2.2: If the questionnaire data is complete, proceed to step S2.3; if the questionnaire data is incomplete, process it by filling in missing values, marking abnormal data, or prompting the user to re-fill the questionnaire, and then proceed to step S2.3;
[0026] Step S2.3: Preprocess the complete questionnaire data, including cleaning, classification, coding and integration.
[0027] Preferably, the method of collecting threat intelligence data from the threat intelligence library includes accessing third-party APIs and data crawling technology;
[0028] The threat intelligence data includes known vulnerabilities, malware families, attack methods, and corresponding protection measures.
[0029] Preferably, each dimension in the feature vector represents an independent security feature or interaction feature;
[0030] For structured data, statistical and data mining techniques are used to extract key features that can reflect the enterprise's network security status from the coded and standardized data;
[0031] For unstructured data, natural language processing (NLP) technology and deep learning models are used to extract key information from the text and convert the key information into feature vectors that can be combined with structured data.
[0032] Preferably, the Transformer-based deep learning model includes a multi-head attention layer, a multi-layer encoder, a linear transformation layer, and an output layer, and a random forest model is introduced as an auxiliary to identify possible evaluation deviations or anomalies;
[0033] The multi-head attention layer can implement the self-attention mechanism by performing multiple parallel linear transformations on the input embedding representation, then calculating the attention score of each feature and using the score for weighted averaging. These scores are then normalized into a probability distribution through the softmax number and applied to the weighted sum of the input vector.
[0034] Preferably, the multi-layer encoder structure, each layer of the encoder includes a self-attention layer and a feedforward neural network layer;
[0035] The self-attention layer is used for feature interaction, and the feedforward neural network layer is used for nonlinear transformation and information extraction;
[0036] The output layer is used to map the high-dimensional risk representation vector processed by the multi-layer encoder to a specific risk score;
[0037] The linear transformation layer is used to output the specific risk score of the enterprise in each GIPDRR dimension.
[0038] Preferably, the feature vector is input into the deep learning model in the form of a multidimensional array or tensor;
[0039] Each dimension of the vector represents an independent security feature, while the threat intelligence data contains external threat information relevant to the current enterprise situation.
[0040] Preferably, an explanation analysis step is also included, combining the use of two explanatory analysis tools, LIME and SHAP, to analyze the output of the model from two levels, local and global, respectively.
[0041] According to the present invention, a cybersecurity insurance questionnaire and risk assessment model adaptive linkage system is provided, comprising:
[0042] Questionnaire management module: This module builds a cybersecurity questionnaire based on the GIPDRR model, sends it to enterprises for completion, and then collects the completed questionnaire to obtain the corresponding questionnaire data.
[0043] Data collection and preprocessing module: collects threat intelligence data from the threat intelligence database and preprocesses the questionnaire data based on the threat intelligence data;
[0044] Feature extraction module: extracts key features from the preprocessed data, integrates all extracted and selected features into a unified feature vector, and uses the feature vector as the input of the risk rating module;
[0045] Risk rating module: Build a Transformer-based deep learning model, integrate the feature vectors and real-time updated threat intelligence to evaluate the network security status, obtain the corresponding risk score, and generate a risk assessment report.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] 1. This paper specifically optimizes the Transformer model and combines it with an unstructured data processing model to enhance the system's performance when processing multi-source heterogeneous data. This multi-model combination significantly improves the system's comprehensiveness and accuracy in cybersecurity risk assessment, ensuring coverage of a wide range of complex risk scenarios.
[0048] 2. This paper combines an attention-based interpretable model with LIME (Local Interpretable Models) and SHAP (Shapley Values) to analyze and interpret assessment results. Through these interpretable tools, this paper clearly demonstrates the source and calculation basis of each risk score, enabling users and security experts to clearly understand the model's decision-making process, thereby increasing system transparency and user trust.
[0049] 3. This invention utilizes a real-time threat intelligence library, combined with the adaptive capabilities of the Transformer model, to enable real-time adjustments to network security risk assessment results. The system dynamically updates risk assessment results based on new threat information and environmental changes, ensuring that enterprises consistently receive accurate risk assessments in an ever-changing network security environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0051] Figure 1 It is a schematic flow chart of the working method of the present invention;
[0052] Figure 2 Schematic diagram of the Transformer model structure of the present invention. DETAILED DESCRIPTION
[0053] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0054] The present invention introduces a multi-dimensional questionnaire design based on the GIPDRR model and a Transformer model for comprehensive evaluation. Existing network security risk assessment methods usually rely on static rules and are difficult to fully cover the actual security needs of enterprises. The present invention designs a comprehensive questionnaire based on the six dimensions of the GIPDRR model (security management, risk identification, security protection, threat detection, incident response, and disaster recovery). The questionnaire covers specific sub-questions in each dimension and comprehensively captures the various security risks that an enterprise may face. By combining the intelligent analysis of questionnaire data with the Transformer model, the present invention can accurately evaluate the network security status of an enterprise and provide risk scores for each dimension. This method goes beyond traditional static rule evaluation and provides more dynamic and accurate risk assessment results.
[0055] Example 1
[0056] According to the present invention, a method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model is provided. Figure 1 Shown, including:
[0057] Questionnaire Management Steps: Build and manage an enterprise cybersecurity questionnaire based on the GIPDRR model. The GIPDRR model, based on the NIST Cybersecurity Framework, encompasses six key dimensions of cybersecurity management: Governance, Identify, Protect, Detect, Respond, and Recover. Each dimension corresponds to a key stage in the enterprise cybersecurity lifecycle, helping enterprises systematically identify and assess potential security risks. Among them, security management focuses on the company's security strategy and management mechanism, including policy formulation, implementation, personnel training, and division of security responsibilities; risk identification involves how the company identifies and assesses security risks that may affect its information system, including asset management, understanding of the business environment, and vulnerability identification; security protection includes the protection measures taken by the company at the technical level, such as access control, data protection, maintenance processes, etc., to ensure the confidentiality, integrity and availability of corporate information; threat detection focuses on the company's ability to detect potential threats, including continuous monitoring, security detection activities and incident reporting; incident response focuses on the company's response capabilities when facing security incidents, ensuring that security incidents can be responded to and resolved quickly and effectively; disaster recovery involves the company's recovery capabilities after encountering major security incidents, including the execution of data backup and recovery plans to ensure rapid recovery of business and minimize losses.
[0058] Detailed questionnaires covering various areas are developed using the six dimensions of the GIPDRR model. These questionnaires are carefully designed, with multiple specific questions for each dimension to ensure they comprehensively capture potential security risks within an enterprise. For example, within the security management dimension, the questionnaire covers questions such as whether the enterprise regularly updates its security policy, conducts employee security awareness training, and has a clear division of security responsibilities. Within the threat detection dimension, the questionnaire inquires about whether the enterprise regularly conducts vulnerability scans and utilizes the latest threat intelligence services. These questions provide enterprises with a comprehensive self-assessment tool and lay the foundation for subsequent risk analysis. To ensure the timeliness and accuracy of the questionnaire, the content is regularly adjusted and optimized based on the latest security intelligence, industry trends, and technological developments. For example, when new cyberattack methods or new security protection technologies emerge, the questionnaire will be updated to ensure it accurately reflects the latest security risks. This dynamic adjustment ensures that the questionnaire remains current with the real-world threats facing the enterprise, providing the most targeted risk assessment information.
[0059] The questionnaires are distributed to businesses through various means, including online submission, email distribution, or direct integration into their security management systems. The system automatically tracks the progress of questionnaire completion, reminding businesses to complete the questionnaires on time to ensure data integrity and timeliness. The collected questionnaire data is initially collated and stored in the system, ready for further analysis and processing.
[0060] Data collection and preprocessing steps: Threat intelligence data is collected from the threat intelligence library, and the network security questionnaire is preprocessed in combination with the threat intelligence data. The preprocessing step provides high-quality input data for subsequent feature extraction and risk assessment. The network security questionnaire includes structured multiple-choice questions, scoring questions, and unstructured open questions. The preprocessing includes cleaning, classification, coding, and integration. Specifically, the received raw data is first preliminarily checked to confirm the integrity of the data. For example, the system will detect whether the data is missing, whether there are format errors, etc. If data is found to be missing, it is processed according to predefined rules. The predefined rules include filling in missing values, marking abnormal data, or prompting the user to re-fill. Then the preprocessing step is performed. The data cleaning is a key step in the data collection and preprocessing module, which directly affects the accuracy of subsequent analysis. Cleaning can remove noise data, process abnormal values and null values, and ensure the accuracy and consistency of the data.
[0061] The method of collecting threat intelligence data from the threat intelligence library includes accessing third-party APIs and data crawling technology. The threat intelligence data includes known vulnerabilities, malware families, attack methods and corresponding protection measures. In the subsequent risk rating step, the threat intelligence is integrated with the company's questionnaire data, and potential security risks are discovered through matching and linkage analysis. For example, if the threat intelligence library shows that a certain attack method is frequent in a certain industry, and the company indicates in the questionnaire that its protection measures in this area are insufficient, the system will include this as an important risk point in the subsequent analysis and evaluation. This integration not only enriches the dimensions of the questionnaire data, but also provides important background information for subsequent risk assessment.
[0062] Feature extraction: Extract key features from the preprocessed data, integrate all extracted and selected features into a unified feature vector, and use this feature vector as input for the risk rating step. Each dimension in this feature vector represents an independent security feature or interactive feature. In other words, the feature extraction step uses a series of algorithms and techniques to transform structured and unstructured data into highly expressive feature vectors. These feature vectors are directly input into the subsequent risk rating machine learning model to perform cybersecurity rating.
[0063] For structured data, we first use statistical and data mining techniques to extract key features that reflect the enterprise's network security status from the coded and standardized data. This data includes options for multiple-choice questions, numerical values for rated questions, and other formatted input data. For example, when processing multiple-choice questions, we calculate the distribution of each option and generate representative features based on the distribution pattern, such as the tendency, frequency, and concentration of choices. For rated questions, we calculate statistical features such as the mean and variance of the scores. These representative and statistical features help quantify the enterprise's security performance in a specific dimension.
[0064] For unstructured data, such as open questions and text data, natural language processing (NLP) technology and deep learning models are used to extract key information from the text and convert the key information into feature vectors that can be combined with structured data. Specifically, the text data is preprocessed, including word segmentation, stop word removal, stemming and other operations. The word segmentation process will perform precise segmentation based on the context to ensure that terms and phrases can be correctly identified. The text data is then converted into a dense vector representation using the word embedding technology BERT. Through word embedding, the semantic information in the text is retained and the potential relationship between words can be captured. These vectorized text data can then be combined with structured data to generate a unified feature vector.
[0065] Risk rating step: Build a deep learning model based on Transformer, integrate the feature vectors and real-time updated threat intelligence to evaluate the network security status, obtain the corresponding risk score, and generate a risk assessment report. Specifically, Figure 2 As shown, first, high-dimensional feature vectors from the feature extraction step are received. These vectors contain all the key information provided by the enterprise in the cybersecurity questionnaire. In addition, data from the threat intelligence library is received and integrated. The data in the threat intelligence library usually contains information such as the latest cybersecurity threats, vulnerabilities, attack methods and response strategies. After cleaning and processing, these intelligence data will be combined with the feature vectors to provide a more comprehensive assessment context. The feature vectors are usually input into the deep learning model in the form of multidimensional arrays or tensors. Each dimension of the vector represents an independent security feature (such as firewall configuration, incident response time, etc.), while the threat intelligence data contains external threat information related to the current enterprise situation, such as the prevalence of a certain attack method in the industry.
[0066] The Transformer-based deep learning model includes a self-attention mechanism and a multi-layer encoder, and introduces a random forest model as an auxiliary. That is to say, the present invention uses the Transformer model to perform the main risk assessment analysis, and combines it with the random forest for auxiliary verification and optimization. The Transformer model is able to capture the complex relationship between features due to its powerful sequence processing capabilities and self-attention mechanism. In the Transformer model, the input vector passes through an input embedding layer to convert the received high-dimensional feature vector into an embedded representation. The embedded representation maps each feature into a high-dimensional space, preserving the relative position and semantic information between the features.
[0067] The self-attention mechanism allows the model to dynamically adjust the weights between different features when processing input features, thereby identifying those feature combinations that are most critical to the final risk assessment. Through this mechanism, the model can efficiently capture the complex interactions between features of different dimensions. In practice, this function is achieved using a multi-head attention layer, which performs multiple parallel linear transformations on the input embedding representation, then calculates an "attention score" for each feature and uses these scores for weighted averaging. These scores are then normalized into a probability distribution using a softmax number and applied to the weighted sum of the input vector.
[0068] The multi-layer encoder structure further enhances the model's representational capabilities. Each encoder layer consists of a self-attention layer for feature interaction and a feedforward neural network layer for nonlinear transformation and information extraction. This stacking of encoder layers allows the model to gradually abstract the representation of input features, ultimately outputting a comprehensive risk representation vector. Finally, the Transformer model's output layer maps the high-dimensional risk representation vector, processed by the multi-layer encoders, to a specific risk score. Through a linear transformation layer, it outputs the company's specific risk score for each GIPDR dimension.
[0069] While the Transformer model excels at capturing complex feature relationships, the system incorporates a random forest model as an auxiliary tool to enhance the robustness and stability of the scoring. The random forest model is a decision tree-based ensemble learning method that reduces model variance and enhances understanding of the importance of different features by constructing multiple decision trees and averaging their results. During the evaluation process, it ensures that the Transformer model's output does not deviate significantly from the actual network security situation and promptly identifies anomalies or inconsistencies in the assessment. Specifically, the random forest model receives as input the high-dimensional risk representation vector output by the Transformer model. By training these decision trees, it understands the contribution of each feature to the final score and uses this contribution to adjust the Transformer model's output. Finally, the high-dimensional representation output by the Transformer model is fed into the random forest for validation, identifying potential biases or anomalies in the assessment. Furthermore, the random forest model receives as input the high-dimensional risk representation vector output by the Transformer model and analyzes the performance of these features across different decision trees to identify potential biases or anomalies. By assessing feature importance, the random forest model can identify features that may contribute to scoring instability, thereby revealing potential miscalculations in the Transformer model. When a random forest makes a prediction, a confidence score is calculated based on the model's stability. If the predictions of multiple random forest trees are consistent, while the output of the Transformer model deviates significantly, this confidence score is used to weight the Transformer output, making the final evaluation more stable.
[0070] After model processing is complete, the system generates risk scores for each dimension of the GIPDRR model. These scores quantify an enterprise's cybersecurity risk into actionable values and are interpreted through a series of classification criteria, allowing enterprises to understand and take appropriate security measures.
[0071] The adaptive linkage method for cybersecurity insurance questionnaires and risk assessment models provided by this invention also includes an explanatory analysis step. Using a variety of explanatory analysis techniques, this step provides transparency into the model's decision-making process and reveals the impact of individual features on the final risk score. This approach aims to provide in-depth analysis and interpretation of the scoring results generated in the risk rating step, enabling enterprise managers and security experts to clearly understand the logic behind the scoring and make more informed security decisions.
[0072] In complex machine learning models (such as the Transformer and random forest models mentioned above), the model's decision-making process is often treated as a "black box," meaning users can only see the inputs and outputs without understanding the intermediate calculations. However, risk assessment is a highly sensitive area, and companies need to understand why certain security measures have low scores or certain risks have been assessed as high. Therefore, the introduction of an explainable analysis step is crucial, as it opens up this "black box" and allows users to clearly understand the source and basis of the scoring results.
[0073] The explanation analysis step primarily utilizes two mainstream explainable analysis tools, LIME and SHAP, to analyze model outputs at both local and global levels, respectively. When explaining the risk score for a specific enterprise, LIME first generates a set of "neighborhood" samples similar to the enterprise's feature vector. These perturbed samples are generated by making slight random modifications to the original features, maintaining semantic similarity to the original input. For example, if the original features indicate an 8 for an enterprise's "firewall configuration," the perturbed samples might modify it to a 7 or 9. After generating the perturbed samples, LIME uses the risk rating model to make predictions for these samples and then fits a simple linear model to these predictions. LIME assumes that within a local neighborhood, the behavior of complex models can be approximated by a linear model. Therefore, this linear model captures the relationship between features and the predicted outcome in the vicinity of the current sample. Using this fitted linear model, LIME calculates the contribution of each feature to the final score. For example, if the score for the "threat detection" dimension is low, LIME might indicate that this is due to the significant negative impact of "vulnerability scanning frequency" on the score. This contribution is visualized to help users understand why certain features contribute to the current score.
[0074] SHAP is a global explanation tool based on game theory that can explain the output of the model from a global perspective. The SHAP value represents the marginal contribution of each feature to the prediction result, and by adding up these contribution values, it explains the output of the entire model. SHAP calculates the marginal contribution of each feature to the final prediction result. For each feature, SHAP considers its role in different feature combinations and calculates its contribution value by weighted average. This process is similar to the "contribution distribution" in game theory. Each feature is regarded as a player, and its contribution to the final prediction result is its "score". SHAP can not only calculate the contribution of a single feature to a single prediction result, but also obtain a global feature importance ranking by cumulatively calculating the average contribution of each feature on the entire dataset. This helps to understand which features have the greatest overall impact on the model's prediction results.
[0075] In the explanation analysis step, LIME and SHAP tools are used in combination to provide a complete explanation. LIME focuses on explaining a single sample (such as the score of a company on a certain dimension), while SHAP provides a global explanation (such as which features have the greatest impact on the score across all dimensions), helping users understand the behavior of the model from different levels. In addition, the results of LIME and SHAP can also be used to improve the performance of the risk rating model. When users find that certain explanation results are not as expected, the system can incorporate this feedback into the model adjustment process. For example, if the SHAP analysis shows that the contribution value of a certain feature is abnormally high or abnormally low, the system can improve the evaluation accuracy by retuning the model or adjusting the feature weights.
[0076] Furthermore, the adaptive linkage method of the cybersecurity insurance questionnaire and risk assessment model of the present invention is described in detail as follows in combination with actual application scenarios:
[0077] Application Scenario: A financial enterprise (hereinafter referred to as "Enterprise A") wishes to conduct a comprehensive assessment of its network security to identify potential security risks and provide a scientific basis for its insurance coverage. Enterprise A's network infrastructure includes multi-layered firewalls, intrusion detection systems, data encryption systems, and regular security review processes. With the increasing frequency of cyberattacks in recent years, Enterprise A hopes to utilize the network security risk assessment system presented in this invention to obtain an accurate assessment of its current network security status and recommendations for improvements.
[0078] Detailed description: Step 1: Questionnaire design and release in the questionnaire management module. A customized cybersecurity questionnaire was designed for Company A based on the six dimensions of the GIPDRR model. The questionnaire covered key areas of Company A's cybersecurity, such as:
[0079] In the “safety management” dimension, the questionnaire asked about Company A’s security policy formulation, the frequency of employee safety training, and the division of security responsibilities among management.
[0080] In the "Threat Detection" dimension, the questionnaire asked about the vulnerability scanning tools used by Enterprise A, the scanning frequency, and the update of threat intelligence.
[0081] The IT department of Company A completed the questionnaire through the online platform provided by the system. The system automatically tracks and ensures the completeness and accuracy of the questionnaire.
[0082] Step 2: Data integration and cleaning in the data collection and preprocessing module. After the questionnaire was submitted, the data collection and preprocessing module began processing the raw data provided by Enterprise A. First, the system performed an integrity check on the questionnaire data and found that the descriptions of some security measures were vague. Based on predefined rules, the system prompted Enterprise A to re-fill these unclear sections. After obtaining the complete data, the system cleaned and standardized the data. Simultaneously, the system accessed multiple threat intelligence repositories via APIs, obtaining the latest threat information related to the financial industry. For example, recent cases of SQL injection attacks that have become common in a certain financial industry, along with their protection recommendations, were included. This threat intelligence data was collated and integrated with Enterprise A's questionnaire data, laying the foundation for subsequent analysis.
[0083] Step 3: Extract and optimize key features of the feature extraction module. The cleaned and integrated data is input into the feature extraction module. This module uses NLP technology to process Company A's open-ended answers and convert the text data into vector representations. For example, Company A's description of its data encryption strategy extracts key features such as "encryption algorithm strength" and "key management process" and generates corresponding vectors. At the same time, the system also extracts important statistical features from the structured data, such as Company A's average security incident response time over the past year and employee security training participation rate. These features are optimized to generate a unified feature vector, which provides input data for the subsequent risk scoring model.
[0084] Step 4: Scoring analysis of the risk rating module. The feature vector is input into the risk rating module, which uses a Transformer-based deep learning model to comprehensively evaluate the network security status of Enterprise A. The system combines real-time updated threat intelligence to score Enterprise A's performance in each dimension of the GIPDRR model. For example, Enterprise A scored high in the "Threat Detection" dimension, indicating that its vulnerability scanning frequency and threat intelligence are relatively good. However, in the "Incident Response" dimension, Enterprise A scored low, indicating that it has insufficient response speed and decision-making efficiency when dealing with sudden security incidents. At the same time, the random forest model is used to verify the scoring results of the Transformer model to ensure the stability of the evaluation. Through the collaboration of the two models, the system finally generated risk scores for Enterprise A in various dimensions and proposed possible risk points.
[0085] Step 5: Interpretation of the scoring results in the explanation and analysis module. In order to enable the management of Company A to better understand the risk score, the system enters the explanation and analysis module. Using LIME technology, the system explains why Company A scored low in the "Incident Response" dimension. For example, the system points out that Company A's "Incident Response Speed" feature contributes significantly to the low score, suggesting that the company should speed up the incident response process. At the same time, the system uses SHAP technology to provide a global perspective, showing which features have the greatest impact on the overall score. The results show that "Data Encryption Strength" and "Employee Training Participation Rate" are the two features that contribute most to the overall score, which helps Company A identify the security areas that are currently relatively stable.
[0086] Step 6: Generate and output a risk assessment report. A detailed risk assessment report was generated, covering Enterprise A's scores on the six dimensions of GIPDRR, detailed explanations of each characteristic, and corresponding improvement suggestions. The report pointed out Enterprise A's weaknesses in the "Incident Response" dimension and recommended that it improve its incident response process and enhance the efficiency of emergency decision-making. At the same time, the report affirmed the high score in the "Threat Detection" dimension and recommended that Enterprise A continue to maintain its existing threat monitoring mechanism. Based on this report, Enterprise A's management developed a further cybersecurity improvement plan and used the report's conclusions in future insurance risk assessments, obtaining risk premium discounts from the insurance company.
[0087] This invention aims to improve the accuracy of cybersecurity risk assessment and the customization level of insurance services through the linkage mechanism of intelligent questionnaires and machine learning models.
[0088] Example 2
[0089] The present invention also provides a network security insurance questionnaire and risk assessment model adaptive linkage system. The network security insurance questionnaire and risk assessment model adaptive linkage system can be realized by executing the process steps of the network security insurance questionnaire and risk assessment model adaptive linkage method. That is, those skilled in the art can understand the network security insurance questionnaire and risk assessment model adaptive linkage method as the preferred implementation method of the network security insurance questionnaire and risk assessment model adaptive linkage system.
[0090] A method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model provided by the present invention includes:
[0091] Questionnaire management module: A cybersecurity questionnaire is established based on the GIPDRR model, and the cybersecurity questionnaire is sent to the enterprise for completion and then collected back to obtain the corresponding questionnaire data. The GIPDRR model is based on the NIST cybersecurity framework and includes six key dimensions of cybersecurity management, namely Governance, Identify, Protect, Detect, Respond, and Recover. A cybersecurity questionnaire is established by raising multiple specific questions for each dimension, and the questionnaire content is regularly adjusted and optimized based on the latest security intelligence, industry trends, and technological developments. The cybersecurity questionnaire includes structured multiple-choice questions, scoring questions, and unstructured open-ended questions.
[0092] Data collection and preprocessing module: collects threat intelligence data from the threat intelligence library, and preprocesses the questionnaire data in combination with the threat intelligence data. The data collection and preprocessing module includes: Module M2.1: performs a preliminary check on the received questionnaire data to confirm the integrity of the data. Module M2.2: If the questionnaire data is complete, module M2.3 is triggered. If the questionnaire data is incomplete, module M2.3 is triggered after processing by methods including filling in missing values, marking abnormal data, or prompting the user to re-fill the data. Module M2.3: Preprocesses the complete questionnaire data, and the preprocessing includes cleaning, classification, encoding, and integration. The method of collecting threat intelligence data from the threat intelligence library includes accessing third-party APIs and data crawling technology. The threat intelligence data includes known vulnerabilities, malware families, attack methods, and corresponding protection measures.
[0093] Feature Extraction Module: Extracts key features from preprocessed data, integrates all extracted and selected features into a unified feature vector, and uses this feature vector as input to the risk rating module. Each dimension in this feature vector represents an independent security feature or interactive feature. For structured data, statistical and data mining techniques are used to extract key features that reflect the enterprise's network security status from the encoded and standardized data. For unstructured data, natural language processing (NLP) technology and deep learning models are used to extract key information from the text and convert this key information into a feature vector that can be combined with the structured data.
[0094] Risk Rating Module: A Transformer-based deep learning model is constructed. By integrating the feature vectors with real-time threat intelligence updates, the model assesses the network security status, generates a corresponding risk score, and generates a risk assessment report. This Transformer-based deep learning model comprises a multi-head attention layer, a multi-layer encoder, a linear transformation layer, and an output layer. A random forest model is also incorporated to assist in identifying potential assessment biases or anomalies. The multi-head attention layer implements a self-attention mechanism by performing multiple parallel linear transformations on the input embedding representation. The attention score for each feature is then calculated and weighted averaged. These scores are then normalized to a probability distribution using a softmax function and applied to the weighted sum of the input vectors. Each encoder layer in the multi-layer encoder structure comprises a self-attention layer and a feedforward neural network layer. The self-attention layer is used for feature interaction, while the feedforward neural network layer performs nonlinear transformation and information extraction. The output layer maps the high-dimensional risk representation vector processed by the multi-layer encoder to a specific risk score. The linear transformation layer outputs the enterprise's specific risk score for each GIPDR dimension. Feature vectors are input into the deep learning model as multidimensional arrays or tensors. Each dimension of the vector represents an independent security feature, while the threat intelligence data contains external threat information relevant to the current enterprise situation.
[0095] The present invention also includes an explanation analysis module, which combines two explanatory analysis tools, LIME and SHAP, to analyze the output of the model from two levels, local and global.
[0096] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.
[0097] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model, characterized in that: include: Questionnaire management steps: Establish a cybersecurity questionnaire based on the GIPDRR model, send the cybersecurity questionnaire to the enterprise for completion, and collect it back to obtain the corresponding questionnaire data; Data collection and preprocessing step: collecting threat intelligence data from the threat intelligence database, and preprocessing the questionnaire data in combination with the threat intelligence data; Feature extraction step: extract key features from the preprocessed data, integrate all extracted and selected features into a unified feature vector, and use the feature vector as input to the risk rating step; Risk rating step: Build a Transformer-based deep learning model, integrate the feature vector and real-time updated threat intelligence, conduct network security status assessment to obtain the corresponding risk score, and generate a risk assessment report.
2. The method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model according to claim 1 is characterized in that: The GIPDRR model is based on the NIST cybersecurity framework, which includes six key dimensions of cybersecurity management: Governance, Identify, Protect, Detect, Respond, and Recover. We propose multiple specific questions for each dimension to establish a cybersecurity questionnaire, and regularly adjust and optimize the questionnaire content based on the latest security intelligence, industry trends and technological developments; The cybersecurity questionnaire includes structured multiple-choice questions, scoring questions, and unstructured open-ended questions.
3. The method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model according to claim 1 is characterized in that: The data collection and preprocessing steps include: Step S2.1: Conduct a preliminary check on the received questionnaire data to confirm data integrity; Step S2.2: If the questionnaire data is complete, then execute step S2.3; if the questionnaire data is incomplete, then execute step S2.3 after processing it by filling in missing values, marking abnormal data, or prompting the user to re-fill the data; Step S2.3: Preprocess the complete questionnaire data, including cleaning, classification, coding and integration.
4. The method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model according to claim 1, characterized in that: The method of collecting threat intelligence data from the threat intelligence library includes accessing a third-party API and data crawling technology; The threat intelligence data includes known vulnerabilities, malware families, attack methods, and corresponding protection measures.
5. The method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model according to claim 1, characterized in that: Each dimension in the feature vector represents an independent security feature or interaction feature; For structured data, statistical and data mining techniques are used to extract key features that can reflect the network security status of the enterprise from the coded and standardized data; For unstructured data, natural language processing (NLP) technology and deep learning models are used to extract key information from the text and convert the key information into feature vectors that can be combined with structured data.
6. The method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model according to claim 1, characterized in that: The Transformer-based deep learning model includes a multi-head attention layer, a multi-layer encoder, a linear transformation layer, and an output layer. A random forest model is introduced as an auxiliary to identify possible evaluation deviations or anomalies. The multi-head attention layer can implement the self-attention mechanism by performing multiple parallel linear transformations on the input embedding representation, then calculating the attention score of each feature and using the score for weighted averaging, and then normalizing these scores into probability distributions through the softmax number and applying them to the weighted sum of the input vector.
7. The method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model according to claim 6 is characterized in that: The multi-layer encoder structure, each layer of the encoder includes a self-attention layer and a feed-forward neural network layer; The self-attention layer is used for feature interaction, and the feedforward neural network layer is used for nonlinear transformation and information extraction; The output layer is used to map the high-dimensional risk representation vector processed by the multi-layer encoder to a specific risk score; The linear transformation layer is used to output the specific risk score of the enterprise in each GIPDRR dimension.
8. The method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model according to claim 1, characterized in that: Feature vectors are input into deep learning models in the form of multidimensional arrays or tensors; Each dimension of the vector represents an independent security feature, while the threat intelligence data contains external threat information relevant to the current enterprise situation.
9. The method for adaptively linking a cybersecurity insurance questionnaire with a risk assessment model according to claim 1, characterized in that: It also includes an explanatory analysis step, combining the two explanatory analysis tools LIME and SHAP to analyze the model output from both local and global levels.
10. A cybersecurity insurance questionnaire and risk assessment model adaptive linkage system, characterized in that: include: Questionnaire management module: Establish a cybersecurity questionnaire based on the GIPDRR model, send the cybersecurity questionnaire to the enterprise for filling out, and then collect it back to obtain the corresponding questionnaire data; Data collection and preprocessing module: collects threat intelligence data from the threat intelligence library, and preprocesses the questionnaire data in combination with the threat intelligence data; Feature extraction module: extracts key features from the preprocessed data, integrates all extracted and selected features into a unified feature vector, and uses the feature vector as the input of the risk rating module; Risk rating module: Build a Transformer-based deep learning model, integrate the feature vector and real-time updated threat intelligence to evaluate the network security status and obtain the corresponding risk score, and generate a risk assessment report.
Citation Information
Patent Citations
Network security risk assessment method, system and device
CN113542279A
Insurance service recommendation method and device, server and storage medium
CN116756424A
A data security risk assessment and management system based on large language model
CN118194358B
Cited By
Automatic questionnaire filling method based on progressive divide-and-conquer and multi-question fusion collaboration
CN121743346A