An enterprise data security capability evaluation method and system
The method and system provide a comprehensive, real-time evaluation of enterprise data security by analyzing data flows, simulating attacks, and performing multi-layered penetration testing to identify and remediate vulnerabilities, enhancing the enterprise's ability to detect and respond to potential threats.
Patent Information
- Application Number
- CN202510293258.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing data security assessment methods rely on manual monitoring inefficient efficiency, cannot analyze and monitor the dynamic flow of enterprise data in real time, lack comprehensive assessments of data security, data access behavior, and data leakage, and cannot promptly detect potential security risks.
By building an enterprise data flow mapping model, abnormal intrusion attack simulation and historical normalized behavior analysis are carried out, combined with multi-time window intrusion trajectory back-pull and multi-level penetration testing, real-time risk assessment and dynamic vulnerability repair of enterprise information systems are achieved, and a comprehensive risk situation assessment report is generated.
It has achieved accurate and comprehensive assessment of enterprise data security capabilities, timely discover abnormal behaviors and potential threats, ensure data security, optimize security protection strategies, and improve defense capabilities and repair efficiency.
Smart Images

Figure CN119808073B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of enterprise security assessment, and particularly to a method and system for evaluating enterprise data security capabilities. Background Art
[0002] With the continuous development of informatization and digitalization, a large amount of data generated and stored by enterprises during operation has gradually become one of the core assets. Enterprise data not only involves all aspects of business operation, but also carries key elements such as market competition and technological innovation. Therefore, ensuring the security and integrity of enterprise data has become an important task in modern enterprise management. With the increasing number of network security threats, the security challenges faced by enterprises are becoming more and more complex. Incidents such as data leakage, hacker attacks, and internal disclosure occur frequently, bringing huge economic losses and reputation crises to enterprises.
[0003] In traditional data security management, enterprises often rely on static security protection measures, such as firewalls, intrusion detection systems (IDS), data encryption, etc. However, these traditional methods focus more on the security protection of the system itself and often ignore the security hazards that occur during the data flow process. The security of data in various links such as transmission, storage, and access is threatened by different behavior patterns and attack means. In the face of constantly changing attack methods, traditional protection measures often seem powerless and cannot detect potential security risks in time.
[0004] In addition, existing data security assessment methods usually rely on manual monitoring and manual detection. This method is not only inefficient, but also often cannot provide a comprehensive security situation awareness. Traditional assessment methods often cannot analyze and monitor the dynamic flow of enterprise data in real time and lack a comprehensive assessment of aspects such as data security, data access behavior, and data leakage. Therefore, how to accurately and comprehensively evaluate the data security capabilities of enterprises in a dynamic and complex network environment has become an urgent problem to be solved. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention proposes a method and system for evaluating enterprise data security capabilities to solve at least one of the above technical problems.
[0006] To achieve the above object, the present invention provides a method for evaluating enterprise data security capabilities, including the following steps:
[0007] Step S1: Obtain enterprise historical security logs based on the enterprise information system, perform real-time data flow analysis and dynamic location of the storage location on the enterprise historical security logs, and construct an enterprise data flow mapping model;
[0008] Step S2: Conduct an abnormal intrusion attack simulation on the enterprise data flow mapping model, perform enterprise security risk assessment, and mark potential risk nodes;
[0009] Step S3: Conduct historical normalization behavior analysis on the enterprise data flow mapping model, and perform access behavior deviation detection on potential risk nodes to obtain enterprise system abnormal intrusion behavior data;
[0010] Step S4: Perform multi-time window intrusion trajectory backtracking on the enterprise system abnormal intrusion behavior data, and conduct in-depth mining of abnormal attack characteristics to obtain intrusion attack characteristics;
[0011] Step S5: Conduct multi-level penetration testing on the enterprise information system based on the intrusion attack characteristics, and perform adaptive vulnerability repair to obtain multi-level vulnerability repair data;
[0012] Step S6: Conduct hierarchical vulnerability repair monitoring on the multi-level vulnerability repair data, and perform comprehensive dynamic risk situation assessment to obtain a comprehensive enterprise risk situation assessment report.
[0013] Through the parsing of historical security logs, the present invention accurately identifies the flow trajectories of various types of data in an enterprise information system, helping the enterprise understand the transfer paths of data in the system. Through the real-time analysis of data flow, the storage location of data is determined in real time to ensure data security, especially the data security in different storage areas or cloud services. By simulating the data flow mapping model, various potential attack methods are simulated, such as SQL injection, malicious code execution, etc. In this way, the performance and response capabilities of the system in the face of various intrusions can be detected. Through the simulation results, the enterprise accurately evaluates security vulnerabilities and determines the weak links of the system by analyzing different attack paths, helping managers make targeted security improvement measures. Through the normalized analysis of historical data flows, the enterprise can identify common normal behavior patterns in the system as a benchmark for subsequent comparison, and perform deviation detection on the access behaviors of potential risk nodes, which helps to detect any access or operation that deviates from normal behavior in a timely manner. This helps to detect abnormal behaviors or potential security incidents at an early stage. By reverse-inferring the intrusion trajectory under multiple time windows, each stage of the attack behavior is revealed, helping security personnel track the source of the attack and the evolution process of the attack method. By deeply mining the characteristics of abnormal attack behaviors, key information such as the attacker's attack means, attack time window, attack mode, etc. can be extracted, providing data support for subsequent defense measures. By performing multi-level penetration testing on the system, the vulnerability of the system under different defense levels can be comprehensively evaluated to ensure that each layer of the protection mechanism can effectively respond to potential attacks. According to the results of the penetration testing, the enterprise quickly discovers and repairs vulnerabilities. The repair process is dynamic and real-time, timely responding to new attack patterns. The layer-by-layer monitoring of the vulnerability repair process ensures that vulnerabilities at each level are effectively repaired and prevents any repair omissions, thereby improving the accuracy of repair. The enterprise can dynamically track the security status of its information system, and adjust the security protection strategy in real time according to the real-time effect of vulnerability repair and the changes in the external threat environment. The comprehensive evaluation report provides a detailed security status analysis for the enterprise management, including the evaluation of existing risks, the effect of repair measures, and the prediction of future potential risks. This helps senior decision-makers understand the overall security situation of the enterprise and make corresponding decisions.
[0014] Preferably, step S1 includes the following steps:
[0015] Step S11: Obtain the enterprise historical security logs based on the enterprise information system, and filter out abnormal outlier data from the enterprise historical security logs to obtain outlier-filtered security logs;
[0016] Step S12: Parse the real-time data flow of the outlier-filtered security logs to generate the enterprise real-time data flow;
[0017] Step S13: Perform user dynamic behavior analysis on the enterprise real-time data flow to generate user dynamic behavior data;
[0018] Step S14: Identify the data flow paths based on the user's dynamic behavior data and extract multiple data flow paths.
[0019] Step S15: Dynamically locate the storage positions of the enterprise's real-time data streams to obtain the dynamic storage positions of each data stream.
[0020] Step S16: Based on the dynamic storage positions of each data stream, perform dynamic data flow mapping on the multiple data flow paths to construct an enterprise data flow mapping model.
[0021] The present invention removes these abnormal data that do not conform to the conventional pattern through outlier data filtering, reducing interference in the subsequent analysis process. Historical security logs contain a large amount of irrelevant information or noisy data. After filtering out the outlier data, the remaining logs will more centrally reflect the real abnormal behaviors in the system, thus enabling more effective subsequent risk assessment and defense. Focusing the data on the abnormal logs significantly improves the efficiency of analysis, helping the security team quickly identify potential threats and vulnerabilities. By performing real-time parsing on the filtered security logs, real-time data streams of the enterprise are generated, helping the enterprise monitor the status of the system instantaneously. This is very important for quickly identifying and responding to security threats. Real-time data stream parsing helps the enterprise visualize the flow and processing paths of data in its information system, contributing to the identification of any abnormal activities or potential security vulnerabilities. By performing user dynamic behavior analysis on the real-time data streams, the enterprise can comprehensively track the operation behaviors of each user in the system and clarify the behavior patterns of the users. By analyzing the user behaviors, abnormal operation behaviors such as unauthorized access or abnormal login patterns can be identified in a timely manner, which are all potential security threats. By identifying the data flow paths in the user behavior data, the enterprise can accurately locate the transfer routes of data in its information system. This is crucial for tracking sensitive data and monitoring data flow. By extracting multiple data flow paths, the enterprise analyzes the data flow from multiple perspectives, helping to discover potential security risks, such as whether there is unnecessary data access or illegal data transmission. Dynamically locating the storage position of each data stream can ensure the enterprise's complete visualization of data storage and clearly understand the changes in the data storage positions. When the storage positions of the data are accurately located, the enterprise can monitor in a timely manner and take corresponding security measures to prevent data loss or illegal access. By mapping multiple data flow paths based on the dynamic storage positions, the enterprise constructs a comprehensive data flow model. This model can accurately reflect the flow, storage, and flow direction of data in the enterprise information system. Through dynamic data flow mapping, the enterprise can identify potential risk points in different paths and implement targeted security protection measures against these risk points, reducing the invasiveness of attackers.
[0022] Preferably, step S2 includes the following steps:
[0023] Step S21: Perform multi-dimensional behavioral scenario semantic analysis on the enterprise historical security logs to obtain various behavioral scenarios of the enterprise system;
[0024] Step S22: Perform multi-attack scenario fitting based on various behavioral scenarios of the enterprise system to obtain multiple enterprise system attack scenarios;
[0025] Step S23: Perform abnormal intrusion attack simulation on the enterprise data flow mapping model according to multiple enterprise system attack scenarios to generate multi-scenario attack simulation data;
[0026] Step S24: Perform enterprise security risk assessment on the multi-scenario attack simulation data and mark potential risk nodes.
[0027] Through multi-dimensional analysis of historical security logs, the present invention deeply understands the operating status of the enterprise information system and the user behavior pattern from different perspectives (such as time, user behavior, data access, network traffic, etc.), which helps to comprehensively master the normal behavior of the system. Through semantic analysis, the interaction behaviors between different users and different system components in the system can be refined, and multiple behavioral scenarios can be identified. These scenarios help to understand the dynamic performance of the system under normal circumstances. Through the fitting of various behavioral scenarios, multiple potential attack scenarios can be generated. These attack scenarios simulate the behaviors of attackers in different situations and help the enterprise predict the attack paths. Through the fitting of multiple attack scenarios, the enterprise can better understand the manifestations of different attack methods and attacker behaviors, which enables the enterprise to strengthen protection measures in advance for specific attack methods and improve the defense prediction ability. By combining multiple attack scenarios with the enterprise data flow mapping model, abnormal intrusion attack simulation is carried out to simulate the process of attackers launching attacks on the enterprise information system through different paths. In this way, the performance of the enterprise system under different attack scenarios can be deeply understood. The simulation attack helps the enterprise comprehensively identify the vulnerabilities or weak links exploited by the attackers, including the risk points in the data flow path and the security problems of system components, etc. This process helps to identify and repair potential security vulnerabilities in advance. Through the assessment of multi-scenario attack simulation data, the potential risk points under different attack scenarios can be accurately identified, and the high-risk areas in the system can be marked. The enterprise focuses on these risk nodes accordingly to strengthen the defense. Through multi-scenario attack simulation and risk assessment, the enterprise evaluates the effectiveness of existing security protection measures, optimizes the security strategy according to the assessment results, fills the gaps in the defense, and improves the comprehensiveness of protection. By comprehensively evaluating multiple attack scenarios, the enterprise quantifies its security situation and provides a detailed security report for the management, which helps the enterprise make scientific security investment decisions, personnel allocation, and security technology selection.
[0028] Preferably, the specific steps of step S24 are:
[0029] Identify the attack data flow nodes for the multi-scenario attack simulation data to obtain multiple attack data flow nodes;
[0030] Analyze the changes in user access permissions for multiple attack data flow nodes to obtain the access permission change data for each node;
[0031] Detect data transmission anomalies for multiple attack data flow nodes and extract the data transmission anomaly characteristics of the nodes;
[0032] Based on the access permission change data of each node and the data transmission anomaly characteristics of the node, conduct intrusion prevention response analysis for each node one by one, so as to generate the intrusion prevention response data for each node;
[0033] Calculate the defense failure time for the intrusion prevention response data of each node and extract the intrusion prevention failure time of each node;
[0034] Quantify the intrusion interception of each node for the multi-scenario attack simulation data to obtain the intrusion interception value of each node;
[0035] Conduct a security risk assessment for each node one by one according to the intrusion prevention failure time of each node and the intrusion interception value of each node, so as to obtain the security risk assessment value of each node;
[0036] Based on the preset multi-scenario attack security assessment threshold, identify the risk weak nodes for the security risk assessment value of each node and mark the potential risk nodes.
[0037] By identifying the flow nodes of the attack data stream, the present invention can clarify the activity path of the attacker and master the data flow direction during the attack process. Understanding this information helps to promptly identify the actual impact scope and intrusion points of the attack. By identifying the flow nodes of the attack data, enterprises can more precisely understand through which paths the attacker enters the system and how the data flows, so as to accurately detect attack behaviors. Analyzing the changes in access permissions for each attack data flow node can help enterprises identify whether users have illegal privilege escalations during the attack process and promptly discover the risk of privilege abuse. By carefully analyzing the privilege changes of the nodes, enterprises can detect privilege escalations or overrides that occur during the attack, preventing attackers from obtaining higher access permissions through privilege expansion and thus penetrating more system areas. By detecting transmission anomalies of the attack data flow nodes, enterprises can discover whether abnormal behaviors occur during the data flow, such as abnormal increases in data volume, abnormal transmission speeds, or data packets that do not conform to the normal pattern. Through anomaly detection, potential data leakage or data tampering behaviors can be discovered, such as abnormal transmission of sensitive data, unauthorized data access, etc., so as to identify and take protective measures as early as possible. Combining the access permission change data and data transmission anomaly characteristics of each node, a defense response strategy for each node is customized. This helps enterprises take the most appropriate protective measures for each node, thereby enhancing the overall protection ability of the system. Through defense response analysis for each node one by one, it is ensured that the defense measures are not blind, but designed according to the actual risks and characteristics of each node. This makes the defense more precise and helps to avoid waste of resources and inappropriate defense strategies. By calculating the intrusion prevention failure time for each node, enterprises can identify the time periods when the defense fails and find the risks of weak defense or bypassed defense during certain time periods. This helps to strengthen the weak links of the defense. Calculating the intrusion prevention failure time can help enterprises evaluate the continuous effectiveness of existing defense measures and ensure that the defense measures are always effective within the activity time window of the attacker. Quantifying the intrusion interception capabilities of each node can clearly understand the defense level of each node, which nodes can effectively prevent attacks, and which nodes have weak defenses. The quantification of the intrusion interception value provides objective data support for security assessment, making the security assessment not only a qualitative judgment but a quantitative analysis based on specific data. Considering both the intrusion prevention failure time and the intrusion interception value comprehensively, a specific security risk value is evaluated for each node, helping enterprises clarify the risk levels of each node and make reasonable arrangements for protection priorities. Through security risk assessment, enterprises can identify which nodes in the system have higher security risks, strengthen protective measures for these nodes, and reduce the vulnerability to attacks. By comparing with the preset security assessment threshold, enterprises can quickly identify those weak nodes that exceed the risk threshold, thus marking them as high-risk areas and giving priority to strengthening security protection.After marking potential risk nodes, the enterprise timely adjusts its defense strategies for these nodes and takes necessary strengthening measures, such as enhancing access control, deploying intrusion detection systems, increasing data encryption, etc.
[0038] Preferably, the specific steps of step S3 are as follows:
[0039] Step S31: Conduct historical normalization behavior analysis on the enterprise data flow mapping model to construct a normalization behavior baseline;
[0040] Step S32: Monitor the real-time access behaviors of potential risk nodes to obtain the real-time access behavior data of risk nodes;
[0041] Step S33: Conduct access behavior deviation detection on the real-time access behavior data of risk nodes based on the normalization behavior baseline, and mark each access behavior data with a baseline deviation;
[0042] Step S34: Identify abnormal intrusion behaviors for each access behavior data with a baseline deviation, so as to obtain the enterprise system abnormal intrusion behavior data.
[0043] By identifying the flow nodes of the attack data stream, the present invention can clarify the activity path of the attacker, master the data flow direction during the attack process, and understanding this information helps to timely identify the actual impact scope and intrusion points of the attack. Analyzing the change of access rights for each attack data flow node can help an enterprise identify whether a user has an illegal privilege escalation during the attack process and timely discover the risk of privilege abuse. By carefully analyzing the privilege changes of the nodes, the enterprise can detect privilege escalations or crossings that occur during the attack and prevent the attacker from obtaining higher access rights through privilege expansion, thereby penetrating more system areas. By detecting transmission anomalies of the attack data flow nodes, the enterprise can discover whether abnormal behaviors occur during the data flow process, such as abnormal increase in data volume, abnormal transmission speed, or data packets that do not conform to the normal pattern. Through anomaly detection, potential data leakage or data tampering behaviors can be discovered, such as abnormal transmission of sensitive data, unauthorized data access, etc., so as to identify and take protective measures as early as possible. Combining the access right change data of each node and the characteristics of data transmission anomalies, a defense response strategy for each node is customized to help the enterprise take the most appropriate protective measures for each node, thereby enhancing the overall protection ability of the system. Through defense response analysis for each node one by one, it is ensured that the defense measures are not blind but designed according to the actual risks and characteristics of each node, which makes the defense more accurate and helps to avoid resource waste and inappropriate defense strategies. By calculating the intrusion prevention failure time of each node, the enterprise can identify the time period of defense failure and find the risk of weak defense or bypassed defense during certain time periods, which helps to strengthen the weak links of the defense. Calculating the intrusion prevention failure time can help the enterprise evaluate the continuous effectiveness of the existing defense measures and ensure that the defense measures are always effective within the activity time window of the attacker. Quantifying the intrusion interception ability of each node can clearly understand the defense level of each node, which nodes can effectively prevent attacks, and which nodes have weak defenses.
[0044] Preferably, the specific steps of step S31 are as follows:
[0045] Identify the normal user access behaviors of the enterprise data flow mapping model and extract the normal user access behavior data;
[0046] Perform data access path analysis on the normal user access behavior data and extract multiple normal data access paths;
[0047] Based on multiple normal data access paths, perform statistical analysis of the transmitted data volume to obtain the transmitted data volume of each access path;
[0048] Perform multi-point data transmission trend analysis on the transmitted data volume of each access path and extract the data transmission trend characteristics of each access path;
[0049] Calculate the data access frequency for the normal user access behavior data to obtain the normal user data access frequency;
[0050] Conduct data flow topology analysis based on the normal user data access frequency to generate normal user data flow topology data;
[0051] Perform enterprise system normal flow state fitting according to the data transmission trend characteristics of each access path and the normal user data flow topology data, thereby constructing a normal behavior baseline.
[0052] Through the normal behavior analysis of the enterprise historical data flow model, the present invention can construct a clear behavior baseline, which represents the normal operation state of the system and the regular behavior pattern of users, provides a reference standard for subsequent deviation detection. By monitoring the real-time access behavior of potential risk nodes, the enterprise can immediately understand the operation status of these nodes and timely detect any abnormal behavior or unusual access requests. By comparing with the normal behavior baseline, it can accurately detect access behaviors deviating from the normal mode, such as abnormal logins, abnormal access times, cross-authority access, etc. This helps to discover potential intrusion behaviors in advance. Based on the access behavior data with baseline deviation, further analyze and identify specific abnormal intrusion behaviors. Through this analysis, it can accurately judge whether there is an intrusion attempt and determine the nature and scope of the attack. Through further analysis of the deviation, abnormal behaviors can be effectively classified, thereby reducing the situation of missed reports and false alarms and ensuring the quality and credibility of the alarms.
[0053] Preferably, the specific steps of step S4 are as follows:
[0054] Step S41: Define the intrusion trace window length, and perform multi-time window intrusion trajectory backtracking on the enterprise system abnormal intrusion behavior data according to the intrusion trace window length, so as to obtain intrusion trajectories in multiple time periods;
[0055] Step S42: Reconstruct the full-cycle intrusion attack chain for the intrusion trajectories in multiple time periods to generate a full-cycle intrusion attack topology graph;
[0056] Step S43: Trace the source of the initial attack point of the full-cycle intrusion attack topology graph to obtain the initial node of the intrusion attack;
[0057] Step S44: Deeply mine the abnormal attack characteristics based on the initial node of the intrusion attack, so as to obtain the intrusion attack characteristics.
[0058] By defining the length of the intrusion traceback window, an enterprise can trace back to the specific time period when an attack occurred, and obtain the attack traces of different time periods through multiple time windows, which makes the tracing of intrusion behavior more accurate and can cover the entire attack process. The full-cycle reconstruction of the intrusion attack chain can help the enterprise comprehensively restore the attacker's intrusion path, intrusion means, attack stages, etc. From the initial attack to the final attack result, the whole picture of the attack is presented. The reconstruction of the attack chain not only helps to identify the attacker's path, but also reveals the tools, strategies, vulnerability exploitation methods, etc. used by the attacker. This helps the enterprise better understand potential attack patterns and provide a basis for future defenses. By tracing the initial attack point in the attack topology graph, the enterprise can identify the initial source of the attack, whether it is a certain vulnerability, a certain user account being stolen, an entrance being attacked, etc. This can help the enterprise strengthen protection from the source and reduce the likelihood of similar attacks in the future. Once the attack source is determined, the enterprise immediately takes defensive measures, such as blocking vulnerabilities, revoking stolen accounts, isolating infected areas, etc., to prevent the attack from spreading or further expanding to other systems. Based on the mining of abnormal attack characteristics of the initial node, the enterprise can understand the specific behavior characteristics of the attacker in the system, such as which attack techniques the attacker used, which vulnerabilities were exploited, and whether there are repetitive attack patterns. By deeply mining intrusion attack characteristics, the enterprise can provide a more detailed attack characteristic library for the future intrusion detection system, improving the intelligence and sensitivity of the detection system, and thus more accurately identifying future attacks.
[0059] Preferably, the specific steps of step S5 are as follows:
[0060] Step S51: Predict the attack diffusion based on the intrusion attack characteristics to obtain the predicted intrusion attack diffusion path;
[0061] Step S52: Conduct multi-level penetration testing on the enterprise information system according to the predicted intrusion attack diffusion path to generate penetration test data of the enterprise information system;
[0062] Step S53: Mark the risk vulnerabilities of the penetration test data of the enterprise information system to obtain enterprise risk vulnerabilities at different levels;
[0063] Step S54: Perform adaptive vulnerability repair on the enterprise risk vulnerabilities at different levels to obtain multi-level vulnerability repair data.
[0064] By analyzing the identified intrusion attack characteristics, the present invention predicts the attack diffusion path to help enterprises understand how attacks spread in the system. In this way, enterprises can take preventive measures to prevent the attacks from spreading on a larger scale and ensure the security of critical resources. The attack diffusion prediction path can reveal the potential spread channels of attacks, thereby helping enterprises optimize the network and system architectures and reduce the possibility of attack diffusion. By adjusting strategies such as network access permissions and strengthening border defenses to cope with the predicted attack paths, and through multi-level penetration testing based on the predicted attack diffusion paths, enterprises can comprehensively simulate the attack paths of attackers, thereby deeply detecting potential security vulnerabilities and weaknesses in the system. This multi-level testing covers multiple attack surfaces and can effectively evaluate the overall security of the system. By labeling risk vulnerabilities for the penetration test data, enterprises can clearly understand the risk levels of different vulnerabilities. For high-risk vulnerabilities, immediate measures are taken for repair, while for low-risk vulnerabilities, subsequent monitoring and evaluation are carried out. This helps optimize the vulnerability management strategy. The adaptive vulnerability repair technology can automatically repair according to the different risk levels and repair priorities of vulnerabilities to ensure the security of the enterprise information system. For high-risk vulnerabilities, enterprises quickly implement repair measures, while for low-risk vulnerabilities, they are gradually processed through automated tools or delayed repair to improve the repair efficiency.
[0065] Preferably, the specific steps of step S6 are as follows:
[0066] Step S61: Conduct hierarchical vulnerability repair monitoring on the multi-level vulnerability repair data to identify the unrepaired vulnerabilities at each level;
[0067] Step S62: Perform global vulnerability iterative repair on the enterprise information system according to the unrepaired vulnerabilities at each level to generate global vulnerability iterative repair data;
[0068] Step S63: Conduct a comprehensive assessment of the dynamic risk situation based on the global vulnerability iterative repair data to obtain a comprehensive assessment report on the enterprise risk situation.
[0069] The present invention effectively identifies the vulnerabilities that remain unrepaired at each level by monitoring the vulnerability repair progress of each level in the enterprise information system level by level. This hierarchical management method can help the enterprise more precisely track each link of the vulnerability repair, avoiding omissions. Through global vulnerability iterative repair, the enterprise can centrally manage and coordinate the vulnerability repairs at different levels to ensure that all levels of vulnerabilities can be repaired in a timely manner, rather than just a single level. This helps to improve the overall security of the system and avoid security problems of the entire system caused by unrepaired vulnerabilities at a certain level. Vulnerability repair is not a one-time task, but requires continuous iterative optimization. The global vulnerability iterative repair method can continuously evaluate the vulnerability repair effect and continuously optimize based on newly discovered vulnerabilities or the feedback of the repaired system to improve the repair quality. Based on the global vulnerability iterative repair data, dynamic risk assessment is carried out to comprehensively analyze the security situation of the enterprise at different repair stages in all aspects. This enables the enterprise to dynamically understand the security effect after vulnerability repair, timely discover newly emerging risks, and adjust security policies according to the actual situation. The dynamic risk situation assessment enables the enterprise to continuously monitor the security status, not only relying on regular vulnerability repairs, but also real-time evaluating the actual impact of vulnerability repair on system security, helping the enterprise to detect potential security problems in a timely manner after repair and avoid new vulnerabilities or attack surfaces caused by the repair. Through the generated comprehensive evaluation report on the enterprise risk situation, the management obtains a more comprehensive feedback on the security status, helping it to make timely security decisions. The report content not only includes the current risk situation, but also can provide the security improvement situation after repair to ensure that the management can make the optimal security decision based on the latest situation.
[0070] In this specification, an enterprise data security capability evaluation system is provided for implementing the enterprise data security capability evaluation method as described above, including:
[0071] A data flow analysis module, configured to obtain enterprise historical security logs based on the enterprise information system, perform real-time data flow analysis and dynamic location of storage locations on the enterprise historical security logs, and construct an enterprise data flow mapping model;
[0072] An intrusion attack simulation module, configured to perform abnormal intrusion attack simulation on the enterprise data flow mapping model, perform enterprise security risk assessment, and mark potential risk nodes;
[0073] A behavior deviation detection module, configured to perform historical normalization behavior analysis on the enterprise data flow mapping model and perform access behavior deviation detection on potential risk nodes to obtain enterprise system abnormal intrusion behavior data;
[0074] An intrusion trajectory reverse inference module, configured to perform multi-time window intrusion trajectory reverse inference on the enterprise system abnormal intrusion behavior data and perform in-depth mining of abnormal attack characteristics to obtain intrusion attack characteristics;
[0075] An adaptive vulnerability repair module for performing multi-level penetration testing on an enterprise information system based on intrusion attack characteristics and performing adaptive vulnerability repair to obtain multi-level vulnerability repair data;
[0076] A comprehensive evaluation module for monitoring vulnerability repair level by level for the multi-level vulnerability repair data and performing a comprehensive evaluation of the dynamic risk situation to obtain a comprehensive enterprise risk situation evaluation report.
[0077] By parsing the historical security logs in the enterprise information system and constructing a data flow mapping model, the present invention can clearly show how data flows, is stored, and interacts with other systems within the enterprise. This is a clear security situation map for enterprise managers, which can effectively guide security decisions. Dynamically locating the storage location can help enterprises quickly identify the storage areas of data and potential security hazards, reducing security risks caused by unclear data storage locations. By simulating different attack scenarios, the intrusion attack simulation module can help enterprises predict potential attack paths and intrusion methods, so as to strengthen defenses in advance. By simulating attacks, potential weak links in the enterprise system can be identified, helping the IT security team to strengthen them in advance to prevent attackers from invading the system through these weak links. By detecting deviation behaviors, the system can identify abnormal access behaviors or data operations that do not conform to the norm, timely discover potential internal or external threats, continuously monitor the access patterns of the enterprise system, and give real-time alarms according to deviation detection, helping the security team to discover and respond to potential intrusion behaviors as early as possible. By reverse inferring the intrusion trajectory, it can help enterprises clearly understand how attackers enter the system, how they move, and how attack behaviors evolve, which is crucial for subsequent defense and repair work. By deeply mining intrusion attack characteristics, it can help the security team understand the behavior patterns of attackers, and even identify potential attack tools and means, enhancing the ability to respond to future attacks. Adaptive vulnerability repair can customize repair strategies according to intrusion attack characteristics to ensure that the most critical vulnerabilities are repaired first. This process adjusts the repair strategy in real time based on attack data, making the repair more in line with the current security threat requirements. Through multi-level penetration testing, different attack surfaces can be comprehensively tested, and vulnerabilities at all levels in the system can be repaired one by one to ensure the comprehensiveness and depth of repair measures. By comprehensively evaluating the progress of vulnerability repair at all levels, a full-scale risk assessment report is provided to help management and the security team comprehensively understand the security status of the enterprise. Through dynamic analysis of the vulnerability repair data, the comprehensive evaluation module can provide real-time risk assessment, helping enterprises to adjust security strategies and repair plans in a timely manner. The comprehensive enterprise risk situation evaluation report not only provides feedback on the repair effect for enterprises, but also helps the decision-making level understand the current security situation of the enterprise and support it in formulating a reasonable security development strategy in the future. Brief Description of the Drawings
[0078] Figure 1 It is a schematic diagram of the step flow of a method for evaluating the enterprise data security capability of the present invention;
[0079] Figure 2 It is a schematic diagram of the detailed implementation steps of step S1;
[0080] Figure 3 It is a schematic diagram of the detailed implementation steps of step S2;
[0081] Figure 4 It is a schematic diagram of the detailed implementation steps of step S3. Specific implementation manners
[0082] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0083] The embodiments of the present application provide a method and a system for evaluating the enterprise data security capability. The execution subjects of the method and system for evaluating the enterprise data security capability include but are not limited to: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc. that carry the system, which can be regarded as general computing nodes of the present application. The data processing platform includes but is not limited to at least one of: audio and image management systems, information management systems, and cloud data management systems.
[0084] Please refer to Figures 1 to 4 , the present invention provides a method for evaluating the enterprise data security capability, and the method for evaluating the enterprise data security capability includes the following steps:
[0085] Step S1: Obtain the enterprise historical security logs based on the enterprise information system, perform real-time data flow parsing and dynamic location of the storage location on the enterprise historical security logs, and construct an enterprise data flow mapping model;
[0086] Step S2: Perform abnormal intrusion attack simulation on the enterprise data flow mapping model, perform enterprise security risk assessment, and mark potential risk nodes;
[0087] Step S3: Perform historical normalization behavior analysis on the enterprise data flow mapping model, and perform access behavior deviation detection on potential risk nodes, so as to obtain enterprise system abnormal intrusion behavior data;
[0088] Step S4: Perform multi-time window intrusion trajectory inversion on the enterprise system abnormal intrusion behavior data, and perform in-depth mining of abnormal attack characteristics, so as to obtain intrusion attack characteristics;
[0089] Step S5: Perform multi-level penetration testing on the enterprise information system based on the intrusion attack characteristics, and perform adaptive vulnerability repair to obtain multi-level vulnerability repair data;
[0090] Step S6: Monitor the vulnerability repair of multi-level vulnerability repair data level by level, and conduct a comprehensive assessment of the dynamic risk situation to obtain a comprehensive assessment report on the enterprise risk situation.
[0091] In the embodiment of the present invention, refer to Figure 1 , which is a schematic diagram of the step flow of a method for evaluating the enterprise data security capability of the present invention. In this example, the steps of the method for evaluating the enterprise data security capability include:
[0092] Step S1: Obtain the enterprise historical security logs based on the enterprise information system, perform real-time data flow parsing and dynamic positioning of the storage location of the enterprise historical security logs, and construct an enterprise data flow mapping model;
[0093] In this embodiment, historical security logs are extracted from enterprise information systems (such as SIEM systems, application servers, network devices, etc.). These logs usually include user activity logs, system event logs, network traffic logs, etc. Set the time range for data collection. It is recommended to collect logs for at least the past six months to ensure the representativeness and comprehensiveness of the data. Standardize the format of the collected logs to ensure that all logs follow a unified format (such as JSON, CSV, or XML) for subsequent parsing and analysis. Use tools (such as Logstash or Fluentd) to preprocess the logs, remove irrelevant information and duplicate records, and retain key information. Define real-time parsing rules based on the log format and key information. For user activity logs, extract fields such as user ID, timestamp, operation type, and access object. Use tools such as regular expressions or JSON parsers to ensure accurate extraction of the required information. Select a suitable real-time stream processing framework (such as Apache Kafka, Apache Flink, or Apache Spark Streaming) for processing real-time data streams. Set the parameters for stream processing, such as batch size (e.g., 1000 records) and window size (e.g., 5 minutes), to optimize performance. Input the real-time log stream into the selected stream processing framework, apply the defined parsing rules, extract key information in real-time, and store it in memory or a temporary database. Monitor performance metrics during the parsing process, such as latency and throughput, to ensure that the system can efficiently process large-scale data streams. Design the enterprise's storage architecture, clarify the locations for data storage, including relational databases (such as MySQL, PostgreSQL), NoSQL databases (such as MongoDB, Cassandra), and data warehouses (such as Amazon Redshift, Google BigQuery). Select a suitable storage solution according to the characteristics of the data. For example, use a relational database for structured data and a NoSQL database for unstructured data. Develop a dynamic location function to automatically identify the data storage location and record the data flow direction through metadata management or data catalog tools (such as Apache Atlas or AWS Glue). Set the monitoring frequency (e.g., every hour) and update mechanism to ensure the dynamism and real-time nature of the data storage location. Design a data flow mapping model based on the parsed log information and storage location, and graphically display the flow path of data in the enterprise information system. Select a suitable modeling tool (such as Graphviz or Neo4j) to visualize the data flow model for subsequent analysis and monitoring. Verify the accuracy and effectiveness of the data flow mapping model by comparing the actual data flow with the model prediction. Randomly select some logs and check whether their flow paths in the model are consistent. Generate a report on the data flow mapping model, detailing the model construction process, key findings, and future optimization directions.
[0094] Step S2: Conduct anomaly intrusion attack simulation on the enterprise data flow mapping model, perform enterprise security risk assessment, and mark potential risk nodes;
[0095] In this embodiment, select a suitable attack model, such as the MITRE ATT&CK framework, as the basis for simulation, which covers a variety of attack techniques and strategies. This framework can help identify attack paths and potential attacker behaviors, define attack scenarios, including targets, attack means, attack timing, and attacker identity characteristics. Design an SQL injection attack scenario targeting the database, determine the parameters of the simulation attack, such as attack frequency (e.g., once per hour), attack duration (e.g., 1 day), and attack intensity (e.g., low, medium, high). Use historical data (e.g., past attack events) to set the attack parameters to ensure the authenticity and rationality of the simulation results. Create an isolated test environment to simulate the actual operation of the enterprise information system, including application servers, databases, and network devices, etc. Use virtualization technologies (such as Docker or VMware) to quickly deploy the test environment, thereby ensuring the flexibility and controllability of the simulation process. Select a suitable simulation tool (such as Cyborg Security or Core Impact), which can simulate various attack behaviors and record the results. Configure the parameters of the simulation tool, such as attack type, target system, and attack path, to facilitate the execution of the attack simulation. Execute the set attack scenario, simulate the behavior of the attacker through the simulation tool, record the results and impacts of each attack step, monitor the system response, including log records, system resource usage, and the effectiveness of security protection measures. Define the criteria for risk assessment, including risk levels (e.g., high, medium, low), scope of impact (e.g., single system or multiple systems), and attack success rate, etc. Develop a risk assessment framework based on industry standards (such as ISO 27001 or NIST SP 800-30) to ensure the professionalism and systematicness of the assessment process. Analyze the simulation results, identify potential risk nodes, including the attacked systems, exploited vulnerabilities, and affected users, etc. Record the detailed information of each risk node, including risk type, potential impact, and repair suggestions, to form a preliminary risk assessment report. Mark the identified potential risk nodes, classify them using color coding (e.g., red for high risk, yellow for medium risk, and green for low risk) for subsequent management. Integrate the risk nodes into the enterprise data flow mapping model to form a complete risk situation map for visual analysis.
[0096] Step S3: Conduct historical normalization behavior analysis on the enterprise data flow mapping model and perform access behavior deviation detection on potential risk nodes to obtain enterprise system abnormal intrusion behavior data;
[0097] In this embodiment, historical behavior data is extracted from the enterprise information system, including user access logs, system event logs, and network traffic data. These data should cover a reasonable time range, preferably at least 6 months, to ensure the representativeness of the samples. Set the parameters for data extraction, such as time period, user type, and access object, to ensure that the acquired data is comprehensive and meaningful. Clean the collected data to remove irrelevant or redundant information, and standardize the data format (such as timestamp format, user ID format, etc.) for subsequent analysis. Use data processing tools (such as Pandas or Apache Spark) to preprocess the data to ensure data consistency and integrity. By analyzing historical data, establish a normalized model of user behavior. Use clustering analysis (such as K-means clustering) to identify the patterns of normal user behavior. Set the parameters of the normalized model, such as the number of clusters (set to 5 clusters based on different types of user behavior) and the similarity threshold, to determine the boundaries of normal behavior. Define the deviation detection criteria for access behavior, including the types of deviations (such as access frequency, access time, access object, etc.), and set the corresponding thresholds (such as the access frequency exceeding 3 times the norm). Consider using multiple detection methods, such as statistical-based methods (such as Z-score analysis) and machine learning-based methods (such as anomaly detection algorithms), to improve the detection accuracy. Monitor the access behavior of potential risk nodes, and record the access patterns of users in real time, including information such as access time, access object, and access frequency. Compare the real-time monitored data with the historical normalized behavior model to identify behavior deviations. Use similarity calculation (such as cosine similarity) to judge the difference between the current behavior and the normalized behavior, identify the access behavior with deviations, and record relevant information, including user ID, timestamp, deviation type, and deviation degree, etc. Generate a deviation detection report, which details the detected abnormal behaviors and their potential impacts, providing a basis for subsequent analysis and response. Integrate the identified abnormal behavior data into a structured dataset for subsequent storage and analysis. The dataset should include fields such as user ID, access time, access object, deviation type, and deviation degree. Use a database management system (such as MySQL or MongoDB) to store the abnormal behavior data to ensure data security and accessibility. Analyze the abnormal behavior data to identify potential intrusion behavior patterns, record the characteristics and impacts of each pattern, and generate an abnormal intrusion behavior data report, which details the identified abnormal behaviors, including the intrusion paths and attack means, to help the security management team formulate countermeasures.
[0098] Step S4: Perform multi-time-window intrusion trajectory reverse inference on the enterprise system abnormal intrusion behavior data, and deeply mine the abnormal attack characteristics to obtain the intrusion attack characteristics;
[0099] In this embodiment, abnormal intrusion behavior data is collected, including timestamps, user IDs, accessed objects, deviation types, etc., to ensure the integrity and accuracy of the data. These data are integrated into a structured database (such as PostgreSQL or MongoDB). Time window parameters are set. For example, the data is divided into multiple time windows (such as 1 hour, 3 hours, 6 hours) to facilitate the analysis of intrusion behaviors in different time periods. Model parameters for trajectory backtracking are set, including backtracking algorithms (such as graph-based backtracking algorithms or sequence pattern mining algorithms) to reconstruct the attacker's behavior path. Graph analysis tools (such as Gephi or NetworkX) are used to construct an intrusion trajectory graph, where nodes represent users or systems and edges represent behavior paths. By analyzing abnormal behavior data, abnormal access sequences are identified and intrusion trajectories are reconstructed. Timestamps and accessed objects are used to concatenate paths to form a complete attack path. The attack activities in each time window are monitored, the intrusion trajectories within each time window are recorded, and corresponding trajectory datasets are generated. Feature engineering is performed on the reconstructed intrusion trajectories to extract key features, such as attack path length, access frequency, target system type, and attack techniques. Data analysis tools (such as Pandas and NumPy) are used to perform statistical analysis on the features, calculate feature distributions and correlations to identify key attack features. Machine learning algorithms (such as random forests, support vector machines, or deep learning models) are applied to train the extracted features to identify abnormal attack features. Model parameters are set, such as the ratio of the training set to the test set (such as 80% training, 20% testing), and cross-validation is used to optimize the model performance. Deep learning frameworks (such as TensorFlow or PyTorch) are used to construct deep learning models for in-depth feature mining. Temporal models such as LSTM or GRU are considered to analyze attack sequence features. A feature importance report is generated, detailing the impact degree of each feature on intrusion behaviors to help understand the attacker's behavior patterns. The identified attack features are integrated into a structured dataset for subsequent analysis. The dataset should include fields such as attack methods, attack paths, target systems, and attack frequencies. Visualization tools (such as Tableau or Matplotlib) are used to visually display the attack features to help the security team understand the attack patterns and features. Feature analysis is performed to identify common attack patterns and strategies, and the description, potential impact, and protection suggestions of each attack feature are recorded. An intrusion attack feature report is generated, detailing the identified attack features and providing a basis for the formulation of enterprise security policies
[0100] Step S5: Conduct multi-level penetration testing on the enterprise information system based on the intrusion attack features and perform adaptive vulnerability repair to obtain multi-level vulnerability repair data;
[0101] In this embodiment, according to the previously identified intrusion attack characteristics, a penetration testing plan is formulated to clarify the scope of testing (such as network layer, application layer, and database layer), and specific testing objectives are set (such as identifying vulnerabilities in critical assets). Determine the time frame for testing. For example, set it for one week to conduct multiple rounds of testing to cover different attack scenarios and time windows. Select appropriate penetration testing tools, such as Metasploit, Burp Suite, and Nessus, to ensure that these tools can cover the testing requirements at all levels. Set up a penetration testing environment to ensure that the testing will not affect the production environment. Use virtualization technologies (such as VMware or Docker) to create an isolated testing environment. Implement multi-level penetration testing according to the formulated penetration testing plan. First, conduct network layer testing to identify network vulnerabilities (such as open ports, unencrypted communications, etc.). Next, conduct application layer testing, use the identified intrusion attack characteristics (such as SQL injection, cross-site scripting, etc.) to attack the application, and record the test results at each step. Record the results of each test session, including information such as the discovered vulnerabilities, the severity of the vulnerabilities, the scope of impact, and exploitability. Classify the test results to generate a preliminary vulnerability report, and detail the characteristics, repair suggestions, and priorities of each vulnerability. According to the penetration testing results, formulate corresponding vulnerability repair strategies. For high-risk vulnerabilities, set the priority of immediate repair; for medium- and low-risk vulnerabilities, formulate a regular repair plan. Set the repair solutions, including patch updates, configuration adjustments, or security policy optimizations, and ensure that the repair measures will not affect the normal operation of the system. Select appropriate repair tools (such as Ansible, Chef) for automated repair to ensure the efficiency and accuracy of the repair process. Implement vulnerability repair step by step according to the formulated repair plan, and record the detailed steps and results of each repair process. Record the repair status of each vulnerability, including information such as repair time, repair measures, and verification results, to form multi-level vulnerability repair data. Integrate the repair data into a database (such as MySQL or MongoDB) to ensure the traceability and accessibility of the data.
[0102] Step S6: Conduct hierarchical vulnerability repair monitoring on the multi-level vulnerability repair data and conduct a comprehensive assessment of the dynamic risk situation to obtain a comprehensive enterprise risk situation assessment report.
[0103] In this embodiment, a vulnerability repair monitoring system is formulated, and the monitoring standards and indicators are set. These indicators include the number of unpatched vulnerabilities, the repair rate, the severity level of vulnerabilities (such as high, medium, low), the repair time, and the vulnerability recurrence rate, etc. The setting of the monitoring frequency is also crucial. It is recommended to monitor the vulnerability status daily and generate a detailed report once a week to promptly discover problems. Select suitable vulnerability monitoring tools (such as Qualys, Nessus, or OpenVAS) to ensure that these tools can cover the vulnerability monitoring requirements at all levels, ensure the integration of the monitoring tools with the enterprise information system, facilitate real-time acquisition of vulnerability data, and conduct automated monitoring. Collect multi-level vulnerability repair data, including the repair status, repair measures, repair time, verification results, etc. at each level. These data should be from the records of the penetration testing and adaptive vulnerability repair phases. Set the data format (such as JSON or CSV) to ensure the readability and consistency of the data for subsequent analysis. Analyze the collected data to identify the unpatched vulnerabilities and their impact scope, evaluate the security status at each level, and use statistical analysis tools (such as the R or Python's Pandas library) for data processing and visualization. The monitoring results should be presented in the form of charts. For example, use bar charts to show the changes in the number of unpatched vulnerabilities and the repair rate to help management quickly understand the current security situation. Select a suitable risk assessment model, such as adopting a quantitative risk assessment model (such as the FAIR model) or a qualitative risk assessment model, comprehensively consider the vulnerability repair status and external threat factors, and set the parameters for risk assessment, including the number of vulnerabilities, threat level, potential impact, and attack probability, etc. for a comprehensive assessment. Based on the monitored vulnerability data, analyze the risk situation at each level and evaluate its impact on the overall information system. Use multi-dimensional assessment methods, such as SWOT analysis or PEST analysis, comprehensively consider internal and external factors to generate risk assessment results. Record the risk level, main risk factors, impacts, and improvement suggestions at each level. Integrate the monitoring results and risk assessment results to generate a comprehensive enterprise risk situation assessment report. The report should include content such as vulnerability monitoring results, risk assessment results, potential risk nodes, and recommended repair measures. Use visualization tools (such as Tableau or PowerBI) to display the risk situation to help management quickly understand the security status of the enterprise.
[0104] In this embodiment, refer to Figure 2 , which is a schematic diagram of the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include:
[0105] Step S11: Obtain the enterprise historical security logs based on the enterprise information system, and filter the abnormal outlier data in the enterprise historical security logs to obtain outlier-filtered security logs;
[0106] Step S12: Perform real-time data stream parsing on the outlier-filtered security logs to generate the enterprise real-time data stream;
[0107] Step S13: Conduct user dynamic behavior analysis on the enterprise real-time data stream to generate user dynamic behavior data;
[0108] Step S14: Identify the data flow paths based on the user dynamic behavior data and extract multiple data flow paths;
[0109] Step S15: Dynamically locate the storage positions of the enterprise real-time data stream to obtain the dynamic storage positions of each data stream;
[0110] Step S16: Perform dynamic data flow mapping on the multiple data flow paths based on the dynamic storage positions of each data stream, and construct an enterprise data flow mapping model.
[0111] In this embodiment, historical security logs are extracted from the enterprise information system to ensure data integrity and coverage. Usually, historical logs include information such as access records, system warnings, and security events, with the data volume reaching millions of records. Set the time range for data collection, such as the past year, for comprehensive analysis. Preprocess the collected historical security logs, including removing duplicate records, standardizing field formats (such as timestamp format, IP address format), etc., to ensure data consistency and availability. Record the basic statistical information of the data, such as the total number of records, the number of fields, and the missing value situation, for subsequent analysis. Select appropriate anomaly detection algorithms, such as Isolation Forest, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), or the Z-score method. These algorithms can effectively identify outliers in the data. Set the parameters of the algorithm, such as the number of trees and the threshold in Isolation Forest, to ensure that the algorithm can adapt to the characteristics of the data. Run the selected outlier detection algorithm to analyze the historical security logs and identify abnormal behaviors or data records. Record the characteristic values of each outlier for subsequent analysis. Remove the detected outlier data from the original logs to form outlier-filtered security logs, ensuring that subsequent analysis is based on valid and reliable data. Generate a statistical report of the outlier-filtered security logs, recording the data changes before and after filtering, including the remaining number of records, the types of abnormal events, and their frequencies. Use visualization tools (such as Matplotlib or Tableau) to display the results of outlier detection for easy analysis and understanding of the characteristics of the outlier phenomenon. Select an appropriate real-time data stream processing framework, such as Apache Kafka, Apache Flink, or Apache Spark Streaming. These tools can handle high-throughput data streams and support real-time data processing and analysis. Configure the data stream system to ensure that it can receive real-time log data from the enterprise information system. Use the outlier-filtered security logs as the input source of the data stream, set the format and transmission protocol of the data stream (such as JSON, XML), to ensure that the data can be correctly parsed and processed. Set the processing frequency of the data stream, usually processing thousands of records per second, to meet the requirements of real-time analysis. Develop a data stream parsing program to parse the input security log data in real time, extract important fields (such as timestamp, user ID, event type, etc.), and convert them into structured data. Implement data cleaning and formatting to remove irrelevant information and noise data, ensuring the accuracy and consistency of the parsing results. Generate an enterprise real-time data stream, recording the results of real-time parsing, including the parsed data structure and its storage location. Store the real-time data stream in a database or data warehouse for subsequent analysis and query.Use a NoSQL database (such as MongoDB) or a relational database (such as PostgreSQL) for storage. Monitor the data stream parsing process in real time to ensure the normal operation of the system and promptly handle any faults that occur. Use visualization tools to display the changes in the real-time data stream, such as real-time user activity charts, to facilitate the monitoring of the enterprise's security status. Extract user-related information from the real-time data stream, including user ID, access time, operation type, accessed module, etc., to form a dataset for user dynamic behavior analysis. Ensure that the integrated data can reflect the user's interaction behaviors and behavior patterns. Select an appropriate user behavior analysis model, such as clustering analysis, association rule analysis, or time series analysis. These models can effectively capture the user's behavior characteristics and trends. Set the parameters of the model, such as the number of clusters and the time window, to better adapt to the characteristics of the user behavior data. Apply the selected behavior analysis model to analyze the integrated user data and identify the user's behavior patterns and dynamic changes. Record the behavior characteristics of each user, including access frequency, operation duration, and access path, etc. By tagging user behaviors, identify normal behaviors and abnormal behaviors and generate user dynamic behavior data. Generate a user dynamic behavior analysis report, which details the behavior characteristics and change trends of each user, including behavior classification and frequency statistics. Use visualization tools to display the results of the user behavior analysis, such as a user behavior heat map, to facilitate enterprise managers to understand user dynamics. Path identification method selection: Select an appropriate data flow path identification method, such as graph algorithms (such as Dijkstra's algorithm, A* algorithm) or sequence pattern mining algorithms (such as PrefixSpan). These methods can effectively identify the user's flow path in the data stream. Determine the key parameters of path identification, such as the minimum path length and the maximum number of hops, to facilitate the screening of valid paths. Based on the user dynamic behavior data, apply the selected path identification method to analyze the user's operation sequence and extract multiple data flow paths. Record the starting point, ending point, and the operation steps passed through for each path. Identify common flow paths and abnormal paths and analyze their impact on enterprise security. Generate a data flow path identification report, which details the characteristics of each path, including access frequency, number of flowing users, and path length, etc. Use visualization tools to display the distribution of the data flow paths to help the enterprise analyze user interaction behaviors. Select an appropriate storage system architecture, such as a distributed file system (such as HDFS) or object storage (such as AWS S3), to support the storage of dynamic data streams. Determine the structure and format of the data storage to facilitate the rapid retrieval and access of data. Implement dynamic positioning of the storage location according to the characteristics of the real-time data stream. Record the storage path of each data stream, including the storage node, storage time, and storage format of the data. Adopt a metadata management system to ensure that the storage information of each data stream can be quickly queried and accessed.Generate a dynamic positioning report of storage locations, which details the dynamic storage locations of each data stream, including information such as storage time, node location, and storage capacity. Use visualization tools to display the distribution of storage locations, facilitating enterprises to understand the storage strategy of data streams. Select a suitable data flow mapping model, such as a graph database (e.g., Neo4j) or a flow graph model, to support the dynamic mapping of multiple data flow paths. Determine the key parameters of the mapping model, such as the definition of nodes and edges, the direction and intensity of data flow. Based on the collected storage location and data flow path information, construct an enterprise data flow mapping model. Ensure that the flow direction, storage nodes, and dynamic characteristics of each path can be clearly represented. Use visualization tools to display the data flow mapping results, facilitating enterprises to analyze the efficiency and security of data flow. Generate a data flow mapping model report, which details the flow paths, storage locations, and flow characteristics of each data stream. Regularly update the mapping model to reflect the changes in enterprise data flow and ensure that the model always remains accurate and effective.
[0112] In this embodiment, refer to Figure 3 , which is a schematic diagram of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include:
[0113] Step S21: Conduct multi-dimensional behavioral scenario semantic analysis on the enterprise historical security logs to obtain various behavioral scenarios of the enterprise system;
[0114] Step S22: Fit multiple attack scenarios based on various behavioral scenarios of the enterprise system to obtain multiple enterprise system attack scenarios;
[0115] Step S23: Conduct abnormal intrusion attack simulation on the enterprise data flow mapping model according to multiple enterprise system attack scenarios to generate multi-scenario attack simulation data;
[0116] Step S24: Conduct enterprise security risk assessment on the multi-scenario attack simulation data and mark potential risk nodes.
[0117] In this embodiment, historical security logs are extracted from the enterprise information system, including user behavior records, system events, security alerts, etc. To ensure the integrity and diversity of the log data, the collection time range is set to the past year, and the data volume reaches millions of records. The data is preprocessed to remove duplicate records, handle missing values, and standardize fields such as timestamps and user IDs to ensure data consistency and quality. Appropriate semantic analysis techniques are selected, such as natural language processing (NLP) techniques, combined with topic models (such as LDA) and sentiment analysis, to extract semantic information and behavior characteristics from the logs. Analysis parameters are set, such as the number of topics and the sentiment analysis threshold, to ensure that the model can effectively identify different behavior scenarios. The selected semantic analysis technique is used to analyze the historical security logs to identify multi-dimensional behavior scenarios, including normal behavior scenarios and potential abnormal behavior scenarios. The behavior scenarios are classified, and the characteristics and key metrics (such as event frequency, user engagement, etc.) of each scenario are recorded. A behavior scenario report is generated, detailing the characteristics of each scenario, including information such as scenario name, description, occurrence frequency, and related user IDs. Based on the analysis results of the multi-dimensional behavior scenarios, the identified attack types, such as SQL injection, DDoS attacks, malware propagation, etc., are determined, and an attack type library is established. The characteristics of the attack types are set, including attack methods, targets, and potential impacts. Appropriate attack scenario fitting methods are selected, such as case-based learning methods or simulation generation methods, to construct attack scenarios that match the user behavior scenarios. A random generation algorithm is used in combination with actual attack cases for fitting. Fitting parameters are set, such as scenario complexity, attack intensity, and attack frequency, to facilitate the generation of multiple attack scenarios. Using the selected fitting method, multiple enterprise system attack scenarios are generated based on the identified attack types and multi-dimensional behavior scenarios, and the characteristics and trigger conditions of each attack scenario are recorded. An attack scenario report is generated, detailing the attack type, path, scope of influence, and potential targets of each scenario. The structure and data flow path of the enterprise data flow mapping model are confirmed to ensure that the model can reflect the dynamic changes and potential risk points of the enterprise data flow. The model should include data storage locations, flow paths, and key nodes. Appropriate attack simulation tools are selected, such as Metasploit or Cyborg, which can simulate different types of attack scenarios and evaluate their impact on the enterprise system. The simulation tool is configured to ensure that it can be combined with the enterprise data flow mapping model. Based on the generated multiple enterprise system attack scenarios, an abnormal intrusion attack simulation is implemented. The data flow and system responses during the attack are simulated, and the impact of the attack on the data flow and storage locations is recorded. An attack simulation data report is generated, detailing the data flow changes, system responses, and potential losses under each attack scenario. Data analysis tools are used to perform statistical analysis on the simulation data to evaluate the impact of different attack scenarios on the enterprise system and identify potential security vulnerabilities and risk points.Select appropriate security risk assessment indicators, such as the attack success rate, data loss volume, system response time, and recovery time, etc., to establish a risk assessment model. Set the weight of each indicator to comprehensively evaluate the security risks of the enterprise. Use the selected risk assessment model to analyze the multi-scenario attack simulation data and evaluate the security risks under different attack scenarios. Calculate the values of each risk indicator and mark potential risk nodes according to the set thresholds. Record the characteristics of each risk node, including the impact scope, risk level, and recommended countermeasures. Generate a security risk assessment report, which details the information of each potential risk node, including the attack scenario, risk level, and recommended defense measures. Use visualization tools to display the risk assessment results to help enterprise management intuitively understand the security risks and formulate corresponding security strategies.
[0118] In this embodiment, the specific steps of step S24 are as follows:
[0119] Identify the attack data flow nodes from the multi-scenario attack simulation data to obtain multiple attack data flow nodes;
[0120] Analyze the changes in user access permissions for multiple attack data flow nodes to obtain the access permission change data for each node;
[0121] Detect abnormal data transmission for multiple attack data flow nodes and extract the abnormal data transmission characteristics of the nodes;
[0122] Based on the access permission change data of each node and the abnormal data transmission characteristics of the nodes, conduct intrusion prevention response analysis for each node one by one, so as to generate the intrusion prevention response data for each node;
[0123] Calculate the defense failure time for the intrusion prevention response data of each node and extract the intrusion prevention failure time of each node;
[0124] Quantify the intrusion interception of each node for the multi-scenario attack simulation data to obtain the intrusion interception value of each node;
[0125] Conduct security risk assessment for each node one by one according to the intrusion prevention failure time of each node and the intrusion interception value of each node, so as to obtain the security risk assessment value of each node;
[0126] Based on the preset multi-scenario attack security assessment threshold, identify risk weak nodes for the security risk assessment value of each node and mark potential risk nodes.
[0127] In this embodiment, a flow analysis method based on graph theory, such as the Dijkstra algorithm or graph traversal algorithm, is adopted to analyze the data flow path and identify the nodes to which the attack data flows. Analysis parameters, such as the maximum path length and minimum data volume, are set to ensure the identification of effective nodes for the attack data flow. Using the selected flow analysis method, the attack simulation data is traversed to identify multiple nodes for the attack data flow, including the attack source, victim nodes, and intermediate forwarding nodes, etc. The characteristic information of each node is recorded, including the node type, flow direction, and data traffic, etc., for subsequent analysis. The relevant user access permission data is extracted from the enterprise user management system to ensure that the data can reflect the changes in access permissions of each node. The permission data is preprocessed to remove redundant information and standardize the permission levels (such as administrator, user, visitor, etc.). A time series analysis method or comparative analysis method is used to analyze the changes in access permissions of each node before and after the attack. Analysis parameters, such as the time window and change threshold, are set to ensure the identification of significant permission changes. Each node for the attack data flow is analyzed, and the changes in its access permissions are recorded to generate a permission change data report, including information such as the permission levels before and after the change, user IDs, and change times. Nodes with increased or decreased permissions are identified, and their potential impacts on data security are analyzed. A suitable anomaly detection algorithm, such as a statistics-based method (such as Z-score) or a machine learning method (such as Isolation Forest), is selected to identify abnormal situations in data transmission. Algorithm parameters, such as the anomaly threshold and training samples, are set to ensure that the model can adapt to the characteristics of the data. The data transmission characteristics, including the packet size, transmission frequency, and transmission time, etc., are extracted from the nodes for the attack data flow as the model input. The feature data is normalized to ensure comparison of different features on the same scale. Using the selected anomaly detection algorithm, the data transmission characteristics of each node are analyzed to identify abnormal transmission behaviors. The abnormal characteristics of each node, including the abnormal type, frequency, and impact degree, etc., are recorded for subsequent analysis. A suitable intrusion prevention strategy, such as traffic monitoring, access control restriction, and intrusion detection system (IDS), is selected to design corresponding defense responses for each node. The triggering conditions of the defense strategy, such as the amplitude of permission change and abnormal transmission frequency, are set to ensure a quick response to potential intrusions. For each node for the attack data flow, the intrusion prevention response effect is analyzed by combining the access permission changes and data transmission abnormal characteristics. The defense response data of each node, including the response time, response measures, and effect evaluation, etc., are recorded. A node intrusion prevention response report is generated, which details the intrusion prevention response of each node, and a visualization tool is used to display the response effect.
[0128] Define the calculation method for the intrusion prevention failure time, such as the time period from the start of an attack to the failure of the defense measure, and analyze it using an event-driven model. Determine the calculation parameters for the failure time, such as the time window and the duration of the attack behavior. For each node, calculate the failure time of its defense measure based on the intrusion prevention response data, and record the failure time of each node and its influencing factors. Identify the nodes with defense failures and analyze their potential impact on enterprise security. Generate a defense failure time report, detailing the failure time and the scope of influence of each node, and use visualization tools to display the distribution of the failure time. Quantify intrusion interception. Determine the metrics for intrusion interception quantification, such as the interception success rate, the number of interceptions, and the interception type. Set the weight for each metric to facilitate the comprehensive evaluation of the defense capabilities of the nodes. Based on the attack simulation data and the defense response data, perform intrusion interception quantification for each node and record the intrusion interception value of each node. Identify the nodes with poor interception effects and analyze the reasons. Generate an intrusion interception quantification report, detailing the interception value and related metrics of each node, and use visualization tools to display the interception effect. Determine the metrics for security risk assessment, including the defense failure time, the intrusion interception value, and the node importance, etc. Set the weight for each metric to facilitate the comprehensive evaluation of the security risks of the nodes. For each node, calculate the security risk assessment value based on the intrusion prevention failure time and the intrusion interception value, and record the risk level of each node. Identify the nodes with higher risks and analyze their potential impact. Generate a security risk assessment report, detailing the risk assessment value and related metrics of each node, and use visualization tools to display the risk assessment results. According to historical data and industry standards, set the threshold for security risk assessment. For example, nodes with a risk value exceeding 0.7 are considered high-risk nodes. Verify the set threshold through experimental data to ensure its rationality and effectiveness. Based on the set threshold, analyze the security risk assessment value of each node and mark the nodes with weak risks. Record the characteristics of each marked node, including information such as the risk assessment value, the intrusion interception value, and the defense failure time. Generate a report on the identification of nodes with weak risks, detailing the information of high-risk nodes, and use visualization tools to display the distribution of nodes with weak risks to help the enterprise formulate corresponding security policies.
[0129] In this embodiment, refer to Figure 4 , which is a schematic diagram of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include:
[0130] Step S31: Conduct historical normalization behavior analysis on the enterprise data flow mapping model to construct a normalization behavior baseline;
[0131] Step S32: Monitor the real-time access behavior of potential risk nodes to obtain the real-time access behavior data of the risk nodes;
[0132] Step S33: Perform access behavior deviation detection on the real-time access behavior data of risk nodes based on the normalized behavior baseline, and mark each access behavior data with a baseline deviation;
[0133] Step S34: Identify abnormal intrusion behavior for each access behavior data with a baseline deviation, so as to obtain enterprise system abnormal intrusion behavior data.
[0134] In this embodiment, key behavior characteristics are determined, such as access frequency, data request volume, access time period, and user role, etc. These characteristics are used as the input for normalized analysis. Statistical analysis is performed on each characteristic to calculate indicators such as mean and standard deviation for subsequent modeling. A suitable modeling method is selected, such as clustering analysis (e.g., K-means) or statistical model (e.g., Gaussian mixture model), to construct the normalized behavior baseline. Model parameters are set, such as the number of clusters and thresholds, to ensure that the model can effectively adapt to the characteristics of the data. The selected modeling method is used to analyze the extracted behavior characteristics to generate the normalized behavior baseline, and record the characteristics and indicators of each behavior pattern. A normalized behavior baseline report is generated, detailing information such as the mean, standard deviation, and occurrence frequency of each behavior pattern. A suitable real-time monitoring tool is selected, such as the ELK Stack (Elasticsearch, Logstash, Kibana) or Splunk, which can collect, analyze, and visualize data streams in real time. The monitoring system is configured to ensure that it can receive real-time access data from risk nodes. The access behavior data of risk nodes is used as the input source for monitoring, and the data stream format (such as JSON) is set for subsequent analysis. The monitoring frequency is set, for example, thousands of records are collected per minute, to ensure that changes in access behavior can be captured in a timely manner. Real-time data collection is implemented to monitor the access behavior of risk nodes, including characteristics such as user access time, access content, and access method. The real-time monitoring data is stored in a database (such as a NoSQL database) for subsequent analysis and query. The data stream is monitored in real time to ensure the normal operation of the system and handle any faults that occur in a timely manner. A visualization tool is used to display the changes in real-time access behavior, such as a user access heat map, to help management personnel understand the access situation of risk nodes in a timely manner. A suitable deviation detection method is selected, such as Z-score analysis, outlier detection algorithms (e.g., Isolation Forest), or control chart methods to identify deviation behaviors. Model parameters are set, such as deviation thresholds and sample sizes, to ensure that the model can effectively identify deviations in access behavior. The selected detection method is used to analyze the real-time access behavior data and calculate the deviation degree of each access behavior from the normalized behavior baseline. Mark each access behavior data with a baseline deviation and record its characteristics, including the deviation value, access time, user ID, and access content, etc.
[0135] Generate an access behavior deviation detection report, which details the characteristics and impacts of each deviation behavior, including the deviation type and security risks. Use visualization tools to display the deviation detection results to help management understand the abnormal access behaviors at risk nodes. Select appropriate abnormal behavior recognition algorithms, such as Support Vector Machine (SVM), Random Forest, or deep learning models (such as LSTM), to identify abnormal intrusion behaviors. Set the training parameters of the model, such as the training sample ratio and cross-validation method, to ensure the accuracy and robustness of the model. Use the labeled baseline deviation access behavior data to train the selected recognition algorithm to generate an abnormal intrusion behavior recognition model. Evaluate the performance of the model through cross-validation to ensure its ability to effectively identify abnormal behaviors. Apply the trained model to analyze the access behavior data of each baseline deviation to identify abnormal intrusion behaviors. Record the characteristics of each identified abnormal behavior, including the behavior type, scope of influence, and potential risks.
[0136] In this embodiment, the specific steps of step S31 are as follows:
[0137] Perform normal user access behavior recognition on the enterprise data flow mapping model to extract normal user access behavior data;
[0138] Conduct data access path analysis on the normal user access behavior data to extract multiple normal data access paths;
[0139] Based on multiple normal data access paths, perform transmission data volume statistics to obtain the transmission data volume of each access path;
[0140] Conduct multi-time point data transmission trend analysis on the transmission data volume of each access path to extract the data transmission trend characteristics of each access path;
[0141] Calculate the data access frequency of normal user access behavior data to obtain the normal user data access frequency;
[0142] Conduct data flow topology analysis based on the normal user data access frequency to generate normal user data flow topology data;
[0143] Fit the normal operation state of the enterprise system according to the data transmission trend characteristics of each access path and the normal user data flow topology data, thereby constructing a normal behavior baseline.
[0144] In this embodiment, user access logs are extracted from the enterprise information system, including user ID, access timestamp, access content (such as files, database tables, etc.), and access methods (such as direct access, API calls, etc.). It is recommended to collect data for at least the past 6 months to ensure the representativeness of the sample. The data is preprocessed to remove duplicate records and incomplete access logs, and standardize the time format and user ID to ensure data consistency. Determine key access behavior characteristics, such as access frequency, access time period, access object type, and user role, etc. These characteristics serve as the basis for the analysis of normal behavior. Conduct statistical analysis on each characteristic, calculate indicators such as mean and standard deviation for subsequent modeling. Select an appropriate modeling method, such as clustering analysis (K-means or DBSCAN) or statistical models (Gaussian mixture model), to construct a normal user access behavior model. Set model parameters, such as the number of clusters (initially recommended to be set to 5 - 10) and thresholds, to ensure that the model can effectively capture the normal access behavior of users. Use the selected modeling method to analyze the extracted behavior characteristics, generate normal user access behavior data, and record the characteristics and indicators of each behavior pattern. Generate a normal behavior report, detailing information such as the mean, standard deviation, and occurrence frequency of each access behavior. Set the identification criteria for access paths, such as a specific access order (such as from the database to the application server), or a specific access type (such as only focusing on file reads). According to the set criteria, extract access paths from the normal user access behavior data, and record the source node, target node, and access order of each access path. Organize the extracted paths into a structured data format (such as a graph structure or a table) for subsequent analysis. Adopt graph theory analysis methods or sequence pattern mining algorithms (such as GSP or PrefixSpan) to analyze the user's access paths and identify multiple representative normal data access paths. Set analysis parameters, such as support and confidence, to ensure that the extracted paths are meaningful. Record the characteristics of each extracted normal data access path, such as access frequency, path length, and participating users, etc. Use visualization tools (such as Graphviz or Cytoscape) to display the normal data access paths to help managers intuitively understand the data flow. Determine the statistical indicators of the transmitted data volume, such as the total data transmission volume (bytes), average transmission volume, and maximum transmission volume of each path, etc. Adopt a data aggregation method to calculate the transmitted data volume of each normal data access path by summarizing the size of each piece of data in the access log. Set a statistical time window, such as hourly, daily, or weekly statistics, to observe the changing trend of data transmission. Conduct statistical analysis on the transmitted data volume of each normal data access path, and record the total transmitted data volume, average transmission volume, and peak transmission volume of each path. Organize the statistical results into an easy-to-analyze format (such as a table or a chart) for subsequent analysis.Select appropriate trend analysis methods, such as time series analysis (ARIMA model) or moving average method, to analyze the changing trend of the transmission data volume of each access path over time. Set analysis parameters, such as the time window size (e.g., 7 days or 30 days) and prediction period, to ensure that the model can capture the changes in trends. Organize the transmission data volume of each access path into time series data according to the timestamps for trend analysis. Record the transmission data volume at each time point, including the date, time, and transmission data volume. Use the selected trend analysis method to analyze the time series data and extract the trend characteristics of each access path, such as upward, downward, or stable trends. Generate a trend characteristics report, which details the transmission trend characteristics of each access path, including the trend type, change amplitude, and duration, etc. Determine the metrics for calculating the access frequency, such as the access frequency of each user, the access frequency of each path, and the total access frequency, etc. Use simple statistical methods to calculate the access frequency of each user and the access frequency of each access path. Record the access frequency of each user and each path and generate a frequency statistics report, which details the access frequency of each access path. Use visualization tools to display the changing trend of the access frequency to help managers understand the access behavior of users. Determine the structure of the data flow topology, including nodes (users, data sources, data targets) and edges (data flow paths). Use graph theory methods, such as network analysis tools (e.g., Gephi), to construct a normalized user data flow topology and record the characteristics of each node and edge. Set analysis parameters, such as node weights (based on access frequency) and edge weights (based on data transmission volume), to ensure the accuracy of the topology structure. Generate a normalized user data flow topology, which records the characteristics of each node and edge, including access frequency, data transmission volume, and connection strength. Generate a topology analysis report, which details the topology structure and the characteristics of each node to help managers understand the overall structure of the data flow. Select an appropriate behavior baseline model, such as a statistical model (e.g., control chart) or a machine learning model (e.g., supervised learning), for fitting the normal transfer state. Integrate the features extracted from the data transmission trend analysis and the data flow topology analysis as the model input for model training. Set training parameters, such as the learning rate and training period, to ensure the accuracy and robustness of the model. Generate a normalized behavior baseline, which records the characteristics of each behavior pattern, including data transmission trends and flow topology characteristics. Generate a normalized behavior baseline report, which details the constructed baseline status and its characteristics.
[0145] In this embodiment, step S4 includes the following steps:
[0146] Step S41: Define the length of the intrusion traceback window, and perform multi-time window intrusion trajectory backtracking on the enterprise system abnormal intrusion behavior data according to the length of the intrusion traceback window, so as to obtain intrusion trajectories for multiple time periods;
[0147] Step S42: Reconstruct the full-cycle intrusion attack chain for the intrusion trajectories in multiple time periods to generate a full-cycle intrusion attack topology graph;
[0148] Step S43: Trace the origin of the initial attack point in the full-cycle intrusion attack topology graph to obtain the initial node of the intrusion attack;
[0149] Step S44: Deeply mine the abnormal attack features based on the initial node of the intrusion attack to obtain the intrusion attack features.
[0150] In this embodiment, the intrusion tracing window length is set according to the enterprise's security policy and data analysis of past attack behaviors. Select time periods such as 30 minutes, 1 hour or 2 hours, and make preliminary settings based on the time characteristics of historical intrusion behaviors. Conduct experiments to analyze the impact of different window lengths on intrusion trajectory identification and select the optimal window length. Calculate the intrusion event recognition rate under different window lengths through statistical analysis tools (such as Python's Pandas library). Collect abnormal intrusion behavior data of the enterprise system, including timestamp, user ID, access path, and abnormal type information to ensure that the data set is complete. Preprocess the data to remove irrelevant information and standardize the time format for subsequent analysis. Use sliding window technology to gradually advance and intercept intrusion behavior data according to the defined intrusion tracing window length to form intrusion trajectories for multiple time periods. Set the sliding step size, such as 1 minute or 5 minutes, to obtain a more fine-grained intrusion trajectory. Use the sliding window to reverse the intrusion trajectory in the abnormal intrusion behavior data and record the intrusion behavior in each time period, including the time when the intrusion event occurred, the participating users and behavior characteristics. Generate an intrusion trajectory data set, record the intrusion behavior and its characteristics in detail for each time period, and ensure the integrity and accuracy of the data. Select a commonly used attack chain model, such as the MITRE ATT&CK model or the Lockheed Martin kill chain model, as the framework for reconstructing the attack chain. Set model parameters, such as the attack phase (initial access, execution, persistence, etc.), to ensure that all aspects of the intrusion behavior are fully covered. Integrate intrusion trajectory data from multiple time periods to identify potential attack paths and attack events, and form a graph structure for analysis. Record the characteristics of each attack event, including information such as the attack type, target node, affected system and user. Use graph theory methods (such as Dijkstra algorithm or depth-first search) to analyze the intrusion trajectory, reconstruct the full-cycle intrusion attack chain, and record the relationship between each attack link. Generate a full-cycle intrusion attack topology diagram, which details each link of the attack chain, the attacker's behavior path and the target node. Generate an attack chain reconstruction report, record the reconstructed attack chain characteristics and topology in detail, and use visualization tools (such as Visio or Lucidchart) to display the attack topology diagram to help managers understand the attack process. Set traceability standards, such as the characteristics of the initial attack point (such as the first access time, the starting point of the attack path, etc.), to ensure that the source of the attack can be effectively identified. Combine historical data to define available attack feature indicators, such as the attacker's IP address, user account, and attack tool. Use the backtracking analysis method, combined with the structure of the attack chain, to identify the initial attack node in the full-cycle intrusion attack topology. Set analysis parameters, such as the tracing depth (usually set to level 1 or 2), to ensure that the initial attack point can be accurately captured. Use data mining technology to analyze the full-cycle intrusion attack topology, trace back to the initial attack source, and record the feature information of the initial attack node.Generate an initial attack point identification report, which details the characteristics of the attack source and its impact, to help managers formulate corresponding defense strategies. Select appropriate feature mining algorithms, such as decision trees, random forests, or deep learning (e.g., convolutional neural networks), for in-depth mining of attack features. Set mining parameters, such as the sample division ratio (e.g., 80% training set, 20% test set), to ensure the reliability and generalization ability of the model. Collect abnormal behavior data related to the initial attack node, including access paths, behavior characteristics, and timestamps, etc., to construct a feature data set. Preprocess the data, remove noise data, and perform standardization and normalization to ensure data quality. Use the selected mining algorithm to train the feature data set to generate an abnormal attack feature identification model, and record the performance metrics of the model, such as accuracy, recall rate, and F1-score. Evaluate the performance of the model through cross-validation methods to ensure that it can effectively identify abnormal attack features. Use the trained model to analyze the initial attack node, identify abnormal attack features, and record the type, frequency, and impact degree of the features. Generate an abnormal attack feature mining report, which details the identified features and their potential risks, to provide a basis for subsequent security protection and response measures.
[0151] In this embodiment, the specific steps of step S5 are as follows:
[0152] Step S51: Predict the attack diffusion based on the intrusion attack features to obtain the intrusion attack diffusion prediction path;
[0153] Step S52: Conduct multi-level penetration testing on the enterprise information system according to the intrusion attack diffusion prediction path to generate penetration test data of the enterprise information system;
[0154] Step S53: Mark the risk vulnerabilities of the penetration test data of the enterprise information system to obtain enterprise risk vulnerabilities at different levels;
[0155] Step S54: Perform adaptive vulnerability repair on the enterprise risk vulnerabilities at different levels to obtain multi-level vulnerability repair data.
[0156] In this embodiment, a suitable attack diffusion prediction model is selected, such as a graph-based propagation model (e.g., SIR model) or a machine learning model (e.g., random forest, support vector machine, etc.). Consider using a deep learning model (e.g., graph neural network) to handle complex network structures and attack characteristics. Collect data related to attacks, including attack paths, attacker behavior characteristics, affected nodes, and timestamps, etc., to ensure the comprehensiveness of the data. Conduct data cleaning and preprocessing, including removing redundant data, standardizing feature values, and extracting key attack features, such as attack type, attack frequency, length of attack path, etc., to form a feature set for model training. Use a correlation analysis method (e.g., Pearson correlation coefficient) to evaluate the relationship between features and attack diffusion, and screen out important features. Divide the prepared feature dataset into a training set and a test set (e.g., 80% for training and 20% for testing). Use cross-validation to optimize model parameters. Select appropriate performance evaluation metrics (e.g., accuracy, F1-score, etc.) to evaluate the performance of the model. Use the trained model to predict the attack diffusion path, generate the predicted path, record the diffusion probability and affected nodes of each path, and generate an attack diffusion prediction report, which details each predicted path and its features, providing a basis for subsequent penetration testing. Determine the scope of penetration testing, including the network architecture, applications, and databases of the enterprise information system, etc. Clearly define the test objectives. Identify key assets and potential vulnerabilities based on the attack diffusion prediction path. Develop a detailed test plan. Select suitable penetration testing tools, such as Metasploit, Burp Suite, and Nessus, etc., to ensure that the tools can cover the test requirements at all levels. Configure the tools according to the test objectives and environment to ensure that they have the latest vulnerability libraries and attack modules. Implement multi-level penetration testing according to the established plan, including the network layer, application layer, and database layer, etc. Record the test results of each step. Use automated tools for preliminary scanning to identify potential vulnerabilities and conduct in-depth verification in combination with manual testing. Record each vulnerability found during the penetration testing and its impact, including information such as vulnerability type, severity score, and exploitability, etc. Generate a penetration testing report, which details the results of each test session and the discovered risk vulnerabilities, providing a basis for subsequent analysis. Determine the criteria for vulnerability annotation, including the severity of the vulnerability (e.g., high, medium, low), scope of impact (e.g., single system, multiple systems), and difficulty of exploitation, etc. Consider using industry standards (e.g., CVSS scoring system) for vulnerability assessment and annotation. Analyze each vulnerability according to the penetration testing results, evaluate its risk level, and conduct annotation. Record the detailed information of each vulnerability, including vulnerability description, occurrence conditions, affected systems, and repair suggestions, etc. Organize the annotated vulnerability data into a structured format (e.g., database table) for easy subsequent query and management. Generate a risk vulnerability report, which details the vulnerabilities at different levels and their characteristics, to help managers conduct risk assessment and decision-making. Develop corresponding repair strategies according to the vulnerability type and risk level.For operations such as patch updates, configuration adjustments, or system isolation, consider using automated repair tools (such as Ansible, Chef, etc.) to improve repair efficiency. For vulnerabilities at different levels, implement corresponding repair measures, record the detailed steps and results of each repair process, conduct verification tests after the repair to ensure that the vulnerabilities have been successfully repaired and no new problems have been introduced, record the repair status of each vulnerability, including information such as repair time, repair measures, and verification results, form multi-level vulnerability repair data, generate a vulnerability repair report, and record the repair process and results in detail to provide a basis for subsequent audits and compliance checks.
[0157] In this embodiment, the specific steps of step S6 are as follows:
[0158] Step S61: Conduct hierarchical vulnerability repair monitoring on the multi-level vulnerability repair data to identify the unrepaired vulnerabilities at each level;
[0159] Step S62: Perform global vulnerability iterative repair on the enterprise information system based on the unrepaired vulnerabilities at each level, thereby generating global vulnerability iterative repair data;
[0160] Step S63: Conduct a comprehensive assessment of the dynamic risk situation based on the global vulnerability iterative repair data to obtain a comprehensive assessment report on the enterprise risk situation.
[0161] In this embodiment, the criteria for vulnerability monitoring are determined, including monitoring frequency (such as daily, weekly) and monitoring metrics (such as the number of unpatched vulnerabilities, repair rate, etc.). The monitoring metrics are set according to industry standards (such as NIST or OWASP) to ensure compliance with best practices. Multilevel vulnerability repair data is collected, including information on both patched and unpatched vulnerabilities, to ensure the integrity and accuracy of the data. The data is preprocessed to remove redundant information and ensure consistent data formats. Vulnerability management tools (such as Qualys, Nessus) are used for vulnerability monitoring. Monitoring parameters and scanning tasks are configured, and a monitoring period is set. For example, a full scan is performed once a week to identify unpatched vulnerabilities. The monitoring task is executed, and the unpatched vulnerabilities at each level are recorded. Their risk levels and impact scopes are evaluated, and a vulnerability monitoring report is generated, which details the number and characteristics of the unpatched vulnerabilities to assist managers in formulating repair plans. Based on the levels and numbers of unpatched vulnerabilities, a global vulnerability repair plan is formulated, including the vulnerabilities to be repaired first, repair strategies, and schedules. Repair priorities are set. For example, high-risk vulnerabilities are given priority. Appropriate repair tools and methods are selected, such as automated repair tools (such as Ansible, Puppet) or manual repair procedures, to ensure the efficiency and accuracy of the repair process. Repair strategies are set, such as step-by-step repair or batch repair, to ensure that the stability of the system is not affected. According to the formulated repair plan, global vulnerability iterative repair is implemented, and the detailed steps and results of each repair process are recorded. Verification tests are performed on each unpatched vulnerability to ensure that the vulnerability has been successfully repaired and no new problems have been introduced. The repair status of each vulnerability is recorded, including information such as repair time, repair measures, and verification results, to form global vulnerability iterative repair data. A global vulnerability repair report is generated, which details the repair process and results, providing a basis for subsequent audits and compliance checks. An appropriate risk assessment model is selected, such as a quantitative risk assessment model (such as the FAIR model) or a qualitative risk assessment model, to comprehensively evaluate the risk situation of the enterprise information system. Evaluation metrics are set, such as the number of vulnerabilities, repair rate, potential impact, and attack probability, etc. Global vulnerability iterative repair data and other relevant data (such as system configuration, user behavior, and network traffic, etc.) are collected and integrated into a comprehensive evaluation dataset. The data is preprocessed to ensure its consistency and integrity. The selected evaluation model is used to analyze the comprehensive evaluation data to evaluate the risk situation of the enterprise information system, obtain the risk level and impact scope, and generate a risk assessment report, which details the evaluation results, including the risk level, main risk factors, and improvement suggestions.
[0162] In this embodiment, an enterprise data security capability evaluation system is provided for performing the enterprise data security capability evaluation method as described above, including:
[0163] A data stream analysis module, which is used to obtain enterprise historical security logs based on an enterprise information system, perform real-time data stream analysis and dynamic location of storage locations on the enterprise historical security logs, and construct an enterprise data flow mapping model;
[0164] An intrusion attack simulation module, which is used to perform abnormal intrusion attack simulation on the enterprise data flow mapping model, conduct enterprise security risk assessment, and mark potential risk nodes;
[0165] A behavior deviation detection module, which is used to perform historical normalization behavior analysis on the enterprise data flow mapping model and conduct access behavior deviation detection on potential risk nodes to obtain enterprise system abnormal intrusion behavior data;
[0166] An intrusion trajectory reverse deduction module, which is used to perform multi-time window intrusion trajectory reverse deduction on the enterprise system abnormal intrusion behavior data and conduct in-depth mining of abnormal attack characteristics to obtain intrusion attack characteristics;
[0167] An adaptive vulnerability repair module, which is used to perform multi-level penetration testing on the enterprise information system based on the intrusion attack characteristics and conduct adaptive vulnerability repair to obtain multi-level vulnerability repair data;
[0168] A comprehensive evaluation module, which is used to conduct hierarchical vulnerability repair monitoring on the multi-level vulnerability repair data and conduct comprehensive evaluation of the dynamic risk situation to obtain a comprehensive evaluation report on the enterprise risk situation,
[0169] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention,
[0170] As described above are only the specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An enterprise data security capability evaluation method, characterized in that, It includes the following steps: Step S1: Obtain the enterprise historical security logs based on the enterprise information system, perform real-time data stream parsing and dynamic location of the storage location on the enterprise historical security logs, and construct an enterprise data flow mapping model; Step S2: Conduct abnormal intrusion attack simulation on the enterprise data flow mapping model, perform enterprise security risk assessment, and mark potential risk nodes; Step S3: Conduct historical normalization behavior analysis on the enterprise data flow mapping model, and perform access behavior deviation detection on potential risk nodes to obtain enterprise system abnormal intrusion behavior data; Step S4: Perform multi-time window intrusion trajectory backtracking on the enterprise system abnormal intrusion behavior data, and conduct in-depth mining of abnormal attack characteristics to obtain intrusion attack characteristics; the specific process of the multi-time window intrusion trajectory backtracking is as follows: set the intrusion traceback window length, analyze the influence of different window lengths on intrusion trajectory recognition, select the optimal window length, calculate the intrusion event recognition rate under different window lengths, adopt the sliding window technology, and gradually advance and intercept the intrusion behavior data according to the defined intrusion traceback window length to form intrusion trajectories in multiple time periods; Step S5: Perform multi-level penetration testing on the enterprise information system based on the intrusion attack characteristics, and conduct adaptive vulnerability repair to obtain multi-level vulnerability repair data; Step S6: Conduct hierarchical vulnerability repair monitoring on the multi-level vulnerability repair data, and perform comprehensive dynamic risk situation assessment to obtain a comprehensive enterprise risk situation assessment report; Among them, the specific steps of Step S2 are as follows: Step S21: Conduct multi-dimensional behavior scenario semantic analysis on the enterprise historical security logs to obtain various behavior scenarios of the enterprise system; Step S22: Fit multiple attack scenarios based on various behavior scenarios of the enterprise system to obtain multiple enterprise system attack scenarios; Step S23: Conduct abnormal intrusion attack simulation on the enterprise data flow mapping model according to multiple enterprise system attack scenarios to generate multi-scenario attack simulation data; Step S24: Perform enterprise security risk assessment on the multi-scenario attack simulation data and mark potential risk nodes; Among them, the specific steps of Step S24 are as follows: Identify the attack data flow nodes in the multi-scenario attack simulation data to obtain multiple attack data flow nodes; Analyze the change in user access rights for multiple attack data flow nodes to obtain the access right change data for each node; Conduct abnormal data transmission detection on multiple attack data flow nodes and extract the abnormal data transmission characteristics of the nodes; Conduct intrusion prevention response analysis for each node based on the access right change data of each node and the abnormal data transmission characteristics of the node to generate intrusion prevention response data for each node; Calculate the defense failure time for the intrusion prevention response data of each node and extract the intrusion prevention failure time for each node; Quantify the intrusion interception for each node in the multi-scenario attack simulation data to obtain the intrusion interception value for each node; Conduct per-node security risk assessment based on the intrusion prevention failure time of each node and the intrusion interception value of each node to obtain the security risk assessment value for each node; Based on the preset security assessment thresholds for multi-scenario attacks, identify risk weak nodes by the security risk assessment values of each node, and mark potential risk nodes.
2. The enterprise data security capability evaluation method according to claim 1, characterized in that The specific steps of step S1 are as follows: Step S11: Obtain the enterprise historical security logs based on the enterprise information system, and filter the abnormal outlier data in the enterprise historical security logs to obtain outlier-filtered security logs; Step S12: Parse the real-time data stream of the outlier-filtered security logs to generate the enterprise real-time data stream; Step S13: Analyze the dynamic behavior of users in the enterprise real-time data stream to generate user dynamic behavior data; Step S14: Identify the data flow paths according to the user dynamic behavior data, and extract multiple data flow paths; Step S15: Dynamically locate the storage locations of the enterprise real-time data stream to obtain the dynamic storage location of each data stream; Step S16: Based on the dynamic storage location of each data stream, perform dynamic data flow mapping on multiple data flow paths to construct an enterprise data flow mapping model.
3. The enterprise data security capability evaluation method according to claim 1, wherein The specific steps of step S3 are as follows: Step S31: Analyze the historical normal behavior of the enterprise data flow mapping model to construct a normal behavior baseline; Step S32: Monitor the real-time access behavior of potential risk nodes to obtain the real-time access behavior data of risk nodes; Step S33: Detect access behavior deviations of the real-time access behavior data of risk nodes based on the normal behavior baseline, and mark the access behavior data with each baseline deviation; Step S34: Identify abnormal intrusion behavior for each access behavior data with a baseline deviation to obtain enterprise system abnormal intrusion behavior data.
4. The enterprise data security capability evaluation method according to claim 3, wherein The specific steps of step S31 are as follows: Identify normal user access behavior in the enterprise data flow mapping model, and extract normal user access behavior data; Analyze the data access paths of the normal user access behavior data, and extract multiple normal data access paths; Based on multiple normal data access paths, perform statistics on the amount of transmitted data to obtain the amount of transmitted data for each access path; Perform multi-time point data transmission trend analysis on the amount of transmitted data for each access path, and extract the data transmission trend characteristics of each access path; Calculate the data access frequency of normal user access behavior data to obtain the normal user data access frequency; Perform data flow topology analysis according to the normal user data access frequency to generate normal user data flow topology data; Fit the normal operation state of the enterprise system according to the data transmission trend characteristics of each access path and the normal user data flow topology data to construct a normal behavior baseline.
5. The enterprise data security capability evaluation method according to claim 1, wherein The specific steps of step S4 are as follows: Step S41: Define the length of the intrusion traceback window, and perform multi-time window intrusion trajectory backtracking on the enterprise system abnormal intrusion behavior data according to the length of the intrusion traceback window to obtain intrusion trajectories for multiple time periods; Step S42: Reconstruct the full-cycle intrusion attack chain for the intrusion trajectories of multiple time periods to generate a full-cycle intrusion attack topology graph; Step S43: Trace the initial attack point of the full-cycle intrusion attack topology graph to obtain the initial node of the intrusion attack. Step S44: Deeply mine abnormal attack features based on the initial intrusion attack node to obtain intrusion attack features.
6. The enterprise data security capability evaluation method according to claim 1, wherein The specific steps of Step S5 are as follows: Step S51: Predict attack diffusion based on intrusion attack features to obtain an intrusion attack diffusion prediction path; Step S52: Conduct multi-level penetration testing on the enterprise information system according to the intrusion attack diffusion prediction path to generate penetration test data of the enterprise information system; Step S53: Mark risk vulnerabilities for the penetration test data of the enterprise information system to obtain enterprise risk vulnerabilities at different levels; Step S54: Perform adaptive vulnerability repair on enterprise risk vulnerabilities at different levels to obtain multi-level vulnerability repair data.
7. The enterprise data security capability evaluation method according to claim 1, wherein The specific steps of Step S6 are as follows: Step S61: Monitor vulnerability repair level by level for the multi-level vulnerability repair data, and identify unrepaired vulnerabilities at each level; Step S62: Perform global vulnerability iterative repair on the enterprise information system according to the unrepaired vulnerabilities at each level to generate global vulnerability iterative repair data; Step S63: Conduct a comprehensive assessment of the dynamic risk situation based on the global vulnerability iterative repair data to obtain a comprehensive enterprise risk situation assessment report.
8. An enterprise data security capability evaluation system, characterized in that For implementing the enterprise data security capability assessment method as described in claim 1, it includes: A data flow analysis module, configured to obtain enterprise historical security logs based on the enterprise information system, perform real-time data flow analysis and dynamic location of storage locations on the enterprise historical security logs, and construct an enterprise data flow mapping model; An intrusion attack simulation module, configured to perform abnormal intrusion attack simulation on the enterprise data flow mapping model, conduct enterprise security risk assessment, and mark potential risk nodes; A behavior deviation detection module, configured to perform historical normal behavior analysis on the enterprise data flow mapping model and conduct access behavior deviation detection on potential risk nodes to obtain enterprise system abnormal intrusion behavior data; An intrusion trajectory reverse inference module, configured to perform multi-time window intrusion trajectory reverse inference on the enterprise system abnormal intrusion behavior data and conduct in-depth mining of abnormal attack features to obtain intrusion attack features; An adaptive vulnerability repair module, configured to perform multi-level penetration testing on the enterprise information system based on intrusion attack features and conduct adaptive vulnerability repair to obtain multi-level vulnerability repair data; A comprehensive assessment module, configured to monitor vulnerability repair level by level for the multi-level vulnerability repair data and conduct a comprehensive assessment of the dynamic risk situation to obtain a comprehensive enterprise risk situation assessment report.
Citation Information
Patent Citations
Enterprise platform safety management method and system based on artificial intelligence
CN118710224A
Enterprise network security test and evaluation method and system
CN119051990A
Defense system for complex collaborative attack in power distribution network
CN119324807A