Big data security risk assessment method and system
By dynamically acquiring big data operation scenarios and business data flows, and combining machine learning models and system environment information, risk assessment parameters are constructed, solving the problems of flexibility and accuracy in big data security risk assessment in existing technologies, and realizing efficient risk identification and assessment in complex environments.
Patent Information
- Application Number
- CN202511289116.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2026-01-27
AI Technical Summary
Existing big data security risk assessment methods lack flexibility and accuracy when facing complex and ever-changing distributed storage and cloud computing environments. They are unable to dynamically model risk propagation characteristics, resulting in incomplete risk identification, delayed assessment, and misjudgment.
By acquiring big data operation scenario types, collecting business data streams and system operation logs in real time or periodically, dynamically selecting risk modeling modes that match the scenarios, and combining real-time business flow logs and historical security event data, risk assessment parameters are constructed. Machine learning models are used to extract threat factor weights and attack link information, calculate multi-dimensional risk intensity, and conduct risk assessment in conjunction with system operation location and network environment information.
It achieves highly accurate and flexible risk assessment of big data systems, can accurately depict complex attack propagation processes, improves the comprehensiveness and adaptability of risk identification, and enhances the reliability and practicality of risk assessment.
Smart Images

Figure CN121412985A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data processing technology, and in particular to a method and system for big data security risk assessment. Background Technology
[0002] In recent years, with the rapid development of big data technology, data has become a crucial asset driving various businesses and social operations. However, the accompanying security risks have also become increasingly prominent. Big data systems typically operate in complex and ever-changing environments, potentially distributed across geographically dispersed data centers or shared by multiple tenants on cloud computing platforms. These operating environments make data transmission links, computing nodes, and storage modules susceptible to attack threats. Most existing big data security risk assessment methods rely on static rule bases or single threat detection methods. For example, some methods assess threats based solely on intrusion detection logs or known attack characteristics, making it difficult to address constantly evolving attack patterns; while some methods incorporate statistical modeling, they often lack the ability to dynamically model the risk propagation characteristics under different business scenarios. Because attack behaviors are characterized by their interconnectedness, complexity, and stealth, single static rule detection can easily lead to incomplete risk identification, delayed assessment, or even misjudgment.
[0003] Furthermore, existing methods often lack flexibility in modeling risks across different scenarios. For example, attack paths, node topologies, and risk propagation patterns differ significantly between distributed storage and cloud computing environments, but traditional methods often use a uniform modeling approach, ignoring the impact of environmental differences on risk assessment results, thus leading to insufficient accuracy in the assessment outcomes.
[0004] Furthermore, while some risk assessment technologies attempt to incorporate historical security incident data, they fail to effectively combine contextual information such as the system's operating location, time, and network environment, making it difficult to comprehensively depict the distribution of potential attack sources and the dynamic intensity of attacks, resulting in insufficient risk prediction capabilities. Summary of the Invention
[0005] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0006] In view of the problems existing in the prior art, the present invention is proposed.
[0007] The system acquires the types of big data operation scenarios that are in operation and collects the corresponding business data streams and system operation logs in real time or periodically through a data acquisition module. The scenario types include at least distributed storage environments and cloud computing environments. The business data streams may include real-time business flow logs and / or historical security event data.
[0008] Determine the corresponding risk modeling mode based on the type of operational scenario; Based on the business data flow and the risk modeling pattern, risk assessment parameters for subsequent analysis are extracted and determined. The original risk indicators are then comprehensively processed and analyzed using the aforementioned risk assessment parameters to generate and output safety risk assessment results.
[0009] As a preferred embodiment of the big data security risk assessment method described in this invention, the process of determining the risk modeling mode includes: When the operating scenario is a distributed storage environment, the risk modeling mode adopted is the first risk modeling mode corresponding to the distributed storage environment; When the operating scenario is a cloud computing environment, the risk modeling mode adopted is the second risk modeling mode corresponding to the cloud computing environment.
[0010] As a preferred embodiment of the big data security risk assessment method of the present invention, wherein: when the risk modeling mode is a first risk modeling mode, the business data stream is a real-time business flow log; the process of determining the risk assessment parameters through the real-time business flow log and the risk modeling mode includes: Real-time business flow logs are input into the risk identification model to extract threat factor weights and attack chain information. The threat factor weights are normalized to obtain the risk impact coefficient. The risk impact coefficient and the attack chain information are used together as risk assessment parameters.
[0011] As a preferred embodiment of the big data security risk assessment method of the present invention, the process of comprehensively analyzing the original risk indicators through risk assessment parameters and generating output results includes: Risk modeling information for constructing original risk indicators based on the attack chain information, wherein the risk modeling information includes an attack node feature matrix and a risk propagation matrix; The attack node feature matrix is input into the risk propagation function to calculate the risk propagation tensor; By performing matrix operations on the risk impact coefficient and the risk propagation tensor matrix, the risk intensity of the attack node in multiple dimensions can be obtained. Then, by combining the risk intensity of each dimension with the risk propagation matrix, the final safety risk assessment result is generated and output.
[0012] As a preferred embodiment of the big data security risk assessment method described in this invention, the method further includes the following steps: Construct big data operation models corresponding to various attack scenarios; For each of the aforementioned operating models, different attack parameters are set and configured, and attack effect logs generated under those attack parameters are obtained respectively; A training set is formed based on the attack effect logs of all running models, the attack parameters corresponding to each log, and the corresponding scenario information. The preset lightweight network is then trained using the training set to obtain the trained risk identification model.
[0013] As a preferred embodiment of the big data security risk assessment method of the present invention, wherein: when the risk modeling mode is the second risk modeling mode, the business data stream is historical security event data; the process of determining risk assessment parameters through historical security event data and risk modeling mode includes: Collect current system operating location data, current time information, and current network environment information; Based on the system's operating location data and time information, the distribution of potential attack sources is calculated; The attack intensity is determined based on the distribution of potential attack sources. The attack intensity, the distribution of potential attack sources, and the network environment information are combined as risk assessment parameters.
[0014] As a preferred embodiment of the big data security risk assessment method described in this invention, the process of comprehensively analyzing the original risk indicators through risk assessment parameters and outputting the results includes: Based on the distribution of potential attack sources and network environment information, risk action information of the original risk indicators is generated. The attack intensity is then used to perform a comprehensive analysis of the original risk indicators, and the final security risk assessment result is output.
[0015] An assessment system applied to the aforementioned big data security risk assessment method includes: The data acquisition module collects the types of big data operation scenarios in operation, as well as the corresponding business data streams and system operation logs; The scene recognition and modeling module selects or determines the corresponding risk modeling mode based on the collected types of operating scenarios. The risk parameter extraction module extracts risk assessment parameters based on business data flow and risk modeling patterns; The risk modeling and analysis module uses risk assessment parameters to conduct a comprehensive analysis of the original risk indicators; The results generation and output module generates and outputs security risk assessment results for subsequent early warning, display, or decision support.
[0016] The present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described big data security risk assessment method.
[0017] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described big data security risk assessment method.
[0018] The beneficial effects of this invention are: 1. This invention provides a big data security risk assessment method that effectively addresses the problems of insufficient realism and limited adaptability in existing methods. By acquiring the types of big data operation scenarios and combining real-time business flow logs and historical security event data, it dynamically selects a risk modeling mode that matches the scenario, thereby ensuring the pertinence and flexibility of risk modeling.
[0019] 2. The method of this invention can construct risk modeling information based on attack chain information, and introduces attack node feature matrices, risk propagation matrices, and risk transmission tensors. Combined with risk impact coefficients, it performs multi-dimensional risk intensity calculations, achieving an accurate characterization of complex attack propagation processes. Furthermore, by constructing multi-scenario, multi-parameter operating models and generating training sets based on attack effect logs, a lightweight risk identification model can be trained, significantly improving the model's ability to identify and generalize novel attack patterns.
[0020] 3. This invention introduces system operating location data, time information, and network environment information as risk assessment parameters, which can calculate the distribution of potential attack sources and attack intensity, and realize dynamic prediction and adaptive assessment of attack risks.
[0021] 4. In summary, this invention enables risk assessment results to better reflect the actual operating state of big data systems, improving both the comprehensiveness and accuracy of risk identification and enhancing adaptability to complex and ever-changing attack environments, ultimately achieving the technical effect of improving the reliability and practicality of big data security risk assessment. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a schematic diagram of the overall process of a big data security risk assessment method proposed in this invention; Figure 2 This is a schematic diagram of the overall process of a big data security risk assessment system proposed in this invention. Detailed Implementation
[0023] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0024] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0025] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0026] Reference Figures 1-2 As an embodiment of the present invention, a big data security risk assessment method is provided, which includes the following steps: Step 1: Obtain the big data operation scenario type that is in operation, and collect the corresponding business data stream and system operation logs in real time or periodically through the data acquisition module.
[0027] The operational scenario types include distributed storage environments and cloud computing environments. The operational scenario type can serve as a classification identifier for the system's operational environment, distinguishing between architectures centered on multi-node redundant storage and those centered on elastic computing and resource virtualization. The operational scenario type can be determined by acquiring current system operational data through data acquisition modules, log pattern recognition, task distribution, network topology analysis, or by analyzing collected data using artificial intelligence models. The operational scenario can also be determined through location information, such as linking system node geographic distribution information with a resource configuration database. Furthermore, the operational scenario can be determined through inter-node temporal consistency analysis. For example, distributed storage environments can include object storage systems and distributed databases, while cloud computing environments can include IaaS services and PaaS platforms.
[0028] Business data flows can be a specific collection of information describing operational characteristics. Business data flows include real-time business flow logs and / or historical security event data. For example, real-time business flow logs may include task execution records, access request chains, and their latency information collected from the system, such as transaction logs, call chain logs, and performance monitoring logs; historical security event data may include attack event records and abnormal behavior data collected through intrusion detection systems, firewalls, or security information and event management systems. Furthermore, historical security event data may be attack sample data at the user session level, such as malicious requests and abnormal traffic, and may also include abnormal connection information on the network side, such as port scanning and traffic flooding.
[0029] Step 2: Determine the corresponding risk modeling mode based on the type of operating scenario.
[0030] The risk modeling mode can be a preset risk parameter processing algorithm used to generate corresponding risk assessment parameters based on data under specific operating scenarios. For example, a first risk modeling mode and a second risk modeling mode can be used respectively to process distributed storage environments and cloud computing environments. It is understood that operating scenarios, in addition to distributed storage environments and cloud computing environments, can also include sub-types of distributed storage environments and sub-types of cloud computing environments; correspondingly, a corresponding risk modeling mode can be used for each sub-type.
[0031] Step 3: Based on business data flow and risk modeling patterns, extract and determine risk assessment parameters for subsequent analysis.
[0032] The risk assessment parameters can be a set of parameters that affect the results of security risk analysis. For example, they can include one or more security attributes such as attack strength and risk propagation rate. In this embodiment, the risk assessment parameters can be obtained by analyzing data under a risk modeling mode.
[0033] Specifically, when the data includes real-time business flow logs, these logs can be input into a risk identification model to determine threat factor weights and attack chain information. The risk impact coefficient can then be obtained by normalizing the threat factor weights, and this coefficient, along with the attack chain information, can be used as risk assessment parameters. When the data includes historical security event data, the distribution of potential attack sources can be calculated based on system location data and time information. Furthermore, the attack intensity can be determined based on the distribution of potential attack sources, thus comprehensively determining the risk assessment parameters.
[0034] In one exemplary embodiment, the risk modeling mode can be based on data to determine the distribution and intensity of attack sources in the current operating environment, thereby obtaining risk assessment parameters for the current operating environment based on the distribution and intensity of attack sources.
[0035] Step 4: Perform comprehensive processing and analysis on the original risk indicators through risk assessment parameters to generate and output the safety risk assessment results.
[0036] Specifically, the original risk indicators can be the initial risk content before it has undergone risk parameter processing. For example, these can be indicators presented in a multi-dimensional form, such as user behavior characteristics, system access records, and network traffic statistics. Furthermore, when the original risk indicator is a single-dimensional indicator, before conducting a comprehensive analysis using risk assessment parameters, this single-dimensional indicator can be mapped into a multi-dimensional risk feature vector before being processed and analyzed comprehensively using risk assessment parameters.
[0037] The original risk indicators can be obtained according to different operating scenarios. That is, under different operating scenarios, the original risk indicators corresponding to the operating scenario can be used to generate the final security risk assessment results.
[0038] In one exemplary embodiment, a comprehensive analysis of the original risk indicators can be performed by adjusting the influence factors in the risk propagation function according to the risk assessment parameters, and obtaining the safety risk assessment results through matrix calculation.
[0039] The results of the security risk assessment can be displayed on a risk monitoring platform. Furthermore, the results can be correlated with real-time collected business data streams for display; for example, they can be output to a security operations center, a situational awareness platform, or a distributed monitoring terminal.
[0040] In summary, this method can adapt the generation process of security risk assessment results to the operating scenario, resulting in an assessment effect that is highly consistent with the actual risk status of the system, thereby improving the accuracy and practicality of risk assessment.
[0041] In one embodiment, the process of determining the risk modeling pattern includes: When the operating scenario is a distributed storage environment, the risk modeling mode adopted is the first risk modeling mode corresponding to the distributed storage environment; When the operating scenario is a cloud computing environment, the risk modeling mode adopted is the second risk modeling mode corresponding to the cloud computing environment.
[0042] The first risk modeling mode corresponds to the operating scenario of a distributed storage environment and is used to generate risk assessment parameters for such environments. It is understandable that when the operating scenario is a distributed storage environment, the potential risks often stem from redundant access by multiple nodes and data consistency issues. Therefore, the first risk modeling mode can determine the corresponding attack chain information and node threat factor weights for the distributed storage environment to obtain risk assessment parameters.
[0043] The first risk modeling approach can be to analyze real-time business flow logs to determine the threat factor weights and attack chain information in the operational scenario. In one exemplary embodiment, the real-time business flow logs can be global or local logs. Statistical distribution analysis and anomaly feature extraction can be performed on the logs to obtain the location of the attack chain in the distributed storage environment. For example, nodes with significantly higher response latency than the normal baseline in the call chain can be identified as potential risk nodes. Alternatively, session profile recognition and correlation feature enhancement processing can be performed on the logs to determine the location of risk chains based on the distribution of abnormal traffic. In another exemplary embodiment, the real-time business flow logs can be input into a pre-trained risk identification model for identification to obtain risk assessment parameters.
[0044] The second risk modeling mode corresponds to the operational scenario of a cloud computing environment and is used to generate risk assessment parameters for that environment. Accordingly, when the operational scenario is a cloud computing environment, the main risk sources in this environment can usually be considered as external attack traffic or resource contention. The distribution of potential attack sources can be calculated using system operation location data and time information, thereby determining the risk assessment parameters.
[0045] Furthermore, the second risk modeling mode can also obtain network environment information based on system operating location data and time information to further adjust the risk assessment parameters. For example, under high load conditions, the risk assessment parameters can exhibit high risk intensity and a rapid propagation trend, while under low load conditions, the risk assessment parameters can exhibit low risk intensity and a slow propagation trend.
[0046] In summary, the big data security risk assessment method provided in this embodiment obtains the corresponding risk modeling mode according to the operating scenario, thereby adapting to the differences between distributed storage environment and cloud computing environment, realizing the accurate determination of risk assessment parameters in different scenarios, reducing misjudgments caused by mismatched risk modeling, and improving the accuracy and practicality of security risk assessment.
[0047] In one embodiment, when the risk modeling mode is the first risk modeling mode, the business data stream is a real-time business flow log; the process of determining risk assessment parameters through the real-time business flow log and the risk modeling mode includes: Real-time business flow logs are input into the risk identification model to extract threat factor weights and attack chain information. The threat factor weights are normalized to obtain the risk impact coefficient. The risk impact coefficient and the attack chain information are used together as risk assessment parameters.
[0048] Specifically, the real-time business flow log in this embodiment can be two-dimensional or multi-dimensional business data obtained through the system acquisition module. The log can contain information such as the task execution chain, access frequency distribution, and abnormal behavior characteristics.
[0049] Risk identification models can be tools built based on convolutional neural networks, graph neural networks, or other machine learning models. They can be trained to learn the mapping relationship between log features and risk factors under different distributed operating scenarios, thus obtaining a pre-trained risk identification model. Furthermore, the risk identification model can take real-time business flow logs, or pre-processed log data, and input them into the model's input layer, then sequentially pass them through convolutional layers to extract local features, fully connected layers, or attention mechanisms to aggregate global information before outputting the result. For example, the risk identification model can also include a hybrid architecture combining anomaly detection branches and classification layers.
[0050] Threat factor weights can be a set of numerical parameters that quantify potential risk characteristics, including parameters such as risk intensity and attack chain probability. In this embodiment, they can be derived by analyzing and deducing the features in the logs through a risk identification model.
[0051] Attack chain information can be parameters describing the system's operational status. For example, it can include semantic tags such as task chain type and call node distribution, and can be output through the model's classification layer or feature segmentation branch. For instance, attack chain information can include chain type tags, such as "database access chain" or "cached call chain."
[0052] Normalizing threat factor weights can be achieved by transforming the original weight parameters in terms of data format. It is understood that during model training, uneven parameter distribution or excessively large numerical ranges can lead to difficulties in model convergence and insufficient generalization. In this embodiment, normalizing the threat factor weights converts the weight values into relatively balanced standardized parameters, thereby improving the stability of model training and inference. In some exemplary embodiments, this can be achieved using min-max normalization or logarithmic scaling.
[0053] This embodiment provides a big data security risk assessment method that extracts threat factor weights and attack chain information by inputting real-time business flow logs into a risk identification model. It utilizes machine learning models to improve the accuracy of risk assessment parameter extraction in complex distributed operating environments. This allows for the simultaneous consideration of risk propagation and operational status characteristics during the risk assessment process, enabling dynamic generation and environmental adaptation of risk assessment parameters. This results in higher accuracy in risk matching and attack chain characterization in distributed scenarios, improving the authenticity and effectiveness of security risk assessment.
[0054] In one embodiment, the process of comprehensively analyzing the original risk indicators using risk assessment parameters and generating output results includes: Risk modeling information is constructed based on attack chain information to build original risk indicators. The risk modeling information includes an attack node feature matrix and a risk propagation matrix. The attack node feature matrix is input into the risk propagation function to calculate the risk propagation tensor; By performing matrix operations on the risk impact coefficient and the risk propagation tensor matrix, the risk intensity of the attack node in multiple dimensions can be obtained. Then, by combining the risk intensity of each dimension with the risk propagation matrix, the final safety risk assessment result is generated and output.
[0055] Risk modeling information can be a set of parameters describing the propagation characteristics of attack behavior. This information includes an attack node feature matrix and a risk propagation matrix. The attack node feature matrix can be a vectorized representation of the attributes of attack nodes in a multi-dimensional space, such as node access frequency, vulnerability level, and service dependency. The risk propagation matrix can be a two-dimensional tensor used to characterize the propagation strength and directionality between different nodes in the attack path, describing the diffusion pattern of risk among different nodes.
[0056] In this embodiment, the risk modeling information generated from the attack chain information can be adaptively modeled based on the attack chain type, adapting the original risk indicators. For example, when the scenario corresponding to the attack chain is a database access chain, the attack node feature matrix can include features such as database query frequency and sensitive table access ratio, and data flow dependencies can be introduced into the risk propagation matrix to form risk modeling information that matches the scenario.
[0057] The risk propagation tensor can be a multidimensional structure generated by mapping the feature matrix of attack nodes through a risk propagation function. This tensor can transform the multidimensional attributes of each attack node into a measurable set of risk propagation coefficients. By substituting the feature matrix of the attack node into the risk propagation function, the static features of the node can be mapped to the dynamic risk propagation space, thereby abstracting the risk transmission process of the attack chain.
[0058] The risk impact coefficient can be a normalized set of weights, including parameters such as node risk intensity and propagation sensitivity. Multiplying the risk impact coefficient by the risk propagation tensor matrix allows us to use the risk impact coefficient as weights to linearly combine the various dimensions of the risk propagation tensor, ultimately calculating the risk intensity of each attack node in different dimensions.
[0059] Each element in the risk propagation matrix can represent the probability of risk transmission between different nodes. By multiplying the risk intensity of each dimension with the risk propagation matrix, the calculated risk intensity can be diffused and superimposed along the propagation path. For example, the risk intensity of node A (0.2) is multiplied by the propagation probability of A→B (0.4) in the propagation matrix to obtain the risk increment intensity of node B, thus realizing the simulation of iterative transmission of risk in the link.
[0060] Furthermore, when outputting security risk assessment results, time series modeling and network topology constraints can be combined to ensure the consistency between risk propagation results and actual attack evolution processes, thereby further improving the authenticity and credibility of risk assessment results in a distributed operating environment.
[0061] This embodiment provides a big data security risk assessment method that generates risk modeling information, including an attack node feature matrix and a risk propagation matrix, based on attack chain information. The attack node feature matrix is converted into a risk propagation tensor to achieve efficient risk propagation modeling. The risk intensity of each dimension is calculated by combining the risk impact coefficient and the risk propagation tensor. Finally, the security risk assessment result is generated and output through iterative calculation of the risk intensity and the risk propagation matrix. This method can further improve the authenticity and effectiveness of risk assessment in characterizing risk propagation paths and attack nodes.
[0062] In one embodiment, the method further includes: constructing big data operation models corresponding to various different attack scenarios; For each operating model, different attack parameters are set and configured, and attack effect logs generated under those attack parameters are obtained respectively; A training set is formed based on the attack effect logs of all running models, the attack parameters corresponding to each log, and the corresponding scenario information. The pre-set lightweight network is then trained using the training set to obtain the trained risk identification model.
[0063] In this embodiment, the runtime model can be a big data runtime environment model constructed using network security simulation technology, used to simulate the service structure and data flow relationships in a real distributed system. The runtime model can include a database runtime model, a distributed computing model, an edge computing model, etc. Furthermore, different runtime models can be generated for different attack scenarios. For example, a database runtime model can be established under the attack scenario of "SQL injection"; a distributed computing runtime model can be established under the attack scenario of "distributed denial-of-service attack". By constructing multiple runtime models with different parameter configurations under the same attack scenario, the training data can be further enriched, and the generalization ability and accuracy of the model can be improved.
[0064] For each operating model, different attack parameters are configured. This can be done within the operating model by setting at least one attack condition according to preset rules. The attack parameters can include, but are not limited to, attack intensity, attack type, attack frequency, and number of attack sources.
[0065] Obtaining attack effect logs corresponding to different attack parameters can be achieved by generating separate attack effect logs for each running model under each attack parameter. Furthermore, since attack events are not fixed in a single network path during actual operation, multiple effect logs can be generated based on different network topologies under the same attack parameters, thereby improving the accuracy of the risk identification model.
[0066] A training set is constructed based on the attack effect logs corresponding to all attack parameters of all running models, the attack parameters corresponding to each log, and the scene information corresponding to each log. This can be achieved by using the attack effect logs as input features and concatenating the attack parameters and scene information as output features, thus forming a set of training data. Multiple training data are then combined to construct the training set.
[0067] A training set is constructed based on the attack effect logs corresponding to all attack parameters of all running models, the attack parameters corresponding to each log, and the scene information corresponding to each log. This can be achieved by using the attack effect logs as input features and concatenating the attack parameters and scene information as output features, thus forming a set of training data. Multiple training data are then combined to construct the training set.
[0068] Furthermore, calculating the loss function based on the predicted output can involve comparing the predicted output with real metrics from the attack effect logs, calculating the differences, and updating the network parameters. For example, metrics such as attack duration, system performance degradation, and packet loss rate can be extracted from the logs as real labels, and the loss function can be constructed by comparing them with the predicted output. By incorporating real log data from various attack scenarios during training, more accurate loss calculation results can be obtained, thereby effectively improving the accuracy and robustness of the post-trained risk identification model.
[0069] This embodiment provides a big data security risk assessment method by constructing multiple operational models corresponding to different attack scenarios; configuring different attack parameters for each operational model and obtaining attack effect logs corresponding to different attack parameters; constructing a training set based on the attack effect logs corresponding to all attack parameters of all operational models, the attack parameters corresponding to each log, and the scenario information corresponding to each log; training a preset lightweight network using the training set to obtain a trained risk identification model. This method can effectively construct the dataset required for training the risk identification model. Furthermore, by introducing real logs to calculate the loss of the prediction results during the training process, the identification accuracy and reliability of the risk identification model under different attack scenarios can be improved.
[0070] In one embodiment, when the risk modeling mode is the second risk modeling mode, the business data stream is historical security event data. The process of determining risk assessment parameters through historical security event data and the risk modeling mode includes: collecting current system operating location data, current time information, and current network environment information. Based on the system operating location data and time information, the distribution of potential attack sources is calculated; the attack intensity is determined according to the distribution of potential attack sources, and the attack intensity, the distribution of potential attack sources, and the network environment information are combined as risk assessment parameters.
[0071] In this embodiment, the system operating location data can be the geographical information of the current operating node of the big data system obtained through sensors or log collection modules, including but not limited to latitude and longitude, data center deployment area, etc. Time information can be the current operating clock or timestamp, including but not limited to date, hour, minute, etc. Network environment information can be operating status information related to the current network environment of the system, which can be obtained through external interfaces or monitoring modules.
[0072] Based on system location data and time information, the distribution of potential attack sources can be calculated. This can be achieved by deriving the distribution of potential attack sources using attack attribution algorithms or geographic mapping models. For example, the distribution of potential attack sources can be represented by a matrix of node density and location coordinates.
[0073] Determining attack strength by analyzing the distribution of potential attack sources is understandable, as the attack strength varies depending on the distribution of potential attack sources. This is due to differences in distribution density, source node bandwidth, and historical attack frequency. Therefore, by analyzing the distribution of potential attack sources, especially in high-density areas, we can infer the current potential attack strength. Furthermore, we can further adjust the attack strength based on network environment information, such as adaptively increasing or decreasing the attack strength value.
[0074] This embodiment provides a big data security risk assessment method that acquires current system operating location data, time information, and network environment information. It then uses the location data and time information to calculate the distribution of potential attack sources, further combines the distribution of potential attack sources to deduce the attack intensity, and finally uses the attack intensity, potential attack source distribution, and network environment information as risk assessment parameters for subsequent risk modeling and analysis. This method can achieve dynamic adaptation to changes in potential attacks under different network scenarios, significantly improving the accuracy and real-time performance of risk assessment.
[0075] In one embodiment, the process of comprehensively analyzing the original risk indicators and outputting results through risk assessment parameters includes: generating risk action information of the original risk indicators based on the distribution of potential attack sources and network environment information, performing comprehensive analysis on the original risk indicators through attack intensity, and finally outputting security risk assessment results.
[0076] In this embodiment, risk action information can be determined based on the distribution of potential attack sources and network environment information, showing the dynamic manifestation of the original risk indicators in a security risk assessment scenario. This is used to adjust the simulated effect of risk propagation, making it more consistent with the real threat characteristics of the current system operation. For example, determining risk action information based on the distribution of potential attack sources can be done by judging whether the original risk indicators are likely to be subject to concentrated attacks based on the concentration of attack sources. In a specific embodiment, when the distribution of potential attack sources shows a dense concentration in a certain area, the risk action information can include the risk manifestation of multi-node coordinated attacks.
[0077] Determining risk action information based on network environment information can be done by considering factors such as bandwidth, packet loss rate, and abnormal traffic characteristics. In one specific embodiment, when bandwidth utilization reaches a preset threshold, the risk action information may include signs of system overload or service denial. In another specific embodiment, when the network environment indicates the presence of large-scale malicious traffic, the risk action information may include signs of propagation chain expansion or multi-point penetration.
[0078] In this embodiment, the original risk indicators are comprehensively analyzed by attack intensity to generate security risk assessment results. This can generate security risk assessment results corresponding to the risk action information, thereby realizing the dynamic nature of the assessment results, improving the matching degree between the risk assessment and the current system environment, and achieving the technical effect of improving the accuracy and authenticity of security risk assessment.
[0079] Furthermore, this invention also discloses an assessment system applied to the aforementioned big data security risk assessment method. The system includes: a data acquisition module for acquiring the types of big data operation scenarios in operation, as well as corresponding business data flows and system operation logs; a scenario identification and modeling module for selecting or determining the corresponding risk modeling mode based on the acquired operation scenario types; a risk parameter extraction module for extracting risk assessment parameters based on business data flows and risk modeling modes; a risk modeling and analysis module for comprehensively analyzing the original risk indicators using the risk assessment parameters; and a result generation and output module for generating and outputting security risk assessment results for subsequent early warning, display, or decision support.
[0080] This embodiment also provides a computer device applicable to a big data security risk assessment method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the big data security risk assessment method proposed in the above embodiment.
[0081] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0082] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements a big data security risk assessment method as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0083] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A big data security risk assessment method, characterized in that, include: The system acquires the types of big data operation scenarios in operation and collects the corresponding business data streams and system operation logs in real time or periodically through a data acquisition module. The scenario types include at least distributed storage environments and cloud computing environments. The business data streams may include real-time business flow logs and / or historical security event data. Determine the corresponding risk modeling mode based on the type of operational scenario; Based on the business data flow and the risk modeling pattern, risk assessment parameters for subsequent analysis are extracted and determined. The original risk indicators are then comprehensively processed and analyzed using the aforementioned risk assessment parameters to generate and output safety risk assessment results.
2. The big data security risk assessment method according to claim 1, characterized in that: The process of determining the risk modeling pattern includes: When the operating scenario is a distributed storage environment, the risk modeling mode adopted is the first risk modeling mode corresponding to the distributed storage environment; When the operating scenario is a cloud computing environment, the risk modeling mode adopted is the second risk modeling mode corresponding to the cloud computing environment.
3. The big data security risk assessment method according to claim 2, characterized in that: When the risk modeling mode is the first risk modeling mode, the business data stream is a real-time business flow log; the process of determining risk assessment parameters through the real-time business flow log and the risk modeling mode includes: Real-time business flow logs are input into the risk identification model to extract threat factor weights and attack chain information. The threat factor weights are normalized to obtain the risk impact coefficient. The risk impact coefficient and the attack chain information are used together as risk assessment parameters.
4. The big data security risk assessment method according to claim 3, characterized in that: The process of comprehensively analyzing the original risk indicators through risk assessment parameters and generating output results includes: Risk modeling information for constructing original risk indicators based on the attack chain information, wherein the risk modeling information includes an attack node feature matrix and a risk propagation matrix; The attack node feature matrix is input into the risk propagation function to calculate the risk propagation tensor; By performing matrix operations on the risk impact coefficient and the risk propagation tensor matrix, the risk intensity of the attack node in multiple dimensions can be obtained. Then, by combining the risk intensity of each dimension with the risk propagation matrix, the final safety risk assessment result is generated and output.
5. The big data security risk assessment method according to claim 3, characterized in that: The method further includes the following steps: Construct big data operation models corresponding to various attack scenarios; For each of the aforementioned operating models, different attack parameters are set and configured, and attack effect logs generated under those attack parameters are obtained respectively; A training set is formed based on the attack effect logs of all running models, the attack parameters corresponding to each log, and the corresponding scenario information. The preset lightweight network is then trained using the training set to obtain the trained risk identification model.
6. The big data security risk assessment method according to claim 2, characterized in that: When the risk modeling mode is the second risk modeling mode, the business data stream is historical security event data; The process of determining risk assessment parameters through historical security incident data and risk modeling patterns includes: Collect current system operating location data, current time information, and current network environment information; Based on the system's operating location data and time information, the distribution of potential attack sources is calculated; The attack intensity is determined based on the distribution of potential attack sources. The attack intensity, the distribution of potential attack sources, and the network environment information are combined as risk assessment parameters.
7. The big data security risk assessment method according to claim 6, characterized in that: The process of comprehensively analyzing the original risk indicators using risk assessment parameters and outputting the results includes: Based on the distribution of potential attack sources and network environment information, risk action information of the original risk indicators is generated. The attack intensity is then used to perform a comprehensive analysis of the original risk indicators, and the final security risk assessment result is output.
8. A big data security risk assessment system, applied to the big data security risk assessment method described in any one of claims 1-7, characterized in that: The system includes: The data acquisition module collects the types of big data operation scenarios in operation, as well as the corresponding business data streams and system operation logs; The scene recognition and modeling module selects or determines the corresponding risk modeling mode based on the collected types of operating scenarios. The risk parameter extraction module extracts risk assessment parameters based on business data flow and risk modeling patterns; The risk modeling and analysis module uses risk assessment parameters to conduct a comprehensive analysis of the original risk indicators; The results generation and output module generates and outputs security risk assessment results for subsequent early warning, display, or decision support.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the big data security risk assessment method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the big data security risk assessment method according to any one of claims 1 to 7.