Information security risk management method and system based on big data
By cleaning and deduplicating multi-source security data, constructing risk nodes and conducting dynamic analysis, the problems of inaccurate risk identification and delayed early warning in information security management under the big data environment are solved, and global perception and real-time protection of complex risks are realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-14
AI Technical Summary
Existing information security management technologies struggle to effectively identify dynamic risks across subjects, objects, and time periods in a big data environment. They lack the ability to process and dynamically analyze multi-source security data in a unified manner, resulting in inaccurate risk identification, delayed early warnings, and a high false alarm rate.
By uniformly cleaning and deduplicating multi-source security data, risk nodes are constructed and risk scores are calculated. Based on a rolling time window, a risk state sequence is generated, and risk evolution indicators are predicted and adaptive early warnings are implemented to achieve closed-loop optimization.
It improves the accuracy and foresight of risk identification, reduces false alarm and false negative rates, enhances the system's ability to protect against complex threats, and supports global risk management across subjects, objects, and time periods.
Smart Images

Figure CN121859342A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, specifically to a method and system for information security risk management based on big data. Background Technology
[0002] With the continuous expansion of information systems and the high degree of digitalization in business models, user access behavior, system operation behavior, and data interaction behavior are becoming more frequent, complex, and multi-sourced. Information security risks have gradually evolved from traditional single-point attacks to dynamic risks that cross subjects, objects, and time. Most existing information security management technologies are based on rule matching, static threshold judgment, or single-event analysis. Risk identification usually relies on predefined high-risk behavior patterns or fixed thresholds. Once the form of attack behavior changes, or the risk appears in a gradual and cumulative manner, existing methods often fail to detect it in a timely manner.
[0003] In a big data environment, security incident data comes from a wide range of sources, including logs, access records, and operation records. These different data sources vary significantly in data structure, sampling frequency, and semantic representation. Existing technologies often employ simple aggregation or single-source analysis methods, lacking the ability to uniformly clean, deduplicate, and correlate multi-source security data. This can easily lead to data redundancy and information distortion, thereby affecting the accuracy of risk assessment.
[0004] Furthermore, existing information security risk assessment methods typically focus on assessing risk intensity at a single point in time or within a short time window, lacking a systematic characterization of the risk's evolution over time. This makes it difficult to reflect the rate of risk growth, changing trends, and potential signs of loss of control. Before a risk reaches a clearly high-risk threshold, the system often fails to identify its potential danger, leading to delayed early warnings.
[0005] On the other hand, traditional risk warning mechanisms often use fixed threshold triggering methods, which cannot be adaptively adjusted according to the system's operating status and historical warning effects. This can easily lead to high false alarm rates or serious missed alarms, and they lack the closed-loop capability to continuously optimize the model based on the warning results.
[0006] Therefore, there is an urgent need for a comprehensive information security risk management method and system that can uniformly process multi-source security data in a big data environment and achieve dynamic analysis, trend prediction, and adaptive early warning based on risk evolution trajectory, so as to improve the foresight and accuracy of risk identification and the overall security protection capability of the system. Summary of the Invention
[0007] This invention provides a method and system for information security risk management based on big data, which helps to solve the problems mentioned in the background art.
[0008] This invention provides the following technical solution: a method for information security risk management based on big data, comprising: Multiple data sources are acquired, repeatedly cleaned, and the historical high-risk behaviors of the subject identifier and the target of operation are statistically analyzed to form an enhanced security dataset. Risk nodes are constructed based on the enhanced security dataset. A risk score is calculated and a risk identifier is set for each behavior record, forming a set of risk nodes. Risk nodes are sorted by subject and time, and a risk status sequence is generated based on a rolling time window to characterize the changes in risk over time. The total risk value, rate of change, and acceleration of change are calculated for the risk state sequence to obtain risk evolution indicators and identify abnormal acceleration states; Predict future risk trends based on risk evolution indicators and identify risk states that are in potentially high-risk stages; Based on the predicted future risk level and potential high-risk phase information, a dynamic risk warning is triggered and the corresponding warning level is output. Based on the comparison between historical early warning results and actual risk conditions, the risk scoring method, risk calculation method and risk threshold are adaptively adjusted to achieve closed-loop optimization of the risk management model.
[0009] Optionally, the step of acquiring multiple data sources, repeatedly cleaning them, and statistically analyzing the historical high-risk behaviors of the subject identifier and the target of the operation to form an enhanced security dataset includes: Identify and remove duplicate records from each data source; Based on predefined high-risk event types, the cumulative amount of historical risk behaviors for each user entity and each operated object is calculated separately. The accumulated historical risk is used as a new feature field and added to each cleaned behavior record to form an enhanced and complete security dataset.
[0010] Optionally, the step of constructing risk nodes based on the enhanced security dataset, calculating a risk score for each behavior record and setting a risk identifier to form a risk node set includes: Based on the enhanced security dataset, risk nodes are defined, and a comprehensive risk score is calculated for each risk node. The score takes into account the historical cumulative risk of the associated user subject, the historical cumulative risk of the associated operation object, and the time interval since the last operation on the same object. Based on the preset high-risk threshold, each risk node is marked as high-risk or ordinary risk; The output contains a set of all risk nodes, their scores, and risk indicators.
[0011] Optionally, the step of sorting risk nodes by subject and time, and generating a risk state sequence based on a rolling time window to characterize the changes in risk over time includes: The risk node set is sorted by user entity and time. The sorted risk nodes are divided using a rolling time window approach to obtain the node set within each time window; For each time window, calculate its total risk value, average risk value, risk standard deviation, and risk peak value to construct the risk state vector for that window; Calculate a dynamic threshold based on the total risk value of all windows, and label each window with its window-level risk identifier; Output a complete window risk state sequence arranged in chronological order.
[0012] Optionally, the step of calculating the total risk value, rate of change, and acceleration of change of the risk state sequence to obtain risk evolution indicators and identify anomalous acceleration states includes: Based on the window risk state sequence, calculate the total risk value, the rate of change of risk value, and the acceleration of risk change for each time window; By setting a high acceleration threshold, the time window state of abnormal acceleration with risk can be identified; The output includes a set of window-level risk evolution indicators, including the total risk value, rate of change, acceleration of change, and acceleration status indicators for each window.
[0013] Optionally, the step of predicting future risk trends based on risk evolution indicators and identifying risk states in potentially high-risk stages includes: Based on window-level risk evolution indicators, a nonlinear formula combining the current risk value, rate of change, acceleration of change, and peak risk correction term is used to predict the total risk value of the next time window. Judge future risk trends based on changes in predicted risk values; By calculating the dynamic intensity index of risk growth and identifying potential high-risk stages in the future based on dynamic thresholds; The output includes predictions and analysis results that include future risk forecasts, trend judgments, dynamic intensity, and risk stage indicators.
[0014] Optionally, the step of triggering dynamic risk warning and outputting the corresponding warning level based on the predicted future risk level and potential high-risk stage information includes: Based on whether the predicted future risk value exceeds the threshold and whether it is in a potentially high-risk stage, a warning trigger signal is dynamically generated. Based on the proportion of predicted risk values exceeding the threshold and the dynamic intensity of the risk, different levels of early warning information are classified and output, including red warning, orange warning, yellow warning, or no warning.
[0015] Optionally, the step of adaptively adjusting the risk scoring method, risk calculation method, and risk threshold based on the comparison between historical early warning results and actual risk conditions to achieve closed-loop optimization of the risk management model includes: Based on the hit rate and false alarm rate of the early warning results, the scores of risk nodes, the total risk value of the window, and the high-risk threshold of the window are adaptively fine-tuned and dynamically adjusted to achieve closed-loop optimization of the risk prediction and early warning model and provide adaptive capabilities for subsequent risk assessment.
[0016] A system for implementing the big data-based information security risk management method includes: The data integration and enhancement module cleans multi-source security data and calculates long-term risk context as enhancement features; The risk node generation and scoring module performs dynamic risk scoring and preliminary classification for each event; The risk status serialization analysis module aggregates discrete risk points into a time window sequence and calculates risk indicators to form a dynamic situation. The Risk Evolution Dynamics and Prediction module analyzes the speed and acceleration of risk changes and predicts future risk trends and potential high-risk stages. The intelligent early warning and adaptive optimization module triggers tiered early warnings based on prediction results and continuously optimizes model parameters based on feedback.
[0017] The present invention has the following beneficial effects: 1. A method for uniformly cleaning, deduplicating, and calculating the historical risk accumulation of subjects and operational objects from multi-source security data. Specifically, this includes defining a set of data sources, removing duplicate records, forming a complete cleaned security dataset, and calculating the historical risk accumulation for each subject identifier and operational object, providing foundational data for subsequent risk scoring. Uniform cleaning and deduplication of multi-source security data effectively eliminates data redundancy and information distortion, ensuring high-quality and complete data input to the risk analysis model. Various data sources differ significantly in structure, sampling frequency, and semantic expression. Traditional methods easily miss potential high-risk events when performing single-source analysis or simple summarization, leading to biased risk assessment. This method systematically integrates multi-source data and marks and removes duplicate data, ensuring that each record is analyzed only once, thereby significantly improving the accuracy and reliability of risk identification. Simultaneously, calculating the historical risk accumulation for each subject identifier and operational object not only reflects the risk characteristics of a single event but also reveals the risk trends of subjects and objects in long-term operations, providing a quantifiable and accessible data foundation for self-created risk scoring and subsequent dynamic analysis. This approach enables information security risk management methods to achieve global perception of complex risk behaviors across subjects, objects, and time in a big data environment. It overcomes the limitations of traditional static analysis methods in dealing with progressive and cumulative risks, and provides forward-looking data support and analytical basis for system security protection.
[0018] 2. Risk node construction, sorting, rolling window division, and window-level risk state vector calculation. This includes defining a risk node set, calculating a self-created risk score, sorting by subject and time to form a dynamic evolution trajectory, and calculating total risk, average risk, standard deviation, and risk peak within each time window, thereby forming a complete risk state sequence. By sorting risk nodes and dividing them into rolling windows, the system can systematically depict the temporal evolution trajectory of risks, enabling continuous monitoring of dynamic risk changes. Each risk node not only includes the time, subject, operation object, and operation type of a specific event, but also quantifies each event through a self-created risk score. Furthermore, a risk state vector is constructed using multi-dimensional indicators such as total risk, average risk, standard deviation, and risk peak, achieving global quantification of risk in space and time. This method can not only identify high-risk events but also analyze the rate of risk accumulation and fluctuation characteristics, revealing potential abnormal behavior patterns, overcoming the limitation of existing technologies in reflecting the evolution of risk over time. Simultaneously, serializing the risk state vector into a complete risk state sequence allows subsequent trend prediction and early warning triggering to be accurately calculated based on quantifiable data, enhancing the system's foresight and intelligence. This structured and continuous risk modeling approach can effectively support the global analysis of cross-subject and cross-object behaviors, providing reliable quantitative evidence and scientific decision support for information security management.
[0019] 3. A method for nonlinear risk prediction, trend determination, and marking of potential high-risk stages based on window-level risk evolution indicators. This method includes calculating the total risk value, rate of change, and acceleration; setting thresholds and tolerances; and marking potential dangerous stages based on dynamic indicators. By performing nonlinear prediction on window-level risk indicators, potential high-risk stages can be identified in advance, enabling proactive risk management. Specifically, by calculating the total risk value, rate of change, and acceleration, not only can the risk level of the current window be quantified, but the growth trend and fluctuation characteristics of the risk can also be reflected, thus accurately capturing potential abnormal behaviors before the risk reaches the explicit high-risk threshold. By using dynamic threshold and tolerance settings, continuous risk changes are mapped to discrete potential dangerous stage markers, effectively avoiding the underreporting and delayed early warning problems that occur in traditional fixed threshold methods. Simultaneously, this method, through its self-developed nonlinear prediction formula and peak correction mechanism, significantly enhances the sensitivity to abnormal peak risks, enabling timely detection of rapidly accumulating or sudden high-risk behaviors. This dynamic analysis method based on risk evolution trajectories upgrades information security risk management from static assessment to dynamic prediction and real-time perception, improving the system's ability to protect against complex and ever-changing threats and its response speed, and providing reliable data for risk intervention in actual operations.
[0020] 4. A method for dynamic early warning and closed-loop optimization based on risk prediction results, including early warning triggering conditions, early warning level classification, calculation of historical hit rate and false alarm rate, and adaptive fine-tuning and threshold adjustment of risk score and total window risk. By introducing a dynamic early warning triggering mechanism and closed-loop optimization, timely early warnings can be issued when information security risks are still in the development stage, preventing potential risks from escalating. The early warning triggering conditions not only consider the predicted total risk but also combine risk evolution dynamics indicators to achieve multi-dimensional judgment, enabling the system to identify risk behaviors that are about to accumulate rapidly. The classification of early warning levels further supports differentiated management, dividing risks into different levels to facilitate maintenance personnel in prioritizing high-risk events. At the same time, by statistically analyzing historical early warning hit rates and false alarm rates, this method can adaptively fine-tune the risk score and total window risk and dynamically adjust the threshold to achieve closed-loop optimization. This mechanism ensures that the early warning model can continuously learn and improve in actual operation, improve the accuracy and sensitivity of risk identification, reduce false alarm and false negative rates, significantly improve the intelligence, reliability, and sustainable operation capabilities of the information security management system in a big data environment, and qualitatively improve the overall security protection capability. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the process of the present invention.
[0022] Figure 2 This is a schematic diagram of the system structure of the present invention. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Example 1, see Figure 1 A big data-based information security risk management method includes: Multiple data sources are acquired, repeatedly cleaned, and the historical high-risk behaviors of the subject identifier and the target of operation are statistically analyzed to form an enhanced security dataset. Risk nodes are constructed based on the enhanced security dataset. A risk score is calculated and a risk identifier is set for each behavior record, forming a set of risk nodes. Risk nodes are sorted by subject and time, and a risk status sequence is generated based on a rolling time window to characterize the changes in risk over time. The total risk value, rate of change, and acceleration of change are calculated for the risk state sequence to obtain risk evolution indicators and identify abnormal acceleration states; Predict future risk trends based on risk evolution indicators and identify risk states that are in potentially high-risk stages; Based on the predicted future risk level and potential high-risk phase information, a dynamic risk warning is triggered and the corresponding warning level is output. Based on the comparison between historical early warning results and actual risk conditions, the risk scoring method, risk calculation method and risk threshold are adaptively adjusted to achieve closed-loop optimization of the risk management model.
[0025] The process involves acquiring multiple data sources, repeatedly cleaning them, and statistically analyzing the historical high-risk behaviors of the subject identifier and the target object to form an enhanced security dataset, including: Identify and remove duplicate records from each data source; Based on predefined high-risk event types, the cumulative amount of historical risk behaviors for each user entity and each operated object is calculated separately. The accumulated historical risk is used as a new feature field and added to each cleaned behavior record to form an enhanced and complete security dataset.
[0026] The process involves constructing risk nodes based on an enhanced security dataset, calculating a risk score for each behavior record, and setting a risk identifier to form a set of risk nodes, including: Based on the enhanced security dataset, risk nodes are defined, and a comprehensive risk score is calculated for each risk node. The score takes into account the historical cumulative risk of the associated user subject, the historical cumulative risk of the associated operation object, and the time interval since the last operation on the same object. Based on the preset high-risk threshold, each risk node is marked as high-risk or ordinary risk; The output contains a set of all risk nodes, their scores, and risk indicators.
[0027] The process of sorting risk nodes by subject and time, and generating a risk status sequence based on a rolling time window to characterize the changes in risk over time, includes: The risk node set is sorted by user entity and time. The sorted risk nodes are divided using a rolling time window approach to obtain the node set within each time window; For each time window, calculate its total risk value, average risk value, risk standard deviation, and risk peak value to construct the risk state vector for that window; Calculate a dynamic threshold based on the total risk value of all windows, and label each window with its window-level risk identifier; Output a complete window risk state sequence arranged in chronological order.
[0028] The calculation of the total risk value, rate of change, and acceleration of change of the risk state sequence to obtain risk evolution indicators and identify anomalous acceleration states includes: Based on the window risk state sequence, calculate the total risk value, the rate of change of risk value, and the acceleration of risk change for each time window; By setting a high acceleration threshold, the time window state of abnormal acceleration with risk can be identified; The output includes a set of window-level risk evolution indicators, including the total risk value, rate of change, acceleration of change, and acceleration status indicators for each window.
[0029] The method of predicting future risk trends based on risk evolution indicators and identifying risk states in potentially high-risk stages includes: Based on window-level risk evolution indicators, a nonlinear formula combining the current risk value, rate of change, acceleration of change, and peak risk correction term is used to predict the total risk value of the next time window. Judge future risk trends based on changes in predicted risk values; By calculating the dynamic intensity index of risk growth and identifying potential high-risk stages in the future based on dynamic thresholds; The output includes predictions and analysis results that include future risk forecasts, trend judgments, dynamic intensity, and risk stage indicators.
[0030] The process of triggering dynamic risk warnings and outputting corresponding warning levels based on predicted future risk levels and potential high-risk phases includes: Based on whether the predicted future risk value exceeds the threshold and whether it is in a potentially high-risk stage, a warning trigger signal is dynamically generated. Based on the proportion of predicted risk values exceeding the threshold and the dynamic intensity of the risk, different levels of early warning information are classified and output, including red warning, orange warning, yellow warning, or no warning.
[0031] The method of adaptively adjusting the risk scoring method, risk calculation method, and risk threshold based on the comparison between historical early warning results and actual risk conditions, in order to achieve closed-loop optimization of the risk management model, includes: Based on the hit rate and false alarm rate of the early warning results, the scores of risk nodes, the total risk value of the window, and the high-risk threshold of the window are adaptively fine-tuned and dynamically adjusted to achieve closed-loop optimization of the risk prediction and early warning model and provide adaptive capabilities for subsequent risk assessment.
[0032] Example 2, see Figure 2 A system for implementing the big data-based information security risk management method includes: The data integration and enhancement module cleans multi-source security data and calculates long-term risk context as enhancement features; The risk node generation and scoring module performs dynamic risk scoring and preliminary classification for each event; The risk status serialization analysis module aggregates discrete risk points into a time window sequence and calculates risk indicators to form a dynamic situation. The Risk Evolution Dynamics and Prediction module analyzes the speed and acceleration of risk changes and predicts future risk trends and potential high-risk stages. The intelligent early warning and adaptive optimization module triggers tiered early warnings based on prediction results and continuously optimizes model parameters based on feedback.
[0033] Example 3: A method for information security risk management based on big data, comprising: Multiple data sources are acquired, repeatedly cleaned, and the historical high-risk behaviors of the subject identifier and the target of operation are statistically analyzed to form an enhanced security dataset. Risk nodes are constructed based on the enhanced security dataset. A risk score is calculated and a risk identifier is set for each behavior record, forming a set of risk nodes. Risk nodes are sorted by subject and time, and a risk status sequence is generated based on a rolling time window to characterize the changes in risk over time. The total risk value, rate of change, and acceleration of change are calculated for the risk state sequence to obtain risk evolution indicators and identify abnormal acceleration states; Predict future risk trends based on risk evolution indicators and identify risk states that are in potentially high-risk stages; Based on the predicted future risk level and potential high-risk phase information, a dynamic risk warning is triggered and the corresponding warning level is output. Based on the comparison between historical early warning results and actual risk conditions, the risk scoring method, risk calculation method and risk threshold are adaptively adjusted to achieve closed-loop optimization of the risk management model.
[0034] The process involves acquiring multiple data sources, repeatedly cleaning them, and statistically analyzing the historical high-risk behaviors of the subject identifier and the target object to form an enhanced security dataset, including: Define data source collection : ; in, This is the Mth data source; Each data source Contains multiple records: ; in, Represents the i-th data source in the data source set. 1 record; Each record Contains multiple raw fields: ; in, For the time of the event, As the main identifier, such as user ID, process ID, etc. For the objects being manipulated, such as files, database tables, network nodes, etc. For operation type, Event type; Tagged data source collection Each data source Duplicate records in the data are denoted as follows: : ; in, for One piece of data in the middle; Remove Duplicate records : ; in, This is the data source after the first cleaning. For the difference of sets, that is, from Remove all from the middle. Elements in; After cleaning, a complete security dataset is formed. : ; in, This is the Nth cleaned record; Calculate the historical cumulative risk for each entity identifier and operational object: , ; in: The set of entity-identified behavior records is expressed as follows: , Representing records The corresponding main identifier; For a collection of records of the actions of the operation object, the expression is: , Representing records The corresponding operation object; Main logo Historical risk accumulation For the operation object The historical accumulation of risks; The set of high-risk events represents a predefined set of high-risk event types in the system, expressed as: , The event type with sequence number k; This is an indicator function used to determine whether a certain behavior record belongs to a high-risk event. (The event type is missing from the original text.) Belongs to the high-risk event set When the condition is met, the function takes the value 1; otherwise, it takes the value 0. Will and Add to the corresponding record; each record contains: .
[0035] The process involves constructing risk nodes based on an enhanced security dataset, calculating a risk score for each behavior record, and setting a risk identifier to form a set of risk nodes, including: Define the set of risk nodes ; in, This is the p-th risk node; Each node specifically includes: ; in, For self-created risk scoring, j is the risk node number; The self-created risk score The specific calculation expression is as follows: ; in, This indicates the time of the currently recorded event. Event time recorded with the same object as last time The time interval between the two ; Define high-risk threshold And set a risk identifier for each node: ; in, This indicates the risk level of a node, classifying it as either high-risk or moderate-risk. Output the complete set of risk nodes ; Each node contains: .
[0036] This method employs a unified approach to cleaning, deduplicating, and calculating the historical risk accumulation of subjects and operational objects from multiple security data sources. Specifically, it involves defining a set of data sources, removing duplicate records, forming a complete cleaned security dataset, and calculating the historical risk accumulation for each subject identifier and operational object, providing foundational data for subsequent risk scoring. Unified cleaning and deduplication of multi-source security data effectively eliminates data redundancy and information distortion, ensuring high-quality and complete data input to the risk analysis model. Various data sources differ significantly in structure, sampling frequency, and semantic expression. Traditional methods often miss potential high-risk events when performing single-source analysis or simple summarization, leading to biased risk assessments. This method systematically integrates multi-source data and marks and removes duplicate data, ensuring that each record is analyzed only once, thereby significantly improving the accuracy and reliability of risk identification. Simultaneously, calculating the historical risk accumulation for each subject identifier and operational object not only reflects the risk characteristics of a single event but also reveals the risk trends of subjects and objects in long-term operations, providing a quantifiable and accessible data foundation for self-created risk scoring and subsequent dynamic analysis. This approach enables information security risk management methods to achieve global perception of complex risk behaviors across subjects, objects, and time in a big data environment. It overcomes the limitations of traditional static analysis methods in dealing with progressive and cumulative risks, and provides forward-looking data support and analytical basis for system security protection.
[0037] The process of sorting risk nodes by subject and time, and generating a risk status sequence based on a rolling time window to characterize the changes in risk over time, includes: By the subject of the event With event time Sort the set of risk nodes: ; in, () is a sorting function, indicating that sorting is done by... The sorting ensures that all risk behavior nodes of the same user or system process are arranged consecutively, facilitating the analysis of the evolution of the main risk. Then, according to... Sorting ensures the correct sequence of risky behaviors, forming a dynamic evolutionary trajectory. After sorting ; Set the scroll window length ; Divide risk nodes according to a rolling window: ; in, For the first A set of time window nodes For the first The starting time point of each scrolling window; Define the risk state vector for each time window node set: ; in, The first node in the time window node set Self-created risk score for each node; Calculate the total window risk for each time window node set. : ; in, This is the fluctuation attenuation coefficient. Indicates variance; Calculate the window average risk for each time window node set. : ; Calculate the standard deviation of window risk for each time window node set. : ; Obtain the peak window risk for each time window node set. : ; Constructing a risk state vector for each time window node set : ; All risk state vectors Form a complete risk state sequence : ; in, The total number of risk state vectors; Define window high-risk threshold : ; in, This is the average value. Standard deviation; Set window risk indicators for each time window node set. : ; Output risk state sequence set : ; Each element specifically contains: .
[0038] The calculation of the total risk value, rate of change, and acceleration of change of the risk state sequence to obtain risk evolution indicators and identify anomalous acceleration states includes: Will according to Arranged in order, denoted as : ; Define a total risk function to calculate the total risk value for each time window node set. : ; Define the rate of change of risk : ; in, For window spacing; Define risk acceleration : ; Define high acceleration threshold : ; Set a high acceleration state for each time window node set. : ; Output window-level risk evolution indicator set : .
[0039] This approach involves constructing, sorting, and dividing risk nodes into rolling windows, as well as calculating window-level risk state vectors. It includes defining a set of risk nodes, calculating a self-created risk score, sorting by subject and time to form a dynamic evolution trajectory, and calculating total risk, average risk, standard deviation, and risk peak within each time window, thus forming a complete risk state sequence. By sorting risk nodes and dividing them into rolling windows, the system can systematically depict the temporal evolution trajectory of risks, enabling continuous monitoring of dynamic risk changes. Each risk node not only includes the time, subject, target, and type of the specific event, but also quantifies each event through a self-created risk score. Furthermore, a risk state vector is constructed using multi-dimensional indicators such as total risk, average risk, standard deviation, and risk peak, achieving global quantification of risk in space and time. This method can not only identify high-risk events but also analyze the rate of risk accumulation and volatility characteristics, revealing potential abnormal behavior patterns and overcoming the limitations of existing technologies in reflecting the evolution of risk over time. Simultaneously, serializing the risk state vector into a complete risk state sequence allows for precise calculation of subsequent trend predictions and early warning triggers based on quantifiable data, enhancing the system's foresight and intelligence. This structured and continuous risk modeling approach can effectively support the global analysis of cross-subject and cross-object behaviors, providing reliable quantitative evidence and scientific decision support for information security management.
[0040] The method of predicting future risk trends based on risk evolution indicators and identifying risk states in potentially high-risk stages includes: Set of window-level risk evolution indicators Arranged by serial number, denoted as : ; Set prediction step size Adaptable: ; in, To predict the step size amplification factor and ensure that the prediction covers the future risk evolution trend; Define a nonlinear prediction formula to predict the total risk of the next time window set of nodes. : ; in, This is a peak correction term, which enhances the predictive sensitivity to the risk of abnormal peaks; Define the trend determination rules: ; in, To determine the tolerance for trend, The trend status of the next time window node set; Define a risk evolution dynamics function and calculate the dynamic strength of risk growth. : ; Mark potentially hazardous stages: ; in, This indicates the threshold for judging potential hazards. This serves as a risk stage marker; Output future risk prediction and dynamic analysis results : .
[0041] The process of triggering dynamic risk warnings and outputting corresponding warning levels based on predicted future risk levels and potential high-risk phases includes: Will Arrange them according to their serial numbers to form a prediction sequence. : ; Define dynamic early warning trigger conditions: ; in, For trigger signal, This indicates that an alert has been triggered. This indicates that no warning has been triggered. Define warning level : ; Output early warning results : .
[0042] This method, based on window-level risk evolution indicators, performs nonlinear risk prediction, trend assessment, and labeling of potential high-risk stages. It involves calculating the total risk value, rate of change, and acceleration; setting thresholds and tolerances; and labeling potential hazardous stages based on dynamic indicators. By performing nonlinear prediction on window-level risk indicators, potential high-risk stages can be identified in advance, enabling proactive risk management. Specifically, by calculating the total risk value, rate of change, and acceleration, not only can the risk level of the current window be quantified, but the growth trend and volatility characteristics of the risk can also be reflected, thus accurately capturing potential abnormal behavior before the risk reaches the explicit high-risk threshold. Using dynamic thresholds and tolerance settings, continuous risk changes are mapped to discrete potential hazardous stage labels, effectively avoiding the underreporting and delayed early warning problems that occur in traditional fixed-threshold methods. Simultaneously, this method, through its self-developed nonlinear prediction formula and peak correction mechanism, significantly enhances the sensitivity to abnormal peak risks, enabling timely detection of rapidly accumulating or sudden high-risk behaviors. This dynamic analysis method based on risk evolution trajectories upgrades information security risk management from static assessment to dynamic prediction and real-time perception, improving the system's ability to protect against complex and ever-changing threats and its response speed, and providing reliable data for risk intervention in actual operations.
[0043] The process of adaptively adjusting the risk scoring method, risk calculation method, and risk threshold based on the comparison between historical early warning results and actual risk conditions to achieve closed-loop optimization of the risk management model includes: Calculate the early warning hit rate With false alarm rate : , ; Based on the early warning hit rate With false alarm rate right and Perform adaptive fine-tuning: ; in, The learning rate can be set from 0.01 to 0.05. and These are the adaptive fine-tuning results. and ; Based on the early warning hit rate With false alarm rate Dynamically adjust the high-risk threshold of the window : ; in, For the adjusted ; Output adjustment This provides a closed-loop adaptive capability for the next round of risk prediction and early warning.
[0044] This method, based on risk prediction results, employs dynamic early warning and closed-loop optimization. It includes early warning triggering conditions, early warning level classification, calculation of historical hit rate and false alarm rate, and adaptive fine-tuning and threshold adjustment of risk scores and total window risk. By introducing a dynamic early warning triggering mechanism and closed-loop optimization, timely early warnings can be issued when information security risks are still in their developmental stages, preventing potential risks from escalating. Early warning triggering conditions not only consider the predicted total risk but also incorporate risk evolution dynamics indicators to achieve multi-dimensional judgment, enabling the system to identify rapidly accumulating risk behaviors. The classification of early warning levels further supports differentiated management, categorizing risks into different levels to facilitate priority handling of high-risk events by operations and maintenance personnel. Simultaneously, by statistically analyzing historical early warning hit rates and false alarm rates, this method can adaptively fine-tune risk scores and total window risk, and dynamically adjust thresholds to achieve closed-loop optimization. This mechanism ensures that the early warning model can continuously learn and improve during actual operation, enhancing the accuracy and sensitivity of risk identification, reducing false alarm and false negative rates, and significantly improving the intelligence, reliability, and sustainable operation capabilities of the information security management system in a big data environment, resulting in a qualitative improvement in overall security protection capabilities.
[0045] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0046] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for information security risk management based on big data, characterized in that, include: Multiple data sources are acquired, repeatedly cleaned, and the historical high-risk behaviors of the subject identifier and the target of operation are statistically analyzed to form an enhanced security dataset. Risk nodes are constructed based on the enhanced security dataset. A risk score is calculated and a risk identifier is set for each behavior record, forming a set of risk nodes. Risk nodes are sorted by subject and time, and a risk status sequence is generated based on a rolling time window to characterize the changes in risk over time. The total risk value, rate of change, and acceleration of change are calculated for the risk state sequence to obtain risk evolution indicators and identify abnormal acceleration states; Predict future risk trends based on risk evolution indicators and identify risk states that are in potentially high-risk stages; Based on the predicted future risk level and potential high-risk phase information, a dynamic risk warning is triggered and the corresponding warning level is output. Based on the comparison between historical early warning results and actual risk conditions, the risk scoring method, risk calculation method and risk threshold are adaptively adjusted to achieve closed-loop optimization of the risk management model.
2. The information security risk management method based on big data according to claim 1, characterized in that, The process involves acquiring multiple data sources, repeatedly cleaning them, and statistically analyzing the historical high-risk behaviors of the subject identifier and the target object to form an enhanced security dataset, including: Identify and remove duplicate records from each data source; Based on predefined high-risk event types, the cumulative amount of historical risk behaviors for each user entity and each operated object is calculated separately. The accumulated historical risk is used as a new feature field and added to each cleaned behavior record to form an enhanced and complete security dataset.
3. The information security risk management method based on big data according to claim 1, characterized in that, The process involves constructing risk nodes based on an enhanced security dataset, calculating a risk score for each behavior record, and setting a risk identifier to form a set of risk nodes, including: Based on the enhanced security dataset, risk nodes are defined, and a comprehensive risk score is calculated for each risk node. The score takes into account the historical cumulative risk of the associated user subject, the historical cumulative risk of the associated operation object, and the time interval since the last operation on the same object. Based on the preset high-risk threshold, each risk node is marked as high-risk or ordinary risk; The output contains a set of all risk nodes, their scores, and risk indicators.
4. The information security risk management method based on big data according to claim 1, characterized in that, The process of sorting risk nodes by subject and time, and generating a risk status sequence based on a rolling time window to characterize the changes in risk over time, includes: The risk node set is sorted by user entity and time. The sorted risk nodes are divided using a rolling time window approach to obtain the node set within each time window; For each time window, calculate its total risk value, average risk value, risk standard deviation, and risk peak value to construct the risk state vector for that window; Calculate a dynamic threshold based on the total risk value of all windows, and label each window with its window-level risk identifier; Output a complete window risk state sequence arranged in chronological order.
5. The information security risk management method based on big data according to claim 1, characterized in that, The calculation of the total risk value, rate of change, and acceleration of change of the risk state sequence to obtain risk evolution indicators and identify anomalous acceleration states includes: Based on the window risk state sequence, calculate the total risk value, the rate of change of risk value, and the acceleration of risk change for each time window; By setting a high acceleration threshold, the time window state of abnormal acceleration with risk can be identified; The output includes a set of window-level risk evolution indicators, including the total risk value, rate of change, acceleration of change, and acceleration status indicators for each window.
6. The information security risk management method based on big data according to claim 1, characterized in that, The method of predicting future risk trends based on risk evolution indicators and identifying risk states in potentially high-risk stages includes: Based on window-level risk evolution indicators, a nonlinear formula combining the current risk value, rate of change, acceleration of change, and peak risk correction term is used to predict the total risk value of the next time window. Judge future risk trends based on changes in predicted risk values; By calculating the dynamic intensity index of risk growth and identifying potential high-risk stages in the future based on dynamic thresholds; The output includes predictions and analysis results that include future risk forecasts, trend judgments, dynamic intensity, and risk stage indicators.
7. The information security risk management method based on big data according to claim 1, characterized in that, The process of triggering dynamic risk warnings and outputting corresponding warning levels based on predicted future risk levels and potential high-risk phases includes: Based on whether the predicted future risk value exceeds the threshold and whether it is in a potentially high-risk stage, a warning trigger signal is dynamically generated. Based on the proportion of predicted risk values exceeding the threshold and the dynamic intensity of the risk, different levels of early warning information are classified and output, including red warning, orange warning, yellow warning, or no warning.
8. The information security risk management method based on big data according to claim 1, characterized in that, The method of adaptively adjusting the risk scoring method, risk calculation method, and risk threshold based on the comparison between historical early warning results and actual risk conditions, in order to achieve closed-loop optimization of the risk management model, includes: Based on the hit rate and false alarm rate of the early warning results, the scores of risk nodes, the total risk value of the window, and the high-risk threshold of the window are adaptively fine-tuned and dynamically adjusted to achieve closed-loop optimization of the risk prediction and early warning model and provide adaptive capabilities for subsequent risk assessment.
9. A system employing the information security risk management method based on big data as described in claim 1, characterized in that, include: The data integration and enhancement module cleans multi-source security data and calculates long-term risk context as enhancement features; The risk node generation and scoring module performs dynamic risk scoring and preliminary classification for each event; The risk status serialization analysis module aggregates discrete risk points into a time window sequence and calculates risk indicators to form a dynamic situation. The Risk Evolution Dynamics and Prediction module analyzes the speed and acceleration of risk changes and predicts future risk trends and potential high-risk stages. The intelligent early warning and adaptive optimization module triggers tiered early warnings based on prediction results and continuously optimizes model parameters based on feedback.