Intelligent risk early warning management method for information system

Through the methods of consistency comparison, simulation detection and knowledge graph construction, the problems of insufficient data correlation mining and difficulty in causal traceability in traditional risk warning methods are solved, and high-precision intelligent risk warning management of the information system is realized, which improves the accuracy and efficiency of risk warning.

CN120263457AInactive Publication Date: 2025-07-04CHINA NAT INST OF STANDARDIZATION
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510384024.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional risk warning methods are difficult to cope with the real-time fusion analysis needs of multi-source heterogeneous data, and there are insufficient data correlation mining, lag in abnormal detection, and difficulty in causal traceability, resulting in low accuracy and credibility of risk warning.

Method used

The methods of consistency comparison, simulation detection, knowledge completion and causal traceability are adopted to build a knowledge graph through a graph attention network to realize the fusion analysis of multi-level data sources, and improve the accuracy and effectiveness of risk warning.

Benefits of technology

It improves the accuracy and efficiency of intelligent risk warning management of information systems, adapts to the needs of intelligent risk warning management of different standards and information systems, and is universal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263457A_ABST
    Figure CN120263457A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent risk early warning management method for an information system, and the method comprises the steps: taking data obtained from an intelligent data source of a preset information system within a specified time as to-be-analyzed data; wherein the intelligent data source comprises a reference data source and the to-be-determined data; abnormal data and fuzzy data are obtained according to the consistency proportion, the reference data source and the to-be-analyzed data, the fuzzy data are screened based on the anomaly degree, and a screening result is added into the abnormal data; constructing a knowledge graph by adopting the intelligent data source based on a graph attention network, performing knowledge completion on the knowledge graph, obtaining a matched entity through the knowledge graph after knowledge completion according to the abnormal data, and performing causal tracing according to the matched entity to obtain tracing data; and constructing an intelligent risk early warning management model according to the traceability data, inputting to-be-managed data into the intelligent risk early warning management model, and outputting a management result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of early warning management, and particularly to an intelligent risk early warning management method for information systems. Background Art

[0002] With the complication of information systems and the acceleration of the digitalization process, the security threats and operation risks faced by them present diversified, concealed and dynamic characteristics. Traditional risk early warning methods mainly rely on static rule libraries, single-dimensional data analysis or manual experience judgment, and are difficult to meet the real-time fusion analysis requirements of multi-source heterogeneous data. Especially when dealing with large-scale system logs, unstructured documents, real-time transaction data and external threat intelligence, there are problems such as insufficient mining of data relevance, lag in anomaly detection, and difficulty in causal tracing. In addition, the collaborative utilization of data sources with different security levels in the prior art is insufficient, resulting in high-value reference data failing to effectively guide the anomaly determination of low-trust data, reducing the accuracy and credibility of risk early warning.

[0003] In recent years, the intelligent analysis technology based on knowledge graphs has provided new ideas for risk early warning, but it still has limitations in aspects such as dynamic data update, complex entity relationship reasoning and fuzzy data determination. For example, the construction of traditional knowledge graphs relies on manually defined rules and is difficult to adapt to real-time changing threat scenarios; the anomaly detection model is not sensitive enough to fuzzy data and is prone to false alarms or missed alarms. At the same time, the existing methods lack in-depth mining of the causal links across data sources, resulting in low efficiency in locating the root causes of risks and unable to meet the requirements of active defense and intelligent decision-making.

[0004] Therefore, there is an urgent need for an intelligent risk early warning management method that can integrate multi-level data sources, realize dynamic knowledge evolution and complementation, and have high-precision anomaly detection and causal tracing capabilities. This method needs to break through the bottleneck of the existing technology, and through the consistency verification of heterogeneous data, the construction of knowledge graphs driven by graph attention networks and causal reasoning mechanisms, improve the risk perception and early warning efficiency in complex scenarios, and provide reliable support for the security protection and operation management of information systems. Summary of the Invention

[0005] The object of the present invention is to provide an intelligent risk early warning management method for information systems.

[0006] To achieve the above object, the present invention is implemented according to the following technical solutions: The present invention includes the following steps: Take the data obtained from the intelligent data sources of the preset information system within the specified time as the data to be analyzed; wherein the intelligent data sources include reference data sources and the data to be determined; the reference data sources are data sources with a security level higher than that of the data sources to be determined; the data includes system logs, network traffic, user behavior, device status, real-time transaction data, unstructured document data, and threat intelligence platform data. Take the reference data source as a control group, perform a consistency comparison on the data to be analyzed to obtain a consistency ratio, obtain abnormal data and fuzzy data based on the consistency ratio, perform a simulation detection on the fuzzy data to obtain an abnormality degree, and add the fuzzy data with an abnormality degree higher than the abnormality threshold to the abnormal data. Construct a knowledge graph using the intelligent data sources based on the graph attention network, complete the knowledge of the knowledge graph, obtain matching entities through the knowledge-completed knowledge graph according to the abnormal data, and perform causal tracing based on the matching entities to obtain tracing data. Construct an intelligent risk warning management model based on the tracing data, input the data to be managed into the intelligent risk warning management model, and output the management result.

[0007] Further, the method for obtaining abnormal data and fuzzy data based on the consistency ratio includes: The consistency ratio is the similarity between the data in the reference data source and the data to be analyzed. Output the data to be analyzed with a consistency ratio less than 0.357 as abnormal data, and output the data to be analyzed with a consistency ratio greater than 0.357 and less than 0.719 as fuzzy data.

[0008] Further, the method for performing a simulation detection on the fuzzy data to obtain an abnormality degree includes: Perform simulation operations on the fuzzy data for poisoning attack, privacy leakage, sample shift, concept drift, backdoor attack, resource competition, cascading failure, decision bias, compliance violation, adversarial simulation, interface abuse, and model hijacking respectively. Use the data drift degree, prediction stability, abnormal recovery time, and compliance violation rate to detect the results after simulation. For the fuzzy data with effective data drift, output the abnormality degree of the fuzzy data with a data drift degree less than 0.049 as 0.95; for the fuzzy data lacking stability, output the abnormality degree of the fuzzy data with a prediction stability less than 0.152 as 0.848; for the fuzzy data causing system anomalies, output the abnormality degree of the fuzzy data with an abnormal recovery time less than 0.98 min as 0.89; for the fuzzy data causing compliance violations, output the abnormality degree of the fuzzy data with a compliance violation rate less than 0.001 as 0.001. For fuzzy data with two or more of data drift, instability, system anomalies, and compliance violations, the data drift degree, prediction stability, anomaly recovery time, and compliance violation rate are weighted and summed, and the weighted sum is output as the anomaly degree of the fuzzy data.

[0009] Further, a method for knowledge completion of the knowledge graph includes: Adding directions according to the connection directions of entities in the knowledge graph to obtain a directed knowledge graph, where the information in the directed knowledge graph flows in three directions: the original relationship, the inverse relationship, and the self-loop relationship; Introducing a multi-head attention mechanism in the directed knowledge graph, with the expression: , , where the triple formed by entity u and entity z through relationship r is ,the attention head is x, and the representation combination of the triple of attention head x is ,the relationship type-specific parameter after being processed by the multi-head attention mechanism is ,the relationship r type-specific parameter after being processed by the multi-head attention mechanism is ,the relationship r type-specific parameter is ,the original relationship set is R, the self-loop relationship set is ,the inverse relationship set is ,the weight matrix of the original relationship is ,the weight matrix of the inverse relationship is ,the weight matrix of the self-loop relationship is ,the entity is u, the relationship is r, Taking the entity as a feature node, aggregating messages using multi-head attention with different learning network parameters, and updating the relationship embedding, with the expression: , , where the hyperbolic tangent function is tanh, the neighbor set of the node is ,the node feature is z, the relationship set connecting entity u and entity i is ,the parameterized weight matrix in the multi-head attention mechanism is ,the normalized attention coefficient of entity u's relationship r is ,the cyclic correlation between node u and node z through relationship r is ,the number of attention heads is ,the non-linear activation function is ,the updated node z feature is ; Adjust the knowledge graph by using the updated node features and output the adjusted knowledge graph.

[0010] Furthermore, the method for obtaining matching entities from the knowledge graph after knowledge completion based on the abnormal data includes: Map the abnormal data to the knowledge graph and mark it as an abnormal node, and calculate the total influence of all neighbor nodes of the abnormal node on the node based on the core-shell value and the node degree: , where the total influence of node a is , the core-shell value of node a is , the core-shell value of node c is , the number of neighbor nodes of node a is , the number of neighbor nodes of node j is , the neighbor node of node a is j, the neighbor node of node j is c, the number of connection edges between node a and node j is , the number of connection edges between node c and node j is , the degree of node a is , the degree of node j is , the initial weight of node a is , the initial weight of node c is , the neighbor node set of node a is ; Sort the abnormal nodes in descending order according to the total influence of the nodes to obtain an abnormal node sequence; Use the restart random walk algorithm to generate a corpus formed by multiple entity relationships based on the abnormal node sequence, and calculate the similarity of the random walk corpus using a graph kernel function; Given the entity for which the corpus needs to be generated, the number of steps of the random walk, and the maximum neighbor order of the random walk, select the out-degree entity as the next point of the walk according to the probability. The probability expression is , where the probability is p, the out-degree neighbor set of entity j is , the in-degree of entity j is , the in-degree of entity w is , the sum of the in-degrees of all out-degree neighbors of entity j is ; Repeat the probability selection. If the order of the next selected entity is greater than the given maximum order, return the entity and continue the above operation until all neighbors are traversed; Calculate the similarity between entities. The expression is: , where the feature vector of entity a is , the feature vector of entity j is , the similarity metric function is , the feature vector and the feature vector The similarity between them is , the similarity between entity a and entity j is , the vector set obtained by entity a through restart random walk and bag - of - words model is , the vector set obtained by entity j through restart random walk and bag - of - words model is , the vector set The number of elements in is , the vector set The number of elements in is ; Use the corpus with similarity greater than 0.512 as the first entity; sort the corpus with similarity less than 0.512 in descending order, and use the top three corpora as the second entity; output the first entity and the second entity as matching entities.

[0011] Further, a method for causal tracing to obtain tracing data based on the matching entity includes: Use a pointer - generator network to generate titles for the entries in the input matching entity that lack titles, find important entities through the document and entry knowledge graphs, find relevant entries through the breadth - first search algorithm, and use the sub - titles of the entries as auxiliary information to expand the search scope of the entries to generate an entry set; Calculate the node importance based on the entry set, select the head entity of the abnormal node sequence of the entry according to the importance, use the breadth - first search algorithm to find the multi - order parent neighbor entities of the abnormal node sequence entity, and obtain the relevant entry knowledge graph based on the multi - order parent neighbor entities; Calculate the similarity between the relevant entry knowledge graph and the target entry through the graph kernel network, perform weighted averaging on the entity similarity of the two entries to obtain the similarity score QC of the two entries, and the similarity score UE of the important entities in the target entry and the important entities in the relevant entry. When QC is less than or equal to 0.796 and UE is greater than 0.8, the relevant entry is in a tracing relationship, and output the relevant entry as tracing data.

[0012] Further, a method for constructing an intelligent risk warning management model based on the tracing data includes: Calculate the average break - in time of each attack damage path based on the tracing data to quantitatively evaluate the security risk and vulnerability of the information system, construct an objective function according to the security risk and vulnerability, and the expression is: , where the objective function at the s - th moment is , the security risk at the s - th moment is , the vulnerability at the s - th moment is , the weight of the security risk is , the weight of the vulnerability is , the loss function is ; Early warning management principle: When the objective function value is greater than 0.937, automatic service fusing is performed; when the objective function value is less than 0.937 and greater than 0.804, traffic cleaning and model heat exchange are performed; when the objective function value is less than 0.804 and greater than 0.569, dynamic blocking of the rule engine is performed; when the objective function value is less than 0.569, manual review and policy optimization are performed; The intelligent risk early warning management model includes an anomaly recognition algorithm, a classification algorithm, and a BP neural network algorithm; The anomaly recognition algorithm identifies anomaly points deviating from the normal pattern by analyzing the law of input data change over time in the information system, and takes the anomaly points as abnormal data; The classification algorithm groups abnormal data points into clusters defined by similarity, making the data points within the clusters highly similar and the data points between the clusters less similar, to obtain classified data; The BP neural network algorithm finds the optimal parameter combination by learning the risk assessment law of the objective function, performs risk assessment on the classified data according to the risk assessment law to obtain a risk value, and performs management based on the risk value according to the early warning management principle to obtain a management result.

[0013] The beneficial effects of the present invention are: The present invention is an intelligent risk early warning management method for information systems. Compared with the prior art, the present invention has the following technical effects: Through steps of consistency comparison, simulation detection, knowledge completion, obtaining matching entities, causal tracing, and model construction, the present invention can improve the accuracy of intelligent risk early warning management of information systems, thereby improving the precision of intelligent risk early warning management of information systems, optimizing the intelligent risk early warning management of information systems, greatly saving resources, improving work efficiency, realizing intelligent management of intelligent risk early warning of information systems, performing data correction and causal tracing on the intelligent risk early warning management of information systems in real time, which is of great significance to the intelligent risk early warning management of information systems, and can adapt to the intelligent risk early warning management of information systems with different standards and the intelligent risk early warning management requirements of different information systems, having a certain universality. Description of the Drawings

[0014] Figure 1 It is a step flow chart of an intelligent risk early warning management method for information systems of the present invention. Detailed Embodiments

[0015] The present invention will be further described below through specific embodiments. The illustrative embodiments of the present invention and the descriptions are used to explain the present invention, but do not limit the present invention.

[0016] An intelligent risk early warning management method for an information system according to the present invention includes the following steps: As Figure 1 shown, in this embodiment, it includes the following steps: Taking the data obtained from the intelligent data source of the preset information system within a specified time as the data to be analyzed; wherein the intelligent data source includes a reference data source and the to-be-determined data; the reference data source is a data source with a security level higher than that of the to-be-determined data source; the data includes system logs, network traffic, user behavior, device status, real-time transaction data, unstructured document data, threat intelligence platform data; In an actual evaluation, taking an information system as the research object, obtaining the data to be analyzed on September 11, 2024, where: The system log is: 2024-09-11 10:00:00, user admin logs in to the system, IP address: 192.168.1.100; 2024-09-11 10:05:00, the system performs a file backup operation, successfully backing up the file: datafile1.txt; 2024-09-11 10:10:00, a system error occurs, error code: 500, indicating a database connection timeout; 2024-09-11 10:15:00, user user1 attempts to log in to the system, fails, reason: incorrect password; 2024-09-11 10:20:00, the system automatically clears temporary files, releasing 100MB of space;

[0017] The network traffic is as follows: Source IP: 192.168.1.101, Destination IP: 10.0.0.1, Port: 80, Traffic size: 500KB, Time: 2024-09-11 10:00:00; Source IP: 192.168.1.102, Destination IP: 192.168.1.100, Port: 22, Traffic size: 100KB, Time: 2024-09-11 10:05:00; Source IP: 192.168.1.103, Destination IP: 192.168.1.100, Port: 3306, Traffic size: 2000KB, Time: 2024-09-11 10:10:00; Source IP: 192.168.1.104, Destination IP: 192.168.1.100, Port: 443, Traffic size: 300KB, Time: 2024-09-11 10:15:00; Source IP: 192.168.1.105, Destination IP: 192.168.1.100, Port: 8080, Traffic size: 400KB, Time: 2024-09-11 10:20:00; The user behaviors are as follows: User user2 accessed the user management module of the system at 2024-09-11 10:00:00; User user3 attempted to modify their personal information at 2024-09-11 10:05:00 and succeeded; User user4 refreshed the system homepage frequently at 2024-09-11 10:10:00, refreshing 20 times within one minute; User user5 downloaded a file named report.pdf at 2024-09-11 10:15:00; User user6 logged out of the system at 2024-09-11 10:20:00; The device status is as follows: Server Device 1: CPU usage rate is 30%, memory usage rate is 60%, remaining hard disk space is 500GB, network connection is normal, time is 2024-09-11 10:00:00; Server Device 2: CPU usage rate is 20%, memory usage rate is 50%, remaining hard disk space is 600GB, network connection is normal, time is 2024-09-11 10:00:00; Server Device 1: CPU usage rate suddenly rises to 90%, memory usage rate is 70%, remaining hard disk space is 490GB, network connection is normal, time is 2024-09-11 10:10:00; Server Device 2: CPU usage rate is 25%, memory usage rate is 55%, remaining hard disk space is 580GB, network connection is normal, time is 2024-09-11 10:10:00; Server Device 1: CPU usage rate is 40%, memory usage rate is 65%, remaining hard disk space is 480GB, network connection is normal, time is 2024-09-11 10:20:00; The real-time transaction data is as follows: Transaction ID: 1001, User ID: user7, transaction amount is 100 yuan, transaction time is 2024-09-11 10:00:00, transaction status is successful; Transaction ID: 1002, User ID: user8, transaction amount is 200 yuan, transaction time is 2024-09-11 10:05:00, transaction status is successful; Transaction ID: 1003, User ID: user9, transaction amount is 5000 yuan, transaction time is 2024-09-11 10:10:00, transaction status is successful; Transaction ID: 1004, User ID: user10, transaction amount is 150 yuan, transaction time is 2024-09-11 10:15:00, transaction status is successful; Transaction ID: 1005, User ID: user11, transaction amount is 80 yuan, transaction time is 2024-09-11 10:20:00, transaction status is successful; The unstructured document data is as follows: Document 1: "The system is running normally today, but there were some minor problems around 10:10 and further investigation is needed."; Document 2: "Users reported some trouble when logging in. Some users entered the correct password but could not log in."; Document 3: "The recent transaction data shows that there was a large transaction and a risk assessment is required."; The threat intelligence platform data is as follows: Suspicious scanning behavior from IP address 192.168.1.106 was detected at 2024-09-11 10:00:00; No other threat intelligence was found; Taking the reference data source as a control group, perform a consistency comparison on the data to be analyzed to obtain a consistency ratio. Based on the consistency ratio, obtain abnormal data and fuzzy data. Perform a simulation test on the fuzzy data to obtain an abnormality degree, and add the fuzzy data with an abnormality degree higher than the abnormality threshold to the abnormal data; In actual evaluation, the abnormal data includes the network traffic source IP: 192.168.1.103, the CPU usage rate of the device status suddenly rises to 90%, the real-time transaction data transaction amount is 5000 yuan, the unstructured document data, and the threat intelligence platform data; the fuzzy data includes the system log 2024-09-11 10:10:00, the user behavior user user4, and the real-time transaction data transaction amount is 200 yuan; the fuzzy data added to the abnormal data includes the system log 2024-09-11 10:10:00 and the user behavior user user4; the abnormality threshold is 0.358; Based on the graph attention network, construct a knowledge graph using the intelligent data source, complete the knowledge of the knowledge graph, obtain matching entities through the knowledge graph after knowledge completion according to the abnormal data, and obtain traceability data through causal tracing based on the matching entities; Construct an intelligent risk early warning management model according to the traceability data, input the data to be managed into the intelligent risk early warning management model, and output the management result.

[0018] In this embodiment, the method for obtaining abnormal data and fuzzy data based on the consistency ratio includes: The consistency ratio is the similarity between the data in the reference data source and the data to be analyzed. The data to be analyzed with a consistency ratio less than 0.357 is output as abnormal data, and the data to be analyzed with a consistency ratio greater than 0.357 and less than 0.719 is output as fuzzy data.

[0019] In this embodiment, the method for obtaining the abnormality degree by performing a simulation test on the fuzzy data includes: Perform simulation operations of poisoning attack, privacy leakage, sample drift, concept drift, backdoor attack, resource competition, cascade failure, decision deviation, compliance violation, adversarial simulation, interface abuse, and model hijacking on the fuzzy data respectively; The simulated results are detected using data drift degree, prediction stability, abnormal recovery time, and compliance violation rate. For fuzzy data with effective data drift, the abnormality degree of fuzzy data with a data drift degree less than 0.049 is output as 0.95; for fuzzy data lacking stability, the abnormality degree of fuzzy data with a prediction stability less than 0.152 is output as 0.848; for fuzzy data causing system anomalies, the abnormality degree of fuzzy data with an abnormal recovery time less than 0.98 min is output as 0.89; for fuzzy data causing compliance violations, the abnormality degree of fuzzy data with a compliance violation rate less than 0.001 is output as 0.001; For fuzzy data having two or more of data drift, instability, system anomalies, and compliance violations, the data drift degree, prediction stability, abnormal recovery time, and compliance violation rate are weighted and summed, and the weighted sum is output as the abnormality degree of the fuzzy data.

[0020] In this embodiment, the method for knowledge completion of the knowledge graph includes: Adding directions according to the connection directions of each entity in the knowledge graph to obtain a directed knowledge graph, and the information in the directed knowledge graph flows in three directions: the original relationship, the inverse relationship, and the self-loop relationship; Introducing a multi-head attention mechanism in the directed knowledge graph, and the expression is: , , where the triple formed by entity u and entity z through relationship r is ,the attention head is x, and the representation combination of the triple of attention head x is ,the relationship type-specific parameter after being processed by the multi-head attention mechanism is ,the parameter specific to relationship r type after being processed by the multi-head attention mechanism is ,the parameter specific to relationship r type is ,the original relationship set is R, the self-loop relationship set is ,the inverse relationship set is ,the weight matrix of the original relationship is ,the weight matrix of the inverse relationship is ,the weight matrix of the self-loop relationship is ,the entity is u, the relationship is r, Taking the entity as a feature node, using multi-head attention with different learning network parameters to aggregate messages and update the relationship embedding, and the expression is: , , where the hyperbolic tangent function is tanh, and the neighbor set of the node is , the node feature is z, and the set of relationships connecting entity u and entity i is , and the parameterized weight matrix in the multi-head attention mechanism is , and the normalized attention coefficient of entity u's relationship r is , and the circular correlation between node u and node z through relationship r is , and the number of attention heads is , and the non-linear activation function is , and the updated node z feature is ; Adjust the knowledge graph using the updated node features and output the adjusted knowledge graph.

[0021] In this embodiment, the method for obtaining matching entities from the knowledge graph after knowledge completion based on the abnormal data includes: Map the abnormal data to the knowledge graph and mark it as an abnormal node, and calculate the total influence of all neighbor nodes of the abnormal node on the node based on the core-shell value and the node degree: , where the total influence of node a is , the core-shell value of node a is , the core-shell value of node c is , the number of neighbor nodes of node a is , the number of neighbor nodes of node j is , the neighbor node of node a is j, the neighbor node of node j is c, the number of connection edges between node a and node j is , the number of connection edges between node c and node j is , the degree of node a is , the degree of node j is , the initial weight of node a is , the initial weight of node c is , the neighbor node set of node a is ; Sort the abnormal nodes in descending order according to the total influence of the nodes to obtain an abnormal node sequence; Use the restart random walk algorithm to generate a corpus formed by multiple entity relationships based on the abnormal node sequence, and calculate the similarity of the random walk corpus using a graph kernel function; Given the entity for which the corpus needs to be generated, the number of steps of the random walk, and the maximum neighbor order of the random walk, select the out-degree entity as the next point of the walk according to the probability, and the probability expression is, , where the probability is p, and the out-degree neighbor set of entity j is , and the in-degree of entity j is The in-degree of entity w is The sum of the in-degrees of all out-degree neighbors of entity j is ; Repeat probability selection. If the order of the next selected entity is greater than the given maximum order, return the entity and continue the above operations until all neighbors are traversed; Calculate the similarity between entities. The expression is: , where the feature vector of entity a is , the feature vector of entity j is , the similarity metric function is , the feature vectors and the feature vector The similarity between them is , the similarity between entity a and entity j is , the vector set obtained by entity a through restart random walk and bag-of-words model is , the vector set obtained by entity j through restart random walk and bag-of-words model is , the vector set The number of elements in is , the vector set The number of elements in is ; Use the corpus with a similarity greater than 0.512 as the first entity; sort the corpus with similarities all less than 0.512 in descending order, and use the top three corpora as the second entity; output the first entity and the second entity as matching entities.

[0022] In this embodiment, the method for obtaining traceability data according to the matching entities includes: Use a pointer-generator network to generate titles for the entries lacking titles in the input matching entities, find important entities through the document and entry knowledge graphs, find relevant entries through the breadth-first search algorithm, and use the sub-titles of the entries as auxiliary information to expand the search scope of the entries to generate an entry set; Calculate the node importance according to the entry set, select the head entity of the abnormal node sequence of the entry according to the importance, use the breadth-first search algorithm to find the multi-order parent neighbor entities of the abnormal node sequence entities, and obtain the relevant entry knowledge graph based on the multi-order parent neighbor entities; Calculate the similarity between the relevant entry knowledge graph and the target entry through the graph kernel network, perform weighted averaging on the entity similarities of the two entries to obtain the similarity score QC of the two entries, and the similarity score UE of the important entities in the target entry and the important entities in the relevant entry. When QC is less than or equal to 0.796 and UE is greater than 0.8, the relevant entry is a traceability relationship, and output the relevant entry as traceability data.

[0023] In this embodiment, the method for constructing an intelligent risk early warning management model based on the traceability data includes: Calculating the average break time of each attack damage path according to the traceability data to quantitatively evaluate the security risk and vulnerability of the information system, and constructing an objective function based on the security risk and vulnerability. The expression is: , where the objective function at the s-th moment is , the security risk at the s-th moment is , the vulnerability at the s-th moment is , the weight of the security risk is , the weight of the vulnerability is , and the loss function is ; Early warning management principle: When the objective function value is greater than 0.937, automatic service fusing is performed; when the objective function value is less than 0.937 and greater than 0.804, traffic cleaning and model heat exchange are performed; when the objective function value is less than 0.804 and greater than 0.569, dynamic blocking of the rule engine is performed; when the objective function value is less than 0.569, manual review and policy optimization are performed; The intelligent risk early warning management model includes an anomaly recognition algorithm, a classification algorithm, and a BP neural network algorithm; The anomaly recognition algorithm identifies anomaly points deviating from the normal pattern by analyzing the law of change of input data in the information system over time, and takes the anomaly points as anomaly data; The classification algorithm groups the anomaly data points into clusters defined by similarity, so that the data points within the clusters have high similarity and the data points between the clusters have low similarity, and obtains classified data; The BP neural network algorithm finds the optimal parameter combination by learning the risk assessment law of the objective function, performs risk assessment on the classified data according to the risk assessment law to obtain a risk value, and performs management based on the risk value according to the early warning management principle to obtain a management result.

[0024] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent risk early warning management method for information systems, characterized in that, Including the following steps: Taking the data obtained from the intelligent data source of the preset information system within a specified time as the data to be analyzed; Wherein the intelligent data source includes a reference data source and the to-be-determined data; The reference data source is a data source with a security level higher than that of the to-be-determined data source; the data includes system logs, network traffic, user behavior, device status, real-time transaction data, unstructured document data, threat intelligence platform data; Taking the reference data source as a control group, performing a consistency comparison on the data to be analyzed to obtain a consistency ratio, obtaining abnormal data and fuzzy data based on the consistency ratio, performing a simulation detection on the fuzzy data to obtain an abnormality degree, and adding the fuzzy data with an abnormality degree higher than the abnormality threshold to the abnormal data; Constructing a knowledge graph using the intelligent data source based on a graph attention network, performing knowledge completion on the knowledge graph, obtaining matching entities through the knowledge-completed knowledge graph according to the abnormal data, and obtaining traceability data through causal traceability based on the matching entities; Constructing an intelligent risk warning management model according to the traceability data, inputting the data to be managed into the intelligent risk warning management model, and outputting a management result.

2. The intelligent risk early warning management method for an information system according to claim 1, characterized in that, The method for obtaining abnormal data and fuzzy data based on the consistency ratio includes: The consistency ratio is the similarity between the data in the reference data source and the data to be analyzed. The data to be analyzed with a consistency ratio less than 0.357 is output as abnormal data, and the data to be analyzed with a consistency ratio greater than 0.357 and less than 0.719 is output as fuzzy data.

3. The intelligent risk early warning management method for an information system according to claim 1, characterized in that The method for performing a simulation detection on the fuzzy data to obtain an abnormality degree includes: Performing simulation operations such as poisoning attack, privacy leakage, sample drift, concept drift, backdoor attack, resource competition, cascading failure, decision deviation, compliance violation, adversarial simulation, interface abuse, and model hijacking on the fuzzy data respectively; Using the data drift degree, prediction stability, abnormal recovery time, and compliance violation rate to detect the simulation results. For the fuzzy data with effective data drift, the abnormality degree of the fuzzy data with a data drift degree less than 0.049 is output as 0.95; for the fuzzy data lacking stability, the abnormality degree of the fuzzy data with a prediction stability less than 0.152 is output as 0.848; for the fuzzy data causing system abnormalities, the abnormality degree of the fuzzy data with an abnormal recovery time less than 0.98 min is output as 0.89; for the fuzzy data causing compliance violations, the abnormality degree of the fuzzy data with a compliance violation rate less than 0.001 is output as 0.001; For the fuzzy data having two or more of data drift, instability, system abnormality, and compliance violation, the data drift degree, prediction stability, abnormal recovery time, and compliance violation rate are weighted and summed, and the weighted sum is output as the abnormality degree of the fuzzy data.

4. The intelligent risk early warning management method for an information system according to claim 1, wherein The method for performing knowledge completion on the knowledge graph includes: Adding directions according to the connection directions of each entity in the knowledge graph to obtain a directed knowledge graph, and the information in the directed knowledge graph flows in three directions: the original relationship, the inverse relationship, and the self-loop relationship; Introducing a multi-head attention mechanism in the directed knowledge graph, and the expression is: , , The triple formed by entity u and entity z through relation r is , the attention head is x, and the representation combination of the triple of attention head x is , the relation type-specific parameter after being processed by the multi-head attention mechanism is , the relation r type-specific parameter after being processed by the multi-head attention mechanism is , the relation r type-specific parameter is , the original relation set is R, and the self-loop relation set is , the inverse relation set is , the weight matrix of the original relation is , the weight matrix of the inverse relation is , the weight matrix of the self-loop relation is , the entity is u, and the relation is r, Taking entities as feature nodes, multi-head attention with different learning network parameters is used to aggregate messages and update relationship embeddings. The expression is as follows: , , where the hyperbolic tangent function is tanh, and the neighbor set of the node is , the node feature is z, and the relation set connecting entity u and entity i is , the parameterized weight matrix in the multi-head attention mechanism is , the normalized attention coefficient of entity u's relation r is , the cyclic correlation between node u and node z through relation r is , the number of attention heads is , the non-linear activation function is , the updated node z feature is ; The knowledge graph is adjusted using the updated node features, and the adjusted knowledge graph is output.

5. The intelligent risk early warning management method for an information system according to claim 1, characterized in that A method for obtaining matching entities from the knowledge graph after knowledge completion according to the abnormal data includes: Mapping the abnormal data to the knowledge graph and marking it as an abnormal node, and calculating the total influence of all neighbor nodes of the abnormal node on the node based on the core-shell value and the node degree: , Among them, the total influence of node a is , the core-shell value of node a is , the core-shell value of node c is , the number of neighbor nodes of node a is , the number of neighbor nodes of node j is , the neighbor node of node a is j, the neighbor node of node j is c, and the number of connecting edges between node a and node j is , the number of connecting edges between node c and node j is , the degree of node a is , the degree of node j is , the initial weight of node a is , the initial weight of node c is , the set of neighbor nodes of node a is ; Sorting the abnormal nodes in descending order according to the total influence of the nodes to obtain an abnormal node sequence; Using the restart random walk algorithm to generate a corpus formed by multiple entity relationships based on the abnormal node sequence, and calculating the similarity of the random walk corpus using a graph kernel function; Given an entity for which a corpus needs to be generated, the number of steps of random walk, and the maximum neighbor order of random walk, the out-degree entity is selected as the next point of the walk according to the probability. The probability expression is , where the probability is p, and the set of out-degree neighbors of entity j is , the in-degree of entity j is , the in-degree of entity w is , and the sum of the in-degrees of all out-degree neighbors of entity j is ; Repeat the probability selection. If the order of the next selected entity is greater than the given maximum order, return the entity and continue the above operation until all neighbors are traversed; Calculate the similarity between entities. The expression is as follows: , Among them, the feature vector of entity a is , and the feature vector of entity j is , the similarity metric function is , the similarity between the feature vector and the feature vector is , the similarity between entity a and entity j is , the vector set obtained by entity a through restart random walk and bag-of-words model is , the vector set obtained by entity j through restart random walk and bag-of-words model is , the number of elements in the vector set is , the number of elements in the vector set is ; Taking the corpus with a similarity greater than 0.512 as the first entity; sorting the corpus with similarities all less than 0.512 in descending order, and taking the first three corpora as the second entity; outputting the first entity and the second entity as matching entities.

6. The intelligent risk early warning management method for an information system according to claim 1, characterized in that A method for obtaining traceability data by causal traceability according to the matching entities includes: Using a pointer-generator network to generate titles for the entries in the input matching entities that lack titles, searching for important entities through the document and the entry knowledge graph, finding relevant entries through the breadth-first search algorithm, and using the sub-titles of the entries as auxiliary information to expand the search scope of the entries to generate an entry set; Calculating the node importance according to the entry set, selecting the head entity of the abnormal node sequence of the entry according to the importance, using the breadth-first search algorithm to find the multi-order parent neighbor entities of the entities in the abnormal node sequence, and obtaining the relevant entry knowledge graph based on the multi-order parent neighbor entities; Calculating the similarity between the relevant entry knowledge graph and the target entry through a graph kernel network, performing a weighted average on the entity similarity of the two entries to obtain the similarity score QC of the two entries, and the similarity score UE of the important entities in the target entry and the important entities in the relevant entry. When QC is less than or equal to 0.796 and UE is greater than 0.8, the relevant entry is a traceability relationship, and the relevant entry is output as traceability data.

7. The intelligent risk early warning management method for an information system according to claim 1, wherein A method for constructing an intelligent risk warning management model according to the traceability data includes: Calculating the average penetration time of each attack damage path according to the traceability data to quantitatively evaluate the security risks and vulnerabilities of the information system, and constructing an objective function according to the security risks and vulnerabilities. The expression is as follows: , where the objective function at the s-th moment is , the security risk at the s-th moment is , the vulnerability at the s-th moment is , the weight of the security risk is , the weight of the vulnerability is , and the loss function is ; Early warning management principle: When the objective function value is greater than 0.937, automatic service fusing is performed; when the objective function value is less than 0.937 and greater than 0.804, traffic cleaning and model heat exchange are performed; when the objective function value is less than 0.804 and greater than 0.569, dynamic blocking of the rule engine is performed; when the objective function value is less than 0.569, manual review and policy optimization are performed; The intelligent risk early warning management model includes an anomaly recognition algorithm, a classification algorithm, and a BP neural network algorithm; The anomaly recognition algorithm identifies anomaly points deviating from the normal pattern by analyzing the pattern of input data changing over time in the information system, and takes the anomaly points as abnormal data; The classification algorithm obtains classified data by grouping abnormal data points into clusters defined by similarity, so that the data points within the cluster have high similarity while the data points between clusters have low similarity; The BP neural network algorithm finds the optimal parameter combination by learning the risk assessment pattern of the objective function, obtains the risk value by performing risk assessment on the classified data according to the risk assessment pattern, and performs management based on the risk value according to the early warning management principle to obtain the management result.

Citation Information

Cited By

  • Information security risk classification method and system

    CN121125247A

  • Land intensive utilization intelligent analysis method and system based on multi-dimensional indexes

    CN121189946A

  • Intelligent Analysis Method and System for Intensive Land Use Based on Multidimensional Indicators

    CN121189946B