Auxiliary decision-making method and system for safety operation center
By building a network topology diagram and risk assessment model, identifying potential attack chains and generating repair decisions, the problem of inefficiency in identifying and repairing complex network attack chains is solved, and the network security defense capabilities and economicality of repair decisions is improved.
Patent Information
- Application Number
- CN202510760105.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-09
AI Technical Summary
Traditional security operation centers find it difficult to efficiently identify complex and diverse network attack chains and perform precise repairs, resulting in insufficient network security defense capabilities.
By acquiring multi-source heterogeneous data, building network topology maps and risk assessment models, identifying potential attack chains, and generating repair reference standards and decisions, performing repair operations.
It realizes accurate identification and efficient repair of potential attack chains, improves network security defense capabilities, reduces missed detection rates and optimizes the economicality of repair decisions.
Smart Images

Figure CN120415879A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the fields of artificial intelligence and network security, and particularly to an auxiliary decision-making method and system for a security operation center. Background Art
[0002] With the rapid development of information technology, network security threats have become increasingly complex and diverse. Traditional security operation centers are facing huge challenges in real-time analysis of massive multi-source heterogeneous data and accurate identification of complex security events. Security events often present as a dynamic attack chain. Attackers gradually intrude, escalate privileges, move laterally, and finally achieve the goal of stealing or destroying data.
[0003] Therefore, it is necessary to propose an auxiliary decision-making method and system for a security operation center that can accurately and efficiently identify these potential attack chains and enhance network security defense capabilities. Summary of the Invention
[0004] One or more embodiments of this specification provide an auxiliary decision-making method for a security operation center. The method includes: obtaining multi-source heterogeneous data of a target network within a preset historical period; determining a potential attack chain based on the multi-source heterogeneous data; determining vulnerability data based on the multi-source heterogeneous data and the potential attack chain; generating at least one repair reference standard based on the vulnerability data, where the repair reference standard includes a repair method and a repair cost; generating a repair decision based on the vulnerability data and the at least one repair reference standard, and performing a repair based on the repair decision.
[0005] One embodiment of this specification provides an auxiliary decision-making system for a security operation center. The system includes: an obtaining module configured to obtain multi-source heterogeneous data of a target network within a preset historical period; an attack chain determination module configured to determine a potential attack chain based on the multi-source heterogeneous data; a vulnerability determination module configured to determine vulnerability data based on the multi-source heterogeneous data and the potential attack chain; a standard determination module configured to generate at least one repair reference standard based on the vulnerability data, where the repair reference standard includes a repair method and a repair cost; a repair module configured to generate a repair decision based on the vulnerability data and the at least one repair reference standard, and perform a repair based on the repair decision.
[0006] One or more embodiments of this specification provide an auxiliary decision-making device for a security operation center, including a processor, and the processor is used to execute the auxiliary decision-making method for a security operation center as described in any embodiment of this specification.
[0007] One or more embodiments of this specification provide a computer-readable storage medium storing computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the auxiliary decision-making method of the security operation center as described in any embodiment of this specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] This specification will be further described by way of exemplary embodiments, which will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where: Figure 1 is an exemplary module diagram of an auxiliary decision-making system for a security operation center shown in some embodiments of this specification; Figure 2 is an exemplary flowchart of an auxiliary decision-making method for a security operation center shown in some embodiments of this specification; Figure 3 is a schematic diagram of an attack chain identification model shown in some embodiments of this specification; Figure 4 is a schematic diagram of a related node prediction model shown in some embodiments of this specification; Figure 5 is a schematic diagram of a future attack prediction model shown in some embodiments of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0009] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some examples or embodiments of this specification. For those of ordinary skill in the art, without creative efforts, this specification can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the drawings represent the same structure or operation.
[0010] It should be understood that the "system", "device", "unit" and / or "module" used herein is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the said words can be replaced by other expressions.
[0011] As shown in this specification and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0012] In this specification, flowcharts are used to illustrate the operations performed by the system according to the embodiments of this specification. It should be understood that the preceding or subsequent operations do not necessarily have to be performed precisely in sequence. On the contrary, the steps can be processed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or more steps can be removed from these processes.
[0013] Figure 1 is an exemplary module diagram of an auxiliary decision-making system for a security operation center shown in some embodiments of this specification.
[0014] In some embodiments, the auxiliary decision-making system 100 of the security operation center may include an acquisition module 110, an attack chain determination module 120, a vulnerability determination module 130, a standard determination module 140, and a repair module 150.
[0015] In some embodiments, the acquisition module 110 is configured to acquire multi-source heterogeneous data of the target network within a preset historical period.
[0016] In some embodiments, the attack chain determination module 120 is configured to determine a potential attack chain based on the multi-source heterogeneous data.
[0017] In some embodiments, the attack chain determination module 120 is further configured to: establish a first network topology diagram based on the connection relationships among the internal servers, databases, and terminals of the target network; establish a hierarchical structure of a stream processing layer and a batch processing layer: the stream processing layer performs a preliminary screening on the multi-source heterogeneous data based on high-frequency data features and risk data features to obtain high-risk data; the batch processing layer determines the risk value of the high-risk data based on high-frequency data features or risk data features; generate a second network topology diagram based on the first network topology diagram and the high-risk data with a risk value greater than a risk threshold; determine a potential attack chain based on the second network topology diagram.
[0018] In some embodiments, the attack chain determination module 120 is further configured to: determine potential attack data corresponding to each high-risk data based on the high-risk data with a risk value greater than a risk threshold, where the potential attack data includes potential attack events, the confidence levels corresponding to the potential attack events, and attack timestamps; determine a second network topology diagram based on the first network topology diagram and the potential attack data, and the node features in the second network topology diagram include potential attack events and the confidence levels and attack timestamps corresponding to the potential attack events.
[0019] In some embodiments, the vulnerability determination module 130 is configured to determine vulnerability data based on the multi-source heterogeneous data and the potential attack chain.
[0020] In some embodiments, the standard determination module 140 is configured to generate at least one repair reference standard based on vulnerability data, and the repair reference standard includes a repair method and a repair cost.
[0021] In some embodiments, the repair module is configured to generate a repair decision based on the vulnerability data and at least one repair reference standard, and perform a repair based on the repair decision.
[0022] In some embodiments, the auxiliary decision-making system 100 of the security operation center may further include an adjustment module (not shown in the figure), and the adjustment module is configured to: determine the intrusion risk value and weak features of each node in the second network topology diagram based on the second network topology diagram, the high-frequency data features, and the risk data features; determine potential attack nodes in a future time period based on the second network topology diagram and the potential attack chain; generate a simulated attack data set with different attack degrees based on the intrusion risk value and the weak features of the potential attack nodes; perform a simulated attack on the target network based on the simulated attack data set to obtain a simulated attack result; and adjust the defense strategy based on the simulated attack result.
[0023] For a detailed description of each module, reference may be made to Figures 2 - 5 and its related description.
[0024] It should be understood that Figure 1 the system and its modules shown can be implemented in various ways. In some embodiments, Figure 1 the system and its modules shown can be implemented by a processor. The processor can process data and / or information obtained from other devices or system components. The processor can execute program instructions based on these data, information, and / or processing results to perform one or more functions described in this application. In some embodiments, the processor may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-chip processing device). By way of example only, the processor may include a central processing unit (CPU), a controller, a microprocessor, etc. or any combination of the above. In some embodiments, the processor may include multiple modules, and different modules may be used to execute different program instructions respectively.
[0025] It should be noted that the above description of the system and its modules is only for convenience of description and does not limit this specification to the scope of the examples given. It can be understood that for those skilled in the art, after understanding the principle of the system, various combinations of the modules may be made, or a subsystem may be formed and connected to other modules without departing from this principle. In some embodiments, Figure 1The acquisition module 110, the attack chain determination module 120, the vulnerability determination module 130, the standard determination module 140, and the repair module 150 disclosed in
[0026] Figure 2 is an exemplary flowchart of an auxiliary decision-making method for a security operation center according to some embodiments of this specification. As Figure 2 shown, process 200 includes the following steps. In some embodiments, process 200 may be executed by the auxiliary decision-making system 100 or the processor of the security operation center.
[0027] Step 210, acquire multi-source heterogeneous data of the target network within a preset historical period.
[0028] The preset historical period refers to a specific time range preset for analysis or comparison. The preset historical period can be preset manually.
[0029] The target network refers to the network environment that needs to be secured. In some embodiments, the target network may include the security operation center.
[0030] Multi-source heterogeneous data refers to a collection of multi-source data that is significantly different in terms of source, structure, or semantics.
[0031] In some embodiments, the multi-source heterogeneous data may include network traffic, terminal logs, external threat intelligence, etc. External threat intelligence refers to structured information about potential or existing network security threats obtained from third-party sources.
[0032] In some embodiments, the processor may acquire multi-source heterogeneous data in various ways. For example, the processor may perform real-time monitoring of the target network within the preset historical period to acquire multi-source heterogeneous data.
[0033] Step 220, determine a potential attack chain based on the multi-source heterogeneous data.
[0034] A potential attack chain refers to a series of concatenated potential attack events.
[0035] A potential attack event refers to a network activity that has not been confirmed but exhibits malicious behavior characteristics. For example, phishing emails, brute-force cracking attempts, data infiltration, etc.
[0036] In some embodiments, a potential attack chain may include a series of network activities such as port scanning - phishing attack - spreading attack - data exfiltration, port scanning - phishing attack - spreading attack, port scanning - vulnerability exploitation - spreading attack - data exfiltration, etc.
[0037] In some embodiments, the processor may determine a potential attack chain based on multi - source heterogeneous data in various ways.
[0038] In some embodiments, the processor may determine a potential attack chain based on multi - source heterogeneous data by matching in a first vector library.
[0039] The first vector library is a pre - set database including a plurality of first feature vectors and corresponding first vector labels. In some embodiments, the processor may convert multi - source heterogeneous data in historical data into structured feature vectors to construct the first feature vectors. In some embodiments, the processor may use the attack chain corresponding to the first feature vectors constructed according to historical data as the first vector labels. For example, in historical attack events, the processor may extract the attack chain and the multi - source heterogeneous data involved, convert this multi - source heterogeneous data into structured feature vectors as input features, and use the attack chain as the corresponding label.
[0040] In some embodiments, the processor may construct a first target vector based on the current multi - source heterogeneous data; match the first target vector in the first vector library to determine a potential attack chain. The processor may convert the current multi - source heterogeneous data into structured feature vectors to construct the first target vector. The processor may search in the first vector library for the first feature vector with the highest similarity (e.g., the smallest vector distance) to the first target vector. In response to the similarity of the first feature vector with the highest similarity being greater than a preset similarity threshold, the first vector label corresponding to the first feature vector is used as the potential attack chain. In response to the similarity of the first feature vector with the highest similarity being less than the preset similarity threshold, it may indicate to a certain extent that the probability of a potential attack is low, and the processor may terminate the process. Among them, the similarity may be calculated by cosine distance or Euclidean distance. The preset similarity threshold may be set according to experience.
[0041] In some embodiments, the processor may establish a first network topology diagram based on the connection relationships among the internal servers, databases, and terminals in the target network; establish a hierarchical structure of a stream processing layer and a batch processing layer: the stream processing layer performs a preliminary screening on multi - source heterogeneous data based on high - frequency data features and risk data features to obtain high - risk data; the batch processing layer determines the risk value of the high - risk data based on high - frequency data features or risk data features; generate a second network topology diagram based on the first network topology diagram and the high - risk data with a risk value greater than the risk threshold; determine a potential attack chain based on the second network topology diagram.
[0042] The first network topology diagram refers to a graph structure that reflects the devices in the target network and their connection relationships.
[0043] In some embodiments, the first network topology diagram includes nodes and edges. The nodes can represent servers, databases, terminals, etc. of the target network. The edges can represent the connection relationships or communication relationships between the nodes.
[0044] In some embodiments, the edges of the first network topology diagram can also include weights. The weights can be set according to the communication frequency between the nodes.
[0045] In some embodiments, the processor can identify the nodes in the target network, collect the communication data between the nodes to determine whether there is a connection relationship or communication relationship between the nodes, and the nodes with the connection relationship or communication relationship are connected based on the edges. The processor can calculate the weights of the edges according to the communication frequency, represent the nodes and the weighted edges in the form of a graph structure, and establish the first network topology diagram.
[0046] The stream processing layer refers to a lightweight processing layer for real-time data streams.
[0047] The batch processing layer refers to a deep analysis layer for historical or cumulative data.
[0048] There is no limitation here on how to construct the hierarchical structure of the stream processing layer and the batch processing layer.
[0049] High-frequency data features refer to the features of data that appear frequently within a certain period of time. In some embodiments, high-frequency data features can include high-frequency IP addresses, high-frequency trigger periods, high-frequency event types, etc. Among them, the high-frequency data features corresponding to data from different sources are different.
[0050] Risk data features refer to data attributes that reflect potential security threats or abnormal behaviors. In some embodiments, risk data features can include unfamiliar IP addresses, unconventional trigger periods, and infrequently used event types, etc. Among them, the risk data features corresponding to data from different sources are different.
[0051] High-risk data refers to data that meets the high-frequency data features or risk data features.
[0052] In some embodiments, the stream processing layer can perform a preliminary screening on multi-source heterogeneous data based on high-frequency data features and risk data features to obtain high-risk data.
[0053] In some embodiments, the stream processing layer may obtain high-frequency data features and risk data features based on the statistical analysis of historical data from different sources. For example, the stream processing layer may acquire historical data from different sources, count the IPs in the historical data of each source, determine the average number of occurrences of the IPs, and identify the IPs with the number of occurrences greater than the average as high-frequency IP addresses.
[0054] In some embodiments, the stream processing layer may classify multi-source heterogeneous data according to different sources to obtain multiple groups of heterogeneous data corresponding to different sources. For each group of heterogeneous data, the stream processing layer may determine the heterogeneous data that meets the high-frequency data features or risk data features as high-risk data.
[0055] For example, for each group of heterogeneous data, the stream processing layer may determine the data in which the IP data, trigger period, and event type in the group of heterogeneous data simultaneously meet the high-frequency IP address, high-frequency trigger period, and high-frequency event type in the high-frequency data features as high-risk data. For another example, for each group of heterogeneous data, the stream processing layer may determine the data in which the IP data, trigger period, and event type in the group of heterogeneous data simultaneously meet the unfamiliar IP address, unconventional trigger period, and infrequently used event type in the risk data features as high-risk data.
[0056] In some embodiments, the processor may obtain the missed detection rate and missed detection attack events after each regular execution of attack events; based on the missed detection rate and missed detection attack events, update the high-frequency data features and risk data features.
[0057] Regular execution of attack events means regularly conducting simulated attack events.
[0058] The missed detection rate refers to the proportion of attack events that are not correctly identified in the total number of actual attack events.
[0059] A missed detection attack event refers to a real attack event that is not detected by the system.
[0060] In some embodiments, the processor may compare the missed detection rate of the missed detection attack event with a preset missed detection threshold to update the high-frequency data features and risk data features. For example, when the missed detection rate of a certain missed detection attack event is greater than the preset missed detection rate threshold, the high-frequency data features and risk data features of the missed detection attack event may be appropriately reduced. For example, the IPs with the number of occurrences greater than the average being determined as high-frequency IP addresses are updated to the IPs with the number of occurrences greater than half of the average being determined as high-frequency IP addresses.
[0061] In the embodiments of this specification, by dynamically adjusting the high-frequency data features and risk data features, the occurrence of missed detection events can be reduced, and the security of the target network can be improved.
[0062] The risk value refers to a numerical value that evaluates the likelihood of a risk event occurring.
[0063] In some embodiments, the batch processing layer may determine the risk value of high-risk data in various ways based on high-frequency data features or risk data features.
[0064] In some embodiments, the batch processing layer may determine the risk value through a first preset table based on high-frequency data features or risk data features. The first preset table may include different data features (including high-frequency data features and risk data features), event types corresponding to different data features, and risk values corresponding to the event types. Among them, the event types of high-frequency data features may include high-frequency data transmission, frequent account logins, etc., and the event types of risk data features may include overseas account logins, large-scale data transmission, etc.
[0065] In some embodiments, the batch processing layer may determine the risk value of an event type by the number of actually occurred attack events corresponding to the event type in historical data. For example, for each event type, the batch processing layer may divide the number of actually occurred attack events corresponding to the event type by the total number of actually occurred attack events to obtain the risk value of the event type.
[0066] In some embodiments, the risk threshold may be preset according to manual experience.
[0067] The second network topology diagram refers to the graph structure obtained by modifying the first network topology diagram.
[0068] In some embodiments, the processor may generate a second network topology diagram in various ways based on the first network topology diagram and high-risk data with a risk value greater than the risk threshold.
[0069] In some embodiments, the processor may generate a second network topology diagram through a second vector library based on the first network topology diagram and high-risk data with a risk value greater than the risk threshold.
[0070] The second vector library is a preset database including multiple second feature vectors and corresponding second vector labels. In some embodiments, the processor may construct the first network topology diagram and high-risk data with a risk value greater than the risk threshold in historical data into second feature vectors. In some embodiments, the processor may use the second network topology diagram corresponding to the second feature vectors constructed according to historical data as the second vector label.
[0071] For example, in historical attack events, the processor can manually label the attack event, the timestamp of the attack event, and the confidence level as the attribute features of the relevant nodes in the first network topology diagram. The labeled diagram can be used as the second network topology diagram corresponding to the first network topology diagram and as the second vector label. The processor can determine the first network topology diagram and the multi-source heterogeneous data corresponding to the attack event as the first network topology diagram and the high-risk data in the second feature vector. The confidence level can be determined based on the frequency of the attack event causing the node failure in historical data.
[0072] In some embodiments, the processor can construct a second target vector based on the current first network topology diagram and high-risk data with a risk value greater than the risk threshold; determine the second network topology diagram based on the similarity between the second target vector and multiple second feature vectors in the vector database.
[0073] The processor can use the second vector label corresponding to the second feature vector with the highest similarity to the second target vector as the second network topology diagram. The similarity includes cosine similarity.
[0074] In some embodiments, the processor can determine the potential attack data corresponding to each high-risk data based on the high-risk data with a risk value greater than the risk threshold. The potential attack data includes potential attack events, the confidence level corresponding to the potential attack events, and the attack timestamp; determine the second network topology diagram based on the first network topology diagram and the potential attack data. The node features in the second network topology diagram include potential attack events and the confidence level and attack timestamp corresponding to the potential attack events.
[0075] In some embodiments, not every node in the second network topology diagram must include node features, and one node can include multiple node features.
[0076] In some embodiments, the potential attack events can include multiple potential attack events belonging to the same attack behavior. For example, the potential attack events can include port scanning, phishing attacks, vulnerability exploitation, spreading attacks, data exfiltration, etc.
[0077] Potential attack data refers to the data related to potential attack events. In some embodiments, the potential attack data can include potential attack events, the confidence level corresponding to the potential attack events, and the attack timestamp.
[0078] In some embodiments, the processor can determine the potential attack data corresponding to each high-risk data based on the high-risk data with a risk value greater than the risk threshold through a large language model.
[0079] A large language model refers to a deep learning model trained based on a vast amount of text data, capable of understanding, generating, and reasoning natural language. In some embodiments, the large language model can be an existing trained model such as the GPT series models, BERT, etc.
[0080] In some embodiments, to determine potential attack data, the input of the large language model can include high-risk data and input text, and the output of the large language model is determined based on the input text. The output of the large language model can be used as potential attack data.
[0081] Among them, the high-risk data can include multi-source heterogeneous data such as network traffic, terminal logs, external threat intelligence, etc. The input text refers to the instructions or questions input by the user into the model, used to guide the model to generate the required output. In some embodiments, the input text can include specific content to be parsed, specific requirements for the output, knowledge domains to be considered, output formats, etc.
[0082] Only as an example, the input text can include: Analyze the input high-risk data, parse the content in the data, comprehensively consider each event data and the occurrence time in different files, combine the relevant knowledge in the field of network attacks, determine the possible network attack events, the timestamp of the event occurrence, and the confidence level of the network attack event in the uploaded data. The output can be in JSON format and needs to contain the following fields (where timestamp, event_type, and confidence_coefficient are required content, and fill in "null" for missing content): - `ip`: IPv4 address, in dotted decimal format.
[0083] - `timestamp`: Timestamp format (milliseconds).
[0084] - `event_type`: Event type (such as "brute-force login", "phishing attack").
[0085] - `confidence_coefficient`: Confidence level of the occurrence of this attack event type.
[0086] - `context`: Briefly describe the event context.
[0087] - `source`: Data source (such as "firewall log").
[0088] In some embodiments, the processor can determine a second network topology map through the large language model based on the first network topology map and potential attack data.
[0089] In some embodiments, to determine the second network topology diagram, the input to the large language model may include the first network topology diagram, potential attack data, and input text, and the output of the large language model is determined based on the input text. Among them, the first network topology diagram and the potential attack data may be in Json format. The output of the large language model may be used as the second network topology diagram.
[0090] By way of example only, the input text may include: 1. Analyze the Json file of the uploaded first network topology diagram and the Json file of the potential attack data, and combine the relevant knowledge in the field of network attacks and the relevant knowledge in the field of network topology diagrams to determine the nodes of the first network topology diagram corresponding to each piece of potential attack data; 2. Determine this piece of potential attack data as the feature of the node. Note that a node may contain multiple features, and the file is output in Json format; 3. The output content includes: Based on the Json file of the first network topology diagram, each node adds a 'characteristic' field, which is used to store the features of the node. Multiple features are separated by semicolons. If a node has no features, "null" is filled in the 'characteristic' field.
[0091] In the embodiments of this specification, constructing the second network topology diagram through the large language model can construct a diagram containing potential attack data, reduce manual participation, and improve the construction efficiency and accuracy.
[0092] In some embodiments, the processor may determine potential attack chains in various ways based on the second network topology diagram.
[0093] In some embodiments, the processor may determine multiple potential attack chains based on the second network topology diagram through an attack chain recognition model.
[0094] The attack chain determination model refers to a model used to identify potential attack chains. In some embodiments, the attack chain recognition model may be a machine learning model. For example, a Graph Neural Network (GNN) model, etc.
[0095] Figure 3 It is a schematic diagram of the attack chain recognition model shown in some embodiments of this specification. In some embodiments, as Figure 3 shown, the input of the attack chain recognition model 320 includes the second network topology diagram 310, and the output includes potential attack chains 330.
[0096] In some embodiments, the processor may train an attack chain recognition model based on multiple first training samples with a first label. For example, the processor may input multiple first training samples into an initial attack chain recognition model, construct a loss function based on the output of the initial attack chain recognition model and the first label, iteratively update the parameters of the initial attack chain recognition model based on the first loss function, and end the iteration when the iteration completion condition is met to obtain a trained attack chain recognition model. Among them, the method of iterative update includes but is not limited to the gradient descent method, and the iteration completion condition may be that the first loss function converges or the number of iterations reaches a threshold.
[0097] In some embodiments, the first training sample may include a sample second network topology diagram. The first label is the attack chain corresponding to the sample second network topology diagram.
[0098] In some embodiments, the processor may determine the first training sample through historical data. For example, the processor may use the second network topology diagram corresponding to the actually occurred attack chain in the historical data as the first training sample, and the corresponding attack chain as the first label.
[0099] In some embodiments of the present specification, screening through the flow processing layer to obtain high-risk data can reduce the amount of multi-source data, thereby improving the efficiency of subsequent identification of potential attack chains and improving real-time performance.
[0100] Step 230, determine vulnerability data based on the multi-source heterogeneous data and the potential attack chain.
[0101] Vulnerability data refers to security defects or errors existing in the target network.
[0102] In some embodiments, the processor may perform a preliminary analysis on the multi-source heterogeneous data through a vulnerability scanning tool to generate an initial vulnerability set. For example, the vulnerability scanning tool may include Nessus, OpenVAS, etc.
[0103] In some embodiments, the processor may utilize the context of the potential attack chain (for example, each attack stage and the techniques utilized in the potential attack chain, etc.) to screen out the vulnerabilities related to the potential attack chain from the initial vulnerability set and determine them as vulnerability data. For example, if SQL injection is involved in the potential attack chain, the processor may preferentially screen the data related to the database as vulnerability data.
[0104] Step 240, generate at least one repair reference standard based on the vulnerability data.
[0105] The repair reference standard refers to the standard for repairing the vulnerability data. In some embodiments, the repair reference standard may include a repair method and a repair cost.
[0106] In some embodiments, the processor may generate at least one repair reference standard based on vulnerability data through a vulnerability processing data set.
[0107] The vulnerability processing data set is a data set including vulnerabilities and corresponding repair methods and repair costs. In some embodiments, the vulnerability processing data set may include multiple pieces of vulnerability processing data, and each piece of vulnerability processing data includes a type of vulnerability and a corresponding repair method and repair cost.
[0108] In some embodiments, a vulnerability may have multiple repair methods.
[0109] In some embodiments, the types of vulnerabilities may include software vulnerabilities, configuration errors, weak password policies, etc. The corresponding repair methods for vulnerabilities of corresponding types may include: for software vulnerabilities, the repair methods include upgrading through official patches, network layer interception, etc.; for configuration errors, the repair methods include tightening permissions, data encryption, monitoring warnings, etc.
[0110] In some embodiments, the vulnerability processing data set may be obtained based on historical data statistics.
[0111] In some embodiments, the repair cost may include computing resource cost and time cost, both of which can be determined through presetting.
[0112] In some embodiments, the processor may query all the corresponding repair costs and repair methods for each vulnerability in the vulnerability data in the vulnerability processing data set and determine them as repair reference standards.
[0113] In some embodiments, the repair cost may further include business interruption cost, and the processor may generate a repair reference standard based on the vulnerability processing data set, the business event data set, and the network topology diagram.
[0114] Regarding the network topology diagram, reference may be made to the first network topology diagram or the second network topology diagram in step 220 and their related descriptions.
[0115] Business interruption cost refers to the direct or indirect economic losses caused by system downtime, service degradation, or function limitation during the process of repairing vulnerabilities. The business interruption cost can be preset.
[0116] The business event data set refers to a data set including multiple pieces of business event data. In some embodiments, each piece of business event data includes a type of business event and a flow path of the business event. The flow path refers to the servers, databases, terminals, etc. that a business will pass through. A type of business event may correspond to multiple flow paths. Each piece of business event data is marked with a business interruption cost. Among them, the flow path can be obtained based on historical data statistics.
[0117] In some embodiments, the processor may, based on vulnerability data, determine, by querying a vulnerability handling data set, the repair methods and repair costs corresponding to each vulnerability in the vulnerability data; determine whether the repair method involves hardware interruption repair; in response to the repair method including hardware interruption repair, determine, by querying a network topology map, the devices to be interrupted and the business events to be interrupted; based on each business event to be interrupted, determine, by querying a business event data set, the business interruption cost corresponding to each business; and use the repair method, the repair cost corresponding to the repair method, and the business interruption cost caused by adopting the repair method as a repair reference standard for the vulnerability data.
[0118] In the embodiments of this specification, by quantitatively evaluating the business interruption cost, the economic analysis of the repair plan is optimized, thereby enhancing the comprehensive benefit of the repair decision and ensuring the efficient handling of security vulnerabilities on the premise of minimizing business impact.
[0119] Step 250: Generate a repair decision based on the vulnerability data and at least one repair reference standard, and perform a repair based on the repair decision.
[0120] A repair decision refers to the finally determined repair strategy.
[0121] In some embodiments, the processor may generate a repair decision based on the vulnerability data and at least one repair reference standard through a large language model. Among them, the repair reference standard may include multiple corresponding repair methods and repair costs for each vulnerability. The processor may determine a suitable repair method for each vulnerability based on the large language model and finally generate a repair decision.
[0122] In some embodiments, in order to generate a repair decision, the input of the large language model may include vulnerability data, a repair reference standard, and input text, and the output of the large language model is determined based on the input text. The output of the large language model may be used as a repair decision. Among them, the repair reference standard includes specific repair methods and repair costs. In some embodiments, the input text may include the required output content, selection conditions, and the format of the output content, etc.
[0123] Merely by way of example, the input text may include: Please give an optimal repair decision based on the uploaded vulnerability data and the repair reference standards that can be adopted for each vulnerability data, and requirements: 1. The repair decision needs to include the repair of each vulnerability; 2. The total repair cost is the lowest; 3. Output in the form of string points, and each point corresponds to a repair method in the repair decision and the vulnerability that the method can solve.
[0124] In the embodiments of this specification, by determining potential attack points and vulnerability data from multi-source heterogeneous data, and then generating repair reference criteria and repair decisions, it is possible to accurately and efficiently identify these potential attack chains and automatically generate repair decisions, thereby enhancing network security defense capabilities.
[0125] In some embodiments, the processor can also determine the intrusion risk value and weak features of each node in the second network topology diagram based on the second network topology diagram, high-frequency data features, and risk data features; determine potential attack nodes in a future time period based on the second network topology diagram and potential attack chains; generate simulated attack data sets of different attack degrees based on the intrusion risk values and weak features of the potential attack nodes; perform simulated attacks on the target network based on the simulated attack data sets to obtain simulated attack results; and adjust the defense strategy based on the simulated attack results.
[0126] The intrusion risk value refers to a quantitative evaluation value of the possibility of this node being attacked.
[0127] Weak features refer to the features of potential security risks in this node that may lead to intrusion. In some embodiments, weak features may include weak passwords, open ports, etc.
[0128] In some embodiments, the processor can obtain different event types included in the high-frequency data features and risk data features; determine the influence weights of different event types on different nodes based on the second network topology diagram and event types through a relevant node prediction model; query the risk values corresponding to the event types based on a first preset table; determine the intrusion risk value of each node in the second network topology diagram; and determine the weak features of the node based on the event types affecting this node through a second preset table.
[0129] For more information about the event types of high-frequency data features and risk data features, please refer to step 220 and its related descriptions.
[0130] A relevant node prediction model refers to a model used to predict the influence weights of different event types on different nodes. In some embodiments, the relevant node prediction model can be a machine learning model. For example, a Recurrent Neural Network (RNN) model, etc.
[0131] Figure 4 It is a schematic diagram of a relevant node prediction model shown in some embodiments of this specification. In some embodiments, as Figure 4 shown, the input of the relevant node prediction model 420 includes the second network topology diagram 410 and the event type 411, and the output can include the influence weights 430 of different event types on different nodes in the second network topology diagram.
[0132] In some embodiments, the processor may train a relevant node prediction model based on multiple second training samples with second tags. The training process of the relevant node prediction model is similar to that of the attack chain recognition model.
[0133] In some embodiments, the second training samples may include a sample second network topology diagram and a sample event type. The second tag is the influence weight of the sample event type on different nodes in the sample second network topology diagram.
[0134] In some embodiments, the processor may determine the second training samples and second tags through historical data. For example, the processor may obtain, from the historical data, for different event types, when an attack of the same event type occurs multiple times, the nodes affected after the attack (such as servers, terminals, etc.), and the number of times the nodes are affected during multiple attacks. The processor may use the second network topology diagram and event type corresponding to the historical data as the second training samples, determine the relevant nodes as the affected nodes, and use the frequency of node influence as the second tag.
[0135] For more content about the first preset table, reference can be made to step 220 and its related descriptions.
[0136] In some embodiments, for each node of the second network topology diagram, the processor may determine the intrusion risk value of the node based on the weight of the event type affecting the node. For example, the processor may determine the intrusion risk value of the node according to the following formula (1): (1) Where, represents the intrusion risk value of the node, , , …… and are respectively the influence weights of event types numbered 1, 2, …… n on the node, , , …… and are respectively the risk values corresponding to event types numbered 1, 2, …… n, and the event types numbered 1, 2, …… n are respectively the event types that will affect the node.
[0137] The second preset table may include different event types and the weak features respectively corresponding to each event type. In some embodiments, different event types may include event types in high-frequency data features and risk data features (such as high-frequency account logins, strange IPs, etc.), and weak features may include lack of data encryption mechanism, unmonitored abnormal traffic, storage or transmission protocol vulnerabilities, weak authentication mechanism, overdevelopment of account permissions, etc. In some embodiments, the second preset table may be constructed based on artificial experience.
[0138] In some embodiments, the processor may look up the weak features corresponding to the event types affecting the node in the second preset table as the weak features of the node.
[0139] In some embodiments, the processor may determine potential attack nodes in a future time period through a future attack prediction model based on the second network topology graph and the potential attack chain.
[0140] The future time period may be preset.
[0141] The future attack prediction model refers to a model used to predict potential attack nodes in a future time period. In some embodiments, the future attack prediction model may be a machine learning model. For example, a GNN model, etc.
[0142] Figure 5 It is a schematic diagram of the future attack prediction model shown in some embodiments of this specification. In some embodiments, as Figure 5 shown, the inputs of the future attack prediction model 520 include the second network topology graph 510 and the potential attack chain 511, and the output may include potential attack nodes 530 in a future time period.
[0143] In some embodiments, the processor may train a future attack prediction model based on multiple third training samples with a third label. The training process of the future attack prediction model is similar to that of the attack chain recognition model. In some embodiments, the future attack prediction model may be trained alone or jointly trained with the future attack prediction model.
[0144] In some embodiments, the third training sample may include a sample second network topology graph and a sample potential attack chain. The third label is the potential attack node corresponding to the sample second network topology graph and the sample potential attack chain.
[0145] In some embodiments, the processor may determine the third training sample and the third label through historical data. For example, after the processor determines the second network topology graph and the potential attack chain from the historical data, when the next attack occurs, it determines the node under attack. The processor may use the second network topology graph and the potential attack chain as the third training sample, and the corresponding node under attack as the third label. Among them, the node under attack, that is, the third label, may be one or more.
[0146] The simulated attack dataset refers to a set of codes for simulating attacks on the system.
[0147] In some embodiments, the processor may screen out nodes whose intrusion risk values of potential attack nodes in a future time period are greater than a risk threshold; based on the screened-out nodes, determine a simulated attack dataset through a third preset table.
[0148] Among them, the risk threshold can be set according to experience. In some embodiments, the risk threshold is related to the number of potential attack nodes in a future time period. The more the number of potential attack nodes, the lower the risk threshold.
[0149] The third preset table may include nodes (such as servers, terminals, etc.), the weak features of the nodes (such as, a single weak feature or a combination of weak features), different attack degrees (such as, mild, moderate, severe) for a certain weak feature of the nodes, and corresponding simulated attack data sets. A node may correspond to multiple different combinations of weak features. In some embodiments, the third preset table may be constructed according to manual experience.
[0150] In some embodiments, the processor may look up the simulated attack data set corresponding to the filtered nodes in the third preset table.
[0151] In some embodiments, the simulated attack result may include the vulnerability type, the number of vulnerabilities, and the nodes where the vulnerabilities occur.
[0152] In some embodiments, the simulated attack may include: vulnerability exploitation attack, lateral movement attack, denial-of-service attack, data leakage simulation, etc. After the simulated attack is completed, the processor may obtain the simulated attack result through network traffic capture, operating system logs, security monitoring tools, etc.
[0153] The defense strategy refers to a series of proactive or passive security measures taken to protect the target network from attacks.
[0154] In some embodiments, when the simulated attack result is that a port is breached, the defense strategy may include restricting access to the IP whitelist, port traffic encryption, etc.; when the simulated attack result is password theft, the defense strategy may include enforcing password complexity, deploying multi-factor authentication, etc.
[0155] In some embodiments, the processor may determine the defense strategy based on the simulated attack result through the fourth preset table.
[0156] The fourth preset table may include the vulnerability type, the number of vulnerabilities, the nodes where the vulnerabilities occur, and the corresponding defense strategy. In some embodiments, the fourth preset table may be constructed by effective measures refined from historical successful defense records. Among them, the historical successful defense record refers to the record that after a vulnerability appears and is defended against, the vulnerability does not appear again within 3 months. The processor may use the defense strategies corresponding to all vulnerability types as the final defense strategy.
[0157] In the embodiments of this specification, by analyzing the second network topology, real-time data, and risk characteristics, actively identifying attack nodes and weak nodes at future time points, and generating simulated attacks, a closed-loop defense of "risk prediction - attack verification - strategy optimization" can be achieved, effectively discovering security blind spots in advance, dynamically optimizing protection strategies, and enhancing the network's active defense capabilities.
[0158] It should be noted that the above description of process 200 is only for illustration and example, and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to process 200 under the guidance of this specification. However, these modifications and changes are still within the scope of this specification.
[0159] One or more embodiments of this specification also provide an auxiliary decision-making device for a security operation center, including a processor, and the processor is used to execute the auxiliary decision-making method for the security operation center as described in any embodiment of this specification.
[0160] One or more embodiments of this specification also provide a computer-readable storage medium, and the storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the auxiliary decision-making method for the security operation center as described in any embodiment of this specification.
[0161] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only an example and does not constitute a limitation to this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are proposed in this specification, so such modifications, improvements, and corrections still belong to the spirit and scope of the exemplary embodiments of this specification.
[0162] At the same time, this specification uses specific terms to describe the embodiments of this specification. Such as "one embodiment", "an embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "an embodiment" or "one embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0163] Further, unless explicitly stated in the claims, the order of process elements and sequences, the use of numerical and alphabetical characters, or the use of other names in this specification are not used to limit the order of the processes and methods in this specification. Although some currently useful embodiments of the invention are discussed through various examples in the above disclosure, it should be understood that such details are for illustrative purposes only. The appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that conform to the essence and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only through software solutions, such as installing the described system on existing servers or mobile devices.
[0164] Similarly, it should be noted that, in order to simplify the presentation of the disclosure in this specification and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of this specification, sometimes multiple features are merged into one embodiment, drawing, or description thereof. However, this method of disclosure does not mean that the features required by the subject matter of this specification are more than those mentioned in the claims. In fact, the features of the embodiments are fewer than all the features of the individual embodiments disclosed above.
[0165] In some embodiments, numbers are used to describe the components and the quantity of attributes. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximate" or "substantially" in some examples. Unless otherwise stated, "about", "approximate" or "substantially" indicate that the said numbers allow a variation of ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, and such approximate values may vary according to the characteristics required by individual embodiments. In some embodiments, the numerical parameters should consider the specified significant digits and adopt the method of retaining the general number of digits. Although the numerical ranges and parameters used in some embodiments of this specification to confirm the breadth of their scope are approximate values, in specific embodiments, such numerical settings are as precise as possible within the feasible range.
[0166] For each patent, patent application, patent application publication, and other materials cited in this specification, such as articles, books, specifications, publications, documents, etc., their entire contents are hereby incorporated into this specification by reference. Except for the application history documents that are inconsistent with or conflict with the content of this specification, and also except for the documents that limit the broadest scope of the claims of this specification (currently or subsequently appended to this specification). It should be noted that if there are inconsistencies or conflicts between the descriptions, definitions, and / or the use of terms in the supplementary materials of this specification and the content described in this specification, the descriptions, definitions, and / or the use of terms in this specification shall prevail.
[0167] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this specification may be regarded as consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly presented and described in this specification.
Claims
1. An auxiliary decision-making method for a security operation center, characterized in that, Including: Obtain multi-source heterogeneous data of the target network within a preset historical period; Based on the multi-source heterogeneous data, determine potential attack chains; Based on the multi-source heterogeneous data and the potential attack chains, determine vulnerability data; Based on the vulnerability data, generate at least one repair reference standard, where the repair reference standard includes a repair method and a repair cost; Based on the vulnerability data and the at least one repair reference standard, generate a repair decision and perform a repair based on the repair decision.
2. The method according to claim 1, characterized in that, The determining of potential attack chains based on the multi-source heterogeneous data includes: Based on the connection relationships among the internal servers, databases, and terminals of the target network, establish a first network topology diagram; Establish a hierarchical structure of a stream processing layer and a batch processing layer: The stream processing layer performs a preliminary screening on the multi-source heterogeneous data based on high-frequency data characteristics and risk data characteristics to obtain high-risk data; The batch processing layer determines the risk value of the high-risk data based on the high-frequency data characteristics or the risk data characteristics; Based on the first network topology diagram and the high-risk data with a risk value greater than a risk threshold, generate a second network topology diagram; Based on the second network topology diagram, determine the potential attack chains.
3. The method according to claim 2, wherein The generating of the second network topology diagram based on the first network topology diagram and the high-risk data with a risk value greater than a risk threshold includes: Based on the high-risk data with a risk value greater than the risk threshold, determine the potential attack data corresponding to each high-risk data, where the potential attack data includes potential attack events, the confidence levels corresponding to the potential attack events, and attack timestamps; Based on the first network topology diagram and the potential attack data, determine the second network topology diagram, where the node characteristics in the second network topology diagram include the potential attack events and the confidence levels and attack timestamps corresponding to the potential attack events.
4. The method according to claim 2, characterized in that, The method further includes: Based on the second network topology diagram, the high-frequency data characteristics, and the risk data characteristics, determine the intrusion risk value and weak characteristics of each node in the second network topology diagram; Based on the second network topology diagram and the potential attack chains, determine potential attack nodes in a future time period; Based on the intrusion risk value and the weak characteristics of the potential attack nodes, generate simulated attack data sets with different attack degrees; Based on the simulated attack data sets, perform a simulated attack on the target network to obtain simulated attack results; Based on the simulated attack results, adjust the defense strategy.
5. An auxiliary decision-making system for a security operation center, characterized in that, Including: An acquisition module configured to obtain multi-source heterogeneous data of the target network within a preset historical period; An attack chain determination module configured to determine potential attack chains based on the multi-source heterogeneous data; A vulnerability determination module configured to determine vulnerability data based on the multi-source heterogeneous data and the potential attack chains; A standard determination module configured to generate at least one repair reference standard based on the vulnerability data, where the repair reference standard includes a repair method and a repair cost; A repair module configured to generate a repair decision based on the vulnerability data and the at least one repair reference standard and perform a repair based on the repair decision.
6. The system according to claim 5, wherein The attack chain determination module is further configured to: Establish a first network topology diagram based on the connection relationships among the internal servers, databases, and terminals in the target network; Establish a hierarchical structure of a stream processing layer and a batch processing layer: The stream processing layer performs a preliminary screening on the multi-source heterogeneous data based on high-frequency data characteristics and risk data characteristics to obtain high-risk data; The batch processing layer determines the risk value of the high-risk data based on the high-frequency data characteristics or the risk data characteristics; Generate a second network topology diagram based on the first network topology diagram and the high-risk data with a risk value greater than a risk threshold; Determine the potential attack chain based on the second network topology diagram.
7. The system according to claim 6, wherein The attack chain determination module is further configured to: Determine potential attack data corresponding to each piece of the high-risk data with a risk value greater than the risk threshold, where the potential attack data includes potential attack events, the confidence levels corresponding to the potential attack events, and attack timestamps; Determine the second network topology diagram based on the first network topology diagram and the potential attack data, where the node characteristics in the second network topology diagram include the potential attack events and the confidence levels and attack timestamps corresponding to the potential attack events.
8. The system according to claim 6, wherein The system further includes an adjustment module, and the adjustment module is configured to: Determine the intrusion risk value and weak characteristics of each node in the second network topology diagram based on the second network topology diagram, the high-frequency data characteristics, and the risk data characteristics; Determine potential attack nodes in a future time period based on the second network topology diagram and the potential attack chain; Generate simulated attack data sets with different attack levels based on the intrusion risk values and weak characteristics of the potential attack nodes; Perform a simulated attack on the target network based on the simulated attack data sets to obtain simulated attack results; Adjust the defense strategy based on the simulated attack results.
9. An auxiliary decision-making device for a security operation center, including a processor, where the processor is used to execute the auxiliary decision-making method for the security operation center according to any one of claims 1 to 4.
10. A computer-readable storage medium, where the storage medium stores computer instructions, and when a computer reads the computer instructions in the storage medium, the computer executes the auxiliary decision-making method for the security operation center according to any one of claims 1 to 4.
Citation Information
Patent Citations
Automatic host computer security configuration vulnerability restoration method and system based on configuration specification
CN104346574A
Network security vulnerability mining method and device
CN112187773A
Composite attack chain completion method and system based on multi-modal data model, and medium
CN115883218A
RASP-based information system vulnerability elimination and control method and system, terminal and storage medium
CN117579305A
Security monitoring alarm device and method for network security vulnerabilities
CN120074857A
Cited By
Method and system for generating open source component repair opinions and storage medium
CN120611388A
Dos attack detection method and defense device for heterogeneous multi-core consistency master node
CN122348864A