Big model-based network security and data security comprehensive analysis method and system

By collecting and processing multi-source heterogeneous security data, a dynamic risk association hypergraph is constructed and a large model is used for risk attribution reasoning. This solves the problem of incomplete risk association in existing technologies, realizes the accurate capture of collaborative risk behavior of multiple entity nodes and the generation of dynamic protection strategies, and improves the accuracy of risk attribution and the adaptability of protection measures.

CN120825344BActive Publication Date: 2025-11-25贵州华谊联盛科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511320198.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2025-11-25
Estimated Expiration
2045-09-16

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively correlate potential risk relationships between data from different sources when processing multi-source heterogeneous security data. This results in an incomplete characterization of collaborative risk behaviors among multiple entity nodes and an inability to accurately trace the primary risk node and complete associated risk paths during risk attribution. Consequently, the generated protection strategies lack specificity and dynamic adaptability.

Method used

Collect multi-source heterogeneous security data from network communication links, data storage nodes, and application interaction interfaces, perform atomized correlation processing of risk behaviors, construct a dynamic risk correlation hypergraph, use a large model to perform multi-round risk attribution reasoning, generate a risk attribution reasoning chain containing principal risk nodes, associated risk paths, and evidence confidence, and construct a risk evolution probability model to calculate the short-term diffusion probability and long-term evolution trend vector of each associated risk path.

Benefits of technology

It achieves accurate capture of collaborative risk behaviors among multiple entity nodes, improves the accuracy of risk attribution, generates dynamic and highly adaptable protection strategies, can cope with incomplete attack chains or new attack patterns, and improves the pertinence and dynamic adaptability of protection measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120825344B_ABST
    Figure CN120825344B_ABST
Patent Text Reader

Abstract

The application provides a network security and data security comprehensive analysis method and system based on a large model, acquires multi-source heterogeneous security data of network communication links, data storage nodes and application interaction interfaces, performs risk behavior atomization correlation processing on the multi-source heterogeneous security data, constructs a dynamic risk correlation hypergraph, calls a large model to perform multiple rounds of risk attribution reasoning based on a preset security field knowledge graph, performs attack chain segment matching and evidence chain completion on a behavior hyperedge set in the dynamic risk correlation hypergraph, generates a risk attribution reasoning chain, constructs a risk evolution probability model according to a time sequence constraint set of the risk attribution reasoning chain and the dynamic risk correlation hypergraph, and calculates a short-term diffusion probability and a long-term evolution trend vector of each correlated risk path based on the risk evolution probability model, to obtain a risk evolution path prediction result. The application can improve the pertinence and dynamic adaptability of protection measures, and solves the hysteresis problem that a static protection strategy is difficult to cope with dynamic changes in risks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data analysis and machine learning, in particular, to a network security and data security comprehensive analysis method and system based on a large model. BACKGROUND

[0002] In the digital era, network security and data security comprehensive analysis plays an important role in ensuring the stable operation of information systems. Its core goal is to identify potential risks and take targeted protection measures by analyzing various security data in the network environment. Currently, common network security and data security comprehensive analysis methods usually collect network traffic, system logs and other security data, analyze the data through rule matching or simple models, identify abnormal behavior and assess risk levels. However, existing technologies have difficulty in effectively associating potential risk relationships between different source data when processing multi-source heterogeneous security data, resulting in incomplete characterization of multi-entity node collaborative risk behavior. In the risk attribution process, due to the fragmentation of attack chains or the emergence of new attack patterns, it is difficult to accurately trace the root cause risk node and completely associate the risk path, resulting in a lack of targetedness and dynamic adaptability of the generated protection strategies. SUMMARY

[0003] The present application provides a network security and data security comprehensive analysis method and system based on a large model.

[0004] In a first aspect, the present application provides a network security and data security comprehensive analysis method based on a large model, comprising: collecting multi-source heterogeneous security data of network communication links, data storage nodes and application interaction interfaces, the multi-source heterogeneous security data including encrypted transmission data packet units, data access audit units and interface call sequence units, each data unit carrying a node identifier and an interaction time stamp; performing risk behavior atomization association processing on the multi-source heterogeneous security data, and constructing a dynamic risk association hypergraph; based on a pre-set security domain knowledge graph, calling a large model to perform multiple rounds of risk attribution reasoning, matching attack chain fragments and completing evidence chains for a set of behavior hyperedges in the dynamic risk association hypergraph, and generating a risk attribution reasoning chain containing root cause risk nodes, associated risk paths and evidence confidence; constructing a risk evolution probability model according to a set of time sequence constraints of the risk attribution reasoning chain and the dynamic risk association hypergraph, and calculating a short-term diffusion probability and a long-term evolution trend vector of each associated risk path based on the risk evolution probability model, to obtain a risk evolution path prediction result.

[0005] In a second aspect, the present application provides a computer system, comprising: a memory having a computer program stored therein; a processor for loading the computer program to implement the network security and data security comprehensive analysis method based on a large model as described above.

[0006] This invention provides a comprehensive network security and data security analysis method based on a large model. By collecting multi-source heterogeneous security data from network communication links, data storage nodes, and application interaction interfaces, the method ensures that the acquired data includes encrypted transmission data packet units, data access audit units, and interface call sequence units, and carries node identifiers and interaction time stamps, providing a comprehensive and relevant data foundation for subsequent security analysis. By performing atomized association processing on the multi-source heterogeneous security data to construct a dynamic risk association hypergraph, and utilizing a set of entity nodes including communication nodes, storage nodes, and interface nodes, a set of behavioral hyperedges connecting multiple entity nodes, and a set of temporal constraints defining the temporal relationships of these hyperedges, the method can accurately capture the collaborative risk behavior associations and temporal dynamics among multiple entity nodes. This solves the problem that traditional graph structures can only express bilateral node associations and cannot reflect multi-node collaborative risks. Finally, by calling a large model based on a security domain knowledge graph to perform multi-round risk attribution reasoning, the method further refines the behavioral hyperedges. This system performs attack chain segment matching and evidence chain completion, generating a risk attribution inference chain that includes the main risk node, associated risk paths, and evidence confidence levels. It effectively handles scenarios with incomplete attack chains or novel attack patterns, improving the accuracy of risk attribution and overcoming the limitations of simple pattern matching that relies on a pre-defined rule base. By constructing a risk evolution probability model based on the risk attribution inference chain and time-series constraint set, it calculates the short-term diffusion probability and long-term evolution trend vector of each associated risk path, achieving a refined and dynamic description of risk evolution. This solves the problem that traditional single-probability prediction cannot reflect the time-dimensional evolution characteristics of risks. Furthermore, by generating a dynamic security protection scheme based on the short-term diffusion probability and long-term evolution trend vector, which includes node protection priorities, path blocking strategies, and interface access control rules, it enables real-time matching of protection strategies with the risk evolution state, improving the targeting and dynamic adaptability of protection measures and solving the lag problem of static protection strategies being unable to cope with dynamic changes in risks. Attached Figure Description

[0007] Figure 1 This is a flowchart of a comprehensive analysis method for network security and data security based on a large model, provided by an embodiment of the present invention.

[0008] Figure 2 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present invention. Detailed Implementation

[0009] Please see Figure 1 , Figure 1 The flowchart illustrates a comprehensive analysis method for network security and data security based on a large model, provided as an embodiment of the present invention. This method can be executed by a computer system and may include the following steps:

[0010] The method for comprehensive analysis of network security and data security based on a large model provided by the embodiments of the present application can be applied to many fields such as e-commerce, finance, medical treatment, industrial control, and government affairs. In order to more clearly and specifically set forth the implementation process and advantages of the method, the e-commerce scenario is taken as an example for detailed description, but the method is not limited to the e-commerce scenario, and can also play an important role in other mentioned or unmentioned fields, effectively cope with security challenges in various complex network environments, and realize comprehensive analysis and protection of network security and data security.

[0011] Step S100: Collecting multi-source heterogeneous security data of network communication links, data storage nodes, and application interaction interfaces, the multi-source heterogeneous security data including encrypted transmission data packet units, data access audit units, and interface call sequence units, each data unit carrying a node identifier and an interaction time stamp.

[0012] The multi-source heterogeneous security data is derived from different components and systems in the network, and its types and formats are different. The encrypted transmission data packet unit is a set of data packets transmitted in encrypted form in the network communication link. These data packets carry basic information of network communication such as source address, destination address, port number, etc. Encryption processing is to ensure the security of data transmission. The data access audit unit records the detailed information of data access operations on the data storage node, including access subject, operation object, operation time, operation type, etc., which is used for auditing and monitoring data access behavior. The interface call sequence unit records the calling process and related information of the application interaction interface, such as the calling sequence of the interface, input parameters, return results, etc., reflecting the interaction behavior between application programs. The node identifier is used to uniquely distinguish each node in the network, such as communication nodes, storage nodes, interface nodes, etc., facilitating the management of data of different nodes. The interaction time stamp records the time information of each data unit in network interaction, which is used to determine the sequence and time interval of data interaction.

[0013] In collecting multi-source heterogeneous security data, for encrypted transmission data packet units, a network packet capturing tool can be deployed in the network communication link, for example, a special packet capturing software is used, which listens to network traffic, captures all data packets passing through the network interface, and performs preliminary analysis and recording on encrypted data packets. For the data access audit unit, audit software is configured on the data storage node, which monitors and records all data access operations in real time, and when a user or program accesses the data storage node, the audit software records detailed information of the access. For the interface call sequence unit, log recording functions are added at the application interaction interface, and when the interface is called, the related information of the call is automatically recorded, including the time of the call, input parameters, return status, etc. For example, in an e-commerce platform, a packet capturing software is deployed on the core switch of its network to capture encrypted transmission data packets between servers; an audit software is installed on the database server to record user access operations on the commodity database; log recording functions are added at the interfaces of the e-commerce application to record the calling conditions of the user ordering, querying commodities and other interfaces. Each collected data unit automatically adds a node identifier and an interaction time stamp, the node identifier can be generated by pre-defined rules, such as coding according to the network area and function where the node is located, and the interaction time stamp is automatically recorded by the system clock.

[0014] Step S200: Perform risk behavior atomization correlation processing on the multi-source heterogeneous security data, and construct a dynamic risk correlation hypergraph, the dynamic risk correlation hypergraph including an entity node set, a behavior hyperedge set and a time sequence constraint set, the entity node set including a communication node, a storage node and an interface node.

[0015] The risk behavior atomization correlation processing is to decompose and correlate various risk behaviors in the multi-source heterogeneous security data, to decompose the complex risk behaviors into atomic level features, and to establish the correlation between the features. Exemplarily, the multi-dimension can involve the following aspects: from the data source dimension, the multi-source heterogeneous security data is derived from network communication links, data storage nodes and application interaction interfaces, which are three different data source dimensions. The data of the network communication links can embody the risk behaviors at the network layer, the data of the data storage nodes can reflect the risks in the data storage and access process, and the data of the application interaction interfaces can show the risks in the application program interaction. From the behavior feature dimension, different types of behavior atomic features are generated after processing the multi-source heterogeneous security data, such as communication behavior atomic features, access behavior atomic features and interface behavior atomic features. The communication behavior atomic features focus on the characteristics of network communication, such as transmission direction identifier, payload length sequence and encryption algorithm type; the access behavior atomic features focus on data access operations, including access subject identifier, operation object path and permission change record; and the interface behavior atomic features focus on interface call conditions, including interface identifier, input parameter digest and return status code. From the time dimension, each data unit carries an interaction time stamp, which enables the analysis of the risk behaviors to consider the sequence and time interval of the behaviors. By dividing the time stamp sequence by time interval, the co-occurrence frequency and time sequence dependence of the behavior atomic features in different time intervals can be counted, thereby establishing the correlation of the risk behaviors in the time dimension. In constructing the dynamic risk correlation hypergraph, these multi-dimension information interacts with each other, the entity node set is constructed based on different data sources and behavior feature dimensions, the construction of the behavior hyperedge set considers the co-occurrence frequency and time sequence dependence of the behavior atomic features between different entity nodes in the time dimension, and the time sequence constraint set further limits the sequence of the behavior hyperedges in the time dimension, thereby realizing the multi-dimension risk behavior atomization correlation processing. The dynamic risk correlation hypergraph is a graph structure for presenting network security risks, which can more comprehensively describe the risk relationships in the network. The entity node set is a node set in the hypergraph, the communication node represents a communication device or a communication link in the network, the storage node represents a data storage device or a data storage location, and the interface node represents an application interaction interface. The behavior hyperedge set is a hyperedge set connecting multiple entity nodes, the hyperedge embodies the complex behavior relationship between multiple entity nodes, and each hyperedge has a corresponding weight. The time sequence constraint set is used to limit the occurrence time sequence of the hyperedges in the behavior hyperedge set, to ensure the time sequence correctness of the risk correlation.

[0016] As an implementation, step S200 can be implemented as steps S210-S270.

[0017] Step S210: Protocol reverse analysis is performed on the encrypted transmission data packet unit in the multi-source heterogeneous security data to extract the transmission direction identifier, the payload length sequence, and the encryption algorithm type, and to generate the communication behavior atomic feature.

[0018] Protocol reverse analysis is a deep analysis and processing of the encrypted transmission data packet to restore the protocol information followed by the data packet. The transmission direction identifier is used to determine the transmission direction of the data packet, i.e., from the source node to the destination node or from the destination node to the source node. The payload length sequence refers to the sequence formed by arranging the length information of the payload in the data packet in the order of the data packet. The encryption algorithm type refers to the encryption algorithm used by the data packet, such as the common symmetric encryption algorithm or asymmetric encryption algorithm. The communication behavior atomic feature is a set of basic features used to describe the network communication behavior, which is generated by extracting the transmission direction identifier, the payload length sequence, and the encryption algorithm type.

[0019] In the protocol reverse analysis, a professional protocol analysis tool can be used. The tool will first identify the header information of the data packet and determine the transmission direction identifier through the fields in the header. For example, there may be fields in the header that specifically indicate the source address and the destination address, and the transmission direction can be determined according to the relationship between these addresses. Then, the tool will extract the payload part of the data packet, calculate its length, and record it in sequence to form the payload length sequence. For the determination of the encryption algorithm type, the tool will analyze the encryption algorithm related fields in the data packet, which may contain the identification information of the algorithm, and determine the encryption algorithm type by comparing with the predefined algorithm identification library. In the e-commerce platform, when the servers transmit the commodity data, the protocol analysis tool is used to analyze these encrypted transmission data packets. The transmission direction identifier (such as from the commodity inventory server to the order processing server), the payload length sequence (such as a sequence of bytes of a certain length), and the encryption algorithm type (such as symmetric encryption algorithm) are extracted to generate the communication behavior atomic feature.

[0020] Step S220: The operation sequence decomposition is performed on the data access audit unit in the multi-source heterogeneous security data to extract the access subject identifier, the operation object path, and the permission change record, and to generate the access behavior atomic feature.

[0021] Operation sequence decomposition is to split and analyze a series of data access operations recorded in the data access audit unit, decomposing complex operation processes into single operation steps. Access subject identification is used to uniquely identify the subject performing the data access operation, which can be a user account, an application program, etc. Operation object path refers to the specific location path of the accessed data object in the data storage system, which is used to accurately locate the data. Permission change record refers to the detailed record of the change in the access subject's permissions during the data access process, such as the promotion or reduction of permissions. Access behavior atomic features are a set of basic features used to describe data access behavior, generated by extracting access subject identification, operation object path, and permission change record, etc.

[0022] When performing operation sequence decomposition, a log analysis tool is used. This tool analyzes the log information in the data access audit unit line by line, first extracting the access subject identification of each operation step from the log, determined by the set field or identification information. Then, according to the relevant information in the operation record, the operation object path is determined, for example, the log may have directory and file name information of data storage, and these information can be combined to obtain the operation object path. Finally, check if there is a record of permission change in the log, determine if the permission has changed through the set identification or flag, and extract the relevant information. In the e-commerce platform, when the user accesses the commodity database, the log analysis tool analyzes the access audit log of the database. Extract the access subject identification (such as the user's account name), operation object path (such as the specific path of a table in the commodity database), and permission change record (such as the user changing from read-only permission to read-write permission), to generate access behavior atomic features.

[0023] Step S230: Call chain tracking is performed on the interface call sequence unit in the multi-source heterogeneous security data, and interface identification, input parameter summary, and return status code are extracted to generate interface behavior atomic features.

[0024] Call chain tracking is to track and analyze the interface call process in the interface call sequence unit, determine the order and path of interface calls. Interface identification is used to uniquely identify an interface, usually the name or number of the interface. Input parameter summary simplifies and summarizes the parameters input during interface call, used to quickly understand the input of interface call. Return status code is the result state information returned after interface call, used to indicate whether the interface call is successful or an error occurs. Interface behavior atomic features are a set of basic features used to describe interface call behavior, generated by extracting interface identification, input parameter summary, and return status code, etc.

[0025] When performing call chain tracking, a distributed tracking tool is used. The tool inserts tracking code at various links of an interface call, and when the interface is called, the tracking code records relevant information of the call. First, the tracking code determines the interface identifier, which can be explicitly identified in the definition file or configuration file of the interface. Then, the input parameters of the interface call are extracted, and these parameters are analyzed and processed to extract key parameter information to form an input parameter digest. Finally, the return status code of the interface call is obtained, which is usually a set field in the return result of the interface. In an e-commerce platform, when a user queries product information through an interface, the distributed tracking tool tracks the interface call process. The interface identifier (such as the interface name for querying product information), the input parameter digest (such as the key parameters of the queried product category), and the return status code (such as the status code indicating a successful query) are extracted to generate interface behavior atomic features.

[0026] Step S240: mapping the communication behavior atomic features, the access behavior atomic features, and the interface behavior atomic features into an entity node set of a dynamic risk association hypergraph, wherein the communication behavior atomic features are mapped into communication nodes, the access behavior atomic features are mapped into storage nodes, and the interface behavior atomic features are mapped into interface nodes.

[0027] In the embodiments of the present application, the communication behavior atomic features, the access behavior atomic features, and the interface behavior atomic features are respectively corresponded to the communication nodes, the storage nodes, and the interface nodes in the entity node set of the dynamic risk association hypergraph. The purpose of this is to represent different types of behavior features in the form of nodes in the hypergraph, which facilitates risk association analysis. When mapping, the features are corresponded according to their types and meanings. For the communication behavior atomic features, the information contained therein is mapped to the communication nodes. For example, the attributes of the communication nodes can be set to the transmission direction identifier, the load length sequence, and the encryption algorithm type in the communication behavior atomic features. For the access behavior atomic features, the corresponding information is mapped to the storage nodes, and the attributes of the storage nodes can include the access subject identifier, the operation object path, and the permission change record. For the interface behavior atomic features, the corresponding information is mapped to the interface nodes, and the attributes of the interface nodes can include the interface identifier, the input parameter digest, and the return status code. In an e-commerce platform, the previously generated communication behavior atomic features (such as the transmission direction from the product inventory server to the order processing server, the load length sequence of a certain length, and the type of a certain symmetric encryption algorithm) are mapped to a communication node; the access behavior atomic features (such as the access of a user account to a certain table in the product database, and the change of permission from read-only to read-write) are mapped to a storage node; and the interface behavior atomic features (such as the interface for querying product information, the input product category parameter, and the return status code indicating a successful query) are mapped to an interface node.

[0028] Step S250: Calculate the co-occurrence frequency and time sequence dependency of the behavior atomic features between different entity nodes based on the interaction time sequence stamps, and construct a set of behavior hyper-edges connecting multiple entity nodes.

[0029] The co-occurrence frequency refers to the frequency of the simultaneous occurrence of the behavior atomic features between different entity nodes within the same time interval. The time sequence dependency refers to the sequential dependency of the behavior atomic features between different entity nodes in time, which is reflected by calculating the conditional probability of the preceding behavior atomic features and the subsequent behavior atomic features. The set of behavior hyper-edges is a set of hyper-edges connecting multiple entity nodes, and each hyper-edge has a corresponding weight, which represents the importance of the relationship.

[0030] As an implementation, step S250 can be implemented as steps S251-S256:

[0031] Step S251: Extract the interaction time sequence stamp sequence of the communication node, the storage node, and the interface node, and determine the occurrence time point of each node behavior atomic feature.

[0032] The interaction time sequence stamp sequence refers to the sequence formed by arranging the interaction time sequence stamps corresponding to the behavior atomic features of the communication node, the storage node, and the interface node in order. The occurrence time point refers to the specific occurrence time of each node behavior atomic feature on the time axis.

[0033] When extracting the interaction time sequence stamp sequence, from the previously collected multi-source heterogeneous security data, according to the association relationship between the node identifier and the behavior atomic feature, the interaction time sequence stamp corresponding to each node behavior atomic feature is extracted, and the sequence is formed by arranging in time order. For example, in an e-commerce platform, the interaction time sequence stamps corresponding to the communication behavior atomic features of the communication node (such as data transmission between servers) are different time points in turn, the interaction time sequence stamps corresponding to the access behavior atomic features of the storage node (such as access to the commodity database) are also different time points in turn, and the interaction time sequence stamps corresponding to the interface behavior atomic features of the interface node (such as the calling of the user interface) are also different time points in turn. By extracting these interaction time sequence stamps, the occurrence time points of each node behavior atomic feature are determined.

[0034] Step S252: Segment the time sequence stamp sequence using a time interval division method, count the number of combinations of the communication node behavior atomic features, the storage node behavior atomic features, and the interface node behavior atomic features that occur simultaneously within each time interval, and obtain the co-occurrence frequency.

[0035] The time interval division method is a method of dividing the time axis into several equal or unequal time intervals. The co-occurrence frequency refers to the frequency of the simultaneous occurrence of the behavior atomic features of different types of nodes within the same time interval.

[0036] In the time interval division, the appropriate time interval length is selected according to the actual situation. The time axis can be divided into time intervals of fixed length, or unequal time intervals can be divided according to the characteristics of the business. Then, the time sequence stamp sequence in each time interval is analyzed, and the combination times of the communication node behavior atomic features, the storage node behavior atomic features and the interface node behavior atomic features appearing at the same time are counted. In the e-commerce platform, for example, the time of a day is divided into time intervals of several hours, and in each hour time interval, the combination times of the communication behavior between servers, the commodity database access behavior and the user interface call behavior appearing at the same time are counted. The more the combination times are, the higher the co-occurrence frequency is.

[0037] Step S253: Time sequence order marking is performed on the behavior atomic feature combination in each time interval, and the conditional probability of the previous behavior atomic feature and the subsequent behavior atomic feature is calculated as the time sequence dependence degree.

[0038] The time sequence order marking refers to marking the behavior atomic feature combination in each time interval according to the order of occurrence time, and clarifying which behaviors are previous behaviors and which behaviors are subsequent behaviors. The conditional probability refers to the probability of the occurrence of the subsequent behavior under the condition that the previous behavior has occurred. The time sequence dependence degree is represented by calculating the conditional probability of the previous behavior atomic feature and the subsequent behavior atomic feature.

[0039] In the time sequence order marking, the order of the behavior atomic features is determined according to the interaction time sequence stamp. For the behavior atomic feature combination in each time interval, the first occurring behavior is marked as the previous behavior, and the second occurring behavior is marked as the subsequent behavior. Then, the number of times of the simultaneous occurrence of the previous behavior and the subsequent behavior and the total number of times of the occurrence of the previous behavior are counted, and the number of times of the simultaneous occurrence of the previous behavior and the subsequent behavior is divided by the total number of times of the occurrence of the previous behavior to obtain the conditional probability, i.e. the time sequence dependence degree. In the e-commerce platform, in a certain time interval, the user first performs the commodity query interface call (previous behavior), and then performs the commodity purchase operation (subsequent behavior). The number of times of the commodity purchase after the commodity query and the total number of times of the commodity query are counted, and the conditional probability of the commodity purchase after the commodity query is calculated as the time sequence dependence degree between the two behavior atomic features.

[0040] Step S254: The co-occurrence frequency is converted into the co-occurrence weight in the preset numerical interval by the maximum value normalization method, and the time sequence dependence degree is converted into the dependence weight in the preset numerical interval by the maximum value normalization method.

[0041] The maximum value normalization manner is to map the data into a preset numerical interval by comparing and converting the original data with the maximum value. The co-occurrence weight is a weight value obtained by normalizing the co-occurrence frequency, and is used to represent the importance of the co-occurrence of the behavior atomic feature between different entity nodes. The dependency weight is a weight value obtained by normalizing the time sequence dependency, and is used to represent the importance of the time sequence dependency of the behavior atomic feature between different entity nodes. The preset numerical interval is usually a fixed range, such as from zero to one.

[0042] In the maximum value normalization, the maximum value of the co-occurrence frequency and the time sequence dependency is found first. For the co-occurrence frequency, each co-occurrence frequency value is divided by the maximum value of the co-occurrence frequency to obtain the normalized co-occurrence weight. For the time sequence dependency, each time sequence dependency value is divided by the maximum value of the time sequence dependency to obtain the normalized dependency weight. In the e-commerce platform, the co-occurrence frequency and the time sequence dependency of the behavior atomic feature between different entity nodes are calculated, and then the maximum value among these values is found. Each co-occurrence frequency value and time sequence dependency value is subjected to maximum value normalization to obtain the corresponding co-occurrence weight and dependency weight.

[0043] Step S255: The co-occurrence weight and the dependency weight are weighted and summed according to a preset proportion to obtain a comprehensive weight of the behavior super edge.

[0044] The preset proportion is a proportion of the co-occurrence weight and the dependency weight in the weighted sum. The comprehensive weight is a weight value obtained by weighting and summing the co-occurrence weight and the dependency weight according to the preset proportion, and is used to represent the importance of the behavior super edge.

[0045] In the weighted sum, according to the preset proportion, the co-occurrence weight is multiplied by the corresponding proportion coefficient, the dependency weight is multiplied by the corresponding proportion coefficient, and then the two results are added to obtain the comprehensive weight of the behavior super edge. In the e-commerce platform, it is assumed that the preset proportion is that the co-occurrence weight occupies a certain proportion and the dependency weight occupies another certain proportion. The co-occurrence weight and the dependency weight obtained before are weighted and summed according to this proportion to obtain the comprehensive weight of the behavior super edge.

[0046] Step S256: According to the type and number of entity nodes in the behavior atomic feature combination, a super edge structure containing two or more entity nodes is constructed, the comprehensive weight is assigned to the super edge structure, and a behavior super edge set is generated.

[0047] The super edge structure is a graph structure used to represent the complex relationship between multiple entity nodes, and unlike ordinary edges, the super edge can connect two or more entity nodes. The behavior super edge set is a set composed of multiple super edge structures, and each super edge structure has a corresponding comprehensive weight.

[0048] In constructing the super-edge structure, according to the entity node types (such as communication nodes, storage nodes, interface nodes) and the number involved in the behavior atomic feature combination, the connection mode of the super-edge is determined. If the behavior atomic feature combination involves multiple entity nodes of different types, a super-edge structure connecting these nodes is constructed. Then, the comprehensive weight calculated before is assigned to the super-edge structure. Repeat this process to generate a behavior super-edge set containing multiple super-edge structures. In the e-commerce platform, when the behavior atomic feature combination involves communication nodes between servers, storage nodes of commodity databases, and interface nodes of user interfaces, a super-edge structure connecting the three nodes is constructed, and the comprehensive weight calculated is assigned to the super-edge structure. By processing multiple such behavior atomic feature combinations, a behavior super-edge set is generated.

[0049] Step S260: generating a time sequence constraint set according to the interaction time sequence stamps of the multi-source heterogeneous security data, the time sequence constraint set being used to limit the occurrence time sequence relationship of each super-edge in the behavior super-edge set.

[0050] The time sequence constraint set is a set of constraint conditions for limiting the occurrence time sequence relationship of each super-edge in the behavior super-edge set. By generating the time sequence constraint set, it can be ensured that the time relationship between the super-edges in the subsequent risk analysis conforms to the actual situation, improving the accuracy of the analysis.

[0051] As an implementation, step S260 can be implemented as steps S261-S264 as follows:

[0052] Step S261: extracting the occurrence time stamps of all behavior super-edges in the dynamic risk correlation hypergraph, establishing a super-edge time stamp sequence, and sorting the super-edge time stamp sequence in ascending order to determine the time sequence of each behavior super-edge.

[0053] The super-edge time stamp sequence is a sequence formed by arranging the occurrence time stamps of all behavior super-edges in the dynamic risk correlation hypergraph in order. Ascending order sorting is to arrange the time stamp sequence in order from small to large. By sorting the super-edge time stamp sequence in ascending order, the time sequence of each behavior super-edge can be determined, providing a basis for subsequent generation of time sequence constraint relationship.

[0054] In the process of extracting the occurrence time stamp of the behavior hyperedge, the information recorded in the process of constructing the dynamic risk correlation hypergraph is obtained. Each behavior hyperedge corresponds to the relevant behavior atomic feature, and the behavior atomic feature is associated with the interaction time stamp, so the occurrence time stamp of the behavior hyperedge can be extracted through the association. The time stamps are arranged in order to form a hyperedge time stamp sequence, and then a sorting algorithm such as the quicksort algorithm is used to sort the sequence in ascending order. In the e-commerce platform, the occurrence time stamps of all communication hyperedges between servers, user interface calling hyperedges, etc. are extracted, the hyperedge time stamp sequence is established and sorted in ascending order, and the time sequence of each behavior hyperedge is determined.

[0055] Step S262: For any two behavior hyperedges with shared entity nodes, if the occurrence time stamp of the previous behavior hyperedge is earlier than that of the subsequent behavior hyperedge, a time sequence constraint relationship is generated.

[0056] The time sequence constraint relationship refers to the constraint condition that the occurrence time of the previous hyperedge must be earlier than that of the subsequent hyperedge between the two behavior hyperedges with shared entity nodes. The shared entity node refers to the same node in the entity nodes connected by the two behavior hyperedges.

[0057] In generating the time sequence constraint relationship, the behavior hyperedge set is traversed to find any two behavior hyperedges with shared entity nodes. Then, the occurrence time stamps are compared. If the occurrence time stamp of the previous behavior hyperedge is earlier than that of the subsequent behavior hyperedge, a time sequence constraint relationship is generated. In the e-commerce platform, behavior hyperedge A connects a server node and a user interface node, and behavior hyperedge B also connects the server node and another user interface node, which are shared entity nodes. If the occurrence time stamp of behavior hyperedge A is earlier than that of behavior hyperedge B, a time sequence constraint relationship is generated, indicating that behavior hyperedge A must occur before behavior hyperedge B.

[0058] Step S263: For behavior hyperedges containing the same set of entity nodes, the time interval of their occurrence time stamps is calculated, and if the time interval is less than the preset minimum interval threshold, a time overlap constraint relationship is generated.

[0059] The time interval refers to the difference between the occurrence time stamps of two behavior hyperedges. The preset minimum interval threshold is a pre-set time value for determining whether the occurrence times of two behavior hyperedges are too close. The time overlap constraint relationship refers to the constraint condition that the occurrence times of the two behavior hyperedges containing the same set of entity nodes overlap.

[0060] When calculating the time interval, for the behavior hyperedge containing the same set of entity nodes, the occurrence timestamp of the latter hyperedge is directly subtracted from the occurrence timestamp of the former hyperedge. If the time interval is less than the preset minimum interval threshold, a time overlap constraint relationship is generated. In the e-commerce platform, the behavior hyperedge C and the behavior hyperedge D are connected to the same server node and user interface node, and the time interval of their occurrence timestamps is calculated. If the time interval is less than the preset minimum interval threshold, a time overlap constraint relationship is generated, indicating that the occurrence time of the behavior hyperedge C and the behavior hyperedge D overlaps.

[0061] Step S264: Structurally encode all generated time sequence constraint relationships and time overlap constraint relationships to generate a time sequence constraint set containing hyperedge identification pairs and constraint types.

[0062] Structural encoding is to encode the time sequence constraint relationship and the time overlap constraint relationship in a structured manner for storage and processing. The hyperedge identification pair refers to a pair of identifications for uniquely identifying two behavior hyperedges. The constraint type refers to the time sequence constraint or the time overlap constraint.

[0063] In the structural encoding, each time sequence constraint relationship and time overlap constraint relationship is represented as a structure containing a hyperedge identification pair and a constraint type. For example, for the time sequence constraint relationship, it is represented as a structure containing a previous hyperedge identification, a subsequent hyperedge identification, and a time sequence constraint type. For the time overlap constraint relationship, it is represented as a structure containing two hyperedge identifications and a time overlap constraint type. All such structures are combined to generate a time sequence constraint set containing hyperedge identification pairs and constraint types. In the e-commerce platform, all generated time sequence constraint relationships and time overlap constraint relationships are structurally encoded to generate a time sequence constraint set.

[0064] Step S270: Integrate the entity node set, the behavior hyperedge set, and the time sequence constraint set into a hypergraph structure to generate a dynamic risk association hypergraph containing node attributes, hyperedge weights, and time sequence constraints.

[0065] Hypergraph structure integration is to combine and associate the entity node set, the behavior hyperedge set, and the time sequence constraint set to form a complete dynamic risk association hypergraph. The node attribute refers to various attribute information possessed by the entity node, such as the transmission direction identification of the communication node, the operation object path of the storage node, etc. The hyperedge weight refers to the comprehensive weight of each hyperedge in the behavior hyperedge set, which is used to represent the importance of the hyperedge. The time sequence constraint refers to the constraint condition in the time sequence constraint set, which is used to limit the occurrence time sequence relationship of the hyperedge.

[0066] In the process of supergraph structure integration, the entity node set is taken as the node of the supergraph, the behavior superedge set is taken as the superedge of the supergraph, and the superedge weight is assigned to the corresponding superedge. Then, the time constraint set is associated with the superedge set to ensure that the time relationship of the superedge meets the constraint condition. For each superedge, check whether it meets the time sequence constraint and the time overlap constraint in the time constraint set. Finally, a complete dynamic risk correlation supergraph containing node attributes, superedge weights and time constraints is generated. In the e-commerce platform, the communication node, the storage node and the interface node are taken as the entity node set, the behavior superedge connecting these nodes is taken as the behavior superedge set, and the comprehensive weight of the superedge calculated before is assigned to the corresponding superedge. At the same time, the generated time constraint set is associated with the superedge set to generate a dynamic risk correlation supergraph.

[0067] Step S300: based on the preset security field knowledge graph, a large model is called to perform multi-round risk attribution reasoning, attack chain segment matching and evidence chain completion are performed on the behavior superedge set in the dynamic risk correlation supergraph, and a risk attribution reasoning chain containing a main cause risk node, a correlation risk path and an evidence confidence is generated.

[0068] The preset security domain knowledge graph is a graph containing various knowledge and information in the network security field, such as common attack patterns, vulnerability characteristics, defense rules, and the like. The large model is an artificial intelligence model with strong computing and reasoning capabilities, such as a language model based on a Transformer architecture. The multi-round risk attribution reasoning is a process of gradually analyzing and attributing risks in the dynamic risk association hypergraph by calling the large model multiple times. The attack chain fragment matching is to compare the behavior hyperedge set in the dynamic risk association hypergraph with the attack patterns in the security domain knowledge graph, and find possible attack chain fragments. The evidence chain completion is to further supplement relevant evidence information after matching the attack chain fragments, and perfect the evidence chain. The main cause risk node refers to the main node that causes the risk, the associated risk path refers to the path formed by a series of nodes and hyperedges associated with the main cause risk node, and the evidence confidence refers to the evidence confidence of each link. Among them, the possible attack chain fragment refers to a combination of a series of behavior hyperedges and entity nodes connected by the behavior hyperedges in the dynamic risk association hypergraph, which matches the key behavior stage, entity node type and core behavior characteristics described in the attack pattern in the security domain knowledge graph. When performing attack chain fragment matching, for example, the attack pattern description can be extracted from the preset security domain knowledge graph and converted into a structured behavior feature template. The template contains information such as key behavior stages, entity node types and core behavior characteristics, for example, the initial stage corresponds to a communication node, and the core behavior characteristic is an unauthorized port connection attempt; the intermediate stage corresponds to an interface node, and the core behavior characteristic is an abnormal parameter construction call; the subsequent stage corresponds to a storage node, and the core behavior characteristic is a permission boundary breakthrough operation. Then, the entity node set, behavior hyperedge set and time sequence constraint set of the dynamic risk association hypergraph are converted into a graph structure text description that can be parsed by the large model, which is used as the input context of the first round of reasoning. The large model is called to perform the first round of reasoning on the graph structure text description in combination with the behavior feature template. The large model parses the graph structure text description to understand the information such as entity nodes, behavior hyperedges and time sequence constraints. For each behavior hyperedge in the behavior hyperedge set, check whether the node type connected by the behavior hyperedge matches the node type in the behavior feature template and whether the behavior of the node conforms to the core behavior characteristic. If it matches, the behavior hyperedge and the node combination connected by the behavior hyperedge are taken as a candidate attack chain fragment. For example, if a communication node connected by a behavior hyperedge has an unauthorized port connection behavior, an interface node has an abnormal parameter construction call behavior, and a storage node has a breakthrough permission attempt, these nodes and hyperedge combination will be taken as a candidate attack chain fragment, and finally the entity node sequence and behavior hyperedge sequence contained in the candidate attack chain fragment are output.

[0069] As an implementation manner, the step S300 can be implemented as steps S310-S380.

[0070] Step S310: Extract attack pattern description, vulnerability feature description and defense rule description from the preset security domain knowledge graph, and convert the attack pattern description into a structured behavior feature template.

[0071] Attack pattern description is a detailed description of network attack behavior, including attack steps, means, target and other information. Vulnerability feature description is a description of the features of the vulnerabilities existing in the network system, such as vulnerability type, impact range, trigger condition, etc. Defense rule description is a description of the rules of network security defense measures, such as prohibiting access to certain ports, limiting the operation of certain users, etc. The structured behavior feature template is to convert the attack pattern description into a structured form, which is convenient for large models to analyze and match.

[0072] In extracting the relevant description, through the query interface of the knowledge graph, the attack pattern description, the vulnerability feature description and the defense rule description are obtained from the preset security domain knowledge graph. Then, the attack pattern description is processed and converted into a structured behavior feature template. Natural language processing technology is used to analyze the attack pattern description and extract the key behavior stages in the attack process, the entity node types corresponding to each stage and the core behavior features. These information is combined into a structured template, which includes stage placeholder area, node type placeholder area and behavior feature placeholder area. In the e-commerce platform, the attack pattern description for the e-commerce system is extracted from the security domain knowledge graph, such as the attacker may first connect to the server through an unauthorized port, then construct an abnormal parameter to call the user interface, and finally break through the database permission to tamper with the data. This attack pattern description is converted into a structured behavior feature template.

[0073] As an implementation, in step S310, the attack pattern description is converted into a structured behavior feature template, which can be implemented as the following steps S311-S314:

[0074] Step S311: Analyze the text content of the attack pattern description and extract the key behavior stages in the attack process, which are a sequence of behaviors arranged in chronological order.

[0075] Analysis is to analyze and understand the text content of the attack pattern description and extract useful information from it. Key behavior stages are the behavior stages that have important significance in the attack process, which are arranged in chronological order to form a behavior sequence.

[0076] In the analysis of attack pattern description, natural language processing techniques such as word segmentation, part-of-speech tagging, named entity recognition, etc. are used. First, the text of the attack pattern description is processed by word segmentation, which splits the text into individual words. Then, through part-of-speech tagging, the part-of-speech of each word is determined, and through named entity recognition, the entity type represented by the word is determined. Next, according to the semantics and logical relationships of the text, the key behavior stages in the attack process are extracted. In the e-commerce platform, the attack pattern description is "the attacker first tries to connect to the e-commerce server through an unauthorized port, then constructs an abnormal parameter to call the commodity query interface, and finally breaks through the database permission to modify the commodity price". Through analysis, the key behavior stages are "unauthorized port connection attempt", "abnormal parameter construction call", and "permission boundary breakthrough operation".

[0077] Step S312: Match the corresponding entity node type for each key behavior stage, where the initial stage corresponds to the communication node, the intermediate stage corresponds to the interface node, and the subsequent stage corresponds to the storage node.

[0078] The entity node type refers to the node type in the dynamic risk association hypergraph, including communication nodes, interface nodes, and storage nodes. By matching the corresponding entity node type for each key behavior stage, the attack pattern can be associated with the nodes in the dynamic risk association hypergraph, facilitating subsequent matching and analysis.

[0079] In the matching process, according to the characteristics and properties of the key behavior stages, they are matched with the corresponding entity node types. In the e-commerce platform, the initial stage of the attack "unauthorized port connection attempt" usually involves network communication, corresponding to the communication node, such as the connection between the server and the external network. The intermediate stage "abnormal parameter construction call" requires calling the server interface for operation, corresponding to the interface node, such as the commodity query interface. The subsequent stage "permission boundary breakthrough operation" may operate on the data in the database, corresponding to the storage node, such as the commodity database.

[0080] Step S313: Extract the core behavior features of each key behavior stage, the core behavior feature of the initial stage is unauthorized port connection attempt, the core behavior feature of the intermediate stage is abnormal parameter construction call, and the core behavior feature of the subsequent stage is permission boundary breakthrough operation.

[0081] The core behavior feature refers to the most representative and characteristic behavior in each key behavior stage. By extracting the core behavior feature, the attack pattern can be more accurately described, facilitating matching with the behaviors in the dynamic risk association hypergraph.

[0082] In extracting core behavior features, according to the description and analysis of key behavior stages, the core behaviors of each stage are determined. In the e-commerce platform, for the initial stage "unauthorized port connection attempt", the core behavior feature is to attempt to connect to the e-commerce server through an unauthorized port. For the intermediate stage "abnormal parameter construction call", the core behavior feature is to construct abnormal parameters to call the interface of the server, such as constructing unreasonable product query parameters. For the subsequent stage "permission boundary breakthrough operation", the core behavior feature is to break through the permission boundary of the server database and perform illegal data operations, such as modifying the price of a product.

[0083] Step S314: Combine the key behavior stages, entity node types and core behavior features into a structured behavior feature template. The behavior feature template includes stage placeholder areas, node type placeholder areas and behavior feature placeholder areas.

[0084] The structured behavior feature template is a structured form for representing attack patterns. By combining key behavior stages, entity node types and core behavior features, it can facilitate processing and matching by large models. The stage placeholder area is used to represent different stages of the attack, the node type placeholder area is used to represent the entity node type corresponding to each stage, and the behavior feature placeholder area is used to represent the core behavior feature of each stage.

[0085] In combining the behavior feature template, the key behavior stages, entity node types and core behavior features are filled into the corresponding placeholder areas. In the e-commerce platform, a behavior feature template can be represented as: {stage placeholder area: initial stage, node type placeholder area: communication node, behavior feature placeholder area: unauthorized port connection attempt; stage placeholder area: intermediate stage, node type placeholder area: interface node, behavior feature placeholder area: abnormal parameter construction call; stage placeholder area: subsequent stage, node type placeholder area: storage node, behavior feature placeholder area: permission boundary breakthrough operation}.

[0086] Step S320: Convert the entity node set, behavior hyperedge set and time sequence constraint set of the dynamic risk association hypergraph into a graph structure text description that can be parsed by the large model, as the input context of the first round of reasoning.

[0087] The graph structure text description is a description of the entity node set, behavior hyperedge set and time sequence constraint set of the dynamic risk association hypergraph in text form, so that it can be understood and processed by the large model. The input context refers to the relevant information provided to the large model during reasoning, which is used to guide the reasoning process of the model.

[0088] When converting, first, the information of the entity node set is sorted, including the attribute information of the node, such as the transmission direction identifier of the communication node, the operation object path of the storage node, etc. Then, the information of the behavior hyperedge set is sorted, including the connection nodes of the hyperedge and the hyperedge weight. Finally, the information of the time sequence constraint set is sorted, including the time sequence constraint and the time overlap constraint. These information is combined into a graph structure text description according to a certain format. In the e-commerce platform, the graph structure text description is expressed in JSON format, which contains the attribute information of the communication node, the storage node and the interface node, the connection node and the weight information of the behavior hyperedge, and the information of the time sequence constraint. Such a graph structure text description is provided as the input context of the first round of reasoning to the large model.

[0089] Step S330: calling the large model, combining the behavior feature template to perform the first round of reasoning on the graph structure text description, identifying the candidate attack chain segment in the behavior hyperedge set that matches the attack pattern description, and outputting the entity node sequence and the behavior hyperedge sequence contained in the candidate attack chain segment.

[0090] The large model is an artificial intelligence model with strong computing and reasoning capabilities, such as a language model based on the Transformer architecture. The behavior feature template is a structured template previously converted from the attack pattern description. The candidate attack chain segment refers to a combination of behavior hyperedges in the dynamic risk association hypergraph that may match the attack pattern description. The entity node sequence refers to the sequence formed by the entity nodes involved in the candidate attack chain segment in order. The behavior hyperedge sequence refers to the sequence formed by the behavior hyperedges involved in the candidate attack chain segment in order.

[0091] When calling the large model for the first round of reasoning, the large model receives the graph structure text description and the behavior feature template as input. The large model first parses the graph structure text description to understand the entity nodes, behavior hyper-edges, and timing constraints. Then, combined with the key behavior stages, entity node types, and core behavior features in the behavior feature template, the behavior hyper-edge set is analyzed one by one. For each behavior hyper-edge, check if the connected node types match the node types in the behavior feature template and if the node behavior conforms to the core behavior features. If it matches, the behavior hyper-edge and its connected nodes are combined as a candidate attack chain segment. Finally, output the entity node sequence and behavior hyper-edge sequence contained in the candidate attack chain segment. In the e-commerce platform, the large model analyzes the graph structure text description of the dynamic risk association hypergraph of the e-commerce platform, combines the behavior feature template for the e-commerce system, and finds possible attack chain segments. For example, it is found that a communication node connected by a behavior hyper-edge has a non-authorized port connection, an interface node has an abnormal parameter construction call, and a storage node has signs of attempted privilege escalation. These nodes and hyper-edges are combined as a candidate attack chain segment, and the corresponding entity node sequence and behavior hyper-edge sequence are output.

[0092] Step S340: Take the candidate attack chain segment output by the first round of reasoning as input for the second round of reasoning, call the large model to combine the vulnerability feature description to match the entity nodes in the candidate attack chain segment, and determine the entity nodes with vulnerability features as potential risk nodes.

[0093] The vulnerability feature description is a detailed description of the network vulnerability features in the security domain knowledge graph, including the impact node type, trigger behavior description, and harm degree description of the vulnerability. The potential risk node refers to the entity node in the candidate attack chain segment that has a behavior or attribute matching the vulnerability feature description. These nodes may cause network security risks due to the existence of vulnerabilities.

[0094] In the secondary round of reasoning, the large model receives the candidate attack chain segment and the vulnerability feature description as input. The large model first analyzes the entity nodes in the candidate attack chain segment in detail to determine the type of each node, such as a communication node, an interface node, or a storage node. Then, according to the node type, the vulnerability-related information matching the type is screened from the vulnerability feature description. For each entity node, its behavior atomic features are extracted, and semantic similarity calculation is performed between these behavior atomic features and the trigger behavior description in the screened vulnerability-related information. If the semantic similarity is high, it means that the behavior of the entity node may trigger the corresponding vulnerability, and the entity node is marked as an entity node with vulnerability features, and its corresponding vulnerability-related information and damage degree description are recorded. In the e-commerce platform, the candidate attack chain segment contains a communication node, an interface node, and a storage node. The large model screens the vulnerability information for communication nodes, interface nodes, and storage nodes from the vulnerability feature description according to the node type. For the communication node, its behavior atomic features (such as unauthorized port connection attempts) are extracted and compared with the trigger behavior description in the communication node vulnerability information. If a high degree of match is found, the communication node is marked as a potential risk node, and the corresponding vulnerability information and damage degree are recorded.

[0095] As an implementation, step S340 can be implemented as steps S341-S345:

[0096] Step S341: Extract vulnerability-related information from the vulnerability feature description, each vulnerability-related information containing an affected node type, a trigger behavior description, and a damage degree description.

[0097] Vulnerability-related information refers to various information related to network vulnerabilities, including the type of nodes affected by the vulnerability, the behavior description that can trigger the vulnerability, and the description of the damage degree that the vulnerability may cause.

[0098] In extracting vulnerability-related information, natural language processing techniques are used to analyze the vulnerability feature description. First, the text of the vulnerability feature description is processed by word segmentation, which splits the text into words. Then, the part-of-speech tagging and named entity recognition are used to determine the part-of-speech and entity type of each word. Next, according to the semantics and logical relationships of the text, the affected node type, trigger behavior description, and damage degree description are extracted. In the e-commerce platform, the vulnerability feature description is "a certain vulnerability affects communication nodes, and when unauthorized port connection occurs and lasts for a period of time, it will trigger, which may cause server paralysis, and the damage degree is high." Through analysis, the affected node type is "communication node", the trigger behavior description is "unauthorized port connection and lasts for a period of time", and the damage degree description is "high".

[0099] Step S342: Type identification is performed on the entity nodes in the candidate attack chain segment to determine the node type of each entity node, and the node type includes a communication node, a storage node, and an interface node.

[0100] Type identification is a classification and judgment of the entity nodes in the candidate attack chain segment to determine whether they belong to a communication node, a storage node, or an interface node. By determining the node type, the nodes that match the vulnerability feature description can be more accurately screened out.

[0101] When performing type identification, the attribute information of the entity node is used for judgment. A communication node usually has attributes related to network communication, such as a transmission direction identifier, a payload length sequence, etc. A storage node has attributes related to data storage, such as an access subject identifier, an operation object path, etc. An interface node has attributes related to interface calling, such as an interface identifier, an input parameter digest, etc. In an e-commerce platform, if the attributes of a node in the candidate attack chain segment include a transmission direction identifier and a payload length sequence, it is determined to be a communication node; if the attributes include an access subject identifier and an operation object path, it is determined to be a storage node; and if the attributes include an interface identifier and an input parameter digest, it is determined to be an interface node.

[0102] Step S343: According to the node type of the entity node, the vulnerability-related information matching the type is screened out from the vulnerability feature description to obtain a node type matching vulnerability information set.

[0103] The node type matching vulnerability information set refers to a set of vulnerability-related information that is matched with the type of the entity node and is screened out from the vulnerability feature description. Through screening, unnecessary matching work can be reduced, and the efficiency of matching can be improved.

[0104] When performing screening, each entity node in the candidate attack chain segment is traversed, and the matching vulnerability-related information is found from the vulnerability feature description according to the node type. In an e-commerce platform, for a communication node, the vulnerability-related information affecting the node type of the communication node is screened out from the vulnerability feature description, and these information is combined into a node type matching vulnerability information set.

[0105] Step S344: The behavior atomic features of the entity nodes in the candidate attack chain segment are extracted, and semantic similarity calculation is performed between the behavior atomic features and the trigger behavior description of each vulnerability-related information in the node type matching vulnerability information set to obtain a behavior matching result.

[0106] Semantic similarity calculation refers to calculating the semantic similarity between two texts. The behavior matching result refers to the result obtained by comparing the semantic similarity between the behavior atomic features of the entity node and the trigger behavior description of the vulnerability-related information, which is used to judge whether the behavior of the entity node can trigger a vulnerability.

[0107] In the semantic similarity calculation, a text similarity algorithm in natural language processing is used. First, the behavior atomic features of the entity node and the trigger behavior description of the vulnerability-related information are processed by word segmentation, and the text is converted into a set of words. Then, the semantic similarity is determined by calculating the similarity between the word sets. If the semantic similarity is high, it means that the behavior atomic features and the trigger behavior description are semantically similar, and the behavior of the entity node may trigger the corresponding vulnerability. In the e-commerce platform, the behavior atomic feature of the entity node is "attempting to connect to the server through an unauthorized port", and the trigger behavior description of the node type matching vulnerability-related information set is "unauthorized port connection attempt". The semantic similarity of the two is calculated using the text similarity algorithm to obtain the behavior matching result.

[0108] Step S345: Mark the entity node corresponding to the highly similar vulnerability-related information in the behavior matching result as an entity node with vulnerability characteristics, and record the vulnerability-related information and the description of the degree of harm of the entity node to generate a potential risk node list.

[0109] The potential risk node list is a list containing information of entity nodes with vulnerability characteristics, including node identification, vulnerability-related information, and description of the degree of harm. By generating the potential risk node list, subsequent risk analysis and processing can be facilitated.

[0110] In generating the potential risk node list, the entity nodes in the candidate attack chain segment are traversed, and for each entity node, the behavior matching result is checked. If the behavior matching result is highly similar, the entity node is marked as an entity node with vulnerability characteristics. Then, the vulnerability-related information (such as the affected node type, trigger behavior description) and the description of the degree of harm corresponding to the entity node are recorded. These marked node information is combined into a potential risk node list. In the e-commerce platform, through semantic similarity calculation, it is found that the behavior atomic feature of a communication node is highly similar to the trigger behavior description of a certain vulnerability, and the communication node is marked as an entity node with vulnerability characteristics, and its corresponding vulnerability information and degree of harm are recorded and added to the potential risk node list.

[0111] Step S350: Take the potential risk nodes and corresponding candidate attack chain segments as inputs for the third round of reasoning, call the large model combined with the defense rule description to perform abnormality evaluation on the behavior hyperedge of the potential risk node, generate an abnormality evaluation result, and mark the potential risk node with high risk in the abnormality evaluation result as a primary cause risk node.

[0112] The defense rule description is a rule description about network security defense measures in the security field knowledge graph, such as prohibiting access to some ports, limiting the operation of some users, etc. The abnormality evaluation is an analysis and judgment on the behavior hyperedge of the potential risk node, evaluating whether it violates the defense rule and whether there is abnormal behavior. The abnormality evaluation result is the result obtained after the abnormality evaluation of the behavior hyperedge of the potential risk node, such as high risk, medium risk, low risk, etc. The main cause risk node refers to the main node that causes the risk. By marking the potential risk node with a high risk abnormality evaluation result as the main cause risk node, the root of the risk can be more accurately located.

[0113] In the third round of reasoning, the large model receives the potential risk nodes and the corresponding candidate attack chain fragments and defense rule descriptions as inputs. The large model first parses the defense rule description to understand the rule content. Then, for each behavior hyperedge of the potential risk node, it checks whether it violates the defense rule. For example, the defense rule may stipulate that connecting through unauthorized ports is prohibited. If the behavior hyperedge of the potential risk node involves connecting through unauthorized ports, the behavior hyperedge may be evaluated as high risk. The potential risk node with a high risk abnormality evaluation result is marked as the main cause risk node. In the e-commerce platform, the defense rule stipulates that unauthorized users are prohibited from calling certain key interfaces. The large model analyzes the behavior hyperedge of the potential risk node and finds that the behavior hyperedge of an interface node involves the operation of unauthorized users calling the key interface. The interface node is marked as the main cause risk node with a high risk abnormality evaluation result.

[0114] Step S360: Based on the main cause risk node, the behavior hyperedge connection path in the dynamic risk association hypergraph is traced back to generate an associated risk path containing the main cause risk node, associated entity nodes and behavior hyperedge sequence.

[0115] The associated risk path refers to a path formed by a series of entity nodes and behavior hyperedges associated with the main cause risk node by tracing back the behavior hyperedge connection path in the dynamic risk association hypergraph. By generating the associated risk path, the propagation and influence range of the risk can be more comprehensively understood.

[0116] As an implementation manner, step S360 can be implemented as steps S361-S365:

[0117] Step S361: Taking the main cause risk node as the starting search point, searching for the behavior hyperedge directly connected to the starting search point in the dynamic risk association hypergraph to obtain the first-level associated hyperedge.

[0118] The starting search point is the main cause risk node, which serves as the starting point of tracing back the behavior hyperedge connection path. The first-level associated hyperedge refers to the behavior hyperedge directly connected to the main cause risk node.

[0119] In the process of retrieval, all hyperedges connected with the primary risk node in the dynamic risk association hypergraph are found. In the e-commerce platform, if the primary risk node is a user interface node, all hyperedges connected with the user interface node in the hypergraph are found, and these hyperedges are taken as the first-level association hyperedge.

[0120] Step S362: For each first-level association hyperedge, other entity nodes connected by the hyperedge are extracted as first-level association nodes, and the weight value of the first-level association hyperedge and the corresponding time sequence constraint relationship are recorded.

[0121] The first-level association node refers to other entity nodes connected by the first-level association hyperedge except the primary risk node. The weight value refers to the comprehensive weight of the first-level association hyperedge, reflecting the importance of the hyperedge. The time sequence constraint relationship refers to the time sequence constraint and time overlap constraint relationship of the first-level association hyperedge.

[0122] In the process of extracting the first-level association node, for each first-level association hyperedge, other entity nodes connected by the hyperedge except the primary risk node are found. At the same time, the weight value of the first-level association hyperedge is recorded, which is the comprehensive weight calculated before. The corresponding time sequence constraint relationship of the first-level association hyperedge also needs to be recorded, which has been determined when the time sequence constraint set is generated. In the e-commerce platform, a first-level association hyperedge connects the primary risk node (user interface node) and a storage node, and the storage node is taken as the first-level association node, and the weight value and the corresponding time sequence constraint relationship of the hyperedge are recorded.

[0123] Step S363: Take the first-level association node as a new starting retrieval point, and repeat the above retrieval process until the boundary node of the dynamic risk association hypergraph is retrieved or the preset retrieval depth limit is reached, and obtain the multi-level association hyperedge and the multi-level association node.

[0124] The boundary node refers to the node at the edge position in the dynamic risk association hypergraph, which may be a node without more connections. The preset retrieval depth limit is a preset limit value of the retrieval depth, which is used to control the range of retrieval. The multi-level association hyperedge refers to the association hyperedge of different levels obtained by multiple retrievals. The multi-level association node refers to the association node of different levels connected by the multi-level association hyperedge.

[0125] When searching with the primary associated node as the new starting search point, the process of steps S361 and S362 is repeated. For each primary associated node, the behavior hyperedge connected thereto is found in the dynamic risk association hypergraph, a new associated hyperedge (secondary associated hyperedge) is obtained, and other entity nodes connected by the hyperedge are extracted as secondary associated nodes. The above process is repeated with the secondary associated nodes as the starting search point until the boundary node is searched or the preset search depth limit is reached. In the e-commerce platform, starting from the primary associated node (storage node), the hyperedge connected thereto is found, the secondary associated hyperedge and the secondary associated node are obtained. This process is repeatedly performed until the termination condition is met, and the multi-level associated hyperedge and the multi-level associated node are obtained.

[0126] Step S364: According to the time sequence constraint relationship in the time sequence constraint set, the multi-level associated hyperedge is time-sequentially sorted to generate a behavior hyperedge sequence arranged in time sequence.

[0127] The time sequence sorting is to arrange the multi-level associated hyperedge according to the time sequence constraint relationship in the time sequence constraint set, so that the hyperedge is arranged in the order of occurrence time. The behavior hyperedge sequence refers to the sequence formed by the multi-level associated hyperedge arranged in time sequence.

[0128] In the time sequence sorting, first, the time sequence constraint relationship of the multi-level associated hyperedge is obtained from the time sequence constraint set. Then, the multi-level associated hyperedge is sorted according to the constraint relationship. A sorting algorithm such as bubble sort or quick sort can be used to sort the hyperedge. In the e-commerce platform, according to the time sequence constraint set, it is known that there is a time sequence relationship between certain multi-level associated hyperedges, such as hyperedge A must occur before hyperedge B. The sorting algorithm is used to sort these hyperedges to generate a behavior hyperedge sequence arranged in time sequence.

[0129] Step S365: Arranging the primary cause risk node and the multi-level associated node in the connection relationship of the behavior hyperedge sequence to generate an associated risk path containing the node sequence and the corresponding behavior hyperedge sequence.

[0130] The associated risk path is a path formed by arranging the primary cause risk node and the multi-level associated node in the connection relationship of the behavior hyperedge sequence, containing the sequence of nodes and the corresponding behavior hyperedge sequence.

[0131] In generating the associated risk path, the main cause risk node, the multi-level associated node are arranged in turn according to the connection relationship of the behavior hyperedge sequence. For each behavior hyperedge, record the order of the connected nodes. In the e-commerce platform, the behavior hyperedge sequence connects the main cause risk node (user interface node), the first-level associated node (storage node) and the second-level associated node (server node). According to the connection relationship of the behavior hyperedge sequence, these nodes are arranged in turn to generate the associated risk path containing the node order (user interface node->storage node->server node) and the corresponding behavior hyperedge sequence.

[0132] Step S370: Calculate the product of the weight value of each behavior hyperedge in the associated risk path and the abnormal evaluation result as the evidence confidence of each link on the path.

[0133] The evidence confidence refers to the evidence confidence of each link on the associated risk path, which is obtained by calculating the product of the weight value of the behavior hyperedge and the abnormal evaluation result.

[0134] In calculating the evidence confidence, for each behavior hyperedge in the associated risk path, its weight value is obtained, which is the comprehensive weight calculated before. At the same time, the abnormal evaluation result corresponding to the behavior hyperedge is obtained, which can be represented by a numerical value or level representing the risk degree. Multiply the weight value and the abnormal evaluation result to get the evidence confidence corresponding to the behavior hyperedge. In the e-commerce platform, the weight value of a behavior hyperedge is a certain weight degree, and the abnormal evaluation result is a high risk degree. Multiplying the two gets the evidence confidence of the behavior hyperedge.

[0135] Step S380: Integrate the main cause risk node, the associated risk path and the evidence confidence of each link in the reasoning order to generate the risk attribution reasoning chain.

[0136] The risk attribution reasoning chain is a chain formed by integrating the main cause risk node, the associated risk path and the evidence confidence of each link in the reasoning order, which is used to clearly show the attribution process and related information of the risk.

[0137] In integration, first determine the order of reasoning, usually from the main cause risk node, along the associated risk path step by step. Put the main cause risk node at the starting position of the reasoning chain, and then list the nodes and behavior hyperedges in the associated risk path in turn, and list the evidence confidence of each link. In the e-commerce platform, first list the main cause risk node (such as the user interface node with abnormal operation), then list the associated nodes (such as storage node, server node, etc.) and behavior hyperedges connecting them in order according to the associated risk path. For each node and hyperedge, attach the corresponding evidence confidence. In this way, a complete risk attribution reasoning chain is formed, which clearly shows how the risk propagates through a series of nodes and hyperedges from the main cause risk node, as well as the credibility of each link. For example, the reasoning chain may be: main cause risk node (user interface node, evidence confidence is high degree)->first level associated hyperedge (connecting user interface node and storage node, evidence confidence is certain degree)->first level associated node (storage node, evidence confidence is certain degree)->second level associated hyperedge (connecting storage node and server node, evidence confidence is certain degree)->second level associated node (server node, evidence confidence is certain degree).

[0138] Step S400: According to the risk attribution reasoning chain and the time sequence constraint set of the dynamic risk associated hypergraph, construct a risk evolution probability model, and based on this, calculate the short-term diffusion probability and long-term evolution trend vector of each associated risk path, and get the risk evolution path prediction result.

[0139] The risk evolution probability model is a model for describing how the risk evolves and propagates in the network, which combines the risk attribution reasoning chain and the time sequence constraint set of the dynamic risk associated hypergraph, and considers the propagation path, time factor and credibility of each link of the risk. The short-term diffusion probability is the possibility of the diffusion of the associated risk path in a short time, which reflects the propagation trend of the risk in the near future. The long-term evolution trend vector is a vector describing the evolution direction and rate of the associated risk path in a long time, which is used to predict the long-term development trend of the risk. The risk evolution path prediction result is the result obtained by combining the short-term diffusion probability and the long-term evolution trend vector, which is used to guide the prevention and response measures of network security.

[0140] In the e-commerce platform, according to the risk attribution reasoning chain, the risk propagates from the user interface node to the storage node and the server node through a series of behavior hyper-edges, and each link has a corresponding evidence confidence. At the same time, the time constraint set specifies the time sequence and time interval of these behavior hyper-edges. Using this information to construct the path state transition relationship, considering the influence of time factors on risk propagation, adjust it with time decay. For example, if the time interval between two behavior hyper-edges is long, the possibility of risk transfer from one hyper-edge to another will decrease accordingly. Based on the adjusted state transition relationship, a risk evolution probability model is constructed, and the short-term diffusion probability and long-term evolution trend vector of each associated risk path are calculated through the model to obtain the risk evolution path prediction result.

[0141] As an implementation, step S400 can be implemented as steps S410-S470:

[0142] Step S410: Extract the associated risk path and the evidence confidence of each link from the risk attribution reasoning chain, and extract the time sequence constraint relationship corresponding to the associated risk path from the time constraint set of the dynamic risk association hypergraph.

[0143] The associated risk path is the risk propagation path displayed in the risk attribution reasoning chain, which contains a series of entity nodes and behavior hyper-edges. The evidence confidence of each link is the confidence information corresponding to each node and hyper-edge in the risk attribution reasoning chain. The time sequence constraint relationship is the constraint condition in the dynamic risk association hypergraph that specifies the time sequence and time interval of the behavior hyper-edges on the associated risk path.

[0144] In extracting this information, the risk attribution reasoning chain is analyzed in detail, and the associated risk path is identified, and the entity nodes and behavior hyper-edges in the path are recorded in turn. At the same time, the evidence confidence corresponding to each node and hyper-edge is extracted. For the time constraint set of the dynamic risk association hypergraph, the time sequence constraint relationship corresponding to the associated risk path is selected, which may include time sequence constraints and time overlap constraints, etc. In the e-commerce platform, the associated risk path from the user interface node to the storage node and then to the server node is extracted from the risk attribution reasoning chain, and the evidence confidence of each node and hyper-edge. From the time constraint set, the time sequence relationship and time interval information of the behavior hyper-edges on the associated risk path are extracted.

[0145] Step S420: Based on the sequence of behavior hyper-edges in the associated risk path and the evidence confidence, construct the path state transition relationship, which represents the possibility value of transferring from one behavior hyper-edge to the next behavior hyper-edge, which is determined by the product of the evidence confidence of the current behavior hyper-edge and the hyper-edge weight.

[0146] The path state transition relationship describes the possibility of risk transition from one behavior hyperedge to the next behavior hyperedge in the associated risk path. The sequence of behavior hyperedges is a sequence formed by arranging the behavior hyperedges in the associated risk path in order. The evidence confidence reflects the credibility of each behavior hyperedge, and the hyperedge weight represents the importance of the behavior hyperedge.

[0147] In constructing the path state transition relationship, for each behavior hyperedge in the associated risk path, the evidence confidence of the behavior hyperedge is multiplied by the hyperedge weight to obtain a possibility value of transition from the behavior hyperedge to the next behavior hyperedge. For example, in the associated risk path of the e-commerce platform, the first behavior hyperedge connects the user interface node and the storage node, and its evidence confidence is a certain degree, and the hyperedge weight is a certain degree. Multiplying the two values obtains the possibility value of transition from the hyperedge to the next behavior hyperedge connecting the storage node and the server node. Such calculation is sequentially performed on each behavior hyperedge in the associated risk path to construct a complete path state transition relationship.

[0148] Step S430: Adjust the path state transition relationship by time decay according to the time interval information in the time sequence constraint relationship, and generate a time-aware state transition relationship.

[0149] The time interval information is the interval between the occurrence times of adjacent behavior hyperedges in the associated risk path specified in the time sequence constraint relationship. The time decay adjustment is an adjustment of the path state transition relationship considering the influence of time factors on the possibility of risk propagation. The time-aware state transition relationship is a state transition relationship that can reflect the influence of time factors on risk transition after time decay adjustment.

[0150] In the e-commerce platform, the time interval of adjacent behavior hyperedges in the associated risk path may be different. For example, there is a certain time interval between the behavior hyperedge from the user interface node to the storage node and the behavior hyperedge from the storage node to the server node. According to the preset time decay coefficient table, the time decay coefficient corresponding to the time interval is determined, and the possibility value in the path state transition relationship calculated before is multiplied by the coefficient to obtain the transition possibility value after time decay. These values are normalized to be within a reasonable range, and are updated to the path state transition relationship to generate the time-aware state transition relationship.

[0151] As an implementation manner, step S430 can be implemented as steps S431-S434:

[0152] Step S431: Extract the time interval information in the time sequence constraint relationship to determine the occurrence time interval of adjacent behavior hyperedges in the associated risk path.

[0153] The time interval information is the difference between the occurrence times of adjacent behavior hyper-edges on the associated risk path recorded in the time constraint relationship. By extracting these information, the time interval of risk propagation between different behavior hyper-edges can be understood.

[0154] In the process of extracting the time interval information, the time constraint relationship corresponding to the associated risk path in the time constraint set is analyzed. The occurrence time stamps of adjacent behavior hyper-edges are found, and the difference between them is calculated to obtain the time interval. In the e-commerce platform, for adjacent behavior hyper-edges on the associated risk path, their occurrence time stamps are obtained from the time constraint set, and the time interval is calculated, for example, the time interval from one behavior hyper-edge connecting the user interface node and the storage node to another behavior hyper-edge connecting the storage node and the server node.

[0155] Step S432: Based on the preset time decay coefficient table, the corresponding time decay coefficient is determined according to the time interval information. The longer the time interval, the smaller the corresponding time decay coefficient.

[0156] The preset time decay coefficient table is a table set in advance, which records the time decay coefficients corresponding to different time intervals. The time decay coefficient is used to adjust the likelihood value in the path state transition relationship, reflecting the influence of time factor on risk propagation.

[0157] In the process of determining the time decay coefficient, according to the time interval information obtained in step S431, the corresponding coefficient is found in the preset time decay coefficient table. Because the longer the time interval, the smaller the possibility of risk propagation in this period of time, the longer the time interval, the smaller the corresponding time decay coefficient. In the e-commerce platform, if the time interval of adjacent behavior hyper-edges is long, the coefficient corresponding to this time interval is found in the time decay coefficient table, which will be relatively small.

[0158] Step S433: Multiply the likelihood value in the path state transition relationship with the corresponding time decay coefficient to obtain the transition likelihood value after time decay.

[0159] The likelihood value in the path state transition relationship is the likelihood of transition from one behavior hyper-edge to the next behavior hyper-edge calculated before. By multiplying it with the corresponding time decay coefficient, the adjusted transition likelihood value considering the time factor can be obtained.

[0160] In the process of multiplication, for each behavior hyper-edge in the associated risk path, the likelihood value in the corresponding path state transition relationship is multiplied by the time decay coefficient determined in step S432. In the e-commerce platform, for the behavior hyper-edge connecting the user interface node and the storage node, the likelihood value of its transition to the next behavior hyper-edge is multiplied by the corresponding time decay coefficient to obtain the transition likelihood value after time decay.

[0161] Step S434: Normalizing the time-decayed transition possibility values, and updating the normalized transition possibility values to the path state transition relationship, to generate a time-aware state transition relationship.

[0162] The normalization process is to adjust the time-decayed transition possibility values to a unified range, so that they have comparability and reasonableness. Updating the path state transition relationship is to replace the original possibility values with the normalized transition possibility values, forming a time-aware state transition relationship that can reflect the influence of time factors.

[0163] When performing normalization processing, a normalization algorithm is used to process the time-decayed transition possibility values. The normalization method is to map these values to a set interval, such as [0, 1]. The normalized transition possibility values are updated to the path state transition relationship, replacing the original possibility values. In the e-commerce platform, the time-decayed transition possibility values of all adjacent behavior super-edges are normalized, and the processed results are updated to the path state transition relationship, to generate a time-aware state transition relationship.

[0164] Step S440: Based on the time-aware state transition relationship, a risk evolution probability model is constructed, which includes a short-term prediction layer and a long-term prediction layer. The short-term prediction layer is used to predict the risk diffusion situation in the future first preset time period, and the long-term prediction layer is used to predict the risk evolution situation in the future second preset time period.

[0165] The risk evolution probability model is a comprehensive model used to predict the evolution and diffusion of risks in different time periods. The short-term prediction layer and the long-term prediction layer are two components of the model, responsible for prediction in different time scales.

[0166] When constructing the risk evolution probability model, the time-aware state transition relationship is used as the basis. The time-aware state transition relationship is used as the input of the model, and through a series of calculations and processing, the short-term prediction layer and the long-term prediction layer are constructed. The short-term prediction layer will predict the diffusion of risks on the associated risk path in the future first preset time period based on the current state and state transition relationship, such as the nodes to which the risk may spread and the size of the diffusion possibility. The long-term prediction layer will consider a longer time to predict the evolution direction and rate of risks in the future second preset time period. In the e-commerce platform, the risk evolution probability model is constructed based on the time-aware state transition relationship, which can help predict whether the risk will spread from the current node to other nodes in the future short time, and the overall development trend of the risk in a longer time.

[0167] Step S450: calculating the diffusion possibility value of each associated risk path in the future first preset time period as a short-term diffusion probability through the short-term prediction layer.

[0168] The short-term diffusion probability refers to the possibility of the diffusion of the associated risk path in the future first preset time period. The short-term prediction layer analyzes and calculates each associated risk path by using the information in the risk evolution probability model and combining the time-aware state transition relationship.

[0169] In the e-commerce platform, for an associated risk path, its current state (such as the current behavior hyperedge connecting the user interface node and the storage node, with a certain degree of evidence confidence) and time-aware state transition relationship are input into the short-term prediction layer. The transition calculation module calculates the transition probability from the current state to the next possible state (such as the behavior hyperedge connecting the storage node and the server node) according to the state transition relationship. The transition probability is multiplied by the evidence confidence of the current behavior hyperedge to obtain a single-step diffusion possibility value. The total diffusion possibility value is obtained by cumulatively summing all single-step diffusion possibility values within the first preset time period in the future. The total diffusion possibility value is converted into a probability value in the preset numerical interval by a preset conversion method, serving as the short-term diffusion probability of the associated risk path. Illustratively, the calculation of the transition probability can be based on the path state transition relationship, which represents the possibility value of transitioning from one behavior hyperedge to the next behavior hyperedge, determined by the product of the evidence confidence of the current behavior hyperedge and the hyperedge weight. For example, the associated risk path and the evidence confidence of each link are extracted from the risk attribution reasoning chain, and the time interval information in the time interval constraint relationship corresponding to the associated risk path is extracted from the time interval constraint set of the dynamic risk association hypergraph. Based on the behavior hyperedge sequence and the evidence confidence in the associated risk path, the path state transition relationship is constructed. For each behavior hyperedge in the associated risk path, the evidence confidence is multiplied by the hyperedge weight to obtain the initial transition possibility value from the behavior hyperedge to the next behavior hyperedge. Then, the time interval information in the time interval constraint relationship is combined to adjust the path state transition relationship for time decay. The time interval information in the time interval constraint relationship is extracted to determine the occurrence time interval of adjacent behavior hyperedges in the associated risk path. Based on a preset time decay coefficient table, the corresponding time decay coefficient is determined according to the time interval information, and the longer the time interval, the smaller the corresponding time decay coefficient. The initial transition possibility value in the path state transition relationship is multiplied by the corresponding time decay coefficient to obtain the time-decayed transition possibility value. The time-decayed transition possibility value is normalized, and the normalized transition possibility value is updated to the path state transition relationship to generate a time-aware state transition relationship. The transition calculation module calculates the transition probability according to this time-aware state transition relationship. When the associated risk path is in the current state, the normalized transition possibility value from the current behavior hyperedge to the next possible behavior hyperedge is found in the time-aware state transition relationship, which is the transition probability from the current state to the next possible state. For example, if the current associated risk path is in the behavior hyperedge connecting the user interface node and the storage node, the transition calculation module will find the normalized transition possibility value from this hyperedge to the behavior hyperedge connecting the storage node and the server node in the time-aware state transition relationship as the transition probability.

[0170] As an implementation, step S450 can be implemented as steps S451-S456 as follows:

[0171] Step S451: input the current state of the associated risk path and the time-aware state transition relationship into the short-term prediction layer, which includes a state input module, a transition calculation module, and a probability output module.

[0172] The current state of the associated risk path describes the position and state of the risk at the current time, including the current behavior hyperedge identifier and the corresponding evidence confidence. The time-aware state transition relationship reflects the possibility of risk transition between different behavior hyperedges considering time factors. The state input module of the short-term prediction layer is used to receive these information, the transition calculation module is used to calculate the transition probability according to the input information, and the probability output module is used to output the final short-term diffusion probability.

[0173] When inputting information, the current state information of the associated risk path is accurately transmitted to the state input module. At the same time, the time-aware state transition relationship is also input into the short-term prediction layer. In the e-commerce platform, the current behavior hyperedge of the associated risk path is in the connection between the user interface node and the storage node, and the evidence confidence of the hyperedge is input into the state input module of the short-term prediction layer together with the time-aware state transition relationship.

[0174] Step S452: receive the current state of the associated risk path through the state input module, which contains the current behavior hyperedge identifier and the corresponding evidence confidence.

[0175] The state input module is a module in the short-term prediction layer specially used to receive the current state information of the associated risk path. It processes the received current behavior hyperedge identifier and corresponding evidence confidence to provide a basis for subsequent transition probability calculation.

[0176] When receiving the current state, the state input module will verify and organize the input information to ensure the accuracy and integrity of the information. In the e-commerce platform, the state input module receives the current behavior hyperedge identifier of the associated risk path (such as the set identifier of the behavior hyperedge connecting the user interface node and the storage node) and the corresponding evidence confidence (such as a certain degree of confidence).

[0177] Step S453: calculate the transition probability of the current state to the next possible state based on the time-aware state transition relationship through the transition calculation module.

[0178] The transition calculation module is a module in the short-term prediction layer responsible for calculating the transition probability. It calculates the probability of transitioning from the current state to the next possible state according to the time-aware state transition relationship combined with the current state of the associated risk path.

[0179] When calculating the transition probability, the transition calculation module looks up the transition possibility value corresponding to the current behavior hyperedge in the time-aware state transition relationship. For example, in an e-commerce platform, according to the behavior hyperedge connecting the user interface node and the storage node, the transition possibility value of the hyperedge transitioning to the next behavior hyperedge connecting the storage node and the server node is looked up in the time-aware state transition relationship, and is taken as the transition probability.

[0180] Step S454: multiplying the transition probability and the evidence confidence of the current behavior hyperedge to obtain a single-step diffusion possibility value.

[0181] The single-step diffusion possibility value reflects the possibility of the risk taking one step of transition in the current state. By multiplying the transition probability and the evidence confidence of the current behavior hyperedge, the possibility of transition and the credibility of the current state can be considered comprehensively.

[0182] When performing multiplication operation, the transition probability obtained in step S453 is multiplied by the evidence confidence of the current behavior hyperedge received in step S452. In an e-commerce platform, the transition probability of the behavior hyperedge connecting the user interface node and the storage node transitioning to the behavior hyperedge connecting the storage node and the server node is multiplied by the evidence confidence of the behavior hyperedge connecting the user interface node and the storage node, to obtain a single-step diffusion possibility value.

[0183] Step S455: cumulatively summing all single-step diffusion possibility values in the future first preset time period to obtain a total diffusion possibility value.

[0184] The total diffusion possibility value is the overall possibility of the risk diffusion in the future first preset time period. By cumulatively summing all single-step diffusion possibility values, the comprehensive possibility of the risk diffusion from the current state in the time period can be obtained.

[0185] When performing cumulative summation, for each possible state transition in the future first preset time period, the single-step diffusion possibility value thereof is calculated, and these values are added. In an e-commerce platform, considering multiple possible state transitions on the associated risk path in the future first preset time period, the single-step diffusion possibility values of each transition are cumulatively added to obtain a total diffusion possibility value.

[0186] Step S456: converting the total diffusion possibility value into a probability value in a preset numerical interval through a sigmoid function, as a short-term diffusion probability.

[0187] The preset numerical interval is, for example, [0, 1], representing a probability range of risk diffusion. When conversion is performed, a preset conversion rule or algorithm is used to process the total diffusion possibility value. In the e-commerce platform, the total diffusion possibility value obtained by cumulative summation is converted into a probability value in the [0, 1] interval by a sigmoid function, as a short-term diffusion probability of the associated risk path in the first preset time period in the future.

[0188] Step S460: Calculate the evolution direction and rate of each associated risk path in the second preset time period in the future by the long-term prediction layer, as a long-term evolution trend vector.

[0189] The long-term evolution trend vector is a vector describing the evolution direction and rate of the associated risk path in the second preset time period in the future. The long-term prediction layer uses information in the risk evolution probability model, combines the time-aware state transition relationship and the historical behavior hyperedge sequence of the associated risk path, and analyzes and calculates each associated risk path.

[0190] In the e-commerce platform, for an associated risk path, its historical behavior hyperedge sequence (such as the order and situation of the behavior hyperedges connecting the user interface node, the storage node and the server node before) and the time-aware state transition relationship are input into the long-term prediction layer. The sequence encoding module encodes the historical behavior hyperedge sequence to generate a historical evolution feature vector. The trend prediction module predicts the appearance probability of each possible behavior hyperedge in the second preset time period in the future according to the vector and the state transition relationship. The top several behavior hyperedges with the highest probability are selected as the main evolution direction, and the average appearance interval of these behavior hyperedges is calculated as the evolution rate. The main evolution direction and the evolution rate are combined into the long-term evolution trend vector.

[0191] As an implementation manner, step S460 can be implemented as steps S461-S466:

[0192] Step S461: Input the historical behavior hyperedge sequence of the associated risk path and the time-aware state transition relationship into the long-term prediction layer, which includes a sequence encoding module, a trend prediction module and a vector generation module.

[0193] The historical behavior hyperedge sequence of the associated risk path records the propagation path and order of the risk in the past, reflecting the historical evolution of the risk. The time-aware state transition relationship reflects the possibility of the risk transferring between different behavior hyperedges considering the time factor. The sequence encoding module of the long-term prediction layer is used to encode the historical behavior hyperedge sequence, the trend prediction module is used to predict the trend according to the encoded information and the state transition relationship, and the vector generation module is used to generate the final long-term evolution trend vector.

[0194] When inputting information, the historical behavior hyperedge sequence associated with the risk path and the time-aware state transition relationship are accurately passed to the long-term prediction layer. In an e-commerce platform, the historical behavior hyperedge sequence associated with the risk path (such as the behavior hyperedge sequence from the connection user interface node to the storage node, and then to the server node) and the time-aware state transition relationship are input to the long-term prediction layer.

[0195] Step S462: The sequence encoding module performs time sequence encoding on the historical behavior hyperedge sequence to generate a historical evolution feature vector.

[0196] The sequence encoding module is a module in the long-term prediction layer that is specifically used to encode the historical behavior hyperedge sequence. It analyzes and processes the behavior hyperedge in the sequence, considers its time sequence and features, and generates a feature vector that can reflect the historical evolution.

[0197] When performing time sequence encoding, the sequence encoding module performs feature extraction and quantization on each behavior hyperedge in the historical behavior hyperedge sequence, and combines these features into a vector. For example, it considers the connection nodes, hyperedge weights, and occurrence times of the behavior hyperedge. In an e-commerce platform, the sequence encoding module encodes the historical behavior hyperedge sequence associated with the risk path, quantizes and combines the relevant features of each behavior hyperedge, and generates a historical evolution feature vector.

[0198] Step S463: The trend prediction module predicts the occurrence probability of each possible behavior hyperedge in the future second preset time period based on the historical evolution feature vector and the time-aware state transition relationship.

[0199] The trend prediction module is a module in the long-term prediction layer that is responsible for trend prediction. It predicts the occurrence probability of each possible behavior hyperedge in the future second preset time period based on the historical evolution feature vector generated by the sequence encoding module and the time-aware state transition relationship.

[0200] When performing prediction, the trend prediction module uses machine learning or statistical methods to combine information from the historical evolution feature vector and the state transition relationship. For example, it uses a probability-based model to predict the probability of each behavior hyperedge appearing in the future based on historical conditions and state transition probabilities. In an e-commerce platform, the trend prediction module predicts the occurrence probability of behavior hyperedges connecting different nodes in the future second preset time period based on the historical evolution feature vector and the time-aware state transition relationship.

[0201] Step S464: According to the occurrence probability, the top N behavior hyperedges with the highest probability are selected as the main evolution direction, N > 0.

[0202] The main evolution direction is the direction in which the associated risk path is most likely to develop in the future second preset time period. By screening the top N behavior hyper-edges with the highest probability according to the occurrence probability, the main evolution direction of the risk can be determined.

[0203] In the screening, the occurrence probabilities of the possible behavior hyper-edges obtained in step S463 are sorted, and the top N behavior hyper-edges with the highest probability are selected. In the e-commerce platform, assuming that N is 3, the occurrence probabilities of the predicted behavior hyper-edges are sorted, and the top 3 behavior hyper-edges with the highest probability are selected as the main evolution direction.

[0204] Step S465: Calculate the average occurrence interval of the behavior hyper-edges corresponding to the main evolution direction in the future second preset time period as the evolution rate.

[0205] The evolution rate reflects the speed of the associated risk path developing along the main evolution direction in the future second preset time period. By calculating the average occurrence interval of the behavior hyper-edges corresponding to the main evolution direction in the time period, the evolution rate can be obtained.

[0206] In calculating the average occurrence interval, the occurrence intervals of the behavior hyper-edges corresponding to the main evolution direction are counted and averaged according to the occurrence times of the behavior hyper-edges predicted by the trend prediction module. In the e-commerce platform, for the behavior hyper-edges corresponding to the determined main evolution direction, the average occurrence interval of them is calculated as the evolution rate according to the predicted occurrence times.

[0207] Step S466: Combine the main evolution direction and the evolution rate into a multi-dimensional vector, the dimension of the multi-dimensional vector corresponds to the number of main evolution directions, and the vector element value corresponds to the evolution rate, to generate a long-term evolution trend vector.

[0208] The long-term evolution trend vector is a vector that comprehensively represents the evolution direction and rate of the associated risk path in the future second preset time period. By combining the main evolution direction and the evolution rate into a multi-dimensional vector, the long-term evolution trend of the risk can be intuitively displayed.

[0209] In combining the vector, the dimension of the multi-dimensional vector is determined according to the number of main evolution directions, and the evolution rate corresponding to each main evolution direction is taken as the element value of the vector. In the e-commerce platform, assuming that there are 3 main evolution directions, the evolution rates corresponding to the 3 main evolution directions are combined into a three-dimensional vector as the long-term evolution trend vector.

[0210] Step S470: Integrate the short-term diffusion probability and the long-term evolution trend vector to generate a risk evolution path prediction result containing the path identifier, the short-term diffusion probability, and the long-term evolution trend vector.

[0211] The risk evolution path prediction result is obtained by comprehensively combining the short-term diffusion probability and the long-term evolution trend vector, and is used to comprehensively describe the evolution of the associated risk path in different time scales. The path identifier is used to uniquely identify each associated risk path.

[0212] In the integration, the path identifier, the short-term diffusion probability and the long-term evolution trend vector of each associated risk path are combined together. In the e-commerce platform, for each associated risk path, the path identifier, the calculated short-term diffusion probability and the generated long-term evolution trend vector are integrated to form the risk evolution path prediction result. For example, the path identifier of an associated risk path is a set number, the short-term diffusion probability is a certain degree of probability value, and the long-term evolution trend vector is a multi-dimensional vector. They are combined into a complete result to guide the network security prevention measures of the e-commerce platform.

[0213] As an implementation manner, the method provided by the embodiment of the present application further includes the step of generating a dynamic security protection scheme including a node protection priority, a path blocking strategy and an interface access control rule based on the short-term diffusion probability and the long-term evolution trend vector in the risk evolution path prediction result, which can be implemented as the following steps S500-S900:

[0214] Step S500: Extract the short-term diffusion probability and the long-term evolution trend vector of each associated risk path from the risk evolution path prediction result.

[0215] The risk evolution path prediction result contains important information of each associated risk path. The short-term diffusion probability reflects the diffusion possibility of the risk in the short term, and the long-term evolution trend vector describes the evolution direction and rate of the risk in the long term.

[0216] In the information extraction, the risk evolution path prediction result is analyzed in detail, and for each associated risk path, its corresponding short-term diffusion probability and long-term evolution trend vector are extracted. In the e-commerce platform, the short-term diffusion probability (such as a certain probability value) and the long-term evolution trend vector (such as a multi-dimensional vector) of each associated risk path are extracted from the risk evolution path prediction result.

[0217] Step S600: Sort the associated risk paths according to the size of the short-term diffusion probability. The larger the short-term diffusion probability of an associated risk path, the higher the node protection priority of the associated risk path.

[0218] The node protection priority refers to the priority order of protecting the nodes on the associated risk path. By sorting the associated risk paths according to the short-term diffusion probability, it can be determined which nodes on the path need to be protected in priority to reduce the possibility of risk diffusion.

[0219] In the sorting process, the short-term diffusion probabilities of all associated risk paths are collected, and these probability values are compared and sorted. The associated risk path with a greater short-term diffusion probability has a higher priority in protection. In the e-commerce platform, the short-term diffusion probabilities of the associated risk paths are sorted, and the nodes on the associated risk path with a greater short-term diffusion probability, such as the path connecting the user interface node and the storage node, are given priority in arranging protection measures.

[0220] As an implementation, step S600 can be implemented as steps S610-S660:

[0221] Step S610: Collect the short-term diffusion probabilities of all associated risk paths and establish a short-term diffusion probability list.

[0222] The short-term diffusion probability list is a list containing the short-term diffusion probabilities of all associated risk paths, which is used for subsequent sorting and analysis.

[0223] In collecting the short-term diffusion probabilities, the short-term diffusion probabilities of each associated risk path are sorted and recorded from the information extracted in step S500 to form a list. In the e-commerce platform, the short-term diffusion probabilities of each associated risk path are recorded in turn to establish a short-term diffusion probability list.

[0224] Step S620: Sort the short-term diffusion probability list in descending order to obtain a sorted associated risk path sequence.

[0225] Descending order sorting is to arrange the probability values in the short-term diffusion probability list from large to small. The sorted associated risk path sequence reflects the arrangement order of each associated risk path according to the size of the short-term diffusion probability.

[0226] In descending order sorting, the short-term diffusion probability list is processed using a sorting algorithm. For example, a quicksort algorithm can be used. In the e-commerce platform, the short-term diffusion probability list is sorted in descending order to obtain a sorted associated risk path sequence, and the path with a greater short-term diffusion probability is arranged in front.

[0227] Step S630: Extract the entity nodes contained in each associated risk path in the sorted associated risk path sequence, and record the node type and position in the path.

[0228] The entity node is a key element on the associated risk path, and recording the node type and position in the path helps to determine the importance and protection needs of the node.

[0229] In the extraction of entity node information, for each path in the sorted associated risk path sequence, the entity nodes contained therein are analyzed to determine the node type (such as communication node, storage node, interface node) and position in the path (such as starting node, intermediate node, ending node). In the e-commerce platform, for each sorted associated risk path, the entity nodes contained therein are extracted, such as user interface nodes, storage nodes, server nodes, etc., and their types and positions in the path are recorded.

[0230] Step S640: For the associated risk paths containing the same entity node, the highest short-term diffusion probability is taken as the comprehensive risk probability of the entity node.

[0231] The comprehensive risk probability is an indicator for evaluating the risk size faced by the entity node. For the associated risk paths containing the same entity node, the highest short-term diffusion probability can more accurately reflect the risk degree of the node.

[0232] In determining the comprehensive risk probability, the sorted associated risk path sequence is traversed to find paths containing the same entity node, their short-term diffusion probabilities are compared, and the maximum value is taken as the comprehensive risk probability of the entity node. In the e-commerce platform, if multiple associated risk paths contain the same storage node, the short-term diffusion probabilities of these paths are compared, and the maximum value is taken as the comprehensive risk probability of the storage node.

[0233] Step S650: Sort the entity nodes according to the size of the comprehensive risk probability, and the entity node with a larger comprehensive risk probability has a higher protection priority.

[0234] Sorting the entity nodes according to the comprehensive risk probability can determine their protection priorities. The entity node with a larger comprehensive risk probability indicates that it faces a higher risk and needs to be protected first.

[0235] In the sorting process, the comprehensive risk probabilities of all entity nodes are compared and sorted. A sorting algorithm is used to arrange the entity nodes in descending order of comprehensive risk probability. In the e-commerce platform, the comprehensive risk probabilities of the entity nodes are sorted, and the entity nodes with larger comprehensive risk probabilities, such as an important server node, are given higher protection priorities.

[0236] Step S660: Assign a priority level to each entity node to generate a node protection priority list containing node identification, node type, and protection priority level.

[0237] The node protection priority list is a list that records the protection priority information of each entity node, facilitating the subsequent development of protection schemes.

[0238] In the step of assigning priority levels, the entity nodes are divided into different priority levels according to the sorted results in the step S650. For example, they can be divided into three levels: high, medium and low. Each entity node is assigned a corresponding priority level, and its node identifier, node type and protection priority level are recorded. In the e-commerce platform, the priority level of each entity node is assigned, and a node protection priority list is generated, such as the node identifier of the user interface node with the set number, the node type of the interface node, and the protection priority level of high.

[0239] Step S700: Based on the main evolution direction in the long-term evolution trend vector, determine the key behavior hyperedge that needs to be blocked, and generate a path blocking strategy. The path blocking strategy contains the behavior hyperedge identifier that needs to be blocked and the blocking execution time point.

[0240] The key behavior hyperedge is a hyperedge that plays a key role in the risk propagation in the associated risk path. By determining the key behavior hyperedge that needs to be blocked, the further spread of risk can be effectively prevented. The path blocking strategy is a strategy formulated to block the risk propagation, which contains the behavior hyperedge identifier that needs to be blocked and the blocking execution time point.

[0241] In the determination of the key behavior hyperedge, the main evolution direction in the long-term evolution trend vector is analyzed, and the behavior hyperedge related to the main evolution direction is found out. According to the importance of these behavior hyperedges and their influence on risk propagation, the key behavior hyperedge that needs to be blocked is determined. At the same time, according to the time sequence constraint relationship in the time sequence constraint set, the expected occurrence time point of the key behavior hyperedge that needs to be blocked is calculated as the blocking execution time point. In the e-commerce platform, the main evolution direction is determined according to the long-term evolution trend vector, and the behavior hyperedge related to the main evolution direction is found out, such as the hyperedge connecting the storage node and the server node. The key behavior hyperedge is determined. According to the time sequence constraint set, the expected occurrence time point of the hyperedge is calculated as the blocking execution time point, and the path blocking strategy is generated.

[0242] As an implementation, the step S700 can be implemented as the following steps S710-S750:

[0243] Step S710: Analyze the main evolution direction in the long-term evolution trend vector, and determine the behavior hyperedge sequence that may appear in the future second preset time period.

[0244] The main evolution direction reflects the main development direction of the associated risk path in the future second preset time period. By analyzing the main evolution direction in the long-term evolution trend vector, the behavior hyperedge sequence that may appear can be determined.

[0245] In the analysis, the long-term evolution trend vector is analyzed, and the behavior hyperedge sequence that may appear in the future second preset time period is determined according to the elements in the vector and the corresponding behavior hyperedge information. In the e-commerce platform, the long-term evolution trend vector is analyzed, and the behavior hyperedge sequence that may appear in the future is determined according to the main evolution direction represented by the vector, such as the behavior hyperedge sequence from the user interface node to the storage node, and then to the server node.

[0246] Step S720: Compare the possible behavior hyperedge sequence with the behavior hyperedge set in the dynamic risk association hypergraph, and identify the behavior hyperedge whose weight value exceeds the preset weight threshold as the candidate key hyperedge.

[0247] The preset weight threshold is a weight value set in advance for screening important behavior hyperedges. The candidate key hyperedge is a hyperedge with a larger weight value screened from the possible behavior hyperedge sequence, which may play a key role in risk propagation.

[0248] In the comparison, the possible behavior hyperedge sequence determined in step S710 is compared with the behavior hyperedge set in the dynamic risk association hypergraph one by one. For each behavior hyperedge, check whether its weight value exceeds the preset weight threshold. If it does, it is considered as a candidate key hyperedge. In the e-commerce platform, the possible behavior hyperedge sequence is compared with the behavior hyperedge set in the dynamic risk association hypergraph, and the hyperedge whose weight value exceeds the preset weight threshold, such as the hyperedge connecting the important server node, is considered as the candidate key hyperedge.

[0249] Step S730: Analyze the position of the candidate key hyperedge in the associated risk path, and determine the candidate key hyperedge located at the starting segment or the middle segment as the key behavior hyperedge to be blocked.

[0250] The key behavior hyperedge to be blocked is the hyperedge that needs to be blocked to prevent risk propagation. The candidate key hyperedge located at the starting segment or the middle segment has a greater impact on risk propagation, so it is determined as the key behavior hyperedge to be blocked.

[0251] In the position analysis, for each candidate key hyperedge, its position in the associated risk path is determined. According to the structure of the path and the characteristics of risk propagation, it is judged whether it is located at the starting segment or the middle segment. In the e-commerce platform, for the candidate key hyperedge, its position in the associated risk path is analyzed, and for the hyperedge located at the starting segment connecting the user interface node and the storage node or the hyperedge located at the middle segment connecting the storage node and the server node, it is determined as the key behavior hyperedge to be blocked.

[0252] Step S740: According to the time sequence constraint relationship in the set of timing constraints, the expected occurrence time point of the critical behavior hyperedge to be blocked is calculated as the blocking execution time point.

[0253] The blocking execution time point is the time point at which the critical behavior hyperedge blocking operation is performed. According to the time sequence constraint relationship in the set of timing constraints, the expected occurrence time point of the critical behavior hyperedge to be blocked can be accurately calculated.

[0254] When calculating the time point, the occurrence time information of the critical behavior hyperedge to be blocked is searched in the set of timing constraints. According to the time sequence constraint relationship, the expected occurrence time point is determined. In the e-commerce platform, for the critical behavior hyperedge to be blocked, the occurrence time stamp and other information are obtained from the set of timing constraints, and the expected occurrence time point is calculated as the blocking execution time point.

[0255] Step S750: The critical behavior hyperedge to be blocked and the corresponding blocking execution time point are combined to generate a path blocking strategy, and the path blocking strategy also includes a blocking method description, and the blocking method description is to interrupt the hyperedge connection or reduce the hyperedge weight.

[0256] The path blocking strategy is a strategy that includes the identification of the critical behavior hyperedge to be blocked, the blocking execution time point, and the blocking method description, which is used to guide the network security protection operation.

[0257] In generating the path blocking strategy, the identification of the critical behavior hyperedge to be blocked and the corresponding blocking execution time point are combined. At the same time, the blocking method description is determined, such as interrupting the hyperedge connection or reducing the hyperedge weight. In the e-commerce platform, the identification (such as number) of the critical behavior hyperedge to be blocked and the blocking execution time point (such as specific time) are combined, and the appropriate blocking method is selected, such as interrupting the hyperedge connection, to generate the path blocking strategy.

[0258] Step S800: Extract the interface nodes involved in the associated risk path, set the interface access frequency threshold according to the evolution rate in the long-term evolution trend vector, generate an interface access control rule, and the interface access control rule includes the interface identification, the access frequency threshold, and the violation handling method.

[0259] The interface access control rule is a rule for controlling interface access, and by setting the access frequency threshold, the interface can be prevented from being accessed too much, reducing the risk.

[0260] In generating the rule, first, the interface nodes involved in the associated risk path are extracted. Then, according to the evolution rate in the long-term evolution trend vector, the access frequency threshold of each interface node is determined. The faster the evolution rate, the faster the risk propagation may be, and a lower access frequency threshold needs to be set. The interface identifier, access frequency threshold and violation handling method, such as access restriction, warning, etc., are recorded for each interface node. In the e-commerce platform, the user interface nodes involved in the associated risk path are extracted, and the access frequency threshold is set for each interface node according to the evolution rate in the long-term evolution trend vector, and the interface access control rule is generated.

[0261] Step S900: The node protection priority, path blocking strategy and interface access control rule are integrated into a scheme, and a dynamic security protection scheme containing scheme effective time, execution subject and verification index is generated.

[0262] The dynamic security protection scheme is a comprehensive network security protection scheme, which contains measures such as node protection, path blocking and interface access control.

[0263] In the scheme integration, the node protection priority list generated in step S660, the path blocking strategy generated in step S750 and the interface access control rule generated in step S800 are combined. The scheme effective time is determined, such as starting to take effect from a specific time. The execution subject is clarified, such as a network security team or a security device. The verification index is formulated, such as the reduction degree of risk diffusion probability, the security improvement of nodes, etc., which is used to evaluate the effectiveness of the scheme. In the e-commerce platform, the node protection priority, the path blocking strategy and the interface access control rule are integrated, the scheme effective time, the execution subject and the verification index are determined, and the dynamic security protection scheme is generated to protect the network security and data security of the e-commerce platform.

[0264] It can be understood that the various algorithms involved in the above introduction of the embodiments of the present application can be known from related contents in the prior art, and in order to save space, the embodiments of the present application are not expanded too much. In addition, those skilled in the art can supplement details according to common knowledge in the art when implementing the scheme of the present application, for example, normalization can be used to eliminate the dimensional conflict before feature fusion, interpolation can be used to eliminate the dimensional difference, historical data, experience or business scene demand can be combined to reasonably set the threshold, the model can be trained based on the general model training method, the number of layers in the model structure can be set based on actual needs, the activation function can be selected, etc. The present application will not introduce the redundant implementation process in too much detail.

[0265] Please refer to Figure 2 , Figure 2A structural schematic diagram of a computer system provided by an embodiment of the present application is shown in FIG. 1. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, the communication interface 102, and the memory 103 can be connected by a bus or other means. The processor 101 (also referred to as a central processing unit (CPU)) is the computing core and control core of the computer system, which can parse various instructions in the computer system and process various data of the computer system. The communication interface 102 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and can be used to transmit and receive data under the control of the processor 101; the communication interface 102 can also be used for the transmission and interaction of data within the computer system. The memory 103 is a memory device in the computer system, used to store programs and data. It can be understood that the memory 103 herein can include a built-in memory of the computer system, and of course can also include an extended memory supported by the computer system. The memory 103 provides a storage space that stores an operating system of the computer system, and the present application does not limit this.

[0266] In an embodiment, the processor 101 executes the computer program in the memory 103 to perform the network security and data security comprehensive analysis method based on a large model provided by the above embodiment of the present application.

Claims

1. A large model-based network security and data security comprehensive analysis method, characterized in that, The method comprises the following steps: Collecting multi-source heterogeneous security data of network communication links, data storage nodes and application interaction interfaces, wherein the multi-source heterogeneous security data comprises encrypted transmission data packet units, data access audit units and interface call sequence units, each data unit carries a node identifier and an interaction time stamp; Performing risk behavior atomization correlation processing on the multi-source heterogeneous security data to construct a dynamic risk correlation hypergraph; Based on a preset security domain knowledge graph, calling a large model to perform multi-round risk attribution reasoning, matching attack chain fragments and completing evidence chains for a set of behavior hyperedges in the dynamic risk correlation hypergraph, and generating a risk attribution reasoning chain comprising a primary risk node, a correlation risk path and an evidence confidence; According to the risk attribution reasoning chain and a set of time sequence constraints of the dynamic risk correlation hypergraph, constructing a risk evolution probability model, and based on this, calculating a short-term diffusion probability and a long-term evolution trend vector of each correlation risk path to obtain a risk evolution path prediction result, which specifically comprises: extracting a correlation risk path and the evidence confidence of each link from the risk attribution reasoning chain, and extracting a time sequence constraint relationship corresponding to the correlation risk path from the set of time sequence constraints of the dynamic risk correlation hypergraph; Based on a sequence of behavior hyperedges and evidence confidences in the correlation risk path, constructing a path state transition relationship, wherein the path state transition relationship represents a possibility value of transitioning from one behavior hyperedge to the next behavior hyperedge, which is determined by the product of the evidence confidence of the current behavior hyperedge and the hyperedge weight; combining time interval information in the time sequence constraint relationship, adjusting the path state transition relationship to generate a time-aware state transition relationship; based on the time-aware state transition relationship, constructing a risk evolution probability model, wherein the risk evolution probability model comprises a short-term prediction layer and a long-term prediction layer, the short-term prediction layer is used to predict risk diffusion within a first preset time period in the future, and the long-term prediction layer is used to predict risk evolution within a second preset time period in the future; calculating a diffusion possibility value of each correlation risk path within the first preset time period in the future as a short-term diffusion probability through the short-term prediction layer; calculating an evolution direction and rate of each correlation risk path within the second preset time period in the future as a long-term evolution trend vector through the long-term prediction layer; integrating the short-term diffusion probability and the long-term evolution trend vector to generate a risk evolution path prediction result comprising a path identifier, a short-term diffusion probability and a long-term evolution trend vector.

2. The method of claim 1, wherein, The risk behavior atomization correlation processing on the multi-source heterogeneous security data to construct a dynamic risk correlation hypergraph comprises: Performing protocol reverse analysis on the encrypted transmission data packet units in the multi-source heterogeneous security data to generate communication behavior atomic features; Performing operation sequence decomposition on the data access audit units in the multi-source heterogeneous security data to generate access behavior atomic features; Performing call chain tracking on the interface call sequence units in the multi-source heterogeneous security data to generate interface behavior atomic features; mapping the communication behavior atomic features, the access behavior atomic features and the interface behavior atomic features into an entity node set of a dynamic risk correlation hypergraph, wherein the communication behavior atomic features are mapped into communication nodes, the access behavior atomic features are mapped into storage nodes, and the interface behavior atomic features are mapped into interface nodes; calculating co-occurrence frequencies and time sequence dependencies of the behavior atomic features between different entity nodes based on interaction time sequence stamps, and constructing a behavior hyperedge set connected to multiple entity nodes; generating a time sequence constraint set based on the interaction time sequence stamps of the multi-source heterogeneous security data, the time sequence constraint set being used to limit time sequence relationships of each hyperedge in the behavior hyperedge set; integrating the entity node set, the behavior hyperedge set and the time sequence constraint set into a hypergraph structure to generate a dynamic risk correlation hypergraph containing node attributes, hyperedge weights and time sequence constraints.

3. The method of claim 2, wherein, The calculation of the co-occurrence frequencies and the time sequence dependencies of the behavior atomic features between different entity nodes based on the interaction time sequence stamps, and the construction of the behavior hyperedge set connected to multiple entity nodes, include: extracting interaction time sequence stamp sequences of the communication nodes, the storage nodes and the interface nodes, and determining time points of occurrence of the behavior atomic features of each node; segmenting the time sequence stamp sequences by using a time interval division method, counting the number of combinations of the communication node behavior atomic features, the storage node behavior atomic features and the interface node behavior atomic features simultaneously appearing in each time interval, and obtaining the co-occurrence frequencies; labeling the behavior atomic feature combinations in each time interval in a time sequence order, and calculating conditional probabilities of the preceding behavior atomic features and the subsequent behavior atomic features as the time sequence dependencies; converting the co-occurrence frequencies into co-occurrence weights in a preset numerical interval by using a maximum value normalization method, and converting the time sequence dependencies into dependency weights in a preset numerical interval by using a maximum value normalization method; weighting and summing the co-occurrence weights and the dependency weights according to a preset proportion to obtain comprehensive weights of the behavior hyperedges; constructing a hyperedge structure containing two or more entity nodes according to the types and the number of the entity nodes in the behavior atomic feature combinations, assigning the comprehensive weights to the hyperedge structure, and generating the behavior hyperedge set.

4. The method of claim 2, wherein, The generation of the time sequence constraint set based on the interaction time sequence stamps of the multi-source heterogeneous security data includes: extracting occurrence time stamps of all the behavior hyperedges in the dynamic risk correlation hypergraph, establishing a hyperedge time stamp sequence, and sorting the hyperedge time stamp sequence in ascending order to determine time sequence orders of the behavior hyperedges; for any two behavior hyperedges sharing an entity node, if the occurrence time stamp of the preceding behavior hyperedge is earlier than the occurrence time stamp of the subsequent behavior hyperedge, a time sequence constraint relationship is generated; for behavior hyperedges containing the same entity node set, a time interval of the occurrence time stamps is calculated, and if the time interval is less than a preset minimum interval threshold, a time overlap constraint relationship is generated; structuring and coding all the generated time sequence constraint relationships and the time overlap constraint relationships to generate a time sequence constraint set containing hyperedge identification pairs and constraint types.

5. The method of claim 1, wherein, The preset security domain knowledge graph is called to perform multi-round risk attribution reasoning by a large model, attack chain fragment matching and evidence chain completion are performed on the behavior hyperedge set in the dynamic risk association hypergraph, and a risk attribution reasoning chain including a main cause risk node, an associated risk path and an evidence confidence is generated, including: Attack pattern descriptions, vulnerability feature descriptions and defense rule descriptions are extracted from the preset security domain knowledge graph, and the attack pattern descriptions are converted into structured behavior feature templates; The entity node set, the behavior hyperedge set and the time sequence constraint set of the dynamic risk association hypergraph are converted into a graph structure text description that can be parsed by the large model, which is used as the input context of the first round of reasoning; The large model is called to perform the first round of reasoning on the graph structure text description in combination with the behavior feature template, identify candidate attack chain fragments in the behavior hyperedge set that match the attack pattern description, and output the entity node sequence and the behavior hyperedge sequence included in the candidate attack chain fragment; The candidate attack chain fragment output by the first round of reasoning is used as the input of the second round of reasoning, the large model is called to perform vulnerability feature matching on the entity nodes in the candidate attack chain fragment in combination with the vulnerability feature description, and the entity nodes with vulnerability features are determined as potential risk nodes; The potential risk nodes and the corresponding candidate attack chain fragments are used as the input of the third round of reasoning, the large model is called to perform abnormality evaluation on the behavior hyperedge of the potential risk node in combination with the defense rule description, an abnormality evaluation result is generated, and the potential risk node with a high risk abnormality evaluation result is marked as a main cause risk node; The behavior hyperedge connection path in the dynamic risk association hypergraph is traced based on the main cause risk node, and an associated risk path including the main cause risk node, associated entity nodes and a behavior hyperedge sequence is generated; The product of the weight value of each behavior hyperedge in the associated risk path and the abnormality evaluation result is calculated as the evidence confidence of each link on the path; The main cause risk node, the associated risk path and the evidence confidence of each link are integrated in the reasoning order to generate a risk attribution reasoning chain.

6. The method of claim 5, wherein, The attack pattern description is converted into a structured behavior feature template, including: The text content of the attack pattern description is parsed, and the key behavior stages in the attack process are extracted, the key behavior stages being a behavior sequence arranged in chronological order; Each key behavior stage is matched with a corresponding entity node type, wherein the initial stage corresponds to a communication node, the intermediate stage corresponds to an interface node, and the subsequent stage corresponds to a storage node; The core behavior features of each key behavior stage are extracted, the core behavior feature of the initial stage being an unauthorized port connection attempt, the core behavior feature of the intermediate stage being an abnormal parameter construction call, and the core behavior feature of the subsequent stage being a permission boundary breakthrough operation; The key behavior stage, the entity node type and the core behavior feature are combined into a structured behavior feature template, and the behavior feature template includes a stage placeholder area, a node type placeholder area and a behavior feature placeholder area.

7. The method of claim 5, wherein, The entity nodes in the candidate attack chain fragment are matched with vulnerability features in combination with the vulnerability feature description by calling the large model, and the entity nodes with vulnerability features are determined as potential risk nodes, including: Extracting vulnerability-related information from the vulnerability feature description, each vulnerability-related information including an impact node type, a trigger behavior description, and a hazard level description; Identifying the type of each entity node in the candidate attack chain segment to determine the node type of each entity node, including a communication node, a storage node, and an interface node; According to the node type of the entity node, filtering out the vulnerability-related information matching the type from the vulnerability feature description to obtain a node type matching vulnerability information set; Extracting the behavior atomic features of the entity nodes in the candidate attack chain segment, and performing semantic similarity calculation on the behavior atomic features and the trigger behavior description of each vulnerability-related information in the node type matching vulnerability information set to obtain a behavior matching result; Marking the entity nodes corresponding to the vulnerability-related information with high similarity in the behavior matching result as entity nodes with vulnerability features, recording the vulnerability-related information and the hazard level description of the entity nodes, and generating a potential risk node list.

8. The method of claim 5, wherein, The behavior hyperedge connection path in the dynamic risk association hypergraph based on the primary risk node is traced to generate an associated risk path containing the primary risk node, the associated entity node, and the behavior hyperedge sequence, including: Taking the primary risk node as the starting retrieval point, retrieving the behavior hyperedge directly connected to the starting retrieval point in the dynamic risk association hypergraph to obtain a first-level associated hyperedge; For each first-level associated hyperedge, extracting other entity nodes connected by the hyperedge as first-level associated nodes, recording the weight value of the first-level associated hyperedge and the corresponding time sequence constraint relationship; Taking the first-level associated nodes as new starting retrieval points, repeating the above retrieval process until the boundary nodes of the dynamic risk association hypergraph are retrieved or the preset retrieval depth limit is reached to obtain multi-level associated hyperedges and multi-level associated nodes; According to the time sequence constraint relationship in the time sequence constraint set, sorting the multi-level associated hyperedges in time sequence to generate a behavior hyperedge sequence arranged in time sequence; Arranging the primary risk node and the multi-level associated nodes according to the connection relationship of the behavior hyperedge sequence to generate an associated risk path containing the node sequence and the corresponding behavior hyperedge sequence.

9. A computer system, characterized by It includes: A memory in which a computer program is stored; A processor for loading the computer program to implement the network security and data security comprehensive analysis method based on a large model according to any one of claims 1-8.

Citation Information

Patent Citations

  • Offshore platform water treatment system balancing method based on knowledge graph

    CN113341859A

  • Network security threat information early warning method and system based on big data

    CN120151080A